Free tool

Agent Insurability Scorecard

Twenty questions, grouped the way an underwriting questionnaire groups them rather than the way an engineer would. The point is to rehearse the conversation in the vocabulary the other side will use.

Score

0%0 / 40

0 of 20 answered.

Inventory

Is there a current list of every agent running in production?

An underwriter cannot price exposure they cannot enumerate. This is the question everything else depends on.

For each agent, is it recorded which tools, APIs, and datasets it can reach?

Reach is the blast radius. An agent with database write access is a different risk from one that can only read a wiki.

Is the inventory updated automatically, rather than by someone remembering?

A manual inventory is accurate on the day it is written. Renewal is annual; drift is continuous.

Identity

Does each agent have its own identity, distinct from a human user's?

Agents running under a developer's credentials are indistinguishable from that developer in every log you own.

Can an action taken by an agent be attributed to a specific agent, version, and triggering request?

Attributability is what turns an incident into a claim you can substantiate.

Are agent credentials rotated and scoped with an expiry?

A long-lived token held by an autonomous process is the failure mode carriers have started asking about by name.

Least privilege

Is each agent's access restricted to what its task actually requires?

Convenience-scoped credentials are the single most common finding in agent security reviews.

Are destructive or irreversible actions gated behind an explicit control?

Deletion, payment, and outbound communication are the three that produce claims.

Is there a deny-by-default policy on tool access rather than allow-by-default?

The default determines what happens when someone adds a tool and forgets to think about it.

Logging and audit

Is every tool call logged with its inputs, outputs, and outcome?

Reconstructing an incident from application logs alone is the reason investigations take weeks.

Are those logs tamper-evident, or at least write-once?

An audit trail an insider can edit is not evidence.

Are logs retained long enough to cover a late-discovered incident?

Thirty-day retention against a breach discovered in month four means the answer is permanently unknown.

Human oversight

Is there a defined set of actions that require human approval before execution?

Both regulators and underwriters ask this one, in almost the same words.

Is the approval enforced in the system, rather than being a documented expectation?

A policy that depends on the operator remembering it is a policy that has already failed once.

Spend and rate control

Are there hard spend caps enforced at runtime?

Runaway spend is the most frequent agent incident and the easiest to bound.

Do circuit breakers stop an agent after a set number of tool calls per task?

A cap that does not depend on anyone noticing is the only one that works at three in the morning.

Incident response

Is there a kill switch that stops a specific agent in production?

Redeploying to disable an agent is not incident response.

Has an agent-specific incident been rehearsed in the last year?

The first time the kill switch is used should not be the first time it is tested.

Assessment

Has an AI risk assessment been completed in the last twelve months?

This appears close to verbatim on renewal questionnaires. It is worth being able to answer yes with a document.

Is that assessment mapped to a recognised framework such as NIST AI RMF?

Framework alignment is treated as a safe harbour under at least one US state statute, and it is the vocabulary assessors already read.

Runs entirely in your browser. Nothing you type here is transmitted, logged, or stored anywhere. There is no server-side copy because there is no transmission. Close the tab and it is gone.

How this works, and what it does not cover

Each answer scores two points for yes, one for in part, zero for no. The score is not a rating anyone recognises — it is a completeness measure, and its only job is to show you which group of questions you cannot answer.

The groups are ordered by dependency rather than by importance. Inventory comes first because every question after it assumes you can enumerate what you are being asked about.

Underwriting depends on your sector, your claims history, and a dozen factors beyond agent controls, so no score can promise a premium. What this one gives you is the part you control: whether every question on the form is answerable with evidence today — and a clear list of what to fix first where the answer is not yet.

The framework language is drawn from NIST AI RMF, which is the vocabulary assessors already read and which at least one US state statute treats as a safe harbour.

The carrier-facing one-pager template: your score turned into the document format an underwriter can actually read.

No cookies, no tracking pixels, never shared or sold. How the data is handled.

If you would rather not do this yourself

Agent Governance & Insurability Package

Three to four weeks: the technical controls plus the evidence pack a carrier, an auditor, or an enterprise customer will actually accept.

3-4 weeks · $18-30K
All free tools