In our last article, we looked at how AI risk is moving out of general liability cover and onto the organizations that deploy AI. The practical consequence arrives at renewal: underwriters writing AI coverage want evidence that your AI is governed before they offer terms.
That raises a narrower question than it first appears. It isn't "do we do AI governance?" It's: what evidence can we actually produce, and can the platforms we rely on produce it?
This article is general commentary for business and risk leaders. It is not insurance or legal advice.
Monitoring is not evidence
Many AI platforms were built to help teams watch what their AI is doing: dashboards, activity feeds, alerts. That's valuable for operations. It's not the same as evidence.
A dashboard tells you what happened. Evidence lets someone else confirm it, independently, months or years later. An underwriter, an auditor or opposing counsel won't rely on a screenshot of your vendor's interface. They need records in a format they can read, with integrity they can check.
The eight questions below are designed to tell the difference quickly. Put them to every AI platform on your shortlist, and to the ones you already use. Including us.
Eight questions
1. Can you export the record of what our agents did in an open, published format? A good answer is a demonstration: an exported record in a standard format such as OpenTelemetry or CloudEvents, loaded into a common security or logging tool without custom work. A weak answer is a proprietary format, a PDF report, or "it's on the roadmap."
2. How would we prove to an auditor that those records haven't been changed? A good answer explains how records are made tamper-evident when they are written, and how an outside party can verify them. "They're stored securely" and "access is role-based" describe access controls, not integrity.
3. If the AI systems being governed were compromised, would the governance records be compromised too? Records produced and stored in the same environment as the AI they describe share its weaknesses. A good answer shows a separate trust boundary: different infrastructure, different credentials, different access paths.
4. Show us an action that was stopped before it happened, and the record of why. Underwriters want to see prevention, not just detection. A good answer walks through a specific action that was blocked or held for approval before it reached a system, with the policy and the decision recorded. "We alert on violations" means the violation already happened.
5. For an agent that ran six months ago, can you reconstruct what it saw, what it did and why, at a specific moment? Claims and investigations regularly look back a year or more. A good answer is a live reconstruction: the inputs, the steps, the policy in force, the model used and the output. A weak answer is a short log retention window, or "we keep everything" without showing how to retrieve it.
6. What happens when an agent's behavior changes between runs? Agents increasingly carry memory, adapt, and in research settings rewrite their own code. A good answer shows that rules are checked against what the agent actually tries to do each time, and that the record is produced somewhere the agent can't alter. A weak answer assumes the agent you approved is the agent that's running.
7. How does your evidence map to what our broker's AI questionnaire asks? Underwriters ask about inventory, controls, incidents, human oversight and change management. A good answer maps each category to specific records the platform produces. "Our audit trail should satisfy any underwriter" leaves the mapping to you.
8. If we leave your platform in year three, what happens to our evidence? Claims and regulatory retention periods often outlast vendor relationships. A good answer is that your records already live in storage you control, in open formats that work without the vendor. "We'll help you migrate" is not the same thing.
Reading the answers
No platform needs to be perfect on all eight. But the pattern of answers is revealing. Questions 1, 2 and 4 are the foundation: without them, there's little an underwriter can rely on. Questions 3, 5 and 6 test design decisions that are very hard to add later. Questions 7 and 8 tell you whether the evidence will still be useful when you need it most.
If your current platforms answer poorly, you have two honest choices: accept the insurance consequences as a documented risk decision, or add evidence production to your requirements and widen your shortlist. What you shouldn't do is discover the gap during a claim.
The same evidence, many audiences
These questions come from the insurance conversation, but the evidence isn't only for underwriters. The same records answer your auditors, your board's risk committee, regulators asking about automated decisions, customers' due-diligence questionnaires, and your counsel responding to a legal hold.
We built Operon so that this evidence is produced as the work runs, in your cloud or ours, in formats you can take anywhere. We're glad to walk any risk or insurance team through our own answers to all eight questions.