Every AI evaluation seems to start with the same number. The model is 95% accurate. The agent gets it right 97% of the time.
That number matters. A system that is often wrong shouldn't be anywhere near a real decision.
But accuracy is where the evaluation should start, not where it ends. In regulated work, and in any work where decisions affect people, a single accuracy figure hides most of what you need to know.
Accuracy is necessary. It is not sufficient.
Not all errors cost the same
An accuracy score treats every mistake as equal. The world doesn't.
Consider a fraud model. A false positive blocks a legitimate payment: an irritated customer and a call to the service desk. A false negative lets a fraudulent payment through: a direct loss, and sometimes a regulatory report. Two errors, very different costs.
Now consider a screening tool that decides which cases get a closer look. A false positive sends a case for review that didn't need it. A false negative means a case that needed attention never gets it. The second error may not surface for months, and when it does, the person harmed wasn't the one who chose the system.
Or an eligibility decision for a benefit. A wrong approval costs money. A wrong denial can cost a family its rent.
In each case, the two kinds of error fall on different people, at different times, with different consequences. A model can improve its overall accuracy while getting worse at the error that matters most. Unless you measure the errors separately, and put a cost on each, you won't see it.
Rare events make accuracy meaningless
When the thing you're looking for is rare, accuracy becomes almost useless as a measure.
If one transaction in a hundred is fraudulent, a model that approves everything is 99% accurate. It is also worthless. The more serious and uncommon the event, the more a headline accuracy figure flatters a system that has learned to ignore it.
A right answer you can't explain is still a problem
In many regulated decisions, being right isn't enough. You have to be able to say why.
Lenders in the US already have to give applicants specific reasons when credit is denied. Public agencies have to explain decisions that affect benefits and services. Employers are increasingly expected to show how automated tools reached a hiring recommendation. None of these obligations is satisfied by "the model is 96% accurate."
A decision you can't explain is a decision you can't defend, can't correct, and can't let the affected person challenge. That remains true even when the decision happens to be correct.
What to measure instead
Accuracy stays on the list. It just needs company. When we evaluate an AI system that makes or shapes decisions, these are the questions we want answered alongside it:
- What does each kind of error cost, and who bears it? Weight the errors, not just the count.
- How does it perform on the cases that matter most? Overall performance can hide poor performance on a specific segment: a region, a product line, a group of customers, the rare but serious case.
- Does its confidence mean anything? When the system says it is 80% sure, is it right about 80% of the time? A well-calibrated system tells you when to trust it.
- Does it know when to stop? The most valuable behavior in an uncertain case is often to hand the decision to a person rather than guess.
- Can each individual decision be explained? Not the model in general, but this decision, for this person, today.
- Does it give the same answer to the same question? We'll come back to this one in the next article.
- Will it still perform next quarter? Accuracy measured at launch describes launch day. Later in this series we'll look at what happens when the world moves underneath a system that doesn't.
A better first question
"How accurate is it?" is a reasonable question to open with. It's a poor one to decide on.
A better one: when it's wrong, who pays, and would we know?
If the answer is clear, accuracy becomes a useful number. If it isn't, a high accuracy score is mostly telling you how confident to feel, not how safe you are.
In the next article: why determinism still matters in AI systems, and why it isn't an argument against AI.