Most failures announce themselves. A service goes down, an error appears, someone gets a page at 2 a.m.
Drift doesn't. Nothing breaks. The dashboards stay green. The outputs still read well, the agent still sounds confident, and the numbers still add up. They're just slowly becoming wrong.
That is what makes drift dangerous. Fluent systems fail fluently.
Three kinds of drift
Models drift. The model behind an API can be updated while its name stays the same. Prompts that behaved one way last quarter behave slightly differently this quarter. Your team didn't change anything, and nothing in your own change log says otherwise.
Data drifts. The population a system serves changes. An upstream application starts filling a field differently. A new product line appears, a new region comes online, a season turns. The system was tuned on a world that has quietly moved on.
Meaning drifts. This is the one that gets the least attention, and in our view it causes the most damage.
Organizations run on definitions. What counts as an "active" customer. What "high risk" means. Which codes belong in which category. Where a policy threshold sits. How a regulation is interpreted this year. What a department is called and what it's responsible for.
Those definitions change all the time, usually for good reasons, and usually by a decision made somewhere else in the organization. Finance redefines a customer segment. A code set is revised. A committee moves a threshold. A regulator publishes new guidance.
An AI system that learned yesterday's meaning keeps applying it, with exactly the same confidence as before. Nobody told it. No accuracy metric catches it, because its answers are perfectly consistent. They're consistently answering a question the business stopped asking.
A fourth kind is arriving: agents that change themselves
Until recently, the thing you approved was the thing that ran. That assumption is weakening.
Agents increasingly carry memory between sessions, adapt their approach based on what worked before, and in research settings rewrite their own code. In 2025, researchers at Sakana AI and the University of British Columbia published work on a coding agent that improved itself by modifying its own code. They also reported that in one experiment the agent faked a log to make it appear that tests had been run and passed, when they hadn't.
That is a research result, run under careful controls, and the researchers caught it because they kept a transparent, traceable record of every change the agent made. The lesson for everyone else is simple: when behavior can change from the inside, "we approved it at launch" stops being a description of what's running today.
Why the cost is so high
Drift isn't expensive because it's hard to fix. Once found, the fix is often small. It's expensive because of what happens before anyone finds it.
It compounds. Decisions made on a drifted meaning become the inputs to other decisions. A misclassified customer receives the wrong offer, then the wrong follow-up, then appears in the wrong report the board relies on.
It's discovered late. Drift tends to surface through a complaint, an audit finding or a quarter-end number that doesn't reconcile, weeks or months after it began.
It's hard to unwind. The first question after discovery is always the same: since when? If you can't answer that precisely, the only safe response is to re-examine every decision the system made, which is exactly the work the system was supposed to save.
The difference between a small correction and a large remediation is almost always whether you can say which decisions were made under which version of the model, the data and the definitions.
Questions worth asking now
We won't go into how drift is detected here. The more important questions come first, and they're organizational, not technical:
- When you approved this system, do you know exactly what you approved: which model version, which data, which definitions?
- Who owns each definition the system depends on, and would that person know the system depends on it?
- If any of those changed tomorrow, who would find out, and how soon?
- For any decision the system made last month, could you say which versions it was made under?
- If the agent can learn or adapt, is there a record of how it has changed that the agent itself can't alter?
If those questions have clear answers, drift becomes a routine maintenance event. If they don't, it becomes an incident, just one that hasn't been discovered yet.
Drift is normal. Not knowing isn't.
Drift is not a failure of AI. It's a property of the world. Models are updated, data changes, and organizations redefine things because they're learning and adapting too.
The failure is building systems that can't tell when the ground has moved beneath them, and can't say afterward which decisions were affected.
In the next article: the agents that never drifted from your approved version, because you never approved them at all.