Drift is quiet
A degrading agent does not throw errors. It keeps answering, in the same confident register, slightly more often wrong. The first signal is usually behavioral: escalation rates creep up, or the team quietly stops trusting a feature and routes around it. By then you have lost months of value and a good deal of goodwill.
Three instruments that catch it early
A client-specific benchmark set. Fifty to two hundred real questions from your domain with verified answers, run on a schedule. Generic benchmarks tell you nothing about whether the agent still knows your refund policy.
A feedback loop with somewhere to go. Users can rate and correct outputs, and those corrections land in a reinforcement pipeline rather than a spreadsheet nobody reads.
Escalation-rate monitoring. The cheapest early warning available, and most organizations already have the data.
The governance question nobody answers up front
Who approves a retrain? When a benchmark drops, someone has to decide whether to retrain, roll back, or accept it. If that person is not named before launch, the decision defaults to nobody and the answer defaults to nothing.
Takeaways
- Drift shows up as rising escalations, not errors.
- A client-specific benchmark set is the only meaningful accuracy check.
- Name the person who approves retraining before launch, not after the first drop.
