The metric improved. The outcome didn't.
Offline accuracy can rise for quarters while the business result stays flat.
By Jay Sharma · September 20, 2026
The claim
A pilot whose evaluation metric improves every quarter is not necessarily approaching production. It may only be approaching its own test set.
How it fails
The team optimizes the number it can measure: accuracy on a labeled set, resolution rate inside a test harness, reviewer agreement in a sample. The business cares about a different number - cost per resolved ticket, claims processed without rework, revenue protected. When the evaluation was never tied to that outcome, the metric becomes the project. The eval set drifts toward what the model does well. The people writing the test are the people shipping the model, and nobody is rewarded for making the test harder to pass.
Warning signs
- The headline metric has improved for two consecutive reviews and the business KPI has not moved.
- Nobody can say which business number the evaluation is a proxy for, or why it is a fair one.
- The evaluation set was written by the same team shipping the model, and it has never been audited against production inputs.
- Evaluation runs on clean inputs that production never sees - perfect documents, polite users, complete data.
What it costs
Quarters of spend against a number that does not pay. The pilot "succeeds" into a larger commitment - more budget, more headcount, more surface area - on evidence that cannot survive contact with the real workflow. When the truth lands, it lands as a write-down instead of a decision.
What would change the answer
An evaluation tied to a business outcome, run against production-like data, with a pass threshold agreed before the next funding decision. If the metric moved and the outcome followed, this note does not apply to you.
A support copilot raised draft-resolution accuracy from 71% to 88% across two releases. Cost per resolved ticket stayed flat, because the evaluation measured the quality of drafts, and agents were still rewriting half of them. The gate that failed was the oracle: the test measured the artifact, not the outcome.
Where this maps
This failure is a Oracle gate failure. The PROOF Scorecard tests this gate in about four minutes; the full review proves it with evidence.
Related notes
The demo works because nobody gave it system access.
A sandbox demo proves the model can talk about the work, not that the system can do it.
Floor'Human in the loop' is not an operating model.
Without a trigger, an authority, and a response time, the human boundary exists on paper only.
Bring one pilot and one decision.
Thirty minutes on one stuck pilot: the decision being avoided, the five production gates, and the first one that lacks evidence. If a full review is not useful, I will say so.
Book a private pilot triage