Failure Note - Oracle gate

The metric improved. The outcome didn't.

Offline accuracy can rise for quarters while the business result stays flat.

By Jay Sharma · September 20, 2026

The claim

A pilot whose evaluation metric improves every quarter is not necessarily approaching production. It may only be approaching its own test set.

How it fails

The team optimizes the number it can measure: accuracy on a labeled set, resolution rate inside a test harness, reviewer agreement in a sample. The business cares about a different number - cost per resolved ticket, claims processed without rework, revenue protected. When the evaluation was never tied to that outcome, the metric becomes the project. The eval set drifts toward what the model does well. The people writing the test are the people shipping the model, and nobody is rewarded for making the test harder to pass.

Warning signs

  • The headline metric has improved for two consecutive reviews and the business KPI has not moved.
  • Nobody can say which business number the evaluation is a proxy for, or why it is a fair one.
  • The evaluation set was written by the same team shipping the model, and it has never been audited against production inputs.
  • Evaluation runs on clean inputs that production never sees - perfect documents, polite users, complete data.

What it costs

Quarters of spend against a number that does not pay. The pilot "succeeds" into a larger commitment - more budget, more headcount, more surface area - on evidence that cannot survive contact with the real workflow. When the truth lands, it lands as a write-down instead of a decision.

What would change the answer

An evaluation tied to a business outcome, run against production-like data, with a pass threshold agreed before the next funding decision. If the metric moved and the outcome followed, this note does not apply to you.

Illustrative composite - method walkthrough, not a client result

A support copilot raised draft-resolution accuracy from 71% to 88% across two releases. Cost per resolved ticket stayed flat, because the evaluation measured the quality of drafts, and agents were still rewriting half of them. The gate that failed was the oracle: the test measured the artifact, not the outcome.

Where this maps

This failure is a Oracle gate failure. The PROOF Scorecard tests this gate in about four minutes; the full review proves it with evidence.