Not for want of a better model. Three scenarios below, each drawn from a pattern that shows up again and again, and in all three the thing that was broken was a measurement, not a model.
Two techniques recur across all three, because they are what separates a demo from a system: multi-model councils, where independent models from different families check each other and their disagreement becomes the signal, and human-in-the-loop designed rather than assumed: deciding explicitly when a person enters, what they see, and crucially where their judgment is written back to.
They had a written bar every AI change had to clear before shipping, and no way to tell whether a change cleared it. For eleven months the gate had approved everything. Building a golden set from real traffic, splitting graders by what they can actually judge, and setting two thresholds instead of one.
Read the case study →The dashboard said the support copilot deflected 42% of contacts. The metric could not distinguish a solved problem from an abandoned customer, and improved whenever the bot got harder to escape. The honest number was 19%, and finding that out is what made the real gains possible.
Read the case study →Fourteen months after launch, operations still reviewed every prediction by hand. The accuracy was real, measured on a balanced test set describing a distribution that does not exist, with the errors concentrated in the two classes the business cared about most. Fixed without retraining anything.
Read the case study →Is the number you are reporting measuring the thing you think it is measuring? The Pilot-to-Production Audit answers that for your pilots in two to three weeks, in writing, with a verdict on each. The triage call before it is thirty minutes and free, and plenty of callers leave it knowing they don't need an audit.
Book a triage call