Depilot.
Depilot · Case studies

How AI pilots actually fail.

Not for want of a better model. Three scenarios below, each drawn from a pattern that shows up again and again, and in all three the thing that was broken was a measurement, not a model.

Two techniques recur across all three, because they are what separates a demo from a system: multi-model councils, where independent models from different families check each other and their disagreement becomes the signal, and human-in-the-loop designed rather than assumed: deciding explicitly when a person enters, what they see, and crucially where their judgment is written back to.

Illustrative scenarios My client work is confidential and the practice is new, so none of these is a named engagement. Each is a composite built from patterns I have seen repeatedly across twenty-five years of shipping systems: at Indeed, Amazon, Agoda, Microsoft, PayPal and Bloomberg. The companies are invented. The failure modes, the diagnostics and the fixes are real, and they are described in enough detail to be useful whether or not we ever speak.
Evaluation · Series C SaaS · 7 weeks

Eleven months of releases. Nothing was checking them.

They had a written bar every AI change had to clear before shipping, and no way to tell whether a change cleared it. For eleven months the gate had approved everything. Building a golden set from real traffic, splitting graders by what they can actually judge, and setting two thresholds instead of one.

Read the case study →
Customer support · 55k contacts/mo · 3 weeks

42% deflection. The honest number was 19%.

The dashboard said the support copilot deflected 42% of contacts. The metric could not distinguish a solved problem from an abandoned customer, and improved whenever the bot got harder to escape. The honest number was 19%, and finding that out is what made the real gains possible.

Read the case study →
Classification · Specialty insurance · 2 weeks

97% accurate. Every prediction still checked by hand.

Fourteen months after launch, operations still reviewed every prediction by hand. The accuracy was real, measured on a balanced test set describing a distribution that does not exist, with the errors concentrated in the two classes the business cared about most. Fixed without retraining anything.

Read the case study →

The same question sits under all three

Is the number you are reporting measuring the thing you think it is measuring? The Pilot-to-Production Audit answers that for your pilots in two to three weeks, in writing, with a verdict on each. The triage call before it is thirty minutes and free, and plenty of callers leave it knowing they don't need an audit.

Book a triage call