The demo works because nobody gave it system access.
A sandbox demo proves the model can talk about the work, not that the system can do it.
By Jay Sharma · September 20, 2026
The claim
Most impressive demos run in a sandbox: curated data, a friendly interface, and a person quietly carrying outputs into the real system. The demo proves the model can talk about the work. It says nothing about whether the system can do the work.
How it fails
Production value lives in the integration layer: reading from systems of record, writing back without a human copying and pasting, holding permissions correctly, surviving latency and rate limits. None of that is visible in a demo. Teams postpone integration because it is slow and political, and each postponed quarter makes the demo more polished and less true.
Warning signs
- The demo runs on a snapshot export, not a live system.
- Every end-to-end run has a person in the middle carrying data between systems.
- "Integration" has been next quarter's work for more than one quarter.
- Nobody has written down the list of systems the agent must read from and write to, with the permissions each requires.
What it costs
The demo earns continued funding on evidence that cannot survive production. The integration bill arrives at the worst moment - after the executive team has already promised the outcome.
What would change the answer
One thin slice running against real systems of record with real permissions, however narrow. A pilot that does one real task end to end beats a demo that pretends to do fifty.
A claims-review agent demoed beautifully on exported CSVs. Connected to the live claims system, it could read 40% of the fields it needed and write none of them. The "pilot" had been a demo for eleven months. The gate that failed was reach: the system could not touch the actual work.
Where this maps
This failure is a Reach gate failure. The PROOF Scorecard tests this gate in about four minutes; the full review proves it with evidence.
Related notes
The metric improved. The outcome didn't.
Offline accuracy can rise for quarters while the business result stays flat.
Floor'Human in the loop' is not an operating model.
Without a trigger, an authority, and a response time, the human boundary exists on paper only.
Bring one pilot and one decision.
Thirty minutes on one stuck pilot: the decision being avoided, the five production gates, and the first one that lacks evidence. If a full review is not useful, I will say so.
Book a private pilot triage