How it works
A durable pilot fixes five things in advance: the specific task, the current baseline cost and quality of that task, the metric and how it will be sampled, the acceptance threshold, and the owner of the evaluation set. Absent any of those, the review meeting becomes a discussion of anecdotes.
The second pattern is treating deployment as installation. Model output usually lands in the middle of a process built around human throughput and human review. If queues, handoffs and approval steps stay as they were, the measurable gain is absorbed by the surrounding workflow even when the model performs well.
Example
A support-summarisation pilot with 92% rater-approved summaries reads as a success until someone asks what changed. If agents still read the full thread because they cannot tell which summaries are the unreliable 8%, handling time is unchanged. The fix is not a better model; it is a confidence signal and a policy for when the summary may be trusted alone.
Why it matters
The gap between announced enterprise AI activity and production deployment is the central fact of this market. Distinguishing genuine production use from pilots and announcements is the only way to read vendor claims sensibly — and it is why we track adoption as a measured series rather than a headline.
