Across 40 enterprise LLM deployments observed over six months, measured output quality declined a median of 14% within twelve weeks of launch, with no corresponding change to the model. Launch-time evaluation detected none of it. Continuous evaluation against a blended golden-plus-production sample detected the decline a median of 9 weeks earlier than user-reported issues.
- Median measured quality fell 14% within 12 weeks of launch, with the model unchanged.
- Launch-time evaluation detected none of the observed decline.
- Continuous evaluation flagged drift a median of 9 weeks before users reported issues.
Why drift is silent
Drift is rarely a model event. It is the slow divergence between the distribution a system was evaluated on at launch and the distribution it actually faces in production as inputs, users, and upstream data change. Because the model file never changes, nothing triggers a re-test — and the decline accrues unobserved.
The result is a system that appears stable on its launch metrics while degrading against the only thing that matters: the live traffic it serves today.
Method
We instrumented 40 deployments with a continuous evaluation harness scoring a blended sample: a fixed golden set plus a rolling sample of real production traffic. Scores were collected weekly and compared against both the launch baseline and user-reported issue rates over a six-month window.
- Output quality against a blended golden-plus-production sample, scored weekly.
- Divergence between launch-baseline metrics and live-sample metrics.
- Lead time between harness-detected decline and user-reported issues.