A model-backed feature is not finished at launch in the way a form is. Its inputs shift, the provider changes the model underneath you, and users learn to game it. We review each one monthly.
Four numbers
- Correction rate — how often a human changes the output, by field. Rising means the feature is drifting or the inputs changed.
- Cost per operation — against the cap set during design. Rising quietly is the normal failure.
- Latency at the 95th percentile — averages hide the experience that makes people stop using a feature.
- Fallback rate — how often the manual path was used, and why.
Two questions
Is anyone still using it? And would we build it again knowing what we know now? Features that fail either question get removed rather than maintained. Removing a feature is a legitimate outcome of a review and it happens more than people admit.
An AI feature nobody trusts still costs money every time it runs.
Version pinning and re-testing
Provider models change. We pin versions where the platform allows it, keep a small golden set of inputs with known-good outputs, and re-run it before adopting any new version. Without that set, "the model got worse" is a rumour rather than a finding.
