Shipping AI Features That Survive Production
Evals first, so quality is measured rather than vibed. The delivery pattern for AI features that keep working after the demo — and the four places they come apart.
Contents8 sections
A demo runs the happy path once, in front of a friendly audience, on an input somebody chose. Production runs every path, thousands of times, on inputs nobody anticipated — and the gap between those two is where most AI features quietly stop being worth their maintenance. This is a delivery pattern for closing it, written around the parts that are usually skipped.
Why do so few AI features show up in the numbers?
Adoption is not the bottleneck. McKinsey's 2026 State of AI survey puts AI use at 88% of organisations and generative AI specifically at 72%, while only 37% report any positive EBIT contribution and roughly 6% attribute more than 5% of EBIT to it — essentially unchanged on 2025 despite another year of spending. That is self-reported survey data and should be read as a direction rather than as anyone's odds, but the direction has not moved in two years.
Adoption is near-universal; measurable value is not
Share of organisations
Self-reported. The gap between the second and third bars is the one worth designing against, and it is mostly a measurement problem before it is a modelling one.
McKinsey, The State of AI in 2026: On the road to ROI (August 2026)
What does evals-first actually mean?
It means the harness exists before the feature does. A small labelled set of realistic inputs with known-good outputs, a way to run the current system against it, and a score you can compare across changes. That is it — it is far less sophisticated than it sounds and it is the difference between engineering and guessing.
The sequence, and why this order
Write down what good looks like, in examples
Twenty to fifty real inputs with the output you would accept. Not a rubric — examples. This is the artefact everything else is measured against, and collecting it usually changes the product spec, which is the first sign it was worth doing.
Build the harness before the feature
Something that runs the set and reports a score. It will be embarrassing at first and that is fine; the value is that the second score is comparable to the first. Once the feature exists, nobody builds this, because now it can only deliver bad news.
Ship the narrowest useful version
One workflow, one user type, real traffic. Breadth before reliability is how features end up quietly switched off.
Instrument cost and quality per feature from day one
Token spend attributed to the thing that caused it, not one monthly number nobody can act on. Retrieval hit rate, retry counts, and the eval score on a schedule rather than on a whim.
Where does the money actually go?
Not where teams expect. The unit price of intelligence has collapsed — Stanford HAI's AI Index tracked the cost of reaching GPT-3.5-level performance on MMLU falling from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a reduction of more than 280 times. Bills still climb, because the things that multiply against that price climb faster.
| Driver | Why it grows | The lever |
|---|---|---|
| Call volume | Success. More users, more calls, often more calls per user | Cache aggressively; not every request needs a model |
| Context resent every turn | Conversations accumulate and the whole history goes back each time | Context compaction, and summarise rather than replay |
| Retries | A failed parse or a timeout silently doubles a request | Budget retries explicitly and log them as their own metric |
| One model for everything | The model chosen on day one for the hardest task serves the easiest ones too | Route by task; most requests do not need the expensive model |
How do agents fail in production?
Four ways, consistently. None appears in a demo, because a demo runs the happy path once. All four are cheap to prevent at design time and expensive to retrofit, which is the argument for deciding them before the first agent ships rather than after the first incident.
The four to engineer against
Loops
The agent calls a tool, gets an unexpected result, and calls it again. With no retry budget and no circuit breaker it continues until the invoice notices. Cap attempts per tool per run, at design time.
Invented tool calls
It calls a tool that does not exist, or invents arguments for one that does. A tool registry plus strict schema validation on every call catches this — but both have to be built, and usually are not until a user finds the gap.
Context bloat
Every step appends messages and every message costs tokens, so a long-running agent gets slower and more expensive the longer it works. Compaction is the difference between a flat cost curve and one that bends upward mid-task.
Goal drift
Long agents optimise the last sub-task and forget the original one. Periodic re-grounding against the user's stated intent fixes it, and almost nobody adds it before watching an agent confidently complete the wrong job.
Who can the feature read data as?
This is the question that gets skipped, and it is an access-control question rather than an AI one. The scope belongs in the query, before retrieval — never in the prompt. An instruction telling a model which records it may mention is not an access control; it is a request, and the model is not the enforcement layer.
When is the answer not to use a model?
More often than the current market implies, and this is worth establishing before any of the above. A model is the right tool when the input is genuinely unstructured, the output tolerates variation, and a wrong answer is recoverable. Change any of those three and a query, a rule or a form is better — faster, cheaper, testable, and explainable to a regulator.
- The rule is knowable. If a person could write the decision down as logic, write it down as logic.
- The output must be exactly right. Anything arithmetic, legal or financial wants a deterministic path with the model kept out of the number.
- You cannot define good. If the eval set cannot be written, the feature cannot be measured, and an unmeasurable feature cannot be improved.
- The failure is not recoverable. Where a wrong answer causes irreversible harm, the model can advise but must not act.
Common questions
How big does an eval set need to be?
Smaller than people expect. Twenty to fifty genuinely representative cases will catch most regressions, and a small set that exists beats a comprehensive one that is still being planned. Grow it from real failures — every production complaint becomes a case, which is also how the set stays representative as the product changes.
Do we need RAG, or just a bigger prompt?
A bigger prompt, if the whole corpus fits and is stable. Retrieval earns its complexity when the corpus is larger than the context window, changes frequently, or must be filtered per user. The third reason is the one people forget, and it is usually decisive for anything multi-tenant — the SaaS agency shortlist goes through how retrieval crosses tenant boundaries.
How do we stop the model making things up?
Ground it in retrieved material, cite what it used, and give it a permitted way to say it does not know — most hallucination in practice is a model with no acceptable escape route. Then measure it: the hallucination rate on your eval set is a number, and if nobody is tracking it the answer is that you do not know.
Should we fine-tune?
Rarely at first. Fine-tuning helps with format and tone, and it is a poor and expensive way to add knowledge. Exhaust prompting and retrieval before taking on the operational burden of a training pipeline, because that burden is permanent.
How do we choose a model?
Against the eval set, per task, and then revisit it — this is a decision with a short shelf life. Route rather than standardise: most requests in a real product are easy, and paying the expensive model to answer them is the largest avoidable line on most inference bills. If you are evaluating an agency on this, the AI development agency shortlist covers what to ask.
The short of it
Write the eval set, build the harness before the feature, ship one narrow workflow to real traffic, and instrument cost and quality per feature. Put the tenant scope in the query rather than the prompt, cap what an agent may do, and keep asking whether this needed a model at all. That sequence is most of the difference between a feature that survives and one that gets quietly switched off.



