Reason Once, Deploy Forever: How to Cut Agent Costs 3–6× Without Losing Accuracy
A practical guide to skill distillation for production AI agents — what it is, when to use it, how to deploy it, and how to know it's working.
The problem: you're paying the reasoning tax on every call
If you run an AI agent in production — a voice bot, a customer service agent, a workflow automation — you've probably noticed a pattern. Reasoning-mode models (GPT-5 with reasoning_effort on, Claude with extended thinking, Qwen with enable_thinking) are noticeably better at multi-step tasks. They're also 3–6× more expensive in output tokens per episode, and that premium is paid every single time the agent runs.
Now read a few of those reasoning traces. What you'll find is uncomfortable: most of the "thinking" isn't solving the specific problem in front of the model. It's the model re-deriving the same domain procedures it derived on the previous call. And the call before that. And the one before that.
The retail support agent reasons its way (again) to "I shouldn't call the account lookup tool until the customer has actually given me an email." The household agent re-discovers (again) that "heat X" is one atomic command, not a sequence of microwave-door operations. Every episode, from scratch, forever.
This is exactly what amortization is for. A recent paper from Microsoft (Reason Wide, Not Deep) shows you can pay that reasoning premium once, offline, by distilling the recurring procedural knowledge into a compact markdown file and injecting it into the cheaper non-reasoning model's system prompt.
The results across four agentic benchmarks: skills recover 55–100%+ of the reasoning gap, sometimes beating the reasoning model outright, at 2.7–6× fewer output tokens and zero reasoning tokens per episode. Distillation cost: roughly $1–$3 in coding-agent time, once per domain.
This post is a practitioner's guide to actually using it.
The process: how skill distillation works
Three steps. No fine-tuning, no RL, no test-time search.
Step 1 — Collect a trajectory corpus
Take 35–50 tasks from a training split and run your non-reasoning model on them. Save everything: observations, actions, tool calls, tool outputs, user turns, final outcome. That's your corpus.
You don't need reasoning-mode traces. The paper's ablation is unambiguous on this: skills distilled from non-reasoning trajectories alone are competitive with skills distilled from paired reasoning/non-reasoning corpora. In some domains (SpreadsheetBench) the non-reasoning-only corpus produces a better skill — plausibly because reasoning narratives anchor the distiller on what the model believed rather than what actually happened in the environment.
For most production teams this step is free. You already store these logs for QA, compliance, or debugging.
Step 2 — Distill with a coding agent
Point an off-the-shelf coding agent (Claude Code, Codex, similar) at the directory containing the corpus. Give it a fixed instruction: analyze the trajectories, identify recurring failure modes, compile a compact rulebook.
The agent then does what a decent engineer would do — writes and runs its own analysis code to compute error-type frequencies, action n-grams, loop detection, and win/loss contrasts between failing and succeeding episodes. It reads individual transcripts where the statistics point somewhere interesting. It compiles findings into 40–130 lines of imperative markdown rules, each traceable to transcript evidence.
The rules are concrete, not platitudes. Example from the paper's retail skill:
Before calling
find_user_id_by_email, check that the customer's message actually contains a real email address. This bug appeared in 13 of 22 rollouts (59%) and accounted for 17 of 18 tool errors (94%).
Total cost of this step: $1.28–$2.44 per domain in the paper's experiments. One pass. No optimization loop.
Step 3 — Deploy
Append the markdown skill verbatim to the non-reasoning model's system prompt. Nothing else changes — same harness, same decoding, same tools, same model.
Because the skill lives in a fixed prompt prefix, prompt caching (Anthropic, OpenAI, and Google all support it) makes the marginal per-call cost effectively zero. You pay for the prefix tokens once when the cache is warm.
When to use it (and when not to)
Use it when:
- Your agent runs a narrow, stable domain — customer service for a specific product line, appointment booking, a specific workflow. Failure modes repeat.
- You're running at volume. Per-episode token savings compound. If you handle 1,000 episodes a day, the arithmetic is trivially favorable. If you handle 10, it isn't.
- Latency matters. Reasoning tokens add hundreds of milliseconds to first-response time. In voice agents this shows up as awkward silence.
- You already have logs. You almost certainly do.
Don't bother when:
- The domain is genuinely open-ended (general research assistant, general coding agent). There's no repeating failure distribution to mine.
- Each instance genuinely requires novel deliberation. The paper's residual gaps on telecom (long instance-specific dependency chains) and SpreadsheetBench (one-off logical structures) show where per-instance reasoning is still doing irreplaceable work.
- Your bottleneck is in a different layer. For voice agents whose failures are dominated by ASR/TTS quality, a smarter system prompt won't help.
- Your model provider changes often. Skills are model-specific and cross-model transfer is untested.
How to deploy: a concrete playbook
For a text/chat agent
- Export 35–50 recent conversations with mixed outcomes (resolved, escalated, abandoned). Include the full trace: system prompt, user turns, model turns, tool calls, tool outputs, final resolution status.
- Drop them in a directory. One file per conversation, structured (JSON or markdown), with the outcome clearly labeled.
- Open a coding agent in that directory. Give it a distillation prompt: "Analyze these agent trajectories. Identify recurring failure modes with frequencies. Produce a markdown rulebook of imperative, concrete rules for the non-reasoning agent to follow. Cite corpus evidence for each rule."
- Review the output. If rules are generic ("be polite," "confirm understanding"), your corpus lacks signal — go back and enrich it with outcome data. If rules are specific and cite frequencies, proceed.
- Append the skill to your production non-reasoning agent's system prompt.
- Ensure prompt caching is enabled and hitting.
- A/B test against your current baseline.
For a voice agent
Same pipeline with three modifications:
- Join transcripts to outcome data before distilling. Voice transcripts alone often don't reveal outcomes — did the caller hang up angry? Did a human have to fix it? Did they call back within 24 hours? Skills distilled from calls-plus-outcomes are dramatically better than skills distilled from transcripts alone.
- Expect voice-specific rules. "Never read a 16-digit number without pausing every 4 digits." "If the caller interrupts, stop mid-sentence and yield." "Confirm spelled-out policy numbers back before calling the lookup tool." These don't appear in text-agent benchmarks but dominate voice quality.
- Verify cache hit rates in your voice stack. Voice orchestrators (LiveKit Agents, Pipecat, Vapi, Retell) don't all pass through caching directives cleanly. A skill that isn't cached is a latency regression.
For an agent embedded in a broader workflow
If your agent is one step in a larger pipeline (e.g., inside a document-processing workflow), the skill goes in the agent step's system prompt. The rest of the pipeline is unchanged.
How to check if it's actually working
The paper measures win rate. In production you should measure more.
The minimum bar (task success)
Run a controlled A/B on held-out episodes:
- Arm A: non-reasoning model, current prompt.
- Arm B: non-reasoning model, current prompt + skill.
- Arm C (reference): reasoning model, current prompt.
Track success rate on your existing evaluation set. You want B to close most of the A→C gap. If B doesn't beat A, either your corpus lacks signal or your domain isn't a good fit.
The economic bar (tokens and latency)
Measure:
- Mean output tokens per episode (B should be at or below A, and 3–6× below C).
- P50 and P95 latency to first meaningful action.
- Cache hit rate on the skill prefix.
If tokens don't drop, something is wrong with your deployment — the model may be ignoring the skill, or your harness may be re-prompting each turn.
The quality bar (business metrics)
Task success on a benchmark isn't the same as customer outcomes. Add:
- Voice/chat: containment rate, average handle time, CSAT, escalation rate, callback rate within 24 hours.
- Workflow agents: end-to-end completion rate, human intervention rate, downstream error rate.
A skill can improve task-level success while making conversations longer, more robotic, or more brittle in ways benchmarks miss. Watch for it.
Guardrails against regression
- Track failure-mode distribution before and after. The paper shows the ALFWorld skill dropped "missed transform" failures from 35.9% to 11.5% and stall loops from 28.7% to 5.3%. You want the same visibility — which failure modes shrank, which grew.
- Watch for new failure modes. Adding rules can over-constrain a model that was previously fine. The paper's Qwen retail result (−4.2 points with the skill on an already-competent baseline) is the warning shot.
- Redistill on a cadence. Traffic shifts. New products, new promotions, seasonal issues. A skill distilled in January describes January's failure distribution. Rerun monthly or quarterly — it costs a couple of dollars.
Commercial impact: what this actually saves
Consider a plausible production scenario: a customer service voice agent handling 100,000 calls per month, averaging 8 turns per call, running on a reasoning-mode frontier model at roughly $10 per million output tokens.
Before: reasoning mode, ~2,000 output tokens per call including reasoning. That's 200M output tokens per month, or ~$2,000/month in output token cost — before considering the compounding effect of latency on abandonment rate.
After: non-reasoning mode + skill, ~500 output tokens per call. That's 50M output tokens per month, or ~$500/month. A 75% reduction in output-side inference cost. One-time distillation cost: ~$3.
Now add the second-order effects. Voice agents lose customers to silence. Every 200ms of first-response latency shaved off improves containment rate measurably. Zero reasoning tokens means faster first action. The token savings are the floor; the retention improvement is often the ceiling.
Scale this up to enterprise volume — millions of calls per month, multiple domains — and the annualized savings run into six or seven figures for what is, mechanically, a couple of hours of engineering work per domain plus periodic redistillation.
The comparison against active prompt optimizers is worth noting too. Techniques like GEPA and MIPROv2 need fresh evaluation rollouts against a metric — often hundreds of calls per optimization step. The paper shows their distillation approach beats GEPA on both τ² domains at ~4× lower production cost, and unlike GEPA it can run on logs you already have, without any live rollouts.
The mental model: wide search vs deep search
The paper's cleanest framing: reasoning is deep search inside one episode — the model explores a tree of considerations before each action, and its cost recurs every time. Distillation is wide search across many episodes — regularities in the failure distribution are extracted once, and the result is reused for free.
Where the required knowledge is episode-invariant (domain procedures, tool-calling discipline, action grammars), width dominates. A rule aggregated over 50 episodes is more reliable than a derivation the reasoning model has to reproduce correctly every time — which is why the skilled non-reasoning model sometimes beats the reasoning model outright.
Where the required knowledge is genuinely per-instance (long dependency chains, one-off logical structures), width can't help. Reasoning stays necessary.
The practical implication is that the two are complements, not substitutes. Use skills to stop re-buying the invariants. Reserve reasoning for the instances that genuinely need it. You can even build hybrid routers: cheap skilled agent by default, reasoning-mode fallback triggered by uncertainty signals.
Getting started this week
If you want to try this on a real agent:
- Pick your narrowest, highest-volume domain.
- Pull 50 recent episodes with mixed outcomes. Join to outcome data.
- Drop them in a folder. Point a coding agent at it with the distillation prompt.
- Review the output. If the rules cite evidence and look concrete, proceed. If they read like a LinkedIn post about "AI best practices," fix the corpus first.
- A/B test the skill in your staging environment against your current non-reasoning baseline and your current reasoning baseline.
- If B closes most of the A→C gap, ship it. If not, iterate on the corpus.
Total time investment on a first attempt: a day or two. Total money: a few dollars. Upside: the reasoning premium becomes a one-time cost instead of a recurring one, for as long as your domain is stable.
That's the whole pitch. Reason once. Deploy forever.
Last updated 2026-08-18
Frequently asked questions
1. "Our voice agent's per-call cost is killing our unit economics. Do we have to switch to a cheaper model and take the accuracy hit?"+
No. Switching from reasoning to non-reasoning mode on the same model, plus a distilled skill, typically recovers most of the accuracy at 3–6× lower output-token cost. You keep the model, drop the reasoning flag, add the skill. That's the whole change.
2. "We tried a cheaper model and quality dropped. Now what?" for voice agenst+
The cheaper model probably lacks specific procedural discipline your reasoning model was compensating for turn by turn. Distill a skill from a corpus of the cheaper model's failed episodes — the failure distribution tells you exactly what discipline is missing. Inject it into the cheaper model's prompt and re-test.
3. "Our reasoning model is fast enough, but we're worried about the token bill scaling with usage. Is this preemptive optimization?"+
If your usage is stable and low, yes — don't bother. If you're on a growth curve where token costs will 10× in the next year, this is a two-hour investment that permanently changes your cost slope. The break-even math is roughly: distillation pays for itself after the skill saves as many tokens as it cost to produce, which for most production agents is a matter of hours to days.
4. "The same agent handles multiple domains (billing, tech support, retention). Does this still work?"+
Distill separately per domain. The paper's skills are per-domain, and mixing domains dilutes the failure distribution. In practice you route incoming conversations to the right domain first (you probably already do), then apply the domain-specific skill in the system prompt for that route.
5. "Our agent's failure modes change every time we ship a product update. Won't the skill get stale?"+
Yes. Redistill on a cadence tied to your release cycle — after a major product update, or monthly for high-churn environments. It's cheap ($1–$3), so err on the side of more frequent redistillation. Track failure-mode distribution as a leading indicator — when the top failure modes shift, it's time to redistill.
7. "How is this different from just writing better prompts by hand?"+
Two differences. First, the distiller mines statistics — which failure modes are actually frequent, not which ones you remember from spot-checking. Humans systematically under-weight boring high-frequency bugs and over-weight vivid rare ones. Second, the distiller writes rules with citations to evidence, which makes review easy and disagreements resolvable ("this rule is here because bug X appeared in 59% of rollouts" — you can check the claim). Hand-written prompts drift into folklore.
8. "How is this different from RAG or a knowledge base?"+
RAG retrieves content at query time — documents, policies, product info. Skills encode procedure at prompt-load time — how the agent should behave, in what order, with what guardrails. They're orthogonal. Most production agents want both: RAG for the facts the model shouldn't hallucinate, skills for the behavior the model should exhibit consistently.
9. "Our compliance team needs to review anything that changes agent behavior. Is this reviewable?"+
Yes, and this is actually a strength versus fine-tuning. The skill is a plain markdown file, human-readable, with each rule citing corpus evidence. Compliance can read it, redline it, approve it, version it in git. Compare to fine-tuning, where behavioral changes are opaque weight updates.
10. "Does this work for agents that use MCP tools or connect to external services?"+
Yes — the skill mainly encodes when and how to call tools, what preconditions to check, what error patterns mean. Tool-heavy agents are actually where the paper's results are strongest (the retail authentication bug is a tool-calling discipline issue). If anything, MCP-heavy agents benefit more from distillation because tool-call errors are exactly the kind of concrete, high-frequency failure mode the distiller mines cleanly.
11. "Can we distill a skill from one model and use it on another?"+
Untested by the paper, and skills are technically model-specific. In practice, skills that encode domain procedure (as opposed to model-specific quirks) tend to transfer, but expect to redistill when you switch model families. Budget a couple of dollars per model per domain.
12. "What if the skill makes the agent worse?"+
It happens — the paper's Qwen retail result is a −4.2 point regression, likely from over-constraining an already-competent model. This is why you A/B test before rolling out. If the skill regresses, three common fixes: (a) shorten it — long skills over-constrain; (b) restrict it to failing task types via a router; (c) accept that this domain doesn't benefit and skip it. Not every domain is a fit.
Reason Once, Deploy Forever: How to Cut Agent Costs 3–6× Without Losing Accuracy
Sign up free and get $0.98 in credit — no card required. Connect your number, pick a template, and go live in minutes.