Why your prototype's defaults are quietly bankrupting you
Every AI agent is a loop around a model. It plans, calls a tool, reads the result, and reasons again. A single completed outcome can take a dozen model requests. That is why the number your CFO cares about is cost per successful outcome, not the price of a token.
Most teams build the same way: pick the strongest model, stuff the prompt with everything the model might need, confirm it works. That instinct is correct for a prototype. The problem is what happens next — those prototype defaults quietly become the production architecture.
Two things break at that point:
- AI workloads are not uniform. A single app mixes intent classification, extraction, formatting, summarization, and genuine multi-step reasoning. Routing all of them to one frontier model means overpaying on the majority of requests that never needed that capability.
- One outcome is many requests. A prototype pays for a single call. An agent pays for the whole loop, so waste gets multiplied. An agent that takes a wrong turn burns tokens on turns that should never have happened — and still lands on a weaker answer.
In production, the goal is not to minimize tokens. It is to reduce the cost of a successful outcome while maintaining quality, safety, and latency. The economics of a request come down to four decisions you control at runtime.

The four levers, side by side
| Lever | What it controls | Foundry capability |
|---|---|---|
| Models and offers | Which model runs, where it runs, how throughput is purchased | Model router, deployment types, provisioned throughput, batch, fine-tuning |
| Caching | What gets reused across turns and sessions | Prompt caching, semantic caching via AI Gateway in Azure API Management |
| Prompt & agent optimization | How much context and how many turns each outcome requires | Prompt optimizer, agent optimizer (instructions, skills, tool descriptions, model selection) |
| Observability & evaluation | Whether the saving is real and quality held | Foundry observability, agent traces, Azure budgets, alerts, cost tagging |
1. Route each request to the right model
Routine requests should not pay frontier-model economics. Complex requests should not sacrifice quality to save tokens. A model router assesses each incoming request and dispatches it to the most suitable underlying model in real time, behind a single endpoint. Routing modes let you prioritize cost, quality, or a balance of the two, and built-in failover buys you resilience for free.
Deployment type matters just as much, and it is the lever teams most often leave untouched:
- Standard — pay-as-you-go, greatest flexibility.
- Priority processing — for interactive apps that need consistent latency.
- Provisioned Throughput Units (PTUs) — for high-volume, predictable demand.
- Batch — up to 50% lower cost for async work like document processing and large-scale classification.
Fine-tuning is the advanced version of this lever. Where routing picks among existing models, fine-tuning changes what a smaller model can do — teaching it your task, tone, or format well enough to match a larger model on that specific job. Reach for it when behavior is stable and volume is high enough to earn back the effort.
2. Stop paying for the same tokens twice
Agents are highly cache-effective. The same system instructions, tool schemas, and policy text are re-sent on every turn — an agent that takes 10 turns pays for that prefix 10 times. Prompt caching lets a previously processed prefix be reused rather than reprocessed. Cache reads are billed at a discount on standard deployments and can be discounted up to 100% on provisioned deployments.
The rule for prompt architecture is straightforward: stable content first, volatile content last.
- Top: system instructions, tool definitions, few-shot examples.
- Bottom: user input, retrieved chunks, turn history.
Caching depends on an exact match at the start of the prompt. Anything that changes per request — a timestamp, a user's name — has to sit below that stable block. Put it at the top and the cache never matches.
3. Optimize the prompt, then optimize the agent
If model choice sets the rate, the instruction sets the volume — and it is the cheapest thing to fix because it ships without touching infrastructure. The practices that cut tokens are the same ones that improve answers: lead with the task, be specific about output format and length, and use a few well-chosen examples instead of paragraphs of explanation. Then keep cross-turn accumulation under control:
- Summarize completed conversations instead of replaying full transcripts.
- Scope tool definitions to only the tools relevant to the task.
- Store working state in external memory and retrieve it only when needed.
Automated optimizers can now run your agent against a dataset of real tasks, generate candidate configurations, score each one, and rank them so you can promote the winner. The dataset can come directly from your own agent traces.
4. Make it visible with observability and evaluation
You cannot tune what you cannot see, and you cannot claim a saving you did not measure. Two numbers matter:
- Cost per request — did the cheaper path still clear the bar?
- Cost per completed outcome — what did the business actually pay, across every turn and retry?
An optimization that lowers the first while raising the number of turns has made things worse, and only the second metric will show it. Keep a standing evaluation set that every optimization has to clear before it ships, and pair it with budgets, alerts, and cost tagging so a regression arrives as a notification rather than a surprise at month end.

Where this goes wrong (and how to avoid it)
A few caveats that rarely make it into vendor blogs:
- Routing is not free. Every routing decision adds a small classification hop. If your latency budget is tight, measure the router's own overhead before assuming it is a pure win.
- Cache invalidation is a real cost. Semantic caching across sessions can serve stale answers when your underlying data changes. Tune TTLs aggressively and never cache anything user-scoped or policy-scoped.
- Fine-tuning is a trap for fast-moving products. If your prompts, tools, or output formats change weekly, a fine-tune will be stale before it pays back. Only fine-tune behavior that has been stable for months.
- Cost-per-outcome is a lagging metric. By the time the number moves, you have already shipped the regression. Pair it with per-request token histograms and cache hit rates so you catch drift early.
- Batch is not a magic discount. It only works for work that genuinely does not need a synchronous response. Forcing an interactive feature onto Batch to save 50% will cost you 10x in user trust.
What to learn next
Once you have the four levers instrumented, the natural next step is to close the loop:
- Build a standing evaluation set from real agent traces, not synthetic prompts. This becomes the ground truth every future optimization must clear.
- Measure cache hit rate as a first-class SLO, not a nice-to-have. A sudden drop usually means someone added a volatile field to the top of the prompt.
- Experiment with fine-tuning on one stable, high-volume task before rolling it out broadly — it is the only lever here that is expensive to reverse.
- Study how inference infrastructure shapes cost. The same workload can cost 3–5x more depending on deployment topology. For a concrete production case of building a multimodal system at scale, see our deep dive on building a production-grade multistage multimodal recommender system on Amazon EKS.

The bottom line
None of these levers is a one-time saving. Together they form a loop that gets cheaper and better every time you go around it:
- Model and offer decides where each request runs, and fine-tuning turns a proven task into a permanently cheaper one.
- Caching lowers the cost of every cycle, which is what lets you run the loop often enough to matter.
- Prompt and agent optimization generates the next candidate and proves it against your evaluation set.
- Observability and evaluation tells you where you are and whether the last change held.
The climb is a cycle with no fixed start, though most teams enter it at measurement. Traces become evaluation datasets. Those datasets drive the optimizer. Optimizer results show which tasks are stable enough to fine-tune. Fine-tuned models change what the router should choose — and the new routing produces fresh traces.
Start by deploying a model router and comparing it against your current baseline. Then instrument cost per completed outcome before you touch anything else. You cannot optimize what you cannot see.
If you are also evaluating how gateway-level infrastructure shapes your AI costs, our breakdown of integrating Recraft's advanced image models with the Vercel AI Gateway is a useful companion piece.
Source: Microsoft Azure Blog — The Economics of Agent Optimization