AI infrastructure • 001 | 5 October 2026 | Inference
From FLOPS to outcomes.
What should production AI actually optimize?
Maddipalli Gopalakrishna · AI / ML Engineer
45 seconds · music only · numbers on screen are illustrative
We spend a lot of time talking about how powerful AI hardware is. Production AI raises a different question: how efficiently are we converting compute into useful work?
The industry is already moving past raw compute. Qualcomm describes its new data center products as optimized for tokens-per-watt “as the key lever to reduce total cost of ownership”, and NVIDIA argues the infrastructure metric is “fast shifting from peak performance to validated agentic tokens per megawatt”. I think the useful move is to go one or two layers further.
The metric ladder
Each of these measures a different layer. None replaces the one below it.
- Level 1 · Compute capabilityFLOPS
- How much arithmetic the hardware can do, peak or sustained depending on how it's reported. Peak FLOPS isn't application performance: memory bandwidth, utilization and the workload decide what you get.
- Level 2 · Inference throughputTokens / second
- How fast a serving setup produces tokens. Only meaningful with its conditions attached: model, hardware, batch size, input and output lengths, serving configuration.
- Level 3 · Energy efficiencyTokens / watt
- Throughput relative to power. A useful efficiency measure that still says nothing about whether the answer was right.
- Level 4 · System economicsCost / task
- What one complete workflow costs: model calls, retrieval, tool and API calls, retries and the supporting infrastructure.
- Level 5 · Outcome efficiencySuccessful tasks / compute
- How much correct, useful work the system produces for the resources it consumes. A conceptual application metric, not a standard benchmark.
Why agents change the equation
A chat request is roughly one model call. An agent request looks more like this:
- Request
- Plan
- Retrieve
- Reason
- Tool
- Observe
- Re-plan
- Tool
- Verify
- Response
One user request can turn into several model calls, retrievals, tool invocations and the occasional retry. Tokens, power, latency and cost accumulate across all of them. If you only measure tokens per second, an agent that loops through four extra large-model calls and a retry looks productive. Cost per task, and how many tasks actually succeed for the compute spent, expose it.
A user request is not one inference call.
Not every decision needs the biggest model
A system that produces fewer tokens is not worse, and a frontier model is not the right tool for every step. Route each decision to the cheapest component that can do it correctly:
| Kind of step | Better handled by |
|---|---|
| Complex reasoning, planning, ambiguity | A capable reasoning model |
| Classification, extraction, routing | A smaller model |
| Knowledge lookup | Retrieval |
| Repeated requests | A cache |
| Strict rules, validation, calculations | Deterministic code |
| High-impact uncertainty | Human review |

Optimize the system for the outcome
Instead of only pushing tokens per second up, it is worth also tracking task success, latency, energy, cost, retries and unnecessary model calls, while holding quality, reliability and safety fixed. The interesting optimization target may not be the fastest model. It may be the architecture that produces the correct outcome using the least unnecessary computation.
Don't optimize the model in isolation. Optimize the system for the outcome.
How are you measuring cost per completed task in your agent systems today, if at all?