9 September, 2026•3 minute read

Where your agent spends its time and tokens

After you get an agent working, you’ll want to get it to run cheaply and quickly. There are four key metrics I track alongside task success rate1 to figure out where to allocate engineering effort.

Cache hit rate

This is probably the single most important metric to watch as it has a big influence on both cost (cached tokens are 10x+ cheaper than regular input tokens) and Time To First TokenHow long a language model takes to emit its first token after a request starts — the latency a user feels before text begins streaming, as distinct from total generation time. (through avoiding prefill). Measured as cached_input_tokens / total_input_tokens over the course of the run you generally want this to be as high as possible, although short-horizon tasks naturally benefit less from caching.

In addition to tracking this across eval runs, you also want to track it in production for anomalies. A sharp decrease immediately following a production deployment likely means you broke something and should investigate.

Note that KV caches are not all equal. An inference provider with lower per-token pricing can end up more expensive than alternatives if their cache management is worse. You need to measure overall run cost as part of your evals to figure out whether a new inference provider will actually be cheaper!

Input tokens per run

At the end of 2023 I noted that agentic workloads at Crimson consumed up to 22.5 input tokens per output token, and this ratio has only become more lopsided over time. Long-horizon agentic workloads can easily exceed a 100:1 input:output token ratio, which means that input tokens are the main cost driver for agent runs—even when you have a very high cache hit rate.

The reason for this is that cumulative input tokens across the run increase quadratically as the agent loops, while output tokens are typically flat per-step. The chart below models an agent with a 50:1 input:output token ratio per step over time, using the current GPT-5.6 Sol list price:

11020304050Completed steps$0.00$0.50$1.00$1.50$2.00$2.50Cost (USD)Input + cache …Cache readsOutputGPT-5.6 Sol — cumulative run cost
Cumulative cost by component over 50 steps, with retained history and full cache-prefix reuse.

Input tokens also have a pretty big impact on success rate, because the input tokens sent to the model are ultimately what it’s using to reason about the workload. For those two reasons, input tokens are what most agent engineers tend to spend most of their time on.

Output tokens per run

Output tokens are the biggest driver of end-to-end inference latency. Prefill can be parallelized, but token generation is completely sequential.

This was actually one of the reasons why we initially started generating output in YAML format at Crimson! We had a lot of interactive surfaces leveraging LLMs, and students disengaged when it took too long to get useful output. YAML being more token-dense than JSON meant we could shave precious seconds off our inference.

Output token counts are particularly useful to have when comparing eval runs. Oftentimes you’re dropping down to flex or batch processing to reduce eval costs, and this makes inference times far less predictable. Token counts tend to be a lot less noisy and are a great proxy measurement.

Inference:tool call time

Most people assume they’re bottlenecked by inference, but this is often not actually the case.

Our main browser use agent spends ~70% of its time running tools. Browsers are complicated pieces of software and the protocol for driving browsers programmatically (CDP) is low level, causing even simple tool calls to require multiple RPC round trips between the agent harness and the web browser. Inference improvements are always helpful, of course, but most of the engineering effort invested at Rye around agent latency is actually focused on optimizing the implementation of our tools.


Footnotes

  1. Maximizing one of these metrics in isolation is counterproductive if it comes at the cost of task success rate, of course. ↩

Don't want to miss out on new posts?

Join 100+ fellow engineers who subscribe for software insights, technical deep-dives, and valuable advice.

Get in touch 👋

If you're working on an innovative web or AI software product, then I'd love to hear about it. If we both see value in working together, we can move forward. And if not—we both had a nice chat and have a new connection.
Send me an email at hello@sophiabits.com