AI in engineering companies · Coding agents

What tokens are, and why they matter

Anyone who uses coding agents seriously soon meets the word "token". It appears on the bill, in the limits, and in every discussion of why an agent slowed down or lost track of what it was doing. Yet I find that many engineers and most managers are not quite sure what a token is, or why it matters so much.

It matters because tokens are the unit in which AI does its work. They determine what it costs, how fast it runs, how much it can consider at once, and, less obviously, how well it performs. Understanding them is the foundation for managing coding agents well, which is the subject of the next two articles in this series.


What a token is

A language model does not read text as letters or words. It reads it as tokens: chunks of text that the model has learned to treat as units. A common short word is usually one token. A longer or less common word may be split into several. Punctuation, spaces and symbols are tokens too.

As a rough guide for English prose, a token is about three-quarters of a word, or around four characters. So a thousand words is somewhere around thirteen hundred tokens. The exact figures vary between models, because each has its own way of dividing text.

Code is usually less efficient than prose. Symbols, brackets, indentation, long identifiers and unusual names all take tokens. A page of code typically uses more tokens than a page of English of the same length.

Everything is counted: the instructions the agent is given, the files it reads, the output of the commands it runs, the conversation so far, and everything it writes in reply.


Input and output

Tokens come in two kinds, and the distinction matters.

Input tokens are everything the model reads: instructions, code, documents, tool results and the history of the task.

Output tokens are everything the model writes: its answer, its code, and, for models that reason before they answer, its internal reasoning.

Providers usually charge for both, per million tokens, and output tokens typically cost several times as much as input tokens. Reasoning is charged as output even though you may never see it.


The context window

Every model has a limit on how many tokens it can consider at once: its context window. Everything the agent needs to know for the current step, including its instructions, the relevant code, what it has done so far and the latest results, has to fit within it.

Context windows are now large, often hundreds of thousands of tokens and sometimes more. That sounds like plenty. A substantial codebase is far larger, and an agent working through a long task accumulates history quickly.

When the window fills, something has to give. Older material is summarised or dropped, and the agent may lose track of a decision it made earlier or a constraint it was told about. This is one reason long agent sessions drift from the agreed design, as I described in Coding agents: strengths, weaknesses and opportunities.

There is a subtler point too. Even within the limit, many models become less reliable at using information buried in a very long context. More context is not always better. The right context is.


Why agents consume so many tokens

A single question to a model is read once and answered once. A coding agent works in a loop, as described in What an agent actually is, and isn't, and at every step it reads again: its instructions, the task, the relevant files, what it has done so far, and the latest results.

A simple illustration, using round numbers rather than any particular model or price. Suppose an agent takes forty steps to complete a task. At each step it reads, on average, sixty thousand tokens of accumulated context, and writes a thousand tokens. Over the task, it reads about 2.4 million input tokens and writes about forty thousand output tokens.

Two things stand out. First, reading dominates: well over ninety per cent of the tokens are input. Second, the total is large for what may feel like a modest piece of work. Multiply by a team of engineers running agents all day and the numbers become significant.


Why this matters

Cost. Tokens are what you pay for. Understanding where they go is the only way to control the bill. I discuss how in Managing a token budget.

Speed. Processing more tokens takes longer. An agent that reads less at each step responds faster.

Quality. An agent given a focused, relevant context usually does better work than one given everything. Clutter competes for attention.

Limits. Providers impose rate limits in tokens. Teams that do not manage consumption hit them at inconvenient moments.

Hardware. If you run models on your own hardware, as discussed in The cost and infrastructure of agents, the time the machine takes to process long inputs is often the limiting factor for agent work.

Privacy. Every token sent to a hosted model is data leaving your organisation. The files an agent reads to do its work are sent to the provider as tokens. That is the subject of Keeping your code local and private.


Two things that reduce the cost

Caching. Many providers offer prompt caching: when the beginning of the input is the same as a recent request, as it usually is in an agent loop, the repeated part is charged at a substantial discount and processed faster. Agent tools that are structured to take advantage of caching cost much less to run.

Choosing the model for the job. Smaller models cost far less per token. For many steps in coding work, such as searching, summarising and routine edits, they are perfectly adequate. I discuss this in Local, cloud and frontier: getting the most from your budget.


Four things worth taking seriously

For engineering leaders: tokens are the unit of cost, speed and capacity for AI. Make sure the people using coding agents understand them.

For engineers: what the agent reads costs more than what it writes. Give it the right context, not all of it.

For finance: expect agent consumption to be dominated by input, and to grow with the length and number of tasks.

For everyone: more context is not always better. Focused context is cheaper, faster and usually produces better results.


I would be interested to hear whether your teams know how many tokens their coding agents consume, and what the biggest consumers are.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.