Skip to content
PD
Claude Code

Claude Code Limits: Context, Spend and How to Save

The three different limits in Claude Code, where context actually goes, which techniques really cut spend, and how to notice a limit approaching before work stops.

All articles in the guide Claude Code · 13

The word “limit” means three different things here, and confusing them costs both money and time.

The three limits

The context window. How much text the model holds in view within one conversation. It runs out from the inside, and shows up as the agent forgetting the beginning.

The plan limit. How much you may use per period on a subscription. It runs out from the outside, and shows up as requests failing until the window resets.

The per-response limit. How much the model can emit at once. It shows up as output cut off mid-way, most visibly when you ask for a large file in one go.

Different symptoms, different fixes: the first is session hygiene, the second is your plan or an API key, the third is splitting the task.

Where context actually goes

Almost none of the spend is your messages. Your messages are a few percent.

  • Files that were read. The agent opened a thousand-line file to fix one function, and the whole file is now in context.
  • Command output. Build logs, test output, repository search results. One badly chosen command with verbose output eats more than your entire conversation.
  • Edit history. Every change made stays in the conversation.
  • Always-on project files. The project description is read at the start of every session. If it has grown to fifteen hundred lines, you pay for it in every task.

Hence the main conclusion: saving context is about managing what the agent reads, not how you phrase things.

Techniques that work

In descending order of effect:

  1. One session, one task. The most effective habit by far. Done means clear the context. Continuing a long conversation costs more and works worse.
  2. Give a precise address. “Look at the parsing function in this file” rather than “figure out why the parser is broken”. The second sends the agent through half the repository to find something you already knew.
  3. Do not ask for large files to be dumped into chat. Have the agent edit the file, not print it.
  4. Keep the project description short. It should hold what is always needed. Move procedures into skills, which load when relevant.
  5. Delegate reconnaissance to a subagent. Searching a large repository returns three pages of output. A subagent reads all of that in its own context and returns a short answer.
  6. Match the model to the task. Routine work does not need the most expensive model - see choosing a model.
  7. Compact deliberately. Compacting beats hitting the wall, but it is worse than starting clean: some detail is lost for good.

Noticing the limit approaching

Three signals that it is time to stop:

  • The context indicator is around three quarters full. That is the moment to finish a step, not to begin one.
  • The agent repeats work it already did, or rereads a file it read ten minutes ago. Early history has been compacted away.
  • Answers have gone generic. The specifics left the context, and the model is answering from general knowledge instead of your project.

The right response to all three is to lock the result into code and a commit, then start a new session. Uncommitted work at the end of a long session is the worst possible combination: context runs out exactly when losing it costs the most.

Next

Model choice and its effect on spend has its own article. What happens to context in agent systems generally is covered on the blog: the agent context budget. The overview is in the Claude Code guide.

FAQ

Why does Claude Code get worse towards the end of a session?

Context filled up and part of the early history was compacted into a summary. Details are lost while the agent keeps working as though it remembers them. This is the main argument for short sessions: a clean context on a new task almost always beats continuing a long conversation.

How do I know how much context is left?

The tool shows how full the context is, and it deserves the same attention as a fuel gauge. With a quarter left, finish the current step and start a new session rather than launching another large task in the same window.

What costs more, one long conversation or several short ones?

The long one. Every request resends the whole accumulated context, so each step gets more expensive as the conversation grows. Three short sessions for three tasks cost less than one session where the same tasks ran back to back.

More on this topic