Claude Code Limits: Context, Spend and How to Save
The three different limits in Claude Code, where context actually goes, which techniques really cut spend, and how to notice a limit approaching before work stops.
All articles in the guide Claude Code · 13
The word “limit” means three different things here, and confusing them costs both money and time.
The three limits
The context window. How much text the model holds in view within one conversation. It runs out from the inside, and shows up as the agent forgetting the beginning.
The plan limit. How much you may use per period on a subscription. It runs out from the outside, and shows up as requests failing until the window resets.
The per-response limit. How much the model can emit at once. It shows up as output cut off mid-way, most visibly when you ask for a large file in one go.
Different symptoms, different fixes: the first is session hygiene, the second is your plan or an API key, the third is splitting the task.
Where context actually goes
Almost none of the spend is your messages. Your messages are a few percent.
- Files that were read. The agent opened a thousand-line file to fix one function, and the whole file is now in context.
- Command output. Build logs, test output, repository search results. One badly chosen command with verbose output eats more than your entire conversation.
- Edit history. Every change made stays in the conversation.
- Always-on project files. The project description is read at the start of every session. If it has grown to fifteen hundred lines, you pay for it in every task.
Hence the main conclusion: saving context is about managing what the agent reads, not how you phrase things.
Techniques that work
In descending order of effect:
- One session, one task. The most effective habit by far. Done means clear the context. Continuing a long conversation costs more and works worse.
- Give a precise address. “Look at the parsing function in this file” rather than “figure out why the parser is broken”. The second sends the agent through half the repository to find something you already knew.
- Do not ask for large files to be dumped into chat. Have the agent edit the file, not print it.
- Keep the project description short. It should hold what is always needed. Move procedures into skills, which load when relevant.
- Delegate reconnaissance to a subagent. Searching a large repository returns three pages of output. A subagent reads all of that in its own context and returns a short answer.
- Match the model to the task. Routine work does not need the most expensive model - see choosing a model.
- Compact deliberately. Compacting beats hitting the wall, but it is worse than starting clean: some detail is lost for good.
Noticing the limit approaching
Three signals that it is time to stop:
- The context indicator is around three quarters full. That is the moment to finish a step, not to begin one.
- The agent repeats work it already did, or rereads a file it read ten minutes ago. Early history has been compacted away.
- Answers have gone generic. The specifics left the context, and the model is answering from general knowledge instead of your project.
The right response to all three is to lock the result into code and a commit, then start a new session. Uncommitted work at the end of a long session is the worst possible combination: context runs out exactly when losing it costs the most.
Next
Model choice and its effect on spend has its own article. What happens to context in agent systems generally is covered on the blog: the agent context budget. The overview is in the Claude Code guide.
FAQ
Why does Claude Code get worse towards the end of a session?
Context filled up and part of the early history was compacted into a summary. Details are lost while the agent keeps working as though it remembers them. This is the main argument for short sessions: a clean context on a new task almost always beats continuing a long conversation.
How do I know how much context is left?
The tool shows how full the context is, and it deserves the same attention as a fuel gauge. With a quarter left, finish the current step and start a new session rather than launching another large task in the same window.
What costs more, one long conversation or several short ones?
The long one. Every request resends the whole accumulated context, so each step gets more expensive as the conversation grows. Three short sessions for three tasks cost less than one session where the same tasks ran back to back.
- Claude Code: A Complete Practical GuideGuide
- Claude Code in Russia: Access, Payment and LimitsWhat actually fails when using Claude Code from Russia, which payment routes remain, how to tell an access problem from a configuration error, and what the service rules mean for your account.
- Claude Code API Key: Where to Get One and How to Connect ItWhere the Claude Code API key is created, where to store it, how your own key differs from a subscription in cost and predictability, and what to do about a 401.
- Claude Code CLI: Commands, Flags and ModesHow Claude Code starts, what the permission modes actually mean, which flags earn their keep in daily work, and what belongs in project configuration instead of being retyped.
Done for you
Need the result, not the agent setup? I will build your MVP
I work in Claude Code every day and take products all the way to users.
from $1,500 · 1 to 2 weeks