Skip to content
PD
ИИ-агенты

AI Agent Memory: Context, Compaction and Artifacts

Why agent context runs out sooner than expected, what to keep and what to discard, how compaction works, and why file-based memory beats retelling the conversation.

All articles in the guide ИИ-агенты · 11

Applied to agents, “memory” means four different mechanisms rather than one. Confusing them produces systems that forget what matters and remember noise.

Why context runs out

The model remembers nothing between calls. Everything it “knows” when answering was sent in that same request.

Two consequences follow, and they shape the whole design:

  • Every agent step resends the entire accumulated history. That is why step twenty costs more than step one rather than the same.
  • The window is finite, and it does not hold everything the agent read along the way.

What actually occupies it: not the task statement but tool results. One data export, one long API response, one error log, and half the window is gone. The task statement is a few percent.

What to keep and what to discard

It helps to split memory into four types and treat each differently.

Working context. What is needed right now, on this step. Lives in the window and disappears with it.

Findings. What the agent established and will need later. Must live outside, in a file or a database. Keeping conclusions only in the conversation means losing them at the first compaction.

State. What has already been done. Not a convenience but a defence: on restart the agent must know the email was already sent, or it sends a second one - see idempotent pipelines.

Long-term memory. Knowledge carried between runs: reference data, preferences, accumulated facts. Stored externally and loaded by relevance rather than wholesale.

The rule that cuts spend more than any other: a raw tool result should not stay in context once the useful part has been extracted. You read a thousand-row table and pulled three numbers; three numbers is what should remain.

Compaction

When the window nears its limit, the early conversation is compressed into a summary. That lets work continue, and it has a price.

Detail is lost. The summary keeps what looked important at compaction time. An exact filename, a specific parameter value, a caveat from the middle of the conversation - all gone.

The agent does not tell you. It carries on as though it remembers, which is more dangerous than an explicit error: an incomplete picture looks complete.

Quality drops gradually, not sharply, as compactions accumulate.

The practical conclusion: compaction is an emergency mode, not a working one. Results should be written outside before it happens. The right habit on a long task is to finish a meaningful chunk, record the result to a file or database, and only then continue.

File-based memory

The most reliable and most underrated mechanism: the agent writes down what it found and reads it back when needed.

Why it beats holding everything in context:

  • It survives compaction and restarts. Context survives neither.
  • It can be read in parts. Need one section, read one section, not the whole file.
  • A human can check it. You can open the file and see what the agent actually found, rather than what it wrote in its report.
  • It transfers between agents. In a multi-agent design this is the natural way to hand over a result.

What usually goes into files: the work plan with completed items marked, findings with links to their source, intermediate exports, and a list of what failed and why.

A separate benefit: a plan in a file is protection against goal drift. An agent on a long task forgets why it started; a plan it rereads brings it back to the original brief.

The practical minimum

A working system needs four things:

  1. State in a database - what has been done.
  2. Findings in files or a database - what has been established.
  3. Compaction treated as an emergency, not as routine.
  4. Raw exports removed from context once the useful part is extracted.

The long version of how to budget context in an agent system is on the blog: agent context budget architecture. The same problem inside a coding agent is in Claude Code limits. The overview is in the AI agents guide.

FAQ

Does an AI agent have memory between runs?

Not by itself. The model remembers nothing: every request resends whatever history it needs. An agent memory is whatever you store outside and feed back into context, and its design is entirely on your side.

What is context compaction?

Compressing the early part of the conversation into a summary to free space. It lets work continue but loses detail irreversibly: whatever looked unimportant at compaction time cannot be recovered. That is why important results get written to files and databases before context nears its limit.

Why should an agent write files if it has context?

A file survives compaction, restarts and session changes; context survives none of them. A file can also be read in parts when needed, instead of sitting in view for the whole conversation.

More on this topic