Your Coding Agent Has Amnesia: Keep Project Memory in the Repo, Not the Chat
Coding agents forget your project every session. Here is the four-file system I keep in every repo so decisions, constraints and unfinished work survive between sessions and between models.
Pavel Duglas
AI Automation & MVP Architect
Every morning I open a fresh session with a coding agent and it knows nothing. It does not know why we chose Postgres advisory locks over Redis, and it does not know that the payment webhook must stay idempotent. It does not know that yesterday I stopped halfway through a refactor of the parser module. So it guesses. And a confident agent that guesses is how you get a clean, well-tested pull request that quietly undoes a decision you made three weeks ago for a very good reason.
I stopped treating this as a model problem. Better models do not fix it, because the missing information was never in the code. It lived in my head and in old chat logs. The fix is boring. Put the project memory in the repository, in plain files, and make the agent read and update them like any other part of the codebase.
Why every session starts from zero
A chat session is a scratchpad. When it ends, everything you explained is gone. Some tools add their own memory features, but I do not rely on them for three reasons:
- They are tied to one tool. I switch between agents and models depending on the task and the price. Memory locked inside one vendor does not travel.
- They are invisible to the team. A contractor or a second agent working in parallel cannot see what the first one learned.
- They are not versioned. I cannot diff them, review them, or roll them back when they drift.
The repository already solves all three. It is shared, portable, versioned and reviewed. The only thing missing is the discipline to write the context down.
Code tells you what, never why
An agent can read your entire codebase in seconds. That is the trap. Reading the code gives it the what: this function retries three times, this table has a soft delete column, this module talks to the Telegram API through a queue. It cannot give it the why.
Why three retries and not five? Because the upstream provider bans IPs after six failures per minute. Why a queue instead of a direct call? Because Telegram rate limits bit us during a broadcast and we lost messages. That reasoning is the most valuable engineering context you have, and it is exactly what an agent throws away when it ‘simplifies’ your code.
So the goal is not to document everything. The goal is to document the things the code cannot explain about itself.
The four files I keep in every repo
I use four small markdown files at the root, sometimes inside a /docs/agent folder. Names do not matter much. Separation does. Each file answers one question, and mixing them is how they turn into a junk drawer nobody reads.
AGENTS.md: how to work here
This is the operating manual. It answers: how do I run, test and ship this project without breaking things?
- The exact commands to install, run, test and lint.
- Stack and versions that matter.
- Conventions that are not enforced by a linter, like ‘all external HTTP calls go through
lib/http_clientso we get retries and logging’. - Hard rules: ‘never edit files in
/generated’, ‘never add a dependency without asking’. - A pointer to the other three files.
Keep it under a page. If the agent has to read 800 lines before it can run the tests, you have written a wiki, not a manual.
DECISIONS.md: the why log
This is a lightweight version of architecture decision records. Every entry is short:
## 2024-11-03 - Advisory locks instead of Redis for job dedup
Context: jobs were running twice after worker restarts.
Decision: use Postgres advisory locks keyed by job id.
Why: one less service to run, we already have Postgres, volume is low.
Revisit if: we exceed ~50 jobs/sec or move off Postgres.
The ‘Revisit if’ line is the most useful part. It tells the agent that the decision is not sacred, and it spells out exactly when it becomes wrong. Without it, agents either treat every old choice as law or ignore all of them.
INVARIANTS.md: things that must never break
This is the shortest file and the most important one. It lists the properties of the system that must hold no matter what change is being made:
- Payment webhooks are idempotent by provider event id.
- A user can never see another tenant’s data, even in admin exports.
- Scraper output is never written to the main table without passing the validation step.
- We never store raw card data or full passport numbers.
When an agent proposes a change, I want it to check that list first. A clean refactor that breaks an invariant is worse than no refactor.
STATE.md: where we left off
This one is disposable and changes constantly. It holds the current task, what is done, what is half-done, known broken things, and the next step. Think of it as the handoff note you would leave a colleague on Friday evening.
Current: migrating parser from BeautifulSoup to selectolax
Done: product page, category page
In progress: search results page - pagination selector still flaky
Broken on purpose: test_search_pagination is skipped, see above
Next: fix pagination, remove skip, run full regression on 200 saved pages
That skipped test is the kind of thing that gets ‘fixed’ by an agent in the worst possible way if nobody explains it.
Keep it short or it rots
The failure mode of project memory is not missing information. It is stale information. A decision log that says we use Redis when we removed Redis two months ago will actively mislead the agent, and it will follow the document over the code with total confidence.
My rules:
- One line of truth per fact. If something is in AGENTS.md, do not repeat it in DECISIONS.md.
- Delete aggressively. STATE.md gets wiped when a task is finished. Superseded decisions get marked as superseded with a link to the new entry, not silently left in place.
- Budget the size. My rough limit is about 150 lines for AGENTS.md and INVARIANTS.md combined. The decision log can grow, but the agent only needs the recent and still-active entries by default.
Make the agent write back
The trick that made this actually work for me: the agent maintains the files, not me. Writing docs by hand after a long session is the first thing I skip when I am tired.
I end every meaningful session with a fixed prompt, something like: ‘Before we finish, update STATE.md with where we stopped. If we made a decision that someone might undo later, add it to DECISIONS.md with a Revisit if line. If anything we did touches an invariant, tell me.’
Then I review the diff like any other change. Usually it takes thirty seconds. Sometimes it catches something important, like the agent recording a ‘decision’ that was actually a temporary hack. That is a signal to fix the hack or to label it honestly.
At the start of a session the reverse happens: ‘Read AGENTS.md, INVARIANTS.md and STATE.md. Summarize the current task in three lines before touching any code.’ If the summary is wrong, I correct it before a single file changes. That three-line summary has saved me more time than any clever prompt.
Enforce, don’t hope
Documents are advice. Agents, like people, ignore advice under pressure. So every invariant that can become a test becomes a test.
- ‘Webhooks are idempotent’ becomes a test that sends the same event twice and asserts one charge.
- ‘No cross-tenant data’ becomes a test that queries as tenant A and asserts zero rows from tenant B.
- ‘Never edit generated files’ becomes a CI check that fails if those files changed without the generator running.
In INVARIANTS.md I put the test name next to each rule. Now the agent knows the rule, knows why it exists, and knows exactly which test will fail if it breaks it. Anything I cannot test stays as a written rule, and I review changes near it with extra care.
A cheap extra: a small CI job that warns when a pull request touches core modules like billing, auth or the migration folder but does not touch DECISIONS.md or STATE.md. It is only a warning. But it makes ‘did we record why?’ a question that gets asked every time.
A real example
On a Telegram bot SaaS I run, an agent once proposed replacing our message queue with direct API calls ‘to reduce complexity’. The code was clean, tests passed, and the reasoning looked sensible. Before the memory files, I might have merged it on a busy day.
With DECISIONS.md in place, the agent flagged the conflict itself: there was an entry explaining that direct calls caused lost messages during broadcasts because of rate limits, with ‘Revisit if: we drop broadcast features’. It asked whether that still applied. It did. Five minutes of reading saved a production incident that would have taken a weekend to diagnose.
What not to put in these files
- Secrets. Obviously. These files get read by every tool you plug in.
- Things the code already says clearly. Do not describe every function. Link to the module instead.
- Wishful architecture. Document what exists and why, not the system you hope to build someday. Put plans in STATE.md or an issue tracker.
- Long prompt tricks. ‘You are a senior engineer who…’ belongs nowhere. Facts and constraints beat personas.
Start today
You do not need a framework for this. In one hour you can:
- Create AGENTS.md with your run and test commands plus five hard rules.
- Write down the three decisions you would be most annoyed to see undone.
- List your top five invariants and link each to a test, or write the missing test.
- Add the end-of-session and start-of-session prompts to your notes.
Code got cheap. The context around it did not. Put that context where your agents can read it, and where your team can review it.
FAQ
Isn't this just the same as a README?
No. A README is written for humans discovering the project: what it is and how to get started. These files are written for whoever is changing the project right now, human or agent. They focus on constraints, past decisions and the current state of unfinished work, which a README usually leaves out on purpose. I keep the README short and point from it to AGENTS.md.
Won't these files eat my context window and raise costs?
Only if you let them grow. I keep AGENTS.md plus INVARIANTS.md around 150 lines and wipe STATE.md after each task, so the default load is a few thousand tokens. The decision log is read on demand, for example only when the agent touches billing or auth. That cost is tiny compared to one session wasted on undoing a change that ignored an old decision.
Do I need a special tool or MCP server for project memory?
Not to start. Plain markdown in the repo works with every agent and model, is versioned by git, and can be reviewed in pull requests. Tools can help with search once the decision log gets large, but I would only add one after the plain-file habit is working. Adding infrastructure before the habit exists just gives you a more expensive place for stale notes.
Related articles
Done for you
I will turn your vibe-coded prototype into a working product
I will review what the AI generated, close the security and data gaps and ship it to production.
from $1,500 · 1 to 2 weeks