Skip to content
PD
Vibe Coding 7 min read

Stop Writing Plans for Your Coding Agent. Write the Tests Instead

Long plans drift and go stale. Failing acceptance tests give a coding agent a spec it can check itself against. Here is the workflow I use to steer agents without plan mode.

PD

Pavel Duglas

AI Automation & MVP Architect

For about a year my workflow with coding agents looked like this: describe the feature, let the agent produce a plan, read the plan, tweak the plan, approve the plan, watch the agent drift away from the plan by step four. I spent more time editing markdown than reviewing code. The plan felt like control, but it was mostly theater. The agent agreed with everything in it and then did whatever the code in front of it suggested.

What actually fixed this was boring: I stopped writing plans and started writing failing tests. A test is a plan the agent cannot misread, cannot quietly skip, and cannot declare finished without proof. Here is how I do it on real client work.

Why plans fail as a steering mechanism

A plan is prose. Prose has three problems when the reader is a model.

It is ambiguous. “Handle rate limits gracefully” means five different things. The agent picks one, usually the one that requires the least code.

It is unverifiable. There is no command you can run that tells you whether step 3 of a plan is done. So the agent tells you it is done, and you either trust it or reread the diff line by line.

It goes stale instantly. The moment the agent discovers the existing code does something unexpected, the plan is wrong. Good agents adapt, but now the plan in your head and the code on disk diverge, and nobody updates the document.

A failing test has none of these problems. It is precise, it is executable, and when reality changes it fails loudly instead of silently becoming fiction.

The core idea: tests are the spec, the agent writes the implementation

The split of responsibilities is simple:

  • I write the acceptance tests. They describe behavior from the outside: inputs, outputs, side effects, error cases.
  • The agent writes the implementation and any internal unit tests it wants.
  • Nobody declares done until a single check command passes.

This is not classic TDD with a red-green cycle for every function. I do not care how the agent structures internals. I care that the feature behaves correctly at its boundary. That boundary is where bugs cost money.

What I actually put in the acceptance tests

Let me use a real shape of task. A client has a Telegram bot for a small booking business. The feature: users can cancel a booking with a command, but only up to 24 hours before the slot, and the admin chat gets notified.

The old me would write a plan with seven bullet points. Now I write something like this:

# tests/acceptance/test_cancel_booking.py

def test_user_can_cancel_more_than_24h_before(bot, db, clock):
    clock.set("2026-03-01 10:00")
    booking = db.create_booking(user_id=42, slot="2026-03-03 10:00")
    reply = bot.send(user_id=42, text=f"/cancel {booking.id}")
    assert "cancelled" in reply.text.lower()
    assert db.get_booking(booking.id).status == "cancelled"

def test_cancel_inside_24h_is_rejected(bot, db, clock):
    clock.set("2026-03-02 11:00")
    booking = db.create_booking(user_id=42, slot="2026-03-03 10:00")
    reply = bot.send(user_id=42, text=f"/cancel {booking.id}")
    assert db.get_booking(booking.id).status == "active"
    assert "24" in reply.text

def test_user_cannot_cancel_someone_elses_booking(bot, db, clock):
    booking = db.create_booking(user_id=42, slot="2026-03-10 10:00")
    bot.send(user_id=99, text=f"/cancel {booking.id}")
    assert db.get_booking(booking.id).status == "active"

def test_admin_is_notified_on_cancel(bot, db, clock, admin_chat):
    clock.set("2026-03-01 10:00")
    booking = db.create_booking(user_id=42, slot="2026-03-05 10:00")
    bot.send(user_id=42, text=f"/cancel {booking.id}")
    assert admin_chat.last_message_contains(str(booking.id))

Four tests, maybe fifteen minutes of work. Look at what they encode that a plan would have left vague:

  • The exact 24 hour boundary and that time comes from an injectable clock, not datetime.now() scattered everywhere.
  • The ownership check, which is the one an agent would most likely forget.
  • That a rejected cancel leaves the booking untouched.
  • The side effect on the admin chat.

I deliberately do not assert exact reply wording. Asserting on “24” appearing in the message is enough to prove the user was told why. Over-specified tests make the agent contort the code to match strings, which is its own kind of drift.

The fixtures matter more than the tests

The bot, db, clock and admin_chat fixtures are the real investment. Once they exist, every future feature gets cheap acceptance tests. On a new project I spend the first session building these harness pieces myself, or with the agent under close supervision. After that, writing a spec for a feature takes minutes.

If your project has no harness at all, that is the first thing to build. Not the feature.

The rules I give the agent

The tests alone are not enough. Agents are very good at making tests pass in ways you did not intend. These rules live in the repo instructions file so they apply to every session:

  1. Files under tests/acceptance/ are read-only. If the agent believes a test is wrong, it stops and explains why. It does not edit it. This single rule prevents the most common failure: the agent “fixing” the test to match its buggy code.
  2. No special-casing test conditions. No checks for test user IDs, no environment flags that bypass logic. I grep for this in review.
  3. One command defines done. In my projects that is make check, which runs lint, type checks, unit tests and acceptance tests. The agent must run it and paste the tail of the output before claiming completion.
  4. Smallest diff that passes. No refactoring unrelated code in the same task. Refactors are their own task with their own green check before and after.

Rule 1 is the one people skip and regret. I also enforce it mechanically: a pre-commit hook that fails if acceptance test files changed in a commit authored during an agent session. Belt and suspenders, because a stated rule is a suggestion and a hook is a fact.

The loop in practice

Here is the actual sequence I run for a feature:

  1. Write acceptance tests. Run them. Confirm they fail for the right reason (missing command, not a broken fixture).
  2. Commit the failing tests on a branch. This is the spec, now versioned.
  3. Give the agent a short prompt: the feature in two sentences, the path to the tests, and “make make check pass without editing acceptance tests.”
  4. Let it work. I do not watch every step anymore. I check back when it says done.
  5. Run make check myself. Agents occasionally report success on a stale run.
  6. Review the diff with a specific question: did it solve the problem, or did it solve the tests? Those are different, and the difference is where I earn my fee.

Notice how short step 3 is. The prompt shrank because the tests carry the detail. My prompts went from half a page to three lines, and the results got better.

When the agent pushes back

Sometimes the agent stops and says a test seems wrong. This is the best outcome of the whole system, because it usually means one of two things:

  • My spec was actually wrong. Maybe the booking model stores slots in UTC and my test assumed local time. Good, I found a real ambiguity before it shipped. I fix the test myself and restart.
  • The existing code makes the spec hard. Maybe cancellations are tangled into a payments module. That is a design conversation, and I would much rather have it now than discover it in a 600 line diff.

With plans, these conflicts stayed invisible. The agent just quietly bent the plan. With read-only tests, the conflict has to surface.

Where plans still earn their place

I have not banned planning entirely. I still write a short plan when:

  • The work is a migration or restructure with no new behavior to test. Here the spec is “all existing tests stay green,” plus an ordered list of steps so nothing is half-moved.
  • The feature crosses several services and the order of deployment matters.
  • I am exploring and do not know the behavior yet. Then I prototype with the agent, throw the code away, and write tests for what I learned.

That last one is important. Vibe coding a throwaway spike is great for discovery. The mistake is keeping the spike. Once you know what you want, encode it in tests and rebuild.

What this changes for founders and small teams

If you run an MVP with one or two developers plus agents, this workflow shifts where human attention goes. You stop being the person who reads every line an agent writes. You become the person who decides exactly what correct means, which is the part a model genuinely cannot do for you because it does not know your business.

The practical payoff I have seen on client projects: fewer regressions when agents touch old code, because the acceptance suite accumulates and protects every feature ever specified. Review gets faster because I am checking intent, not syntax. And onboarding a new developer is easier because the acceptance tests read like a list of what the product promises its users.

Start small. Pick the next feature on your list, write three to five acceptance tests before you open the agent, make them read-only, and give the agent a three line prompt. Compare the result to your last planned feature. I would be surprised if you go back.

FAQ

Isn't this just test-driven development with extra steps?

It borrows the test-first idea but at a different level. I only write tests at the feature boundary, from the user's point of view, and I leave internal design and unit tests to the agent. Classic TDD drives every function through red-green cycles. Here the point is to give the agent an unambiguous, executable spec and to have a hard definition of done that it cannot talk its way around.

What stops the agent from writing hacky code that just passes the tests?

Three things. The acceptance files are read-only and enforced with a hook, so it cannot weaken the spec. The repo rules forbid special-casing test conditions. And my review focuses on one question: does this solve the real problem or only the tests? Writing tests that check behavior rather than exact strings also leaves less room for gaming.

My project has no tests at all. Where do I start?

Start with the harness, not the tests. Build fixtures that let you drive the app from the outside: a fake database or test database, a controllable clock, and a way to send input and capture output, such as a fake Telegram client or an HTTP test client. That takes a session or two. After that, each feature spec costs minutes, and you can add acceptance tests for existing critical flows before letting an agent touch them.

Related articles