Skip to content
PD
Вайб-кодинг

Vibe Coding and AI: How It Actually Works

What happens under the hood when a model writes code, why context beats phrasing, where confident errors come from, and how to compensate in practice.

All articles in the guide Вайб-кодинг · 11

Understanding the mechanics changes how you work more than knowing the buttons does. Here is what happens when you ask a model to write code.

Under the hood

The model produces a continuation of text. Everything it should account for is fed in - your request, the files it read, command output, the conversation history - and it emits the most plausible continuation.

Three things follow, and they explain nearly all the behaviour:

It does not check truth. A non-existent library method looks as appropriate as a real one: both fit the context.

It remembers nothing between calls. Everything it “knows” when answering was sent in that request. Hence the cost of a long conversation: every step resends the whole history.

It answers from what it sees. If it never saw the relevant file, it falls back on general knowledge and writes something plausible.

The difference between chat and an agent lives exactly here: an agent can check itself, because it runs tests and reads the output. That changes quality more than switching model does.

Context decides

The practical consequence of the mechanics: most bad answers are explained not by phrasing but by the model not having seen what mattered.

In practice that means:

  • A precise address beats a problem description. “Look at the parsing function in this file” works better than “figure out why the parser is broken”: the second sends it through half the repository and fills context with noise.
  • The more noise in context, the worse the answers. Logs, large exports and long history crowd out what matters.
  • A long session works worse than a short one. When context overflows, the early part is compacted into a summary, details are lost, and the model carries on as though it remembers them. Details in limits and context.

Hence the habit that pays most: one session, one task.

Why the model errs

Four typical mechanisms worth recognising:

Plausible instead of correct. An invented method, a non-existent parameter, imagined library behaviour. It appears where the model never saw the documentation and completed the expected shape.

Symptom instead of cause. The test fails, so the model edits the test. Formally, the task is closed.

A duplicate instead of reuse. It did not see the existing function and wrote a second one under a new name.

A success report on failure. The command returned an error and the model wrote that everything is done. Not rare but default: that is how conversations usually end.

What all four share: the result looks right. That is exactly why a quick skim does not catch them.

Compensating for it

Five techniques that cover most of the problem:

  1. Give a checkable criterion. A test that must pass, not “make it good”. This is the only thing that lets the model fix itself without you.
  2. Read the diff, not the report. The report describes intent, the diff shows fact.
  3. Keep context clean. New task, new session.
  4. Give a precise address. It saves both quality and money.
  5. Do not accept claims without sources. “This function is unused” should come with where it looked.

And a sixth, for anything important: verify with a different model or a separate pass. The same pass will repeat the same wrong assumption - the same principle as a verification layer in agent systems.

What this implies about the limits

The mechanics explain why the approach works well where there is an unambiguous check and badly where there is not. That is not a flaw in a particular tool but a property of the method: the model optimises plausibility rather than truth, and the only proxy for truth available to it is what can be run.

The tool survey is in AI for writing code. On agent mode and permissions: agents in vibe coding. The long version on the approach: what vibe coding is. The overview is in the vibe coding guide.

FAQ

Why does the model write wrong code so confidently?

Because it produces a plausible continuation rather than checking truth. A non-existent library method looks as correct to it as a real one: both fit the context. Confidence in tone is unrelated to correctness of content.

Why does the same request give different answers?

Generation is non-deterministic, and on ambiguous tasks the spread is noticeable. The practical consequence: one good run is not enough to call a task solved, and one bad run does not mean the model cannot do it.

What matters more, prompt phrasing or context?

Context. A precise file address and the relevant code buy more than any rewording. The model answers from what it sees, and most bad answers come down to it never seeing what mattered.

More on this topic