Skip to content
PD
ИИ-агенты

How to Build an AI Agent: A Step-by-Step Real Case

One agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.

All articles in the guide ИИ-агенты · 11

Abstract designs translate badly into code. Here is one concrete task end to end, including the parts that did not become an agent.

The case and the requirements

The task: free-form incoming messages have to become tickets with fields filled in. They arrive from several channels and are written by people rather than forms: typos, no structure, sometimes several questions in one message.

Requirements stated before any work:

  • Assign a category from a fixed list.
  • Set priority by explicit rules.
  • Find the customer in the database if they exist.
  • Create the ticket.
  • When confidence is low, hand it to a human rather than guessing. This requirement outranks the others.

Definition of done: a ticket exists with mandatory fields filled, or the message is flagged as needing a human. Verifiable by a program.

The design

The first question is whether an agent is needed. The answer: partly.

  • Parsing the text is a single model call with a response schema. Not an agent.
  • Finding the customer is a database query. Not an agent, just code.
  • Priority by rules is conditionals. Not an agent.
  • Resolving ambiguity is where a loop appears: the model can notice that data is missing, pull more from the database and revise its decision.

The resulting proportion: the agentic part is roughly a quarter of the system. The rest is code. That is not a compromise, it is normal architecture; making everything agentic yields a system that costs three times more and debugs three times worse.

Tools and their descriptions

The agent got four tools:

  • Find customer by phone, email or name. Read only.
  • Read customer ticket history. Read only.
  • Create ticket with mandatory fields. The only tool with an external effect.
  • Hand to a human with a stated reason. Also terminal.

What the descriptions taught us. The first version of the customer lookup description read “finds a customer in the database”. The agent called it constantly, including when there was nothing to look up. Rewritten to say when it applies, when it does not, and what comes back when nothing is found, the redundant calls dropped by a large factor.

Second lesson: the handover tool initially required no reason. The agent used it as a way to end work whenever the task got hard. Making the reason mandatory, and checking that it was substantive, brought it back into bounds.

The first run and the fixes

We ran it over a hundred real messages. What surfaced:

Reported success after an error. Ticket creation failed validation, the tool returned an error, the agent wrote “ticket created”. A classic. Fixed not by prompting but by checking the fact: after completion the system confirms the ticket exists, and if it does not, the run failed regardless of what the agent said.

Confused two adjacent categories. Not a model problem: the category definitions themselves had no boundary. We wrote the boundary explicitly, with examples on both sides.

Too many steps on simple messages. The agent pulled customer history where the message text was sufficient. An explicit ordering fixed it: try to decide from the text, go to the database only when data is missing.

Lost part of multi-topic messages. A message with three questions became one ticket. We added an explicit multi-topic check and a split before the main loop.

What it became

A system where the model owns parsing and the ambiguous decisions, and everything else is deterministic code. The practical outcomes:

  • The share of messages routed to a human is a deliberate setting, not a side effect. The confidence threshold lives in one place.
  • Per-step cost is predictable, because most of the work never touches the model.
  • Debugging is possible. The logs show which tool was called and what it returned, not just the final text.

The transferable lesson: first work out which part of the task genuinely requires decisions, and give only that part to an agent. The rest is cheaper to write.

What this looks like on shipped projects: the portfolio. The general shape of the work is in building an agent, and what to check before launch is in rollout. The overview is in the AI agents guide.

FAQ

How long does a working agent take to build?

A first version that does something: an evening. A version you can trust with real volume: weeks. Almost all the difference goes into handling the cases where something went wrong, and those cases decide whether the system saves work or creates it.

Does everything have to be the agent?

No, and usually it should not be. In working systems the model owns the one step that requires interpreting unstructured input, and the rest is ordinary code. That proportion is cheaper, more predictable and far easier to debug than an agent that does everything.

What goes wrong most often on the first run?

The agent picks the wrong tool because the descriptions look alike, and it reports success where the tool returned an error. Neither is fixed by prompting: both are fixed in the tool descriptions and in what the system returns on failure.

More on this topic