How to Build an AI Agent: A Step-by-Step Real Case
One agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.
All articles in the guide ИИ-агенты · 11
Abstract designs translate badly into code. Here is one concrete task end to end, including the parts that did not become an agent.
The case and the requirements
The task: free-form incoming messages have to become tickets with fields filled in. They arrive from several channels and are written by people rather than forms: typos, no structure, sometimes several questions in one message.
Requirements stated before any work:
- Assign a category from a fixed list.
- Set priority by explicit rules.
- Find the customer in the database if they exist.
- Create the ticket.
- When confidence is low, hand it to a human rather than guessing. This requirement outranks the others.
Definition of done: a ticket exists with mandatory fields filled, or the message is flagged as needing a human. Verifiable by a program.
The design
The first question is whether an agent is needed. The answer: partly.
- Parsing the text is a single model call with a response schema. Not an agent.
- Finding the customer is a database query. Not an agent, just code.
- Priority by rules is conditionals. Not an agent.
- Resolving ambiguity is where a loop appears: the model can notice that data is missing, pull more from the database and revise its decision.
The resulting proportion: the agentic part is roughly a quarter of the system. The rest is code. That is not a compromise, it is normal architecture; making everything agentic yields a system that costs three times more and debugs three times worse.
Tools and their descriptions
The agent got four tools:
- Find customer by phone, email or name. Read only.
- Read customer ticket history. Read only.
- Create ticket with mandatory fields. The only tool with an external effect.
- Hand to a human with a stated reason. Also terminal.
What the descriptions taught us. The first version of the customer lookup description read “finds a customer in the database”. The agent called it constantly, including when there was nothing to look up. Rewritten to say when it applies, when it does not, and what comes back when nothing is found, the redundant calls dropped by a large factor.
Second lesson: the handover tool initially required no reason. The agent used it as a way to end work whenever the task got hard. Making the reason mandatory, and checking that it was substantive, brought it back into bounds.
The first run and the fixes
We ran it over a hundred real messages. What surfaced:
Reported success after an error. Ticket creation failed validation, the tool returned an error, the agent wrote “ticket created”. A classic. Fixed not by prompting but by checking the fact: after completion the system confirms the ticket exists, and if it does not, the run failed regardless of what the agent said.
Confused two adjacent categories. Not a model problem: the category definitions themselves had no boundary. We wrote the boundary explicitly, with examples on both sides.
Too many steps on simple messages. The agent pulled customer history where the message text was sufficient. An explicit ordering fixed it: try to decide from the text, go to the database only when data is missing.
Lost part of multi-topic messages. A message with three questions became one ticket. We added an explicit multi-topic check and a split before the main loop.
What it became
A system where the model owns parsing and the ambiguous decisions, and everything else is deterministic code. The practical outcomes:
- The share of messages routed to a human is a deliberate setting, not a side effect. The confidence threshold lives in one place.
- Per-step cost is predictable, because most of the work never touches the model.
- Debugging is possible. The logs show which tool was called and what it returned, not just the final text.
The transferable lesson: first work out which part of the task genuinely requires decisions, and give only that part to an agent. The rest is cheaper to write.
What this looks like on shipped projects: the portfolio. The general shape of the work is in building an agent, and what to check before launch is in rollout. The overview is in the AI agents guide.
FAQ
How long does a working agent take to build?
A first version that does something: an evening. A version you can trust with real volume: weeks. Almost all the difference goes into handling the cases where something went wrong, and those cases decide whether the system saves work or creates it.
Does everything have to be the agent?
No, and usually it should not be. In working systems the model owns the one step that requires interpreting unstructured input, and the rest is ordinary code. That proportion is cheaper, more predictable and far easier to debug than an agent that does everything.
What goes wrong most often on the first run?
The agent picks the wrong tool because the descriptions look alike, and it reports success where the tool returned an error. Neither is fixed by prompting: both are fixed in the tool descriptions and in what the system returns on failure.
- What an AI Agent Is and How It Differs from a ChatbotGuide
- Building an AI Agent: From Prompt to ProductionThe path from idea to a working agent: framing the task, choosing tools, the execution loop, testing before launch, and what has to be true before it reaches production.
- AI Agents for Business: Where They Pay Off and Where They Do NotWhich processes an AI agent genuinely makes cheaper, where ordinary automation or a hire wins, how to calculate payback honestly, and the risks that rarely make it into the model.
- AI Agent Development: Stack, Timelines and PitfallsWhat the work of building an AI agent actually consists of, the stack used in practice, where the time goes, and how to accept the result so you get a system rather than a demo.
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks