Skip to content
PD
n8n

AI Agents in n8n: LLMs, Tools, Memory and When You Do Not Need an Agent

How the AI nodes in n8n work: the difference between a plain model call and an agent, wiring up tools, conversation memory, working with your own documents and RAG, cost control, and what breaks in production.

All articles in the guide n8n · 18

The AI nodes are what draws many people to n8n right now, and the tooling genuinely is ahead of its neighbours: model, tools, memory and a vector store assemble on one canvas. But this is also the easiest place to build an expensive, unpredictable system where a single request would have done.

Model call or agent

The distinction is fundamental and worth settling before you start building.

A plain model call is one step: send the prompt and the data, get an answer. Predictable in time, cost and behaviour. It covers most tasks: triaging tickets, extracting fields from text, summarising, generating a templated reply.

An agent is a model handed tools and the right to decide what to call. It can search, hit your API, look in a database, do it again, then answer. One agent run means several model calls.

The selection rule is simple: if the sequence of steps is known in advance, you do not need an agent. Build it with nodes and get something cheaper, faster and reproducible. An agent earns its place when the steps depend on the content of the request.

Tools

A tool is anything the agent can invoke: a search, an HTTP call to your service, a database read, another workflow.

The practice that determines agent quality is descriptions. The model picks a tool by its description, not its name. “Get data” is a poor description; “Find a customer by email and return subscription status and last payment date” is a good one. Half of all “the agent does not call the right tool” problems are fixed by rewriting descriptions rather than switching models.

The second rule: give it the fewest tools you can. Fifteen options degrade decision quality and multiply pointless calls.

And a third, on safety: a tool is a real action. An agent holding a “delete record” tool will eventually call it. Either do not expose dangerous operations at all, or put a check between the agent and the action.

Memory

By default each run is independent - the model does not remember the previous message. Conversational scenarios, Telegram bots and support flows need memory tied to a participant identifier.

The detail people miss: memory is keyed by a session id. Get the key wrong and every user lands in one shared context and reads someone else’s conversation. In a bot the key should be the chat identifier - that is both functionality and a privacy matter.

The second point: memory grows, and the cost of each request grows with it, because the whole history goes to the model. For long conversations, cap the window.

Your own documents and RAG

The stock request is “make it answer from our knowledge base”. The pattern is standard: documents are chunked, embedded and stored; at query time the relevant chunks are retrieved and sent to the model alongside the question.

What actually determines the result: chunking quality (too small loses context, too large dilutes relevance) and index freshness - documents change while the vectors stay stale until you reindex. Build reindexing in as a scheduled workflow from the start, rather than discovering six months later that the bot answers from a superseded policy.

Cost

An agent is a variable cost, and that changes how you operate it. One run may cost a cent, or considerably more if the agent wanders through its tools.

What to do before production:

  • Cap the agent’s step count, or the “call a tool, think, call again” loop can run a long way.
  • Match the model to the task. Classification does not need your most expensive model; the saving is a multiple, not a percentage.
  • Log spend per run. Without it the month-end invoice is a surprise.
  • Test on junk input. An empty message, emoji, text in another language - the agent should not spiral on nonsense.

What breaks in production

Wrong response shape. The model returned prose instead of JSON and the next node failed. If the answer feeds downstream processing, require structured output and validate it - data of the wrong shape remains the leading cause of failures here too.

Non-determinism. The same input yields different answers. Normal for a model, unacceptable for a process that needs reproducibility.

Silent degradation. The agent keeps answering, just worse - the prompt drifted, the index went stale, the model changed. Nothing fails, and the error workflow stays quiet. The only defence is a human spot-checking answers, even once a week.

FAQ

How does an n8n AI agent differ from a plain model call?

A plain node sends a request and returns an answer - one step, predictable cost. An agent is given a set of tools and decides which to call and how often, so a single run can mean several model calls. Reach for an agent only when the sequence of actions is not known in advance.

Do I need an AI agent to extract data from text?

No. Field extraction, classification and summarisation are a single model call with a structured response. An agent adds cost, latency and unpredictability while improving nothing. Agents earn their place when tools and in-flight decisions are required, not for one transformation.

More on this topic

Done for you

I will build the automation in n8n or in code

Leads, sheets, CRM and Telegram connected, so nobody moves data by hand again.

from $300 · 3 to 7 days

Similar caseBAS Script License Issuing Automated on MakeA Make scenario that turns one Telegram message into a full licence handover: generated login and password, a licence for the requested term, FingerprintSwitcher Business enabled, and a row written to Google Sheets.

"Thanks to Pavel, the task is done. Always reachable, gave me detailed instructions and a guide, I will come back and I recommend him to everyone."

MarkBorisov · KworkTranslated from Russian