Local AI Agents: Self-Hosted Models and Stack
Why you would run an agent locally, what hardware it takes, where local models hit their ceiling on agentic work, and why a hybrid split usually beats either extreme.
All articles in the guide ИИ-агенты · 11
A local agent is not “the same thing but free”. It is a different set of trade-offs: you gain control over data and predictable cost, and you pay in quality on hard tasks and in your own operations time.
Why run locally
Three reasons, each sufficient on its own.
The data cannot leave. Personal data, healthcare, banking confidentiality, a client contract that forbids passing anything to third parties. This is the most common and most honest reason: price is not the question here.
Volume. When the token bill is comparable to the cost of a server, local starts winning. That threshold arrives sooner than expected if the agent runs continuously.
Independence. A provider can change prices, restrict access or retire a model. A local model stays put, along with all its shortcomings.
What is not on that list: “local is faster” and “local is simpler”. Usually the opposite.
What hardware it takes
The governing resource is video memory, and it decides which model will run at all.
- A consumer GPU. Small models at acceptable speed. Fine for narrow work: classification, field extraction, short templated answers.
- A professional card. Mid-size models. This is the level at which an agent loop starts behaving sensibly.
- Several cards or a server. Large models. At this point the question is whether paying a provider is cheaper.
- CPU only. Formally works, practically unusable for an agent: a twenty-step loop stretches into hours.
The second resource that matters is generation speed. Every agent step needs a model response, so a slow model turns a one-minute task into half an hour. Invisible on a single answer, multiplied across a loop.
Where local models fall short
Honestly about the ceiling: the gap is wider on agentic work than on text generation.
The reason is that an agent does not need a well-written paragraph; it needs to hold a long chain together: remember what has been done, pick correctly among ten tools, build call arguments precisely, recognise an error in a response and change plan. That is the most demanding use there is, and it is exactly where smaller models slip.
What shows up in practice:
- Wrong tool choice when descriptions look similar.
- Malformed calls - extra prose around the structure, so the call cannot be parsed.
- Goal drift over a long chain: the agent forgets why it started.
- Looping on a repeated step.
Some of this is treatable: a strict response schema, fewer tools, a shorter loop, more deterministic code around it. But a local agent has to be designed for a weaker executor, and that affects the architecture rather than just the model choice.
The hybrid split
The best answer is usually not an extreme but a division.
Keep locally whatever touches the data. Classification, field extraction, search over internal documents - anything where the text is sensitive.
Send to the cloud whatever needs depth and can be de-identified: hard reasoning, unusual cases, planning.
A router decides what goes where. A simple predicate: if the text contains personal data, stay local, otherwise it may go out. The mechanics of that routing are on the blog: the model router.
This gives privacy where it is required and quality where it is required. The price is a slightly more complex system, which is still usually simpler than forcing everything through a local model.
The surrounding infrastructure
The local model is only half of it. The rest is the same as in the cloud case: a database for state, a queue, per-step logs, monitoring. If the stack goes on your own server, the self-hosting in Docker article covers the same principles: backups, upgrades, keys.
And the main point: running locally waives none of the requirements on an agent. Step limits, verification and idempotency are needed exactly as much - see rollout.
The overview is in the AI agents guide.
FAQ
Can an AI agent run fully locally?
Yes, technically everything exists: a local model server with a compatible API, your own agent loop, your own database. The limit is not infrastructure but quality: agentic work requires holding a long chain of steps together, and local models lose more ground there than on plain text generation.
What hardware does a local agent need?
It comes down to video memory. Small models run on a consumer card, mid-size models need a serious one, and the largest need several. CPU-only is formally possible, but the speed makes a twenty-step agent loop impractical.
When is a local model justified?
When data cannot leave your perimeter by law or contract, when volume is high enough that token spend exceeds hardware cost, and when you need independence from provider availability. In all three cases the decision is made by constraints, not by quality.
- What an AI Agent Is and How It Differs from a ChatbotGuide
- Building an AI Agent: From Prompt to ProductionThe path from idea to a working agent: framing the task, choosing tools, the execution loop, testing before launch, and what has to be true before it reaches production.
- How to Build an AI Agent: A Step-by-Step Real CaseOne agent walked through end to end: requirements, design, tools and their descriptions, the first run and the fixes it forced, and which parts deliberately stayed ordinary code.
- AI Agents for Business: Where They Pay Off and Where They Do NotWhich processes an AI agent genuinely makes cheaper, where ordinary automation or a hire wins, how to calculate payback honestly, and the risks that rarely make it into the model.
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks