Skip to content
PD
AI Automation 8 min read

Your Agent Doesn't Need a Smarter Model, It Needs Fewer Tools

Most agent failures I debug are wrong-tool failures, not reasoning failures. Here is how I design a tool surface: intent-based naming, coarse tools, schemas that block bad calls, progressive disclosure, and a tool-selection eval you can run in CI.

PD

Pavel Duglas

AI Automation & MVP Architect

Last month I got a familiar message from a client: “the agent got dumber after we added the new integration.” Nothing about the prompt had changed. The model had not changed. What changed was that they plugged in a third MCP server and went from 12 tools to 31. The agent was now spending its attention choosing between update_record, patch_row, set_field and write_cell, three of which touched the same database through different services.

That is not a model problem. That is a UX problem, and the user is the model. I have debugged enough of these to say it plainly: in production agent systems, wrong-tool calls are more common than bad reasoning. And almost nobody measures them.

The failure mode nobody logs

When an agent fails, teams look at the final answer and conclude “hallucination” or “needs a better prompt.” Then they add 400 words of instructions telling the model which tool to use when. That is a patch on top of a badly designed interface.

Go look at your raw tool-call logs for a week. You will find four distinct patterns:

  • Wrong tool, right intent. The agent wanted to find a customer and called search_documents instead of find_customer, because both descriptions say “search”.
  • Right tool, garbage arguments. It passed project: "Acme Corp" where the schema wanted a slug like acme-corp, got an empty result, and then confidently reported that the project does not exist.
  • Tool thrash. Four calls to list endpoints in a row because no single tool answered the actual question, so the agent tried to assemble it by hand and ran out of context.
  • Silent no-op. The tool returned {"ok": true} for a write that matched zero rows. The agent reported success. Nobody noticed for three days.

Every one of those is fixable at the interface level. None of them is fixed by a bigger model.

Tool definitions are not free

There is a second cost that is easy to miss. Those 31 tools cost about 8,400 tokens of JSON schema in every single request of the loop. On a 20-step task that is 168,000 tokens of pure menu, before any actual work. You pay for it in money, in latency, and in attention: the more tools you show, the flatter the model’s preference between them.

Rule 1: name tools by intent, not by endpoint

Most tool sets are generated from an API surface. That is why they read like get_v2_customers_by_id and post_orders_bulk. The model does not care about your REST layout. It is matching the user’s intent against your tool names and descriptions.

Name for the job to be done:

  • find_customer_by_email instead of get_v2_customers
  • refund_order instead of post_transaction_reversal
  • draft_reply_for_review instead of create_message

And make descriptions do real work. My template is three lines: what it does, when to use it, and when not to use it. The negative clause is the highest-leverage sentence you can write.

{
  "name": "find_customer_by_email",
  "description": "Look up exactly one customer by their email address. Use this first whenever the user mentions an email. Do NOT use this for names, company names or partial matches - use search_customers for that.",
  "input_schema": {
    "type": "object",
    "required": ["email"],
    "properties": {
      "email": { "type": "string", "format": "email" }
    }
  }
}

If two tools could both plausibly satisfy the same sentence, either merge them or write an explicit disambiguation clause into both descriptions. Ambiguity between tools is a bug you can fix in a text editor.

Rule 2: one tool per decision, not per API call

The biggest quality jump I get is from making tools coarser. Agents are bad at multi-step plumbing and good at making one decision at a time. So push the plumbing into the tool.

A real example. A support agent had these five tools:

list_projects, get_project, list_subscriptions, get_invoice, get_payment_status

To answer “is this client paid up?” it needed four chained calls, each with an ID it had to carry forward. Selection and argument errors compounded. I replaced all five with one:

get_account_billing_summary(email_or_slug) returning a flat object: plan, status, last invoice, amount due, days overdue, payment method validity.

One call, one decision, no ID juggling. The rule I use: if the agent will almost always call B right after A, make A+B one tool. Deterministic orchestration belongs in your code, not in a probabilistic loop. That is also cheaper, because you are not paying for a reasoning round trip between each hop.

Rule 3: build schemas that make bad calls impossible

Treat every tool schema like a public form. Free-text fields are where agents go to die.

  • Use enum for anything with a fixed set of values. status: "open" | "pending" | "closed" beats status: string every time.
  • Never accept a human-readable label where you need an ID. If you must, accept both and resolve server-side.
  • Set "additionalProperties": false. Invented parameters should fail loudly, not get silently dropped.
  • Keep required params to three or fewer. If you need seven, you have a workflow, not a tool.
  • Put the dangerous flag in its own tool. delete_records(confirm: true) will eventually get confirm: true. A separate delete_records behind an approval gate will not.

And validate before execution with the same schema you advertise. I return validation errors back into the loop as text rather than throwing, which leads directly to the next rule.

Rule 4: error messages are prompts

Whatever your tool returns on failure is now part of the model’s context and will shape the next call. {"error": "not found"} teaches it nothing, so it retries the same thing or gives up and invents an answer.

Write errors that contain the fix:

{
  "error": "unknown_project",
  "message": "No project with slug 'Acme Corp'. Slugs are lowercase with hyphens. Closest matches: acme-corp, acme-corp-eu. Call list_projects if you need the full list.",
  "retryable": true
}

Same for empty results on writes. A write that matched zero rows is not a success. Return {"updated": 0, "warning": "no rows matched filter; verify the id"} and your silent no-op class of bug largely disappears.

Rule 5: progressive disclosure over one giant menu

You do not need to show all 31 tools on every turn. Two approaches that work:

Phase-scoped tool sets. Most real workflows are a small state machine: triage, research, act, report. Expose only the tools legal in the current phase. A triage step gets read-only lookups. The act step gets writes, and only after your guardrails passed. This cuts token overhead and removes whole categories of wrong-tool calls structurally, not by persuasion.

Tool search as a meta-tool. For large surfaces, expose 5 to 8 core tools plus find_tool(intent: string) that returns the 3 best matching definitions from a registry. The agent asks for capability, you inject the schemas just-in-time. It costs one extra hop but scales to hundreds of integrations without flattening the model’s preferences.

I default to phase-scoping for anything under ~20 tools and tool search above that.

Rule 6: measure tool-selection accuracy in CI

This is the part almost everyone skips. You cannot improve what you do not score.

Build a small labeled set: 50 to 80 realistic first-turn user messages, each with the expected tool name and the expected key arguments. Run them through the model with your real tool definitions, no execution needed for the first call. Score three numbers:

  1. Top-1 tool accuracy - did it pick the right tool?
  2. Argument validity - does the payload pass the schema?
  3. Confusion pairs - which tool got picked instead? This is your redesign backlog.

On the 31-tool client system, the baseline was 71 percent top-1. After merging chained tools into four coarse ones, adding negative clauses to nine descriptions, and phase-scoping writes, it went to 96 percent with a 12-tool surface. No model change, no prompt rewrite. That eval now runs on every PR that touches a tool definition, and a drop below 92 percent fails the build.

The confusion matrix is where the gold is. If search_documents steals 8 calls from find_customer_by_email, you know exactly which sentence to rewrite.

The checklist I run before shipping an agent

  • Count your tools. If the number is above 15 without progressive disclosure, fix that first.
  • Every tool name describes an intent, not an endpoint.
  • Every description has a “do NOT use this for” line.
  • No two tools can answer the same user sentence equally well.
  • No free-text where an enum or ID belongs, and additionalProperties: false everywhere.
  • Chains the agent always walks are collapsed into single tools.
  • Destructive actions live in separate, gated tools with no boolean escape hatch.
  • Errors name the fix and say whether a retry makes sense.
  • Zero-row writes report zero rows, not success.
  • A tool-selection eval exists and gates merges.

None of this is glamorous. It is API design with a very literal, very fast, slightly overconfident consumer. But every hour I have spent tightening a tool surface has bought more reliability than any hour spent rewriting a system prompt.

Before you upgrade the model, go read your tool definitions as if you were the model. If you would hesitate, it already did.

FAQ

How many tools is too many for one agent?

In my experience quality starts degrading noticeably past roughly 15 to 20 tools in a single request, and it degrades faster when several tools overlap in purpose. The count matters less than the ambiguity: 25 crisply separated tools can work better than 10 that all say "search" or "update". If you need more capability, do not flatten everything into one menu. Scope tools to the current workflow phase, or expose a small core set plus a tool-search meta-tool that injects definitions just in time.

Should I just use MCP servers as-is, or wrap them?

Wrap them. Third-party MCP servers are built for generality, so they expose fine-grained endpoints, generic names and permissive schemas. I treat them as a transport layer and put my own adapter in front: rename tools to match the intent in my domain, merge the chains my agent always walks, tighten the schemas with enums and required fields, and rewrite error responses so they tell the model how to recover. The wrapper is usually a couple of hundred lines and pays for itself immediately in fewer wrong-tool calls and fewer tokens spent on schema overhead.

How do I build a tool-selection eval without a big labeled dataset?

Start with 50 real user messages from your logs or from the client's support inbox. For each one write down the tool you would expect a competent human to call first plus the key arguments. That is an afternoon of work. Then run those messages through the model with your real tool definitions and only inspect the first tool call, no execution required, which keeps the whole run cheap. Track top-1 accuracy, schema validity of the arguments, and which tool was picked wrongly instead. The confusion pairs tell you exactly which descriptions to rewrite, and you can grow the set every time production surprises you.

Related articles