Your LLM Classifier Needs a Taxonomy, Not a Better Prompt
Most LLM classification work fails because the labels are undefined, not because the prompt is weak. Here is the workflow I use: build the taxonomy, extract facts instead of verdicts, and let code decide.
Pavel Duglas
AI Automation & MVP Architect
Every few weeks a client shows me the same thing: a prompt that ends with “Return one of: urgent, normal, spam” and a complaint that “the model is inconsistent”. Then I ask two people on their team to label the same twenty messages by hand and they disagree on six of them. The model is not inconsistent. The labels are undefined. Nobody can beat a prompt into agreeing with a rule that does not exist yet.
Classification is the most common real use of LLMs in business automation - ticket triage, lead scoring, moderation, categorizing scraped listings, routing inbound messages in a Telegram bot. It is also where people skip all the boring engineering and then blame the model. Here is the workflow I actually use.
Step 1: A label must map to a different action
Before you write a single line of prompt, write down what happens downstream for each label. If two labels lead to the same action, they are one label. If one label leads to three different actions depending on context, it is three labels, or you are missing a field.
Bad taxonomy for support tickets:
billing,technical,question,complaint,other
Those overlap. A complaint about a failed charge is billing and complaint and technical. The model has to guess your priority order, and it will guess differently on Tuesday.
Better taxonomy, defined by action:
refund_request-> goes to the finance queue with a 4 hour SLAaccess_blocked-> goes to on-call with a 30 minute SLAhow_to-> auto-reply with docs link, no humansales_inquiry-> pushed to CRMunclear-> human triage queue
Rules I hold to:
- Mutually exclusive. If a message can honestly be two labels, add an explicit tie-breaker rule: “if the user cannot log in AND asks for a refund, label
access_blocked.” - Exhaustive, with an escape hatch. Always include
unclearorother. Without it the model invents confidence it does not have. - Max 7-9 labels per level. More than that, go hierarchical: coarse label first, then a second pass within the branch. Flat taxonomies of 40 categories are where accuracy goes to die.
- Write a rubric, not a word. Each label gets a one-line definition, two positive examples, one near-miss negative example, and any tie-breaker. This rubric is your prompt’s core and also your onboarding doc for human reviewers.
If two humans on your team cannot agree using the rubric, fix the rubric. Ambiguity is a taxonomy bug, and no model upgrade will fix it.
Step 2: Ask the model for facts, let your code decide
This is the single change that has improved more of my pipelines than any model swap. Instead of asking for a verdict, ask for observable facts, then apply business rules in code.
Compare. Verdict style:
Classify this ticket as urgent / normal / low.
Fact style:
{
"mentions_payment": true,
"mentions_cannot_access": false,
"has_order_id": true,
"asks_for_money_back": true,
"customer_tone": "frustrated",
"language": "en",
"quoted_error_code": null
}
Then:
def route(facts, account):
if facts["mentions_cannot_access"] and account.plan == "paid":
return "access_blocked"
if facts["asks_for_money_back"]:
return "refund_request"
if facts["customer_tone"] == "frustrated" and account.mrr > 500:
return "escalate_human"
return "how_to"
Why this wins in production:
- Policy changes do not touch the prompt. Finance decides refunds over $200 need manager approval? That is a Python
if, not a prompt rewrite plus a re-eval of the whole taxonomy. - Debuggable. When a ticket routes wrong you can see whether the model misread the text or your rule was wrong. With a single verdict you get one word and no explanation you can trust.
- Cheaper and more stable. Fact extraction is a narrower task. A small model does it reliably. Judgement calls that need a big model are where you spend tokens.
- You can blend in data the model never sees. Account MRR, signup date, prior ticket count. The model does not need them, your rules do.
This is what people mean when they say classification is feature engineering. The LLM is a feature extractor for messy text. The decision stays in code where you can test it.
Always force structured output: JSON schema with enums and booleans, no free-text label field. If your provider supports strict schema mode, use it. If not, validate and retry once, then fall back to unclear rather than accepting garbage.
Step 3: Build a golden set before you tune anything
150 to 300 real examples, hand-labeled by the person who owns the downstream action. Not synthetic. Not the twelve examples you remember.
How I compose it:
- Stratified across labels, including the rare ones. Random sampling from production will give you 90%
how_toand teach you nothing about the labels that matter. - 20% deliberately hard cases: the ambiguous ones, the multi-intent ones, the ones your team argued about.
- 5% garbage: empty strings, emoji only, wrong language, a pasted stack trace, HTML from a broken scrape.
- Store it as a file in the repo with an
id, the raw input, the expected label, and expected facts. Version it.
Then run it in CI. Any change to the prompt, the model, the temperature or the schema runs the golden set and prints per-label precision and recall. Not accuracy. Accuracy on an imbalanced set is a comfort metric: if 85% of items are how_to, a classifier that only ever says how_to scores 85% and fails at the job.
What I look at:
- Recall on the expensive-to-miss labels. Missing
access_blockedcosts you a churned customer. Missinghow_tocosts nothing. - Precision on the labels that trigger automated actions. If
refund_requestauto-creates a refund ticket, false positives cost money. - The confusion matrix. One pair of labels usually causes most of the errors. That pair is a taxonomy problem, not a model problem, 80% of the time.
Set explicit thresholds in CI: recall on access_blocked must stay above 0.95, precision on refund_request above 0.90. Fail the build below that. Now your prompt is a tested artifact, not a lucky string.
Step 4: Cheap model, then escalate
A three-tier routing pattern I reuse everywhere:
- Deterministic pre-filter. Regex and keyword rules catch the obvious 20-40%. An exact match on “unsubscribe” needs zero tokens. Blank input needs zero tokens.
- Small model for fact extraction on everything else. This handles the bulk.
- Large model only for items the small model flagged as
unclear, or where facts conflict, or where the item crosses a value threshold (enterprise account, order above X).
For confidence, do not trust a self-reported "confidence": 0.93 field. Models will happily produce that number for a total hallucination. Two things that actually work:
- Self-consistency. Run the small model three times at temperature 0.7. Three matching answers is a strong signal, a split is your escalation trigger. Costs 3x on the cheap tier, still cheaper than always using the big model.
- Token logprobs on the label token, if your provider exposes them. Cheaper than self-consistency and reasonably calibrated after you bucket it against your golden set.
And always keep a human queue. A classifier that sends 5% of items to a person and gets the other 95% right is a real system. One that pretends to handle 100% is a liability.
Step 5: Log like you will be audited, monitor distribution
For every classification, store: input hash, raw input (or a pointer), extracted facts, final label, prompt version, model name and version, latency, cost, and the tier that produced it. Later, join the downstream outcome: did the refund get approved, did the lead convert, did a human overturn the label.
That overturn rate is the best production metric you have. It is free supervision. Every human correction is a new golden set candidate. Once a week I pull the overturns plus a 1% random sample and review them in a spreadsheet for fifteen minutes.
Monitor the label distribution per day, not just error rates. If how_to was 60% of volume for three months and drops to 30% overnight, either your product broke, your traffic source changed, or your upstream scraper started returning different HTML. In parsing pipelines this is usually the last one: the site changed layout, your extracted text is now navigation menus, and the classifier is dutifully labeling menus. Distribution shift alerts catch this days before anyone notices bad routing.
One more thing: pin the model version. “Latest” is not a version. A silent provider update can shift a borderline label pair by several points, and you will spend a day looking for a bug in your code.
The short version
Write the taxonomy by action. Make the model extract facts, not verdicts. Keep decisions in code. Build a 200-row golden set and run per-label metrics in CI. Escalate instead of guessing. Log the overturns and watch the distribution.
None of that is exciting and none of it depends on which model is winning this month. That is exactly why it still works six months after you ship it.
FAQ
Should I fine-tune a model instead of prompting for classification?
Only after you have a stable taxonomy and a few thousand labeled examples from production. Fine-tuning locks in your label definitions, so doing it before the taxonomy settles means paying to bake in your mistakes. In practice, a good rubric plus fact extraction plus a small model handles most business classification well enough, and the money is better spent on the golden set. Revisit fine-tuning when latency or per-item cost becomes the actual bottleneck.
How do I handle items that legitimately belong to two categories?
Decide whether multi-label is real or whether your taxonomy is wrong. If a message truly contains two independent intents that trigger two independent actions, return an array of labels and let each action fire once, with idempotency keys so nothing runs twice. If the two labels lead to the same queue and just feel different, merge them. If one should win, write the tie-breaker rule explicitly into the rubric and add a test case for it in the golden set.
What accuracy is good enough to automate an action?
It depends on the cost of being wrong, not on a universal number. For an action that is cheap to reverse, like tagging or sorting, precision around 0.85 is usually fine. For anything that touches money, sends an external message, or deletes data, I want precision above 0.97 on that specific label plus a reversible path, and everything else goes to a human queue. Measure per label, set the threshold per label, and enforce it in CI.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks