Stop Scoring AI Output by Vibes. Measure How Much Humans Fix It
Most AI features are judged by demos and gut feeling. I track one thing instead: how much a human has to change the output before it ships. Here is how to capture it, compute it and use it to pick models and prompts.
Pavel Duglas
AI Automation & MVP Architect
Every AI feature I have shipped looked great in the demo. Then real users touched it, and the question changed from “is the answer good?” to “how long did someone spend fixing it?” That second question is the only one that maps to money. If a support agent rewrites 70% of every AI draft, your feature is a slower way to type. If a parser output needs two fields corrected per record, your automation is a data entry job with extra steps.
So in every system where a human reviews AI output, I log one thing religiously: the difference between what the model produced and what the human actually shipped. I call it correction burden. It replaced almost every other quality metric I used to track.
Why the usual metrics lie
The standard ways of judging AI output all have the same flaw: they measure something adjacent to value, not value itself.
- Manual spot checks tell you what you noticed on a Tuesday afternoon. They do not scale and they drift with your mood.
- LLM-as-judge scores are useful for regression tests, but a judge happily gives 8/10 to a polished answer that a domain expert would throw away in two seconds.
- Thumbs up / thumbs down buttons get clicked by maybe 2% of users, and mostly when they are angry.
- “Accuracy” on a test set is frozen in time. Your inputs in month three look nothing like the 50 examples you wrote in week one.
Correction burden is different because it is a byproduct of work people already do. You are not asking anyone to rate anything. You are just watching what they change.
What to capture
The whole idea rests on storing two versions of every output: the AI draft and the human final. Most systems throw the draft away the moment the user edits it. That is the mistake.
Here is the minimal record I keep:
CREATE TABLE ai_outputs (
id uuid PRIMARY KEY,
task_type text NOT NULL, -- 'reply_draft', 'invoice_extract', ...
customer_id uuid,
build_id text NOT NULL, -- prompt + model + tools version
input_hash text NOT NULL,
draft jsonb NOT NULL, -- what the model produced
final jsonb, -- what the human shipped
outcome text, -- 'accepted', 'edited', 'rejected', 'abandoned'
shown_at timestamptz NOT NULL,
resolved_at timestamptz
);
Four details matter more than they look:
build_idties every output to the exact prompt, model and tool configuration that produced it. Without it you can compute burden but you cannot explain it.outcomeseparates “edited” from “rejected”. A user who deletes the draft and writes from scratch is a very different signal from one who fixes a typo.shown_atandresolved_atgive you time-to-accept, which catches cases where the text barely changed but the human spent four minutes verifying it.abandonedis its own state. If people open the draft and walk away, that is a failure the diff will never show.
The metrics I actually compute
From that table I derive a small set of numbers per task type, per build and per customer segment.
Accept-as-is rate
The share of outputs shipped with zero meaningful changes. This is the headline number. For a reply drafting bot I built for a Telegram support channel, it started at 18% and got to 61% over six weeks, almost entirely through prompt changes driven by the edit data.
Normalized edit distance
For free text, I compute a token-level distance between draft and final, divided by the length of the longer one. Character-level diffs overreact to reformatting, so I normalize whitespace and punctuation first.
import difflib, re
def normalize(text: str) -> list[str]:
text = re.sub(r"\s+", " ", text.strip().lower())
return re.findall(r"\w+|[^\w\s]", text)
def edit_burden(draft: str, final: str) -> float:
a, b = normalize(draft), normalize(final)
if not a and not b:
return 0.0
ratio = difflib.SequenceMatcher(None, a, b).ratio()
return round(1 - ratio, 3) # 0 = untouched, 1 = fully rewritten
This is not academically perfect. It does not need to be. It is consistent, cheap and good enough to rank builds against each other.
Field-level correction rate
For structured output, text distance is the wrong tool. If a parser extracts invoices into JSON, I diff field by field and count which fields humans changed.
def field_corrections(draft: dict, final: dict) -> dict[str, bool]:
keys = set(draft) | set(final)
return {k: draft.get(k) != final.get(k) for k in keys}
Aggregate that and you get something like: vendor_name corrected 3% of the time, due_date 22%, tax_amount 9%. Now you know exactly where to spend effort. In my experience it is almost always dates, currencies and anything involving time zones, not the “hard” semantic fields.
Time-to-accept
Median seconds between showing the output and resolving it. When accept-as-is rate is high but time-to-accept is also high, users do not trust the output. They are reading every line. That is a trust problem, and the fix is often showing sources or confidence, not changing the model.
Using the numbers to make decisions
Collecting metrics is pointless unless they change what you ship. Here is where correction burden earns its place.
Picking models on cost per useful output
The classic debate is “should we use the expensive model?” Correction burden turns it into arithmetic. Say the expensive model costs $0.02 per draft with 12% average edit burden, and a cheap one costs $0.002 with 19%. If your reviewer earns $30 an hour and each point of burden adds roughly 3 seconds of editing, the extra 7 points cost about 21 seconds, or $0.17 of human time. The expensive model wins by a mile.
Flip the scenario: a batch extraction job where 95% of records are never reviewed. Burden only matters on the reviewed slice, and the cheap model wins. Same metric, opposite answer, and neither is a guess.
Comparing prompt versions without arguing
Because every row has a build_id, I roll out a new prompt to 20% of traffic and compare burden after a few hundred outputs. No opinions, no “I feel like the new one is better.” If burden does not drop, the change does not ship. This has killed a lot of my own clever prompt ideas, which is exactly the point.
Finding the segments that break
Averages hide the pain. I always slice burden by customer, input language and input source. On one project the global accept-as-is rate looked fine at 55%, but for Vietnamese-language messages it was 9%. Nobody had complained because those users had quietly stopped using the feature.
Turning edits into training data
The most underrated part of this setup: every human edit is a free labeled example. The draft is the wrong answer, the final is the right one, and you did not pay anyone to annotate.
What I do with it:
- Build the regression set. Every week I pull the 20 highest-burden outputs and add their input plus human final to the eval suite. The test set grows in the direction of real failures instead of my imagination.
- Mine few-shot examples. Edited pairs make excellent examples for the prompt, especially for tone and formatting rules that are hard to describe in words.
- Spot new categories of failure. When a cluster of edits all do the same thing, like removing a greeting or fixing a currency symbol, that is a rule you should write down, not a model you should replace.
The traps
This metric is not free of noise. The ones that bit me:
Rubber-stamping. Some reviewers accept everything because they are busy. Their burden is zero and meaningless. I track per-reviewer accept rates and treat anyone at 99% with suspicion. A few random audits per week keep the numbers honest.
Cosmetic edits. People change “Hi” to “Hello” out of habit. Normalization helps, and for text I ignore edits below a small threshold (around 0.03) when computing accept-as-is rate.
Silent rejection. If your UI lets users ignore the draft and type in a separate box, you will never see the rejection. Design the flow so the draft is the starting point of the final, or log when it is discarded.
Privacy. You are storing human-written finals, which may contain more sensitive data than the drafts. Apply the same retention and access rules as the rest of your customer data, and do not ship finals to a third-party eval tool without thinking about it.
Goodhart. If the team is rewarded for low burden, someone will eventually make edits harder. Keep it a product metric, not a performance review metric.
Where to start this week
You do not need a platform. For an existing AI feature with a review step:
- Stop overwriting the draft. Store it next to the final.
- Add
build_idandoutcometo every record. - Write the two functions above and run them nightly.
- Put accept-as-is rate and median burden per task type on one dashboard.
- Pull the worst 20 examples every Friday and read them.
That last step is where most of the value lives. The dashboard tells you something is wrong. The worst 20 tell you what. After a month you will know more about your AI feature than any benchmark could tell you, and every decision about models, prompts and cost will have a number behind it instead of a feeling.
FAQ
What if my AI automation has no human review step?
Then you need a sampled one. Route 2-5% of outputs to a reviewer, or to yourself, and log corrections the same way. For fully automated pipelines you can also use downstream corrections as the signal, like records later edited in the CRM or refunds issued after an automated decision. Without some ground truth from humans you are flying blind.
How many examples do I need before comparing two prompt or model versions?
For a rough ranking, a few hundred outputs per build is usually enough if the difference is meaningful, say 5 or more points of accept-as-is rate. For small differences you need more, and you should compare within the same task type and customer segment. I would rather run a test a week longer than ship a change based on 40 examples.
Is edit distance fair for outputs where many different answers are acceptable?
Not perfectly. Two valid replies can look very different, so a human rewrite does not always mean the draft was wrong. That is why I pair edit distance with outcome and time-to-accept, and why I read the worst examples manually. Use edit burden to compare builds against each other on the same traffic, not as an absolute truth about quality.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks