You Don't Know What Each Customer Costs You: AI Cost Attribution for SaaS
A practical metering and cost-attribution setup for SaaS with LLM features: usage event schema, a money ledger, showback vs chargeback, and margin alarms that catch a loss-making customer before your invoice does.
Pavel Duglas
AI Automation & MVP Architect
Every AI SaaS I have audited in the last year had the same blind spot. The founder can tell me monthly revenue to the cent. Ask them what their top ten accounts cost to serve and the answer is a shrug plus a screenshot of a single OpenAI invoice. That is not an accounting problem. It is a product problem, because you cannot price, rate limit, or fire a customer you cannot measure.
This is the setup I put in place, usually in two or three days of work, before touching prompt optimization or model routing. Cutting costs without attribution is guessing. Attribution first, then optimization.
The question you can’t answer today
Here is the test. Pick your five biggest accounts and answer these for last month:
- Inference spend per account, split by model.
- Gross margin per account after inference, storage, and third-party API calls.
- The single most expensive feature in your product per unit of use.
- How many dollars you burned on requests that failed and got retried.
- How much your free tier cost you.
If any of those takes more than a SQL query, you are flying blind. And AI costs are not like server costs. Servers scale with users. LLM costs scale with user behavior, which varies wildly. One power user pasting 60-page PDFs into your summarizer can cost 200 times more than the median user on the same $49 plan. That does not average out. It concentrates.
Meter at the boundary, not in the prompt
The mistake I see most often is logging usage inside business logic, scattered across a dozen call sites. Six weeks later half of them are missing a tenant id and the numbers do not reconcile.
Put one wrapper around every paid external call. Not just the LLM. Every call that has a price tag: transcription, embeddings, search APIs, scraping proxies, OCR, image generation. That wrapper is the only place allowed to talk to a vendor SDK, and it always emits a usage event.
// the only function allowed to call a paid vendor
async function meteredCall<T>(ctx: CallContext, fn: () => Promise<VendorResult<T>>) {
const startedAt = Date.now();
try {
const res = await fn();
await usage.emit({
request_id: ctx.requestId, // idempotency key
tenant_id: ctx.tenantId,
user_id: ctx.userId,
feature: ctx.feature, // "doc_summary", "lead_enrich"
vendor: ctx.vendor, // "openai", "deepgram"
model: res.model, // what actually ran, not what you asked for
input_tokens: res.usage?.input ?? 0,
cached_input_tokens: res.usage?.cachedInput ?? 0,
output_tokens: res.usage?.output ?? 0,
units: res.usage?.units ?? 0, // minutes, pages, images
attempt: ctx.attempt, // 1, 2, 3...
outcome: "ok",
latency_ms: Date.now() - startedAt,
});
return res.value;
} catch (e) {
await usage.emit({ ...ctx, outcome: classify(e), attempt: ctx.attempt });
throw e;
}
}
Two details that matter more than they look.
Record the model that actually ran. If you have a router, a fallback, or a provider that silently serves a different snapshot, the model you requested is not the model you pay for. Read it off the response.
Record the attempt number and the outcome. Failed calls still cost money in most setups, and retries are where phantom spend hides. If you count every attempt as one unit of customer usage, you overcharge. If you count none of the failures, your internal cost numbers come in low and your margin looks better than it is. Keep both: billable units and incurred cost are different columns.
Don’t trust your own token math
Estimating tokens with a local tokenizer and multiplying by a price list is where attribution quietly breaks. Things that wreck naive math:
- Prompt caching. Cached input tokens are often billed at a fraction of the normal rate. If you ignore the distinction you will overstate cost for your heaviest, most cache-friendly customers, which is exactly the group whose pricing you care about.
- Reasoning tokens. On reasoning models, output tokens include tokens you never see in the response text. Estimating from the visible string undercounts badly.
- Tool loops. One user action can trigger seven model calls. If you meter per user action instead of per vendor call, agents will make you look profitable right until the invoice lands.
- Double counting on retry. I once traced a 30% discrepancy to an inner retry inside an SDK plus an outer retry in the job queue, both emitting events. The request_id idempotency key fixed it in one commit.
At the end of each month, reconcile your own ledger totals against each vendor invoice. If the gap is above 3%, do not go build pricing on top of it. Find the leak first. A ledger nobody has reconciled is a rumor.
From events to money: the ledger
Usage events are raw and high volume. Money lives in a separate, small, append-only table.
create table cost_ledger (
id bigserial primary key,
day date not null,
tenant_id uuid not null,
feature text not null,
vendor text not null,
model text not null,
billable_units numeric not null, -- what you charge for
incurred_cost_usd numeric not null,-- what you actually paid, retries included
price_version text not null, -- which rate card was applied
unique (day, tenant_id, feature, vendor, model, price_version)
);
The price_version column is the one people skip and regret. Vendor prices change. Your own margin analysis from March must stay reproducible in September. Store the rate card as data with an effective date range, never as constants in code, and stamp every ledger row with the version used.
Roll up nightly from events into the ledger. Keep raw events for 30 to 90 days for debugging, then drop them. The ledger is tiny and you keep it forever.
Once that exists, the interesting queries are one liner territory:
select tenant_id,
sum(incurred_cost_usd) as ai_cost,
sum(incurred_cost_usd) / nullif(mrr, 0) as cost_ratio
from cost_ledger join subscriptions using (tenant_id)
where day >= date_trunc('month', current_date)
group by tenant_id, mrr
order by cost_ratio desc
limit 20;
That ordered list is the single most useful report in an AI SaaS. It tells you who to talk to, who to rate limit, and which plan is mispriced.
Showback before chargeback
Attribution has two different destinations and mixing them up causes a lot of pain.
Showback means you tell the customer what they consumed but the bill does not change. Internally, it also means you tell each team or feature owner what they spent without moving budget around.
Chargeback means consumption directly drives the invoice. Credits, overage, per-seat-with-limits, metered billing.
Start with showback. Always. Ship a usage panel that shows the customer: documents processed, minutes transcribed, agent runs, and how far into their plan allowance they are. No dollars yet. You get three wins immediately. Customers self-regulate when they can see a bar filling up. Your support team stops guessing. And you find out whether your metering is even correct before money depends on it.
Then watch for the signals that it is time for chargeback:
- Your cost ratio distribution has a long tail, with a handful of accounts above 40% of their own MRR.
- Support is repeatedly negotiating exceptions to invisible limits.
- Sales cannot answer “what happens if we triple volume next quarter?”
When you move to chargeback, move on a unit the customer understands and can predict. Tokens are a terrible billing unit for a business buyer. “Pages analyzed”, “calls transcribed”, “leads enriched” are good ones. Internally you convert units to cost using the ledger; externally you sell whole actions. Keep a healthy conversion buffer, because your cost per action will move as models change and you do not want to reprice every quarter.
One migration rule I insist on: run showback and chargeback in parallel for at least one full billing cycle. Show the customer the invoice they would have received. Nothing burns trust faster than a surprise metered bill built on metering that has never been audited.
The margin alarm
Budgets that are only checked monthly are not controls, they are postmortems. Two alarms worth wiring up on day one:
- Per-tenant daily cost anomaly. If a tenant’s daily AI cost exceeds 3x its trailing 14-day median and is above a floor of a few dollars, notify a human. Most of the time it is a legitimate bulk import. Occasionally it is a loop, a leaked API key, or someone running your free tier as a batch pipeline.
- Feature level cost ceiling. Each feature gets a hard daily spend cap. Hitting it degrades the feature gracefully: queue the work, drop to a cheaper model, or return a clear “we are catching up” state. Never let one feature silently consume the whole month’s budget.
Degrade, do not crash. A queued job with an honest ETA is a fine product experience. A 500 error is not.
Common mistakes
- Attributing at user level only. In B2B you need tenant, user, and feature. Without feature you know who is expensive but not why.
- No cost on background work. Nightly re-embedding, scheduled agent runs, and cache warmers belong to a tenant too. Tag them with a system actor and the tenant they serve.
- Free tier with no ceiling. Free users should have a hard unit cap enforced in code, not a polite note in the docs.
- Averages in the pricing deck. Median cost per account is almost useless. Look at p90 and p99, because that is where churn-inducing rate limits and margin holes live.
- Instrumenting after the growth spurt. Attribution added during a crisis is attribution nobody trusts. It takes two days now and two weeks later.
Start this week
One wrapper around paid calls. One usage event table with an idempotency key. One nightly rollup into a ledger with a versioned rate card. One query sorting tenants by cost-to-MRR ratio. One showback panel in the UI. One anomaly alert into Slack.
That is the whole thing. Do it before you tune a single prompt, because once you can see per-customer margin you will usually discover the fix is a plan change or a limit, not a cheaper model.
FAQ
Can't I just use my LLM provider's dashboard instead of building this?
Provider dashboards show spend by API key and model, not by customer, feature, or retry attempt. They also cannot join to your MRR, so they can never tell you gross margin per account. Provider dashboards are useful for reconciliation - you compare your ledger totals against the invoice each month - but they cannot answer the questions that drive pricing and rate limiting decisions.
What is the difference between showback and chargeback in practice?
Showback means you report consumption without changing the bill: the customer sees a usage panel, you see cost per account, but the invoice stays flat. Chargeback means consumption drives the invoice through credits, overage, or metered billing. Start with showback so you can validate that your metering actually reconciles with vendor invoices, then move to chargeback once the long tail of expensive accounts justifies it. Run both in parallel for one full billing cycle before the first metered invoice goes out.
Should I bill customers in tokens?
Almost never. Tokens are unpredictable for the buyer, they change meaning when you swap models, and they invite arguments you cannot win. Bill in units the customer can forecast, like documents processed, minutes transcribed, or leads enriched. Internally, keep the token-level ledger so you know your cost per unit and can protect margin as model prices shift. The conversion between the two is your business, not the customer's.
Related articles
Done for you
I will build a platform with accounts, roles and payments
A user area, an admin area, payment and CRM integrations, and a structure that survives the second version.
from $3,000 · 3 to 5 weeks