Every AI feature has a marginal cost that scales with usage, unlike every other software feature you have shipped. Token costs are real, and a popular AI feature can quietly turn a profitable product unprofitable. The trap is measuring success in usage instead of in resolved tasks per dollar.
The trap has a mechanism worth naming: for twenty years, SaaS teams were trained that usage is the leading indicator of retention, so more usage is always good. AI features break that training. A user who runs your assistant nine times to get one useful answer generates nine times the cost and one unit of value, and your usage dashboard reports it as engagement. The dashboard is not lying; it is answering the wrong question.
The right unit of measurement
Usage is a vanity metric for AI features. The honest metric is cost per resolved task: the total inference cost divided by the number of user tasks the AI actually completed. A feature with high usage and low resolution is a money pit. A feature with modest usage and high resolution is a moat.
Defining "resolved" is the real work, and it should be behavioural before it is survey-based: the user accepted the draft, applied the suggestion, did not retry the same request, did not escalate to support. Retries are the most underused signal in AI products; a burst of near-identical queries is a user telling you the first answers failed, and it inflates both your usage metric and your bill simultaneously. Instrument resolution loosely at first (accepted or not-retried within an hour is a fine v1) and tighten it as you learn what your product's real completion looks like.
What changes when you measure it
You stop adding tokens and start adding judgement. You route easy queries to cheap models and reserve the expensive model for the queries where it makes a measurable difference. You add a "did this resolve your task" feedback loop, because without it you have no denominator. You renegotiate provider contracts based on real usage instead of forecasted usage.
The distribution matters as much as the average. Pull the histogram and you will almost always find the same shape: most tasks resolve for fractions of a cent, and a thin tail of pathological ones (huge contexts, retry loops, agent runs that wander) consumes a third of the bill. That tail is where the engineering goes: context caps, retry budgets, early-exit conditions, and routing the genuinely hard cases to the expensive model once instead of the cheap model five times. Fixing the tail routinely cuts the bill by double digits without touching the experience of the median user.
The pricing implication
Once you know cost per resolved task, you can price the feature properly. Outcome-based or task-based pricing aligns with the unit economics. Per-seat pricing for a feature with variable marginal cost is a slow leak.
This is the quiet reason task-priced AI products defend margin better than seat-priced ones: the unit of revenue and the unit of cost are the same object, so heavy usage grows revenue instead of eroding it. It is the same logic behind Wheels-style pricing, where the thing metered is the thing that costs. If you keep seats for simplicity, at least model your P95 user's task volume and set the fair-use line before launch, not after the first invoice surprises you. A weekly cost-per-resolved-task report, split by feature and by model, is one dashboard and it pays for itself the first month.