AI in SaaS

The eval suite is the new product surface

Shipping AI features without evals is shipping software without tests. What the eval discipline actually looks like.

For traditional software, a regression is when a test that used to pass now fails. For AI features, a regression is when a model update or prompt change quietly makes the experience worse for a slice of users you have not tested. Without an eval suite, you will not know until churn tells you.

The structural problem is that AI regressions are invisible at the point of change. An engineer tweaks one sentence of a prompt to fix a complaint, the fix works, everyone moves on, and three unrelated behaviours degraded silently, because prompts are global variables with non-local effects. The same is true when a provider updates a model underneath you on their schedule, not yours. Code diffs show what changed; behaviour diffs are something you have to build. That is what an eval suite is.

What an eval suite is

An eval suite is a structured set of representative inputs, expected behaviours, and a way to score the actual behaviour against them. It is not unit tests; it is closer to a regression test for a fuzzy system. It runs every time you change a prompt, model, retrieval pipeline, or routing rule, and it tells you whether you are improving or breaking things.

Scoring comes in three tiers, and good suites mix them deliberately. Assertions where determinism allows: the output parses, cites a source, contains no forbidden claims, stays under length. Model-graded checks for fuzzy qualities: a second model scores tone, faithfulness to the source, or completeness against a rubric, which is imperfect but consistent, and consistency is what regression detection needs. Human review for the small golden set where judgement is the product. Cheap assertions catch most breakage; the graded tiers catch the drift assertions cannot see.

The discipline that makes them work

Evals are only as good as the inputs in them. They have to be drawn from real user interactions, including the edge cases that hurt. New failure modes from the wild get added back to the suite; passing tests stay there forever. The suite is a living product surface, not a one-off.

The operational rule that keeps the suite alive: every AI bug report becomes an eval case before it becomes a fix, exactly as regression tests work in mature codebases. The fix is only believed when the new case passes and every old case still does. Teams that skip this re-fix the same failure quarterly and each fix silently trades against the last one; the suite is what turns that circle into a ratchet, where solved stays solved.

What to ship first

Start with twenty cases covering the most common user intents and the three or four edge cases that already burned you. That suite, run automatically on every change, prevents most of the regressions that quietly erode AI products. Expand from there.

Wire it into the same muscle memory as CI: prompt and pipeline changes ship through pull requests, the suite runs on each one, and a score drop blocks the merge with a diff of exactly which cases moved. Keep the blocking tier fast (a hundred assertion cases run in a couple of minutes) and push the slower graded tiers to a nightly run. When a model provider announces an update, run the full suite against the new model before switching, and you have converted the scariest event in AI product maintenance (the involuntary upgrade) into a reviewable diff. That is also the quiet competitive moat: two teams using the same models differ mostly in who can change things without breaking things, and the eval suite is that capability, made durable.

Takeaways

What to do with this

Related

Keep reading

Put the playbook to work.

Cafiyn Lens tells you which market is worth the effort, and Cafiyn FlyWheel runs the acquisition loop against it. Two products, one shared Blueprint, from $14.99/mo.