Blueprint mode is on — see the design decisions behind this site

Evaluating AI quality: from "looks good to me" to systematic measurement

An AI product can't be approved with a handful of manual tries. An evaluation set, clear criteria and a quality gate before each release are the only dependable way to change models and prompts without breaking the product.

In ordinary software the same input gives the same output, and you can write a test for it. In an AI product the same question may get a slightly different answer each time, and "correct" is not always black and white.

So many teams measure quality by feel: ask a few questions, the answers look fine, ship. That approach lasts until the first change.

Why feel isn't enough

  • Hand-picked samples. You tend to try the questions you know will work.
  • Invisible regression. Fixing one case breaks three others and nobody notices.
  • Model changes. When the model changes you have nothing to compare against.

Build an evaluation set

An evaluation set is a list of real inputs together with what a good output should look like.

Where samples come from

  • Real conversations. The best source is what users actually sent.
  • Failures. Every time the assistant gets something wrong, that case joins the set.
  • Edge cases. Vague, out of scope, ambiguous, angry.
  • Manipulation attempts. Inputs that try to push the assistant out of its role.

A few dozen good samples are enough to begin. What matters is that they represent reality, not that the set is large.

What to measure

Criteria depend on the product. For a sales assistant:

  • Correctness: does the information match the store's data?
  • Faithfulness to the source: did it invent anything?
  • Completeness: did it answer every part of the question?
  • Tone: does it match the defined persona?
  • Right behaviour when it doesn't know: did it hand over to a person when it should?
  • Right action: if it is an agent, did it call the right tool with the right input?

Three ways to score

Rule-based checks. For what is definite: is the price right? Was the right tool called? Is the answer within the allowed length? These are fast, cheap and reliable.

Model grading. For qualities such as tone or completeness, another model can score against a clear rubric. Useful, not infallible, and it should be compared with human judgement from time to time.

Human review. For a sample of cases, especially sensitive ones, nothing replaces a person reading the output.

The three together give a picture you can rely on.

A quality gate before release

Evaluation is worth something when it stops a worse version from shipping. The rule is simple:

  1. Any change to the prompt, the model or the retrieval method runs the evaluation set.
  2. The result is compared with the previous version.
  3. If the key scores have dropped, the change doesn't ship.

It is the same logic as the approval gates in my process, applied to the intelligent part of the product.

Measuring after release

An evaluation set won't catch everything. In production, watch:

  • Handoff rate. A sudden rise means something broke.
  • Abandoned conversations. The user left mid-task.
  • Direct feedback. A way to say "this didn't help" gives you the most valuable data.
  • Cost per outcome. Look at quality and cost together.

Each interesting case found in production goes back into the set. That loop makes the product steadily more robust.

Common mistakes

  • Testing only easy samples.
  • One overall number. A good average can hide total failure in one important category. Break results down.
  • Forgetting privacy. Real conversations must be stripped of personal data before joining the set.
  • One-off evaluation. A set that isn't updated slowly drifts away from reality.

The takeaway

What separates a demo from an AI product is evaluation. With real samples, clear criteria and a gate that stops regressions, you can change models and prompts with confidence.

Related: AI in the product: core, not decoration.

Written by

Mohammad Ali Eslamipour

Mohammad Ali Eslamipour is a digital product architect who defines, architects and — with his own engineering team — ships AI products, SaaS platforms and web apps.