How to evaluate an AI agent (and why most teams skip it)
The demo works. Everyone’s impressed. You ship it — and three weeks later it’s quietly giving wrong answers to real customers and nobody noticed until one complained. This is the most common way AI projects fail, and it has one root cause: no evaluation. Here’s how to do the thing most teams skip.
”It works” is not a measurement
LLMs are non-deterministic — the same input can give different outputs, and a tiny prompt change can quietly break a case that used to work. Testing by eyeballing five prompts feels fine and tells you almost nothing. At any real scale, “it seems to work” is a guess, not a fact.
An eval harness turns that guess into a number. It’s the single highest-leverage thing you can build for a production AI system — and the clearest signal of whether the team building your agent actually knows what they’re doing.
What an eval harness actually is
It’s a set of test cases — real inputs paired with a definition of what a good response looks like — that you run automatically and score. Run it on every change. If the score drops, you caught a regression before your users did.
That’s the whole idea. The craft is in how you score.
The four ways to score an output
- Rule-based / deterministic. Does the output contain the required fields? Valid JSON? The right status? Cheap, fast, and perfect for structured tasks.
- Reference-based. Compare the output to a “gold” answer — exact match, or semantic similarity for free-text. Good when there’s a known-correct response.
- LLM-as-judge. Use a strong model to grade outputs against a written rubric (“is this accurate, grounded, and on-tone?”). Scales to fuzzy quality judgments that rules can’t capture — just validate the judge against human ratings first.
- Human review. For the highest-stakes cases, nothing beats a person. Use it sparingly, on a sample, to keep the automated scorers honest.
Most real systems use a mix: rules for the cheap checks, LLM-as-judge for quality, humans on a sample.
How to build one (step by step)
- Collect real examples. Pull actual queries — from logs, support tickets, or your own test set. Real inputs beat invented ones.
- Define “good” per case. For each, write down what a correct/acceptable answer must do. This is the work; do it honestly.
- Automate the scoring. Wire up the rule/reference/judge checks so the whole set runs with one command.
- Set a threshold. Decide the passing bar (e.g., 90% task success, zero hallucinations on the safety set).
- Gate every change behind it. No prompt, model, or RAG change ships unless the eval passes. This is the part that actually protects you.
The metrics that matter
- Task success rate — did it actually accomplish the goal?
- Groundedness / hallucination rate — are claims supported by retrieved context? (This is where RAG earns its keep.)
- Format compliance — valid structure every time, for anything downstream.
- Latency & cost — quality you can’t afford to run isn’t quality. (See AI agent cost.)
The loop is the point
Evals aren’t a one-time gate — they’re a loop. Run the suite, find the failures, fix the cause (a better prompt, improved retrieval, or — only if you’ve proven you need it — a fine-tune), then re-run and ship behind the gate. Each turn of the loop makes the agent measurably better instead of mysteriously different.
If you can’t answer “how do you know it got better?” with a number, you’re not improving your agent — you’re just changing it.
The mistakes we see most
- Testing on a handful of prompts. Five happy-path examples hide every edge case.
- No regression gate. Without one, every fix risks silently breaking something else.
- Evaluating once. Real-world edge cases keep arriving; the eval set has to grow with them.
- Optimising the wrong metric. A higher “helpfulness” score is worthless if hallucinations went up too. Measure the things that actually matter to your users.
How we approach it
At Malgary Labs, every agent we build ships with an eval harness — tied to real examples, with scoring automated and changes gated behind it. It’s not glamorous, but it’s the difference between an agent you can trust in production and a demo that embarrasses you later. It’s also why our AI agent builds hold up after launch instead of drifting.
Building something where “mostly right” isn’t good enough? That’s exactly where evals pay for themselves — book a free call and we’ll show you what a real eval setup looks like for your use case.