AI, Software Development
LLM evaluation: what to measure and how to trust the score
TL;DR: LLM evaluation measures if your AI feature gives right, grounded and safe answers, and keeps measuring after launch. Split it into model, system, RAG and agent tests. Pick two or three metrics with a pass mark, check any LLM judge against answers you scored by hand, and lock down prompts, versions and scoring rules so you can rerun the test and trust the number.
LLM evaluation is how you prove your AI feature works. You measure if the answers are right, if they stick to your data, and if they're safe. Then you keep measuring, so you know the moment something slips.
Most teams get stuck on the wrong part. They hunt for the perfect metric. The bigger win is a setup you can rerun next month and get a number you can compare.
So here's what to measure, which methods hold up, and how to make the score something you can trust.
What LLM evaluation covers
LLM evaluation isn't one job. It's four, and mixing them up is the most common mistake.
- Model evaluation tests the raw model on its own. Can it reason? Does it make things up? You're testing the engine, not the car.
- System evaluation tests your app. That's the prompts, the retrieval, the guardrails and the code around the model. A great model in a weak system still gives bad answers.
- RAG evaluation is a slice of system evaluation. It checks two things: did retrieval find the right documents, and did the answer stick to them?
- Agent evaluation checks a whole task. Did the agent pick the right tools and reach the right outcome over many steps?
Each one needs its own test data. Model tests lean on public benchmarks. System tests need real user questions. RAG tests need your documents paired with questions that expose gaps. Agent tests need task scripts with a clear pass or fail. Our AI Orchestrators guide to AI agent testing walks through ten scenarios for that last one.
Get the scope wrong and you'll tune the model when the real fault sits in your retrieval layer. That's weeks gone on the wrong fix.
Offline, online and in your pipeline
Timing matters as much as metrics.
- Offline. Run the system against a fixed test set before you ship. It's cheap and repeatable. The catch? No test set covers every weird thing a real user types.
- Online. Watch live traffic. Thumbs up and down, escalations to a human, repeat questions. This is where you find the failures you didn't think of.
- A/B tests. Split users between two versions. It's the only way to prove a change moved a business number, not just a test score. Run it once offline checks pass.
- In CI. Every prompt change, model update or new index runs the test set before merge. That's where regression checks stop being a chore and become part of the build.
We've already covered the CI side in detail. Our post on LLM unit testing sets out what to run on each commit, each pull request and each night. And CI/CD for AI shows where the eval gates sit and how to wire rollback to the same score. This post is about the step before that: what the gate should measure.
The metrics that tell you something
No single number sums up quality. You need a small set, and you need to know what each one answers.
Reference-based metrics compare the output to a known right answer.
- Exact match and accuracy work for short, closed answers like labels. They break when a right answer is worded differently.
- F1 gives credit for partial overlap. Handy for pulling names or fields out of text.
- BLEU and ROUGE count shared words and phrases. They punish a correct answer that uses different words.
- BERTScore compares meaning instead of exact words, so it catches paraphrases the others miss.
Reference-free metrics don't need a right answer at all. Perplexity tells you how fluent the text is, not if it's true. Instruction-following checks if the model did what you asked.
Safety metrics cover what accuracy can't see. Toxicity checks. Bias checks across different groups of users. Hallucination rate, which is the share of claims in an answer that nothing in your sources backs up.
RAG metrics get their own group, because retrieval and generation fail on their own. The Ragas paper defines three:
- Faithfulness. Does the answer stick to what was retrieved?
- Answer relevance. Does it answer the question that was asked?
- Context relevance. Did retrieval pull useful documents, or a pile of noise?
Now here's the important bit. The authors tested those scores against human judges, and the results weren't even. Ragas matched the humans 95% of the time on faithfulness. It hit 78% on answer relevance and 70% on context relevance, which they called the hardest to judge.
So if your faithfulness score and your spot checks disagree, look at context relevance first. That's where automated scoring is weakest. If you're still building the pipeline itself, our guide to building a RAG pipeline covers chunking and retrieval, and AI-Led's walkthrough on how to build a RAG application adds the eval step at the end.
LLM-as-a-judge, without fooling yourself
You can't have a person read every answer. So most teams use a second model as the judge. It reads the output, checks it against a rubric, and gives a score.
There are two ways to run it.
- Pairwise. Show the judge two answers and ask which is better. Good for taste calls, like which summary reads clearer. Comparing is easier than scoring from scratch, for models and for people.
- Pointwise. Score one answer against a fixed rubric. This scales better, because you're not comparing every pair.
Ask the judge for its reasons, not just a number. LangChain's openevals judges return a comment with every score by default. When a score looks off, that comment tells you if the rubric or the answer is the problem.
One problem though. Judges have blind spots. Databricks warns that an LLM judge can carry the same biases as the model it's grading, and that people still need to check the subtle calls. Judges also tend to reward answers that are longer or sound more sure of themselves.
Tip: before you trust a judge at scale, run it over 30 or so answers you've already scored by hand. If it disagrees with you more than now and then, fix the judge prompt first. It's almost always the prompt, not the model.
Build a test set that looks like real use
A test set that doesn't match real use gives you false comfort. Building a good one takes some care.
- Start with your main use cases. List the jobs users bring to the system, weighted by how often they come up. Every big one gets its own examples.
- Mix made-up and real cases. An LLM can write question and answer pairs from your own documents, fast. Hugging Face's RAG evaluation notebook does this, then has critique agents score each question on three tests: can it be answered from the text, is it useful, and does it make sense on its own. Anything under 4 out of 5 on any test gets dropped. Real questions from users catch the odd ones a model won't dream up.
- Add trouble on purpose. Vague questions. Prompt injection attempts. Questions with no answer in your data. Test only the easy cases and you'll only catch the easy failures. Our post on prompt injection defence has attacks worth adding.
- Write down how you built it. Where each example came from, how it was filtered, who labelled it. That's what lets you trust the numbers six months on.
Treat the test set like code. Version it. Add a case every time a new failure shows up in production.
Make human review count
People still catch what automated scores miss. But only with structure. "Does this look okay?" from whoever's free that afternoon isn't evaluation. It's a vibe check.
- Write a rubric with real examples of a 1, a 3 and a 5. Not just "good" and "poor".
- Run a practice round before any score counts, so reviewers read the rubric the same way.
- Measure how often reviewers agree, with a stat like Cohen's kappa. Low agreement usually means a vague rubric, not lazy reviewers.
- Settle disagreements with a written process, not by picking the score you like.
- Bring in experts where fluent text can hide a wrong answer. Health, legal and finance all need this.
Make the numbers repeatable
A score you can't rerun isn't a result. It's a story.
That's why OLMES exists. It's an open standard for running LLM tests, and it starts from a simple finding: small choices in how you run a test can shift the score a lot. Prompt format, which examples you show the model, how you normalise the answers. Change any of them and the same model on the same benchmark gets a different number.
So lock these down:
- Prompts. Save the exact format, and which examples go in and in what order.
- Scoring rules. Case, spacing, how you normalise answers. Store them with the results, not in a notes file nobody opens.
- Versions. Pin the model version and your libraries. A provider can update a model under you, and your score moves with no code change on your side.
- Everything else. Log the prompts sent, the raw outputs, the scoring script version and the test set version for every run.
A plain metric you run the same way every time beats a clever one nobody can repeat.
Which tools to start with
You don't need to build scoring from scratch.
- Ragas for RAG scoring. It works out faithfulness and relevance without hand-labelled answers, and plugs into LangChain and LlamaIndex.
- openevals for ready-made judges you can drop into an existing app.
- **The Hugging Face evaluation guidebook** if you're new to this. It covers benchmarks, judges and fixing results that won't reproduce.
Start with one of these, get a baseline, then add more only when you hit a question the first one can't answer. Picking the model in the first place is its own call. AI-Led's guide on how to choose an LLM covers that side.
Your first eval run, in six steps
- Pick two or three metrics and set a pass mark for each, before you build anything. "Good enough" should be a number.
- Build a small test set with normal cases and trouble cases.
- Pick pairwise or pointwise judging, and check the judge against answers you scored by hand.
- Add a regular human review, mostly for safety calls.
- Wire the test set into CI, so every prompt or model change runs it.
- Save every run: prompts, outputs, scores, versions. Re-check the test set as usage changes.
Where Devwiz fits
Most teams don't need more theory. They need faithfulness scoring, judge prompts and regression gates built into a product people use every day, without breaking it.
That's the work we do. We build AI platforms as production software, with the RAG pipeline, the evals and the monitoring underneath. It's the same thinking behind our LLM guardrails architecture: checks built in from day one, not bolted on after launch.
Teams spend weeks picking the perfect metric and five minutes writing down how they ran it. Flip that. A plain score you can rerun beats a clever one you can't.
*James*
Devwiz has shipped 200+ apps, including work for NSW Government (Justice and Corrective Services), Briometrix, Vivid and Huskee.
Building an AI feature and not sure your tests would catch a bad change before a customer does? Start with AI app development. Worth a chat.
Frequently asked questions
What is LLM evaluation?
It's how you measure if a language model, or the app built on it, gives answers that are right, grounded in your data and safe. It mixes automated metrics like faithfulness with human review, and it runs before and after launch so you catch regressions.
What are the best LLM evaluation tools?
Ragas is the common pick for scoring RAG apps without hand-labelled answers. LangChain's openevals gives you ready-made LLM judges. The Hugging Face evaluation guidebook is the best free starting point if you're new to it.
What is an LLM evaluation rubric?
It's a written scale, say 1 to 5, with a real example of what each score looks like. Human reviewers and LLM judges both use it. Clear examples at each level are what keep scores steady from one run to the next.
How do RAG metrics differ from standard LLM metrics?
Standard metrics like BLEU or accuracy compare the output to a fixed right answer. RAG metrics like faithfulness and context relevance check if the answer sticks to the documents retrieval found, and if retrieval found the right ones in the first place.
Is LLM-as-a-judge reliable?
It's useful at scale, but it has blind spots. Judges can share the biases of the model they grade and often favour longer, more confident answers. Check the judge against a sample you scored by hand before you trust it.
About James Killick
10+ years building digital products · 200+ apps shipped since 2015
James is a co-founder of Devwiz and an AI product specialist. Since 2015 he has helped ship 200+ apps for founders, businesses and government, including work for NSW Government, Briometrix and Huskee. He builds AI-first platforms and writes about turning a proven program into software. He also hosts the Up in the AI podcast.
More articles by James · James's personal site · LinkedIn · AI Orchestrators
Tags: AI, LLM, Testing, RAG


