
Capability · RAG evaluation
RAG evaluation on production traffic
A RAG pipeline can return a fast, confident answer that none of its retrieved documents support. Evaluation scores tell you when that happens, and the trace tells you whether retrieval or generation was at fault.
Last reviewed by the Perimattic AI Suite team
In short
What is RAG evaluation?
RAG evaluation measures how well a retrieval-augmented generation pipeline answers questions. The standard metrics check two halves of the pipeline: retrieval (did it fetch the right context, ranked well?) and generation (is the answer supported by that context, and does it address the question?). Scores run from 0 to 1 and are most useful when computed on live traffic as well as on a fixed test set.
Green dashboards can hide bad answers
A request can return 200 OK in 800 ms and still contradict every document it retrieved. Latency and error rate can’t see that; faithfulness can.
Find which half failed
Low context recall points at retrieval. High recall with low faithfulness points at the generator. The fix is different in each case.

Capabilities
RAG evaluation in Perimattic AI Suite
Scores on live traffic
Faithfulness and relevance scored after each answer, or in batches against stored traces.
Signal it produces: Faithfulness score per request
Retrieval spans
Embedding, vector search, re-ranking and chunk selection traced with timing and scores.
Signal it produces: Context precision per query
LLM-as-judge for domain rules
Judge prompts for checks the standard metrics miss, such as citation format or required disclaimers.
Signal it produces: Custom quality scores
Regression tracking
Compare scores across model versions, chunk sizes and retrieval settings before and after a change.
Signal it produces: Score change per release
The metrics
Four RAG evaluation metrics, with worked examples
These definitions follow the open-source RAGAS framework. Each example uses one invented question: “What is the waiting period for dental cover?”
To compute these from your own graded samples, use the RAG and LLM evaluation metrics calculator. To check the guardrails around a pipeline first, try the hallucination risk assessment.
| Metric | What it measures | How it is computed | Worked example |
|---|---|---|---|
| Faithfulness | Whether the answer is supported by the retrieved context | Claims in the answer supported by the context ÷ all claims in the answer | The answer makes 4 claims; 3 appear in the policy text. Faithfulness = 3/4 = 0.75 |
| Answer relevancy | Whether the answer addresses the question asked | An LLM generates questions from the answer; the score is their average similarity to the original question | An answer about dental limits, not waiting periods, scores low even if it is accurate |
| Context precision | Whether relevant chunks are ranked near the top | Average of precision at each rank where a relevant chunk appears | Relevant chunks at ranks 1 and 3 of 3: (1/1 + 2/3) ÷ 2 = 0.83 |
| Context recall | Whether retrieval found everything needed to answer | Claims in a reference answer that the context supports ÷ claims in the reference | The reference has 2 claims; the context supports 1. Context recall = 0.5 |
Context recall needs a reference answer, so it is usually measured on a graded sample. Faithfulness and answer relevancy can be scored on every request.
Offline vs production
Offline evaluation and production evaluation do different jobs
You need both. Offline runs gate a release; production scoring catches what the test set never covered.
| Aspect | Offline evaluation | Production evaluation |
|---|---|---|
| Data | A fixed, graded test set | Live requests, sampled or scored in full |
| When it runs | Before a release, in CI | Continuously, after each answer or in batches |
| Best metrics | All four, including context recall | Faithfulness and answer relevancy; recall on graded samples |
| Catches | Regressions from a model, prompt or chunking change | Drift as documents, users and questions change |
| Typical action | Block or approve the release | Alert, review flagged answers, add them to the test set |
Setting thresholds
Choose thresholds from your own data
There is no universal good score. A legal research tool may need faithfulness close to 1.0; an internal search assistant can tolerate less. Start by grading 100 to 200 real answers, compare human judgements with metric scores, and set the alert threshold where scores and people start to disagree.
- Score a baseline before you change anything
- Alert on sustained drops, not single answers
- Send low-scoring answers to a review queue
- Add reviewed failures to the offline test set
FAQ
Common questions
Short answers to the questions engineering and platform teams ask most often.
What is RAGAS?
RAGAS is an open-source framework for evaluating retrieval-augmented generation. It defines metrics such as faithfulness, answer relevancy, context precision and context recall, most of which use an LLM to judge claims rather than needing human-written answers for every question.
Can RAG evaluation run in production?
Yes. Faithfulness and answer relevancy need only the question, the retrieved context and the answer, so they can be scored on live requests. Scoring every request adds judge-model cost, so many teams score a sample and every answer that users flag.
What causes RAG hallucinations?
Two things, usually. Retrieval misses the passage that holds the answer (low context recall), or the generator goes beyond the retrieved text and fills gaps from what it learned in training. Re-ranking and hybrid search help the first; tighter instructions and smaller, cleaner chunks help the second.
How is LLM-as-judge different from RAGAS?
RAGAS metrics are fixed definitions that are comparable across systems. LLM-as-judge is a general technique: you write the criteria, for example “cites a clause number” or “never gives dosage advice”, and a separate model scores each answer. Most teams use both.
How many graded samples do I need?
Enough to cover your main question types. A hundred to two hundred graded answers is a practical starting point for setting thresholds, and you can grow the set from flagged production answers over time.
Go further
Related tools, guides and services
- Free toolRAG and LLM evaluation metrics calculatorCompute faithfulness, context precision and recall, and answer relevancy from graded samples.Open
- Free toolHallucination risk assessmentScore the guardrails, retrieval quality and review process around your LLM system.Open
- ArticleHow to evaluate an LLM before deploying itA pre-production evaluation plan you can reuse.Open
- ArticleHow to choose a vector databaseThe retrieval layer decisions that shape context precision and recall.Open
- Case studyAI knowledge system for a consulting firmA production RAG system Perimattic built for document search and Q&A.Open
- ServiceAI data annotation servicesGraded evaluation sets and labelled data for testing RAG quality.Open

See what your AI systems are doing, with evidence to back it up
Perimattic AI Suite is in early access. Tell us what you are building and a Perimattic engineer will follow up to scope your first instrumented system.
Prefer email? sales@perimattic.com