Perimattic

Capability · RAG evaluation

RAG evaluation on production traffic

A RAG pipeline can return a fast, confident answer that none of its retrieved documents support. Evaluation scores tell you when that happens, and the trace tells you whether retrieval or generation was at fault.

Last reviewed by the Perimattic AI Suite team

Illustrative view with sample values, not customer data.

In short

What is RAG evaluation?

RAG evaluation measures how well a retrieval-augmented generation pipeline answers questions. The standard metrics check two halves of the pipeline: retrieval (did it fetch the right context, ranked well?) and generation (is the answer supported by that context, and does it address the question?). Scores run from 0 to 1 and are most useful when computed on live traffic as well as on a fixed test set.

  • Green dashboards can hide bad answers

    A request can return 200 OK in 800 ms and still contradict every document it retrieved. Latency and error rate can’t see that; faithfulness can.

  • Find which half failed

    Low context recall points at retrieval. High recall with low faithfulness points at the generator. The fix is different in each case.

Glass beakers and a pipette on a white laboratory bench

Capabilities

RAG evaluation in Perimattic AI Suite

  • Scores on live traffic

    Faithfulness and relevance scored after each answer, or in batches against stored traces.

    Signal it produces: Faithfulness score per request

  • Retrieval spans

    Embedding, vector search, re-ranking and chunk selection traced with timing and scores.

    Signal it produces: Context precision per query

  • LLM-as-judge for domain rules

    Judge prompts for checks the standard metrics miss, such as citation format or required disclaimers.

    Signal it produces: Custom quality scores

  • Regression tracking

    Compare scores across model versions, chunk sizes and retrieval settings before and after a change.

    Signal it produces: Score change per release

The metrics

Four RAG evaluation metrics, with worked examples

These definitions follow the open-source RAGAS framework. Each example uses one invented question: “What is the waiting period for dental cover?”

To compute these from your own graded samples, use the RAG and LLM evaluation metrics calculator. To check the guardrails around a pipeline first, try the hallucination risk assessment.

RAG evaluation metrics with definitions and worked examples
MetricWhat it measuresHow it is computedWorked example
FaithfulnessWhether the answer is supported by the retrieved contextClaims in the answer supported by the context ÷ all claims in the answerThe answer makes 4 claims; 3 appear in the policy text. Faithfulness = 3/4 = 0.75
Answer relevancyWhether the answer addresses the question askedAn LLM generates questions from the answer; the score is their average similarity to the original questionAn answer about dental limits, not waiting periods, scores low even if it is accurate
Context precisionWhether relevant chunks are ranked near the topAverage of precision at each rank where a relevant chunk appearsRelevant chunks at ranks 1 and 3 of 3: (1/1 + 2/3) ÷ 2 = 0.83
Context recallWhether retrieval found everything needed to answerClaims in a reference answer that the context supports ÷ claims in the referenceThe reference has 2 claims; the context supports 1. Context recall = 0.5

Context recall needs a reference answer, so it is usually measured on a graded sample. Faithfulness and answer relevancy can be scored on every request.

Offline vs production

Offline evaluation and production evaluation do different jobs

You need both. Offline runs gate a release; production scoring catches what the test set never covered.

Offline and production RAG evaluation compared
AspectOffline evaluationProduction evaluation
DataA fixed, graded test setLive requests, sampled or scored in full
When it runsBefore a release, in CIContinuously, after each answer or in batches
Best metricsAll four, including context recallFaithfulness and answer relevancy; recall on graded samples
CatchesRegressions from a model, prompt or chunking changeDrift as documents, users and questions change
Typical actionBlock or approve the releaseAlert, review flagged answers, add them to the test set

Setting thresholds

Choose thresholds from your own data

There is no universal good score. A legal research tool may need faithfulness close to 1.0; an internal search assistant can tolerate less. Start by grading 100 to 200 real answers, compare human judgements with metric scores, and set the alert threshold where scores and people start to disagree.

  • Score a baseline before you change anything
  • Alert on sustained drops, not single answers
  • Send low-scoring answers to a review queue
  • Add reviewed failures to the offline test set

FAQ

Common questions

Short answers to the questions engineering and platform teams ask most often.

What is RAGAS?

RAGAS is an open-source framework for evaluating retrieval-augmented generation. It defines metrics such as faithfulness, answer relevancy, context precision and context recall, most of which use an LLM to judge claims rather than needing human-written answers for every question.

Can RAG evaluation run in production?

Yes. Faithfulness and answer relevancy need only the question, the retrieved context and the answer, so they can be scored on live requests. Scoring every request adds judge-model cost, so many teams score a sample and every answer that users flag.

What causes RAG hallucinations?

Two things, usually. Retrieval misses the passage that holds the answer (low context recall), or the generator goes beyond the retrieved text and fills gaps from what it learned in training. Re-ranking and hybrid search help the first; tighter instructions and smaller, cleaner chunks help the second.

How is LLM-as-judge different from RAGAS?

RAGAS metrics are fixed definitions that are comparable across systems. LLM-as-judge is a general technique: you write the criteria, for example “cites a clause number” or “never gives dosage advice”, and a separate model scores each answer. Most teams use both.

How many graded samples do I need?

Enough to cover your main question types. A hundred to two hundred graded answers is a practical starting point for setting thresholds, and you can grow the set from flagged production answers over time.