Perimattic

RAG & LLM Evaluation Metrics Calculator

Compute faithfulness, context precision, context recall and answer relevancy from your own graded samples, then check them against the thresholds you set. Definitions follow the open-source RAGAS framework.

  • Free, no sign-up
  • 4 RAGAS-style metrics
  • Worked example included
Engineer reviewing generative AI model outputs on a monitor
Calculator

Grade your samples

Three worked examples are filled in. Replace them with your own question, context and answer grades.
Sample 1
F 0.75 · AR 0.75
CP 0.83 · CR 0.50
Sample 2
F 1.00 · AR 1.00
CP 1.00 · CR 1.00
Sample 3
F 0.40 · AR 0.50
CP 0.33 · CR 0.33

Pass thresholds

Starting points only. Set them from your own graded data.

What You Get

Four metrics, two halves of the pipeline

Retrieval metrics tell you whether the right context was found. Generation metrics tell you what the model did with it.

Faithfulness

Share of claims in the answer that the retrieved context supports. The main hallucination signal for RAG.

Answer relevancy

Whether the answer addresses the question, from your 1 to 5 grade mapped to 0 to 1.

Context precision

Whether relevant chunks are ranked near the top, averaged over the ranks where they appear.

Context recall

Share of the reference answer’s claims that the retrieved context supports.

Pass or fail per metric

Averages compared with thresholds you set, so you can use the result as a release gate.

Which half failed

A short diagnosis that separates retrieval problems from generation problems.
Worked example

Faithfulness = supported claims ÷ all claims

Take the question “What is the waiting period for dental cover?”. The answer makes four claims and three of them appear in the retrieved policy text, so faithfulness is 3 ÷ 4 = 0.75. The relevant chunks sit at ranks 1 and 3 of 3, so context precision is (1/1 + 2/3) ÷ 2 = 0.83. The reference answer has two claims and the context supports one, so context recall is 0.5.

Low recall with low faithfulness usually means retrieval missed the evidence and the model filled the gap. Read more about RAG evaluation or check the safeguards around your system with the hallucination risk assessment.

Abstract visualization of a large language model
FAQ

Frequently asked questions

What are the main RAG evaluation metrics?

Four metrics cover most RAG pipelines. Faithfulness checks whether the answer is supported by the retrieved context. Answer relevancy checks whether the answer addresses the question. Context precision checks whether the relevant chunks were ranked near the top. Context recall checks whether retrieval found everything needed to answer. Each runs from 0 to 1.

How is faithfulness calculated?

Split the answer into individual claims, then count how many of them can be inferred from the retrieved context. Faithfulness is the supported claims divided by all claims. An answer with 4 claims, 3 of them supported, scores 0.75.

How is context precision calculated?

Look at the retrieved chunks in rank order. At each rank where a relevant chunk appears, compute precision so far (relevant chunks up to that rank divided by the rank). Context precision is the average of those values. Relevant chunks at ranks 1 and 3 out of 3 give (1/1 + 2/3) / 2 = 0.83.

Why does this calculator ask me to grade answer relevancy?

RAGAS computes answer relevancy by having an LLM generate questions from the answer and comparing them with the original question using embeddings. That needs a model, so this calculator takes a 1 to 5 human grade instead and converts it to 0 to 1. Use the same rubric for every sample so scores are comparable.

What thresholds should I use?

There is no universal pass mark. The defaults here (0.80 faithfulness, 0.70 for the others) are starting points. Grade 100 to 200 real answers, compare the metric scores with human judgement, and set thresholds where they start to disagree. High-stakes domains such as legal or clinical use usually need higher faithfulness.

Can these metrics run on production traffic?

Faithfulness and answer relevancy only need the question, the retrieved context and the answer, so they can be scored on live requests, usually on a sample to control cost. Context recall needs a reference answer, so it is measured on graded test sets.

Need Expert Help?

Need an evaluation set for your RAG system?

Perimattic's data annotation team builds graded question and answer sets, and Perimattic AI Suite keeps scoring your pipeline once it is live.

Or email sales@perimattic.com