RAG & LLM Evaluation Metrics Calculator
Compute faithfulness, context precision, context recall and answer relevancy from your own graded samples, then check them against the thresholds you set. Definitions follow the open-source RAGAS framework.
- Free, no sign-up
- 4 RAGAS-style metrics
- Worked example included
Grade your samples
Pass thresholds
Starting points only. Set them from your own graded data.
Four metrics, two halves of the pipeline
Faithfulness
Answer relevancy
Context precision
Context recall
Pass or fail per metric
Which half failed
Faithfulness = supported claims ÷ all claims
Take the question “What is the waiting period for dental cover?”. The answer makes four claims and three of them appear in the retrieved policy text, so faithfulness is 3 ÷ 4 = 0.75. The relevant chunks sit at ranks 1 and 3 of 3, so context precision is (1/1 + 2/3) ÷ 2 = 0.83. The reference answer has two claims and the context supports one, so context recall is 0.5.
Low recall with low faithfulness usually means retrieval missed the evidence and the model filled the gap. Read more about RAG evaluation or check the safeguards around your system with the hallucination risk assessment.
Frequently asked questions
What are the main RAG evaluation metrics?
Four metrics cover most RAG pipelines. Faithfulness checks whether the answer is supported by the retrieved context. Answer relevancy checks whether the answer addresses the question. Context precision checks whether the relevant chunks were ranked near the top. Context recall checks whether retrieval found everything needed to answer. Each runs from 0 to 1.
How is faithfulness calculated?
Split the answer into individual claims, then count how many of them can be inferred from the retrieved context. Faithfulness is the supported claims divided by all claims. An answer with 4 claims, 3 of them supported, scores 0.75.
How is context precision calculated?
Look at the retrieved chunks in rank order. At each rank where a relevant chunk appears, compute precision so far (relevant chunks up to that rank divided by the rank). Context precision is the average of those values. Relevant chunks at ranks 1 and 3 out of 3 give (1/1 + 2/3) / 2 = 0.83.
Why does this calculator ask me to grade answer relevancy?
RAGAS computes answer relevancy by having an LLM generate questions from the answer and comparing them with the original question using embeddings. That needs a model, so this calculator takes a 1 to 5 human grade instead and converts it to 0 to 1. Use the same rubric for every sample so scores are comparable.
What thresholds should I use?
There is no universal pass mark. The defaults here (0.80 faithfulness, 0.70 for the others) are starting points. Grade 100 to 200 real answers, compare the metric scores with human judgement, and set thresholds where they start to disagree. High-stakes domains such as legal or clinical use usually need higher faithfulness.
Can these metrics run on production traffic?
Faithfulness and answer relevancy only need the question, the retrieved context and the answer, so they can be scored on live requests, usually on a sample to control cost. Context recall needs a reference answer, so it is measured on graded test sets.
Related Tools
Hallucination Risk Assessment
Score the guardrails and review process around your LLM system.
Try it free AssessmentAI Observability Maturity Assessment
See whether you evaluate on live traffic yet, and what to add next.
Try it free CalculatorLLM Cost Calculator
Estimate token cost per request, per agent task and per user.
Try it freeNeed an evaluation set for your RAG system?
Perimattic's data annotation team builds graded question and answer sets, and Perimattic AI Suite keeps scoring your pipeline once it is live.
Or email sales@perimattic.com