
Guide · AI observability
What is AI observability? A practical guide
How to see what LLMs, agents and RAG pipelines are doing in production, and why it matters for reliability, cost and compliance. Written for platform, ML and risk teams who have to run AI, not just demo it.
Last reviewed by the Perimattic AI Suite team
In short
What is AI observability?
AI observability is the practice of instrumenting LLMs, AI agents and RAG pipelines so teams can see what the system did and why. It works through five signals: traces, metrics, logs, evaluation scores and user feedback. Teams use that evidence to debug failures, control cost and show regulators and customers how the system behaves in production.
Why it is a separate discipline
Traditional monitoring assumes a correct system either works or errors. AI systems can be up, fast and wrong. Observability adds quality signals that tell you when an answer is bad, not just slow.
Why now
In May 2026 Gartner predicted that by 2028, 40% of organisations deploying AI will use AI observability tools to monitor model performance, bias and outputs.

The five signals
The five signals of AI observability
Cloud-native observability rests on traces, metrics and logs. AI observability keeps all three and adds two that only make sense for probabilistic systems.
Traces
Every step of a request, from the first prompt through retrievals, tool calls and agent handoffs to the final answer, linked as parent and child spans.
Answers: what happened, in what order?
Metrics
Latency, tokens, cost, error rate and throughput, broken down by model, agent, feature and user.
Answers: how much, how fast, how often?
Logs
Structured records of inputs, outputs, guardrail decisions and system events, with personal data redacted.
Answers: what exactly was said?
Evaluation scores
Faithfulness, relevance, hallucination and custom LLM-as-judge scores, computed on test sets and on live traffic.
Answers: was the answer any good?
User feedback
Ratings, corrections, retries and escalations linked to the trace that produced the response.
Answers: did it help the person?
How it works
How AI observability works on OpenTelemetry
Most AI observability tools now accept OpenTelemetry, the open standard for telemetry. Its GenAI semantic conventions define how model calls, agent invocations, tool executions and retrievals are described, so the same data works in any backend. The conventions are still marked as in development and moved to their own repository in 2026, so check which version your tools support.
Instrument
Add the OpenTelemetry SDK or an instrumentation library to your LLM client, agent framework and retriever.
Collect
An OpenTelemetry Collector receives spans, redacts personal data, samples and forwards them.
Evaluate and store
The backend stores traces, prices tokens and runs evaluation on a sample or on every request.
Alert and report
Thresholds trigger alerts; dashboards and exports serve engineers, finance and compliance.
How it differs
AI observability vs APM and LLM monitoring
The three overlap, and most teams run more than one. The difference is what each can explain.
| Aspect | APM | LLM monitoring | AI observability |
|---|---|---|---|
| Unit of analysis | HTTP request, service call, query | Single model call | Whole session across agents, tools and retrievals |
| Typical signals | Latency, errors, saturation | Latency, tokens, cost, errors | All of those plus quality scores, feedback and decision context |
| Can it tell a wrong answer from a right one? | No | No | Yes, through evaluation scores |
| Cost view | Infrastructure | Per model | Per model, agent, feature and user |
| Audit use | Uptime and incident history | Usage records | Decision records mapped to obligations |
Reference architecture
Design choices that matter in production
Keep instrumentation open. If your code emits standard OpenTelemetry spans, you can change or add backends later by editing collector configuration, not application code.
Redact before export. Prompts and outputs often contain personal or confidential data. Removing it in the collector, inside your network, keeps traces useful while limiting what any backend stores.
Sample deliberately. Keep every trace that errored, scored badly or was flagged by a user. Sample the healthy majority. Decide retention per project, because regulated systems may need records for months or years.
Fan out instead of replacing. An OpenTelemetry Collector can send the same spans to your existing APM for on-call engineers and to an AI observability tool for quality and compliance work.
Maturity model
Five stages of AI observability maturity
Most teams move through these stages in order. Use the AI observability maturity assessment to see where you are today.
| Stage | What you have | What you can answer |
|---|---|---|
| 1. Logging | Prompts and responses in application logs | What did the model say? |
| 2. Monitoring | Dashboards for latency, tokens, cost and errors | Is it up, and what does it cost? |
| 3. Tracing | End-to-end traces across agents, tools and retrieval | Which step caused this failure? |
| 4. Evaluation | Quality scores on live traffic with alerts and review queues | Are answers getting worse, and where? |
| 5. Governance | Records mapped to policies and regulations, with retention and access control | Can we prove how the system behaved? |
Buyer’s checklist
Ten questions to ask any AI observability platform
- Does it ingest OpenTelemetry, including the GenAI conventions?
- Can it trace a multi-agent session end to end?
- Does it score quality on production traffic, not only on test sets?
- Can it attribute cost to agents, features and users?
- Can personal data be redacted before it leaves your network?
- Can you self-host it or pin data to a region?
- Does it alert on quality as well as latency and errors?
- Can it export evidence in a form your auditors accept?
- Does it work with the frameworks and providers you use?
- What are its own security attestations?
Regulation
Where regulation expects observability
Several frameworks now expect records of how an AI system behaved. The compliance hub covers each in detail; in brief:
| Framework | What it expects |
|---|---|
| EU AI Act | Automatic logging, log retention, human oversight and post-market monitoring for high-risk systems (from 2 Dec 2027 or 2 Aug 2028) |
| HIPAA | Audit controls for systems that contain or use electronic PHI |
| SOC 2 | Monitoring (CC7) and change management (CC8) evidence for in-scope systems |
| DORA | ICT risk management, incident reporting and third-party monitoring for EU financial entities |
Background
From DevOps observability to AI observability
Observability came out of site reliability engineering, where teams learned that dashboards alone can’t explain new failures. AI systems repeat that lesson with a twist: the failure is often in the content, not the infrastructure. If your DevOps practice is still building its foundations, the DevOps maturity assessment is a useful first step, and MLOps vs DevOps explains what changes when models enter the pipeline.
Sources checked on 1 October 2026.
FAQ
Common questions
Short answers to the questions people ask most often about this topic.
What is the simplest definition of AI observability?
It is being able to answer “why did our AI system produce this output?” from production data, without having to reproduce the problem in a test environment.
How is AI observability different from MLOps?
MLOps covers the model lifecycle: data, training, versioning and deployment. AI observability covers what happens after deployment, at inference time: traces, quality, cost and decision records. They meet at model performance monitoring.
What is an OpenTelemetry span for an LLM call?
A record of one unit of work with a start time, end time and attributes. Under the GenAI conventions, a model call span carries attributes such as gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, and sits inside a parent span such as an agent invocation.
Is there an open-source option for AI observability?
Yes. The OpenTelemetry SDKs and collector are open source, and Langfuse publishes an MIT-licensed core you can self-host. Commercial platforms build on the same standard, so your instrumentation stays portable.
Is AI observability the same as AI governance?
No. Governance is the set of policies, roles and controls an organisation uses to manage AI risk. Observability is the evidence those controls depend on: logs, monitoring records and incident data. Governance without observability can’t show that controls work.
Where should a team start?
Trace one production system end to end, add cost per user, then add a faithfulness or relevance score on a sample of live traffic. That covers the three questions people ask first: what happened, what did it cost, and was it right.
Go further
Related tools, guides and services
- Free toolAI observability maturity assessmentFind your stage across tracing, evaluation, cost and governance, with next steps.Open
- Free toolDevOps maturity assessmentBenchmark the delivery and monitoring practices AI observability builds on.Open
- ArticleMLOps vs DevOpsWhat actually changes when you deploy ML models.Open
- ServiceDevOps and platform engineeringPerimattic builds the pipelines and monitoring stacks AI systems run on.Open

See what your AI systems are doing, with evidence to back it up
Perimattic AI Suite is in early access. Tell us what you are building and a Perimattic engineer will follow up to scope your first instrumented system.
Prefer email? sales@perimattic.com