Perimattic

Guide · AI observability

What is AI observability? A practical guide

How to see what LLMs, agents and RAG pipelines are doing in production, and why it matters for reliability, cost and compliance. Written for platform, ML and risk teams who have to run AI, not just demo it.

Last reviewed by the Perimattic AI Suite team

Illustrative view with sample values, not customer data.

In short

What is AI observability?

AI observability is the practice of instrumenting LLMs, AI agents and RAG pipelines so teams can see what the system did and why. It works through five signals: traces, metrics, logs, evaluation scores and user feedback. Teams use that evidence to debug failures, control cost and show regulators and customers how the system behaves in production.

  • Why it is a separate discipline

    Traditional monitoring assumes a correct system either works or errors. AI systems can be up, fast and wrong. Observability adds quality signals that tell you when an answer is bad, not just slow.

  • Why now

    In May 2026 Gartner predicted that by 2028, 40% of organisations deploying AI will use AI observability tools to monitor model performance, bias and outputs.

Overhead view of a team working at a shared table covered with laptops and notebooks

The five signals

The five signals of AI observability

Cloud-native observability rests on traces, metrics and logs. AI observability keeps all three and adds two that only make sense for probabilistic systems.

  • Traces

    Every step of a request, from the first prompt through retrievals, tool calls and agent handoffs to the final answer, linked as parent and child spans.

    Answers: what happened, in what order?

  • Metrics

    Latency, tokens, cost, error rate and throughput, broken down by model, agent, feature and user.

    Answers: how much, how fast, how often?

  • Logs

    Structured records of inputs, outputs, guardrail decisions and system events, with personal data redacted.

    Answers: what exactly was said?

  • Evaluation scores

    Faithfulness, relevance, hallucination and custom LLM-as-judge scores, computed on test sets and on live traffic.

    Answers: was the answer any good?

  • User feedback

    Ratings, corrections, retries and escalations linked to the trace that produced the response.

    Answers: did it help the person?

How it works

How AI observability works on OpenTelemetry

Most AI observability tools now accept OpenTelemetry, the open standard for telemetry. Its GenAI semantic conventions define how model calls, agent invocations, tool executions and retrievals are described, so the same data works in any backend. The conventions are still marked as in development and moved to their own repository in 2026, so check which version your tools support.

  1. Instrument

    Add the OpenTelemetry SDK or an instrumentation library to your LLM client, agent framework and retriever.

  2. Collect

    An OpenTelemetry Collector receives spans, redacts personal data, samples and forwards them.

  3. Evaluate and store

    The backend stores traces, prices tokens and runs evaluation on a sample or on every request.

  4. Alert and report

    Thresholds trigger alerts; dashboards and exports serve engineers, finance and compliance.

How it differs

AI observability vs APM and LLM monitoring

The three overlap, and most teams run more than one. The difference is what each can explain.

Application performance monitoring, LLM monitoring and AI observability compared
AspectAPMLLM monitoringAI observability
Unit of analysisHTTP request, service call, querySingle model callWhole session across agents, tools and retrievals
Typical signalsLatency, errors, saturationLatency, tokens, cost, errorsAll of those plus quality scores, feedback and decision context
Can it tell a wrong answer from a right one?NoNoYes, through evaluation scores
Cost viewInfrastructurePer modelPer model, agent, feature and user
Audit useUptime and incident historyUsage recordsDecision records mapped to obligations

Reference architecture

Design choices that matter in production

Keep instrumentation open. If your code emits standard OpenTelemetry spans, you can change or add backends later by editing collector configuration, not application code.

Redact before export. Prompts and outputs often contain personal or confidential data. Removing it in the collector, inside your network, keeps traces useful while limiting what any backend stores.

Sample deliberately. Keep every trace that errored, scored badly or was flagged by a user. Sample the healthy majority. Decide retention per project, because regulated systems may need records for months or years.

Fan out instead of replacing. An OpenTelemetry Collector can send the same spans to your existing APM for on-call engineers and to an AI observability tool for quality and compliance work.

Maturity model

Five stages of AI observability maturity

Most teams move through these stages in order. Use the AI observability maturity assessment to see where you are today.

AI observability maturity stages
StageWhat you haveWhat you can answer
1. LoggingPrompts and responses in application logsWhat did the model say?
2. MonitoringDashboards for latency, tokens, cost and errorsIs it up, and what does it cost?
3. TracingEnd-to-end traces across agents, tools and retrievalWhich step caused this failure?
4. EvaluationQuality scores on live traffic with alerts and review queuesAre answers getting worse, and where?
5. GovernanceRecords mapped to policies and regulations, with retention and access controlCan we prove how the system behaved?

Buyer’s checklist

Ten questions to ask any AI observability platform

  • Does it ingest OpenTelemetry, including the GenAI conventions?
  • Can it trace a multi-agent session end to end?
  • Does it score quality on production traffic, not only on test sets?
  • Can it attribute cost to agents, features and users?
  • Can personal data be redacted before it leaves your network?
  • Can you self-host it or pin data to a region?
  • Does it alert on quality as well as latency and errors?
  • Can it export evidence in a form your auditors accept?
  • Does it work with the frameworks and providers you use?
  • What are its own security attestations?

Regulation

Where regulation expects observability

Several frameworks now expect records of how an AI system behaved. The compliance hub covers each in detail; in brief:

Regulations that expect AI monitoring records
FrameworkWhat it expects
EU AI ActAutomatic logging, log retention, human oversight and post-market monitoring for high-risk systems (from 2 Dec 2027 or 2 Aug 2028)
HIPAAAudit controls for systems that contain or use electronic PHI
SOC 2Monitoring (CC7) and change management (CC8) evidence for in-scope systems
DORAICT risk management, incident reporting and third-party monitoring for EU financial entities

Background

From DevOps observability to AI observability

Observability came out of site reliability engineering, where teams learned that dashboards alone can’t explain new failures. AI systems repeat that lesson with a twist: the failure is often in the content, not the infrastructure. If your DevOps practice is still building its foundations, the DevOps maturity assessment is a useful first step, and MLOps vs DevOps explains what changes when models enter the pipeline.

FAQ

Common questions

Short answers to the questions people ask most often about this topic.

What is the simplest definition of AI observability?

It is being able to answer “why did our AI system produce this output?” from production data, without having to reproduce the problem in a test environment.

How is AI observability different from MLOps?

MLOps covers the model lifecycle: data, training, versioning and deployment. AI observability covers what happens after deployment, at inference time: traces, quality, cost and decision records. They meet at model performance monitoring.

What is an OpenTelemetry span for an LLM call?

A record of one unit of work with a start time, end time and attributes. Under the GenAI conventions, a model call span carries attributes such as gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, and sits inside a parent span such as an agent invocation.

Is there an open-source option for AI observability?

Yes. The OpenTelemetry SDKs and collector are open source, and Langfuse publishes an MIT-licensed core you can self-host. Commercial platforms build on the same standard, so your instrumentation stays portable.

Is AI observability the same as AI governance?

No. Governance is the set of policies, roles and controls an organisation uses to manage AI risk. Observability is the evidence those controls depend on: logs, monitoring records and incident data. Governance without observability can’t show that controls work.

Where should a team start?

Trace one production system end to end, add cost per user, then add a faithfulness or relevance score on a sample of live traffic. That covers the three questions people ask first: what happened, what did it cost, and was it right.