Perimattic

Capability · Agent observability

AI agent observability for multi-agent systems

An agent that runs for two minutes may make forty model calls and a dozen tool calls before it answers. Agent observability records every step as one connected trace, so you can see why it did what it did.

Last reviewed by the Perimattic AI Suite team

Illustrative view with sample values, not customer data.

In short

What is AI agent observability?

AI agent observability is tracing applied to autonomous and multi-agent systems. Each user session becomes a tree of spans: the agent invocation at the root, then planning steps, model calls, tool calls, retrievals and handoffs to other agents, each with timing, tokens, cost and outcome. With that tree you can debug failures, attribute cost per agent and reconstruct any decision later.

  • Why chatbot tracing isn’t enough

    A chatbot trace is one prompt and one answer. An agent loops, calls tools, spawns sub-agents and retries. Without parent-child spans you see the failure but not the step that caused it.

  • Built on an open standard

    The OpenTelemetry GenAI conventions define invoke_agent, chat and execute_tool spans, so agent traces look the same whichever framework produced them.

Two engineers reviewing code together on a laptop in an open-plan office

Capabilities

What agent observability should capture

These are the signals that answer the questions engineers and reviewers ask after an agent misbehaves.

  • Tool calls

    Every API call, database query or code execution, with its arguments, result, latency and errors.

    Signal it produces: Tool error rate, slowest tools

  • Handoffs between agents

    When one agent delegates to another: what context was passed, which agent took over and what it returned.

    Signal it produces: Handoff path per session

  • Loops and retries

    Repeated tool calls or planning cycles, flagged before they burn through a token budget.

    Signal it produces: Loop count, retries per session

  • Cost per agent and per user

    Token spend split by agent role, feature and end user.

    Signal it produces: Cost per session, per agent, per user

  • Decision context

    Model version, prompt version and retrieved sources for each step, kept with the trace for later review.

    Signal it produces: Replayable session record

  • Prompt injection attempts

    Inputs and tool results that try to override the agent’s instructions, classified and linked to the session.

    Signal it produces: Injection attempts by source

How a trace is shaped

One agent session as an OpenTelemetry trace

Span names follow the OpenTelemetry GenAI semantic conventions, which are still marked as in development. A session for a claims-handling agent might look like this.

Simplified span tree for one session
invoke_agent claims-agent                 2.9s   $0.041
├─ chat <model>   (plan next step)        0.8s
├─ execute_tool lookup_policy             0.3s
├─ invoke_agent fraud-check-agent         1.1s
│  ├─ chat <model>                        0.6s
│  └─ execute_tool score_claim            0.2s
└─ chat <model>   (draft reply)           0.6s   faithfulness 0.88

Illustrative. Attribute names such as gen_ai.agent.name, gen_ai.operation.name and gen_ai.usage.input_tokens come from the GenAI conventions.

Frameworks

Works with the agent framework you already use

Perimattic AI Suite reads OpenTelemetry, so it doesn’t need a plugin per framework. Agent frameworks emit OpenTelemetry spans either natively or through open-source instrumentation libraries such as OpenLLMetry and OpenInference. Custom agent loops can add spans with the standard OpenTelemetry SDK.

  • LangChain and LangGraph
  • CrewAI
  • AutoGen
  • LlamaIndex agents
  • OpenAI Agents SDK
  • Amazon Bedrock Agents
  • Custom agent loops in Python, TypeScript or Java

Agent security

Prompt injection is an agent problem first

Agents read documents, web pages and tool results, and any of them can carry instructions. Because an agent can act, an injected instruction can turn into a real action. Tracing shows where the instruction came from and what the agent did next. Read more on prompt injection monitoring, or score your exposure with the prompt injection risk assessment.

FAQ

Common questions

Short answers to the questions engineering and platform teams ask most often.

How is agent observability different from LLM observability?

LLM observability covers individual model calls: prompt, completion, latency, tokens and cost. Agent observability covers how those calls fit together: which tools ran, in what order, which agent made each decision and what happened as a result. Multi-agent systems need both layers.

Which agent frameworks does Perimattic AI Suite support?

Any framework that emits OpenTelemetry spans, including LangChain, LangGraph, CrewAI, AutoGen, LlamaIndex and the OpenAI Agents SDK through native or open-source instrumentation. Custom agent loops can be instrumented with the standard OpenTelemetry SDK.

How long does it take to start tracing an agent?

As a typical estimate, a single-agent app can send its first trace in under 30 minutes if it already uses OpenTelemetry. Multi-agent systems with dashboards and alerts usually take one to two days, depending on the stack.

What is the difference between agent observability and agent monitoring?

Observability is the data: the full trace of what each agent did. Monitoring is what you watch and alert on from that data, such as error rates, loop counts and cost per session. Our guide to AI agent monitoring covers the alerts worth setting.

Can agent traces be used as audit evidence?

Yes, when they are complete, timestamped and kept for the required period. Regulated teams use agent traces to show how a decision was reached, which supports EU AI Act logging, HIPAA audit controls and DORA incident records.