Perimattic

Capability · AI agent monitoring

AI agent monitoring for production systems

Agents fail in ways uptime checks miss: they loop, call the wrong tool, run up a bill or finish confidently with the wrong answer. Here are the metrics and alerts that catch those failures, and what to do when one fires.

Last reviewed by the Perimattic AI Suite team

Illustrative view with sample values, not customer data.

In short

How do you monitor AI agents in production?

Monitor outcomes, behaviour and cost together. Track whether sessions complete their task, how many steps and retries they take, tool error rates, cost per session, quality scores on the final answer and how often a human has to take over. Alert on sustained changes in those numbers, and link every alert to the traces behind it so the on-call engineer can see the failing step.

  • Monitoring vs observability

    Observability is the data: the full trace of each agent session. Monitoring is the small set of numbers you watch and alert on from that data.

  • Start with five numbers

    Task success rate, steps per session, tool error rate, cost per session and escalation rate cover most production incidents.

Capabilities

Agent monitoring in Perimattic AI Suite

  • Alerts on agent behaviour

    Thresholds on errors, latency, steps, cost and quality, built on the same traces your engineers debug with.

    Signal it produces: Alert history linked to sessions

  • Cost per agent and per user

    Token spend split by agent role, feature and end user, with budget alerts.

    Signal it produces: Cost per session and per user

  • Quality on live traffic

    Faithfulness and custom LLM-as-judge scores on sampled sessions.

    Signal it produces: Quality trend per agent

  • Incident records

    Alerts, affected sessions and resolution kept together for later review.

    Signal it produces: Incident file for audits

Metrics

Agent monitoring metrics and example alerts

Thresholds below are starting points. Set your own from a few weeks of baseline data.

AI agent monitoring metrics, what they reveal, and example alert rules
MetricWhat it revealsExample alert (starting point)
Task success rateWhether sessions reach a correct end stateDrops 10 points below its 7-day average for 30 minutes
Steps per sessionLoops, dead ends and planning failuresp95 steps doubles against baseline
Tool error rateBroken integrations, bad arguments, permission problemsAny single tool above 5% errors for 15 minutes
Cost per sessionRunaway loops and expensive model choicesp95 cost per session above budget for an hour
Latency per sessionSlow tools, slow models, queueingp95 session time above your SLO for 15 minutes
Answer qualityWrong or unsupported final answersFaithfulness score average falls below your threshold
Escalation rateHow often a human has to take overRises 50% week on week
Injection attemptsAttacks through inputs, documents or tool resultsAttempts from one source above normal levels

Service levels

Set service level objectives for agents

Classic SLOs cover availability and latency. Agents need two more: an outcome objective and a cost objective. Writing them down gives on-call engineers and product owners the same definition of “broken”.

  • Availability

    Share of sessions that start and finish without a system error.

    Example: 99.5% of sessions per 30 days

  • Latency

    Time from request to final answer for most sessions.

    Example: p95 under 20 seconds

  • Outcome

    Share of sessions that complete the task correctly, measured by evaluation or review.

    Example: 90% task success on sampled sessions

  • Cost

    Spend per session or per resolved task, against a budget.

    Example: p95 under your per-task budget

Runbook

When an agent alert fires

  1. Open the traces

    Go from the alert to a sample of affected sessions and find the step where they diverge.

  2. Classify

    Is it a tool outage, a model or prompt change, bad input data, or an attack?

  3. Mitigate

    Roll back the change, switch to a fallback model, disable the tool or route to a human.

  4. Record

    Keep the incident with its traces. Regulated teams reuse it as SOC 2, DORA or EU AI Act evidence.

FAQ

Common questions

Short answers to the questions engineering and platform teams ask most often.

What is the difference between AI agent monitoring and agent observability?

Agent observability captures the full trace of each session: every model call, tool call and handoff. Agent monitoring is the set of metrics and alerts you run on top of that data. You need observability to explain an alert; you need monitoring to know when to look.

Which metrics matter most for AI agents?

Task success rate, steps per session, tool error rate, cost per session and escalation rate. Add answer quality scores and injection attempts once the basics are in place.

How do I measure task success for an agent?

Define what “done” means for each task, such as a ticket resolved or a form completed correctly. Measure it from the final state where you can, and from evaluation or human review on a sample where you can’t.

Can I use my existing APM for agent monitoring?

For latency and errors, yes. APM tools don’t usually understand agent steps, tool calls or answer quality, so teams send the same OpenTelemetry data to their APM and to an AI observability tool. Read Perimattic vs Datadog for how that works in practice.