Every enterprise has an AI agent demo that works. It books the meeting, resolves the ticket and drafts the reply, all in a sandbox, on a good day, with someone from the vendor driving. The hard part starts the week after, when the agent is in production and calling your real APIs, writing to your real databases and opening tickets in your real queue at three in the morning. That's when you find out whether anyone can see what it's doing.
Most teams can't. They've spent a year evaluating models, and almost no time asking how an agent will be observed once it's live. This post covers that gap: what's different about agents in production, the ways they fail quietly, and what good observability for them looks like.
An agent isn't another service
For twenty years, monitoring has rested on one assumption: software does what it was written to do. A service receives a request, runs the code path its engineers wrote and returns a response. When something breaks, you trace the request, find the slow span and fix the code. Behavior is fixed at deploy time.
Agents break that assumption. An agent decides at runtime which tools to call, in what order and how many times. The same prompt can produce a different plan tomorrow. The logic isn't only in the code any more. It lives partly in the model, partly in the context the agent was given and partly in whatever it found along the way.
For your infrastructure, that means an agent is a new kind of workload that writes its own call patterns. Your stack sees the effects: API calls, database queries, network traffic, authentication events. What it can't see is the decision that caused them.
Three ways agents fail quietly
Agents don't usually fail with a stack trace. They fail by looking like something else. Here are the three patterns we see teams struggle with most.
- The runaway loop. An agent gets an ambiguous result from a tool, decides to try again, gets the same result and tries again. Each call is valid on its own. Together they look like a traffic spike on an internal API, and your APM flags it as load. Nobody connects it to a single agent session stuck in a retry loop until the rate limits trip or the cloud bill arrives.
- The fan-out. One agent action (say "reconcile this customer's account") turns into dozens of downstream calls across billing, CRM and the data warehouse. The database team sees slow queries. The network team sees east-west traffic climbing. The security team sees a service account touching tables it rarely touches. Three teams, three alerts, one cause, and none of their tools show the connection.
- The slow drift. A model update, a prompt change or a new document in the knowledge base shifts how the agent behaves. It isn't broken. It's just taking a different path: an extra lookup here, a more expensive tool there. Nothing crosses a threshold. Latency creeps up, costs creep up, and by the time anyone notices, it's hard to say when it started.
None of these trigger the classic "service is down" alert. All of them are the kind of incident that eats a whole afternoon on a bridge call.
Why traditional monitoring misses it
The evidence for an agent problem is almost never in one place. The retry loop shows up in APM. The fan-out shows up in network flows and database metrics. The unusual data access shows up in the SIEM. The drift shows up as a slow change spread across all of them.
If those signals live in separate tools, each team sees a piece and none of them sees the pattern. This is the same tool-sprawl problem that has always driven up mean time to resolution. Agents just make it worse, because they generate cross-layer activity faster and less predictably than any human-written service.
There's a second blind spot. Most LLM monitoring products watch the model: tokens, prompt latency, evaluation scores. That's useful, but it stops at the edge of the model. It won't tell you that the agent's third tool call triggered a lock on your orders table, or that its service account just authenticated from an unexpected subnet. Watching the model isn't the same as watching what the model does to your systems.
What good agent observability looks like
Moving agents from demo to deployment safely comes down to one idea: correlate what the agent decided with what happened in your infrastructure. In practice, that means four things.
- Treat agent sessions as first-class traces. Every agent run should carry an ID that follows it through each tool call, API request and database query. When something goes wrong, you should be able to start from the infrastructure symptom and walk back to the agent session and step that caused it, or start from the session and see everything it touched.
- Baseline behavior, not just thresholds. "More than 1,000 calls a minute" is a useless alert for a workload that plans its own work. What you need is a learned picture of normal for each agent (typical tool mix, call depth, data touched) so drift shows up as a change in pattern, not a broken limit.
- Put security in the same view. Agents act through service accounts and API keys. That makes them an identity-and-access question as much as a performance one. Unusual data access by an agent should land on the same timeline as the latency spike it caused, not in a different team's queue.
- Close the loop. When a high-confidence pattern appears (a loop, a runaway fan-out), the platform should be able to act: throttle the session, pause the agent, page the owner with the root cause already attached. An alert that says "API latency elevated" is not a root cause.
Before you ship: five questions to answer
If you're moving an agent from pilot into production this quarter, get clear answers to these first:
- Can we trace any database query or API call back to the agent session and step that caused it?
- Do we know what normal looks like for this agent: its usual tools, call volume and data scope?
- Will unusual data access by the agent show up next to the performance impact, or in a separate security tool?
- What happens automatically when the agent loops or fans out: does anything throttle it, or do we find out from the bill?
- Who gets paged, and what do they see? A root cause, or twelve unrelated alerts?
If the honest answer to most of these is "we'd figure it out on the bridge call," the agent isn't ready for production. More accurately, your observability isn't. The model is the part everyone is evaluating. What decides whether it survives contact with production is the layer that watches everything else.
