Observability for AI Agents: Tracking Token Costs, Latency, and Spans with OpenTelemetry
Distributed tracing, token cost accounting and tool performance monitoring for autonomous AI agents, using OpenTelemetry.
The Black Box Challenge in Autonomous AI Systems
Building production applications with generative AI and Autonomous Multi-Agent Orchestrators introduces unprecedented opacity into software engineering. In standard web systems, request lifecycles are linear and deterministic: HTTP request in, SQL query executed, JSON response returned.
In multi-agent systems, however, a single inbound message on WhatsApp can trigger:
- Intent classification by a local Small Language Model (SLM).
- Multiple hybrid vector searches against PostgreSQL with pgvector.
- Two or three parallel Model Context Protocol (MCP) tool calls.
- A final synthesis call to a frontier model consuming thousands of tokens.
If the request takes 4.5 seconds or costs $0.05 instead of $0.001, where was the bottleneck? Legacy APM platforms (New Relic, Datadog) were not architected to track prompt token economics or reasoning traces.
At MSC Company, we eliminate this blind spot by standardizing on OpenTelemetry (OTel).
The OpenTelemetry Standard for Multi-Agent AI Systems
OpenTelemetry is the vendor-neutral, open-source standard maintained by the Cloud Native Computing Foundation (CNCF). In our architecture, every AI workflow is instrumented as a Distributed Trace with nested Spans:
+-----------------------------------------------------------------------------------+
| DISTRIBUTED TRACE OF AN AI AGENT (OTEL) |
| |
| [ Trace: WhatsApp_Sales_Booking - Total: 1,420ms | Cost: $0.0014 USD ] |
| |
| |-- Span: Webhook_Ingest (Bun/Elysia) .......................... [ 12ms ] |
| | |
| |-- Span: Prometheus_Semantic_Router ........................... [ 95ms ] |
| | |-- Attributes: model="gemini-3.7-flash", intent="quote_request" |
| | |
| |-- Span: MCP_Tool_Parallel_Execution ......................... [ 85ms ] |
| | |-- Sub-span: Postgres_pgvector_Search ..................... [ 14ms ] |
| | |-- Sub-span: Postgres_Pricing_Query ....................... [ 8ms ] |
| | |
| |-- Span: LLM_Synthesis_Completion ............................. [ 1,150ms ] |
| | |-- Attributes: |
| | | * gen_ai.model = "gemini-3.7-flash" |
| | | * gen_ai.usage.prompt_tokens = 1,420 (Cache Hit: 75%) |
| | | * gen_ai.usage.completion_tokens = 180 |
| | | * gen_ai.cost_usd = 0.00018 |
| | |
| +-- Span: WhatsApp_Delivery .................................... [ 58ms ] |
+-----------------------------------------------------------------------------------+
Four Critical Metrics Tracked in Real Time
- Exact Token Costs (Prompt, Completion, and Cached): By recording exact token metrics per span, we track the financial ROI of each customer conversation and monitor prompt efficiency.
- Tool Failure & Retry Rates: If an agent encounters a Zod schema validation error on a tool call, OTel records an error event, alerting engineers to recalibrate prompts.
- Time-to-First-Token (TTFT) and Latency: We ensure that voice and chat pipelines consistently deliver responses under 1.5 seconds.
- Data Privacy Compliance: Traces record anonymized user IDs and token hashes without logging sensitive customer PII, guaranteeing compliance with global privacy standards.
TypeScript Code Example: Instrumenting an Agent Call
import { trace, SpanStatusCode } from "@opentelemetry/api";
const tracer = trace.getTracer("msc-agent-orchestrator");
export async function executeAgentWithTelemetry(params: {
tenantId: string;
userPrompt: string;
agentRole: string;
}) {
return tracer.startActiveSpan(`Agent:${params.agentRole}`, async (span) => {
try {
span.setAttribute("msc.tenant_id", params.tenantId);
span.setAttribute("msc.agent_role", params.agentRole);
const result = await callModelWithTools(params.userPrompt);
span.setAttribute("gen_ai.system", "google_vertex_ai");
span.setAttribute("gen_ai.model", result.model);
span.setAttribute("gen_ai.usage.prompt_tokens", result.promptTokens);
span.setAttribute("gen_ai.usage.completion_tokens", result.completionTokens);
span.setAttribute("gen_ai.cost_usd", result.costUsd);
span.setStatus({ code: SpanStatusCode.OK });
return result.data;
} catch (error: any) {
span.recordException(error);
span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
throw error;
} finally {
span.end();
}
});
}
Frequently Asked Questions (FAQ AEO)
Does OpenTelemetry add overhead to real-time AI responses?
No. The OpenTelemetry TypeScript SDK operates asynchronously with in-memory ring buffers and batch export workers, adding less than 0.5 milliseconds of latency to request lifecycles.
Can OpenTelemetry data be exported to self-hosted tools?
Yes. As an open CNCF standard, OTel traces can be exported directly to open-source collectors (Grafana Tempo, Prometheus, Jaeger, SigNoz, ClickHouse) with zero vendor lock-in.
Related Articles & Next Steps:
- Learn about our orchestrator in Prometheus: Multi-Agent Orchestration via MCP.
- Discover our backend speed in Elysia vs. FastAPI on Bun Runtime.
- Explore model fine-tuning at Cendar Lab.
- Need to gain full observability over your AI agent infrastructure? Talk to MSC Company.