Infrastructure & Sovereign Cloud10 de março de 2026· Leitura: 3 min

Observability for AI Agents: Tracking Token Costs, Latency, and Spans with OpenTelemetry

Distributed tracing, token cost accounting and tool performance monitoring for autonomous AI agents, using OpenTelemetry.

Observability for AI Agents: Tracking Token Costs, Latency, and Spans with OpenTelemetry

The Black Box Challenge in Autonomous AI Systems

Building production applications with generative AI and Autonomous Multi-Agent Orchestrators introduces unprecedented opacity into software engineering. In standard web systems, request lifecycles are linear and deterministic: HTTP request in, SQL query executed, JSON response returned.

In multi-agent systems, however, a single inbound message on WhatsApp can trigger:

  1. Intent classification by a local Small Language Model (SLM).
  2. Multiple hybrid vector searches against PostgreSQL with pgvector.
  3. Two or three parallel Model Context Protocol (MCP) tool calls.
  4. A final synthesis call to a frontier model consuming thousands of tokens.

If the request takes 4.5 seconds or costs $0.05 instead of $0.001, where was the bottleneck? Legacy APM platforms (New Relic, Datadog) were not architected to track prompt token economics or reasoning traces.

At MSC Company, we eliminate this blind spot by standardizing on OpenTelemetry (OTel).


The OpenTelemetry Standard for Multi-Agent AI Systems

OpenTelemetry is the vendor-neutral, open-source standard maintained by the Cloud Native Computing Foundation (CNCF). In our architecture, every AI workflow is instrumented as a Distributed Trace with nested Spans:

+-----------------------------------------------------------------------------------+
|                        DISTRIBUTED TRACE OF AN AI AGENT (OTEL)                    |
|                                                                                   |
|  [ Trace: WhatsApp_Sales_Booking - Total: 1,420ms | Cost: $0.0014 USD ]           |
|                                                                                   |
|  |-- Span: Webhook_Ingest (Bun/Elysia) .......................... [ 12ms ]        |
|  |                                                                                |
|  |-- Span: Prometheus_Semantic_Router ........................... [ 95ms ]        |
|  |   |-- Attributes: model="gemini-3.7-flash", intent="quote_request"             |
|  |                                                                                |
|  |-- Span: MCP_Tool_Parallel_Execution ......................... [ 85ms ]        |
|  |   |-- Sub-span: Postgres_pgvector_Search ..................... [ 14ms ]        |
|  |   |-- Sub-span: Postgres_Pricing_Query ....................... [ 8ms ]         |
|  |                                                                                |
|  |-- Span: LLM_Synthesis_Completion ............................. [ 1,150ms ]     |
|  |   |-- Attributes:                                                              |
|  |   |   * gen_ai.model = "gemini-3.7-flash"                                      |
|  |   |   * gen_ai.usage.prompt_tokens = 1,420 (Cache Hit: 75%)                    |
|  |   |   * gen_ai.usage.completion_tokens = 180                                   |
|  |   |   * gen_ai.cost_usd = 0.00018                                              |
|  |                                                                                |
|  +-- Span: WhatsApp_Delivery .................................... [ 58ms ]        |
+-----------------------------------------------------------------------------------+

Four Critical Metrics Tracked in Real Time

  1. Exact Token Costs (Prompt, Completion, and Cached): By recording exact token metrics per span, we track the financial ROI of each customer conversation and monitor prompt efficiency.
  2. Tool Failure & Retry Rates: If an agent encounters a Zod schema validation error on a tool call, OTel records an error event, alerting engineers to recalibrate prompts.
  3. Time-to-First-Token (TTFT) and Latency: We ensure that voice and chat pipelines consistently deliver responses under 1.5 seconds.
  4. Data Privacy Compliance: Traces record anonymized user IDs and token hashes without logging sensitive customer PII, guaranteeing compliance with global privacy standards.

TypeScript Code Example: Instrumenting an Agent Call

import { trace, SpanStatusCode } from "@opentelemetry/api";

const tracer = trace.getTracer("msc-agent-orchestrator");

export async function executeAgentWithTelemetry(params: {
  tenantId: string;
  userPrompt: string;
  agentRole: string;
}) {
  return tracer.startActiveSpan(`Agent:${params.agentRole}`, async (span) => {
    try {
      span.setAttribute("msc.tenant_id", params.tenantId);
      span.setAttribute("msc.agent_role", params.agentRole);

      const result = await callModelWithTools(params.userPrompt);

      span.setAttribute("gen_ai.system", "google_vertex_ai");
      span.setAttribute("gen_ai.model", result.model);
      span.setAttribute("gen_ai.usage.prompt_tokens", result.promptTokens);
      span.setAttribute("gen_ai.usage.completion_tokens", result.completionTokens);
      span.setAttribute("gen_ai.cost_usd", result.costUsd);

      span.setStatus({ code: SpanStatusCode.OK });
      return result.data;
    } catch (error: any) {
      span.recordException(error);
      span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
      throw error;
    } finally {
      span.end();
    }
  });
}

Frequently Asked Questions (FAQ AEO)

Does OpenTelemetry add overhead to real-time AI responses?

No. The OpenTelemetry TypeScript SDK operates asynchronously with in-memory ring buffers and batch export workers, adding less than 0.5 milliseconds of latency to request lifecycles.

Can OpenTelemetry data be exported to self-hosted tools?

Yes. As an open CNCF standard, OTel traces can be exported directly to open-source collectors (Grafana Tempo, Prometheus, Jaeger, SigNoz, ClickHouse) with zero vendor lock-in.


Related Articles & Next Steps:

Engineering Radar & Technical Inquiries

Scale your operations with audited AI and backend architecture

Subscribe to our technical briefing or submit your system requirements directly to MSC Company's lead architects. Responses within 1 business day.

Applied Engineering & AI

Scale Your Operations with Custom AI Systems

From autonomous WhatsApp agents to sovereign data architecture and fine-tuned SLMs. Talk directly to the MSC Company engineering team.

Contact MSC →
Related Articles

Continue Reading

View all articles →