What you will be able to do
- Choose between metrics, log events and traces for a given agent monitoring question
- Enable OpenTelemetry export for Agent SDK runs at process level or per query() call without breaking the environment
- Read the claude_code span tree to find where time went inside an agent turn
- Tag telemetry with service and tenant attributes so a fleet of agents can be sliced in one backend
Key concept
Three telemetry signals, three jobs — Agent telemetry comes as three OpenTelemetry signals. Metrics are aggregate counters, log events are one record per occurrence, and traces are nested spans that keep the structure of a run. Picking a monitoring strategy means matching the question you have to the signal that can answer it, then turning on only what you need.
1.Why agent runs are hard to observe, and the three signals that help
An agent run is not one request. One prompt turns into a loop of model calls and tool calls. When a run misbehaves in production, the questions are specific: which tools the agent called, how long each model request took, how many tokens were spent, and where the failure happened. The Agent SDK answers these by exporting OpenTelemetry data to a collector you already operate. It splits that data into three signals, and each one answers a different kind of question.
| Signal | What it contains | Enable with | Best suited to |
|---|---|---|---|
| Metrics | Counters for tokens, cost, sessions, lines of code, and tool decisions | OTEL_METRICS_EXPORTER | Dashboards and alerts on spend and volume |
| Log events | Structured records for each prompt, API request, API error, and tool result | OTEL_LOGS_EXPORTER | Searching individual occurrences during an investigation |
| Traces | Spans for each interaction, model request, tool call, and hook (beta) | OTEL_TRACES_EXPORTER plus CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 | Seeing where time went inside one turn |
The last column is where strategy comes in. A counter can tell you that tool decisions spiked this afternoon, but it can't tell you which step of a particular run stalled. A trace can, because it keeps the parent-child structure of the turn. Traces are also the only signal still in beta, and they need an extra switch.
Sources1
2.Turning export on: process environment or per-call options
The answer to the prediction: the exporter selector alone isn't enough for traces. CLAUDE_CODE_ENABLE_TELEMETRY switches telemetry on, and traces also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1. The SDK's own example says this in a comment: the beta flag is required for traces, and metrics and log events do not need it.
OTEL_ENV = {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
# Required for traces, which are in beta. Metrics and log events do not need this.
"CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
# Choose an exporter per signal. Use otlp for the SDK; see the Note below.
"OTEL_TRACES_EXPORTER": "otlp",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
# Standard OTLP transport configuration.
"OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4318",
"OTEL_EXPORTER_OTLP_HEADERS": "Authorization=Bearer your-token",
}You can deliver these variables in two ways. For production, the documentation recommends setting them in the process environment (your shell, container or orchestrator), because every query() call then picks them up with no code change. The second way is per-call options, through ClaudeAgentOptions.env in Python or options.env in TypeScript. Use it when agents running in the same process need different telemetry settings. The two languages behave differently here. In Python, env is merged on top of the inherited environment. In TypeScript, env replaces the inherited environment entirely, so you must spread process.env first or you lose variables such as PATH and your API key.
A SaaS vendor runs one Agent SDK deployment that serves many customer organizations from a single Anthropic credential. Support asks for a per-user, per-tenant audit trail of tool calls so they can answer 'which customer triggered this MCP action' during incident review. What should the engineering team implement?
Correct answer: A — Inject percent-encoded enduser.id and tenant.id into OTEL_RESOURCE_ATTRIBUTES per call so tool_decision, tool_result, and mcp_server_connection events carry the user and tenant
- A. Correct. Attaching enduser.id and tenant.id as resource attributes on each call makes the security-relevant events (tool_decision, tool_result, mcp_server_connection, permission_mode_changed) attributable per end user, which is exactly the documented pattern for building a per-user audit trail from a shared credential.
- B. Per-customer keys would require restructuring authentication and billing for every tenant and still would not tag individual tool-call events with a user identity within a shared key's traffic.
- C. The telemetry stream is not partitioned by tenant at the transport layer, so there is no mechanism to split a shared export stream per customer without the attribute tagging described in the correct option.
- D. Tool content logging exposes payload bodies but does not attach an identifiable user or tenant attribute to the emitted events, so it does not solve the attribution requirement.
3.Reading the span tree: where did the time go?
With tracing on, each turn of the agent loop becomes a tree of spans. The root is claude_code.interaction, which covers one turn from receiving the prompt to producing the response. Under it, claude_code.llm_request wraps each Claude API call and carries the model name, latency and token counts as attributes. claude_code.tool wraps each tool invocation.
claude_code.interaction
├── claude_code.llm_request
├── claude_code.hook (requires detailed beta tracing)
└── claude_code.tool
├── claude_code.tool.blocked_on_user
├── claude_code.tool.execution
└── (Agent tool) subagent claude_code.llm_request / claude_code.tool spansclaude_code.tool has two children. claude_code.tool.execution covers the tool actually running. claude_code.tool.blocked_on_user covers the wait for permission. If most of the 40 seconds sits in blocked_on_user, the tool is fine and the delay comes from approval, which you fix with permissions, not performance tuning.
Two more details matter at scale. First, subagents launched through the Agent tool nest their own llm_request and tool spans under the parent tool span, so a multi-agent run stays in one tree. Second, the interaction span records a parent.source attribute. It shows env when the span was parented under an inbound TRACEPARENT and none when it started its own trace. That lets the agent's spans join a trace your application already started. Hook spans are an exception: they need detailed beta tracing (ENABLE_BETA_TRACING_DETAILED=1 and BETA_TRACING_ENDPOINT), and that pair also changes where your logs and traces are sent.
An enterprise is rolling out centrally managed OpenTelemetry export for Claude Code across 400 engineers on Kubernetes, and the observability team wants configuration that scales without per-developer setup and keeps backend storage costs predictable. Which combination of practices should they adopt? (Select all that apply)(Select 3)
Correct answers: A, B, C — Set the telemetry environment variables in the administrator-managed settings.json rather than relying on each developer's shell profile; Add custom OTEL_RESOURCE_ATTRIBUTES such as department and cost_center so spend and usage can be sliced by team in the backend; Turn off OTEL_METRICS_INCLUDE_SESSION_ID and similar cardinality flags that are not needed for the team's dashboards to limit unique time-series growth
- A. Correct. Centralizing telemetry variables in administrator-managed settings.json applies the same collector endpoint and exporter choices to every developer without relying on individual shell configuration, which is the documented enterprise pattern.
- B. Correct. Custom resource attributes like department, team.id, and cost_center are the documented mechanism for multi-team attribution, letting the backend break down token and cost metrics per group.
- C. Correct. High-cardinality attributes such as session.id or account UUIDs multiply the number of unique time series a metrics backend must store; disabling the ones a dashboard does not need is the documented way to control that cost as headcount grows.
- D. The console exporter writes telemetry to the same standard output channel the SDK/CLI uses for other output and is documented as unsuitable for anything beyond local debugging; it is not a scalable production export target.
- E. Requiring per-developer manual shell exports is exactly the fragile, non-scalable setup that centrally managed settings.json is meant to replace, and it invites drift across 400 machines.
4.Tagging a fleet: service names and per-tenant attributes
Once many agents report to the same backend, raw spans are only useful if you can filter them. The SDK supports the standard OpenTelemetry resource variables for this. OTEL_SERVICE_NAME names the agent, for example support-triage-agent. OTEL_RESOURCE_ATTRIBUTES adds labels such as service.version and deployment.environment. Because the per-call env option exists, you can also stamp each query with the end user and tenant it serves. The values are taken from the incoming request and URL-encoded:
from urllib.parse import quote
options = ClaudeAgentOptions(
env={
# ... exporter configuration from the Enable telemetry export example ...
# request is the incoming request object from your web framework.
"OTEL_RESOURCE_ATTRIBUTES": f"enduser.id={quote(request.user_id)},tenant.id={quote(request.tenant_id)}",
},
)These labels also reach your metrics, not only your spans. OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTES defaults to true, so the keys you put in OTEL_RESOURCE_ATTRIBUTES appear on metric datapoints. That means a cost counter can be grouped by tenant.id directly. OTEL_METRICS_INCLUDE_SESSION_ID also defaults to true. OTEL_METRICS_INCLUDE_VERSION and OTEL_METRICS_INCLUDE_ENTRYPOINT are off unless you enable them.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Setting OTEL_TRACES_EXPORTER=otlp is enough to get agent spans, just as the metrics and logs exporters are enough for their signals.Why is that wrong?
Traces are in beta and also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1. Metrics and log events work without that flag.
Covered in Turning export on: process environment or per-call options
2.Passing telemetry variables through options.env adds them to the existing environment in both the Python and TypeScript SDKs.Why is that wrong?
Only Python merges env on top of the inherited environment. TypeScript replaces it, so you must spread process.env into the object you pass.
Covered in Turning export on: process environment or per-call options
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Structured records for each prompt, API request, API error, and tool result”
↩︎ Why agent runs are hard to observe, and the three signals that help“Counters for tokens, cost, sessions, lines of code, and tool decisions”
↩︎ Why agent runs are hard to observe, and the three signals that help“Every query() call picks them up automatically with no code change. This is the recommended approach for production deployments.”
↩︎ Turning export on: process environment or per-call options“claude_code.llm_request: wraps each call to the Claude API, with model name, latency, and token counts as attributes.”
↩︎ Reading the span tree: where did the time go?“wraps each tool invocation, with child spans for the permission wait (claude_code.tool.blocked_on_user) and the execution itself (claude_code.tool.execution).”
↩︎ Reading the span tree: where did the time go?“Spans for each interaction, model request, tool call, and hook (beta)”
↩︎ Key concept“Required for traces, which are in beta. Metrics and log events do not need this.”
↩︎ Exam trap 1“In TypeScript, env replaces the inherited environment entirely, so include ...process.env in the object you pass.”
↩︎ Exam trap 2 - 2.https://code.claude.com/docs/en/monitoring-usageOfficial docs
“Claude Code has no default protocol, so set this or the signal-specific protocol variable for each otlp exporter you enable”
↩︎ Turning export on: process environment or per-call options“env when it parented under an inbound TRACEPARENT, none when it started its own trace”
↩︎ Reading the span tree: where did the time go?“Include keys from OTEL_RESOURCE_ATTRIBUTES as attributes on metric datapoints”
↩︎ Tagging a fleet: service names and per-tenant attributes