CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 15/38

    Agent SDK Observability: OpenTelemetry Signals, Spans and Tenant Tagging

    Analyze observability challenges and select monitoring strategies at scale

    9 min read
    2.38% of exam
    2 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Choose between metrics, log events and traces for a given agent monitoring question
    • Enable OpenTelemetry export for Agent SDK runs at process level or per query() call without breaking the environment
    • Read the claude_code span tree to find where time went inside an agent turn
    • Tag telemetry with service and tenant attributes so a fleet of agents can be sliced in one backend

    Key concept

    Three telemetry signals, three jobs — Agent telemetry comes as three OpenTelemetry signals. Metrics are aggregate counters, log events are one record per occurrence, and traces are nested spans that keep the structure of a run. Picking a monitoring strategy means matching the question you have to the signal that can answer it, then turning on only what you need.

    1.Why agent runs are hard to observe, and the three signals that help

    An agent run is not one request. One prompt turns into a loop of model calls and tool calls. When a run misbehaves in production, the questions are specific: which tools the agent called, how long each model request took, how many tokens were spent, and where the failure happened. The Agent SDK answers these by exporting OpenTelemetry data to a collector you already operate. It splits that data into three signals, and each one answers a different kind of question.

    The three Agent SDK telemetry signals and how each is switched on
    SignalWhat it containsEnable withBest suited to
    MetricsCounters for tokens, cost, sessions, lines of code, and tool decisionsOTEL_METRICS_EXPORTERDashboards and alerts on spend and volume
    Log eventsStructured records for each prompt, API request, API error, and tool resultOTEL_LOGS_EXPORTERSearching individual occurrences during an investigation
    TracesSpans for each interaction, model request, tool call, and hook (beta)OTEL_TRACES_EXPORTER plus CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1Seeing where time went inside one turn

    The last column is where strategy comes in. A counter can tell you that tool decisions spiked this afternoon, but it can't tell you which step of a particular run stalled. A trace can, because it keeps the parent-child structure of the turn. Traces are also the only signal still in beta, and they need an extra switch.

    Sources1

    2.Turning export on: process environment or per-call options

    The answer to the prediction: the exporter selector alone isn't enough for traces. CLAUDE_CODE_ENABLE_TELEMETRY switches telemetry on, and traces also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1. The SDK's own example says this in a comment: the beta flag is required for traces, and metrics and log events do not need it.

    A complete exporter configuration for all three signals, sent over OTLPpython
    OTEL_ENV = {
        "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
        # Required for traces, which are in beta. Metrics and log events do not need this.
        "CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
        # Choose an exporter per signal. Use otlp for the SDK; see the Note below.
        "OTEL_TRACES_EXPORTER": "otlp",
        "OTEL_METRICS_EXPORTER": "otlp",
        "OTEL_LOGS_EXPORTER": "otlp",
        # Standard OTLP transport configuration.
        "OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
        "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4318",
        "OTEL_EXPORTER_OTLP_HEADERS": "Authorization=Bearer your-token",
    }

    You can deliver these variables in two ways. For production, the documentation recommends setting them in the process environment (your shell, container or orchestrator), because every query() call then picks them up with no code change. The second way is per-call options, through ClaudeAgentOptions.env in Python or options.env in TypeScript. Use it when agents running in the same process need different telemetry settings. The two languages behave differently here. In Python, env is merged on top of the inherited environment. In TypeScript, env replaces the inherited environment entirely, so you must spread process.env first or you lose variables such as PATH and your API key.

    A SaaS vendor runs one Agent SDK deployment that serves many customer organizations from a single Anthropic credential. Support asks for a per-user, per-tenant audit trail of tool calls so they can answer 'which customer triggered this MCP action' during incident review. What should the engineering team implement?

    Sources12

    3.Reading the span tree: where did the time go?

    With tracing on, each turn of the agent loop becomes a tree of spans. The root is claude_code.interaction, which covers one turn from receiving the prompt to producing the response. Under it, claude_code.llm_request wraps each Claude API call and carries the model name, latency and token counts as attributes. claude_code.tool wraps each tool invocation.

    The span hierarchy the SDK emits for one interactiontext
    claude_code.interaction
    ├── claude_code.llm_request
    ├── claude_code.hook                    (requires detailed beta tracing)
    └── claude_code.tool
        ├── claude_code.tool.blocked_on_user
        ├── claude_code.tool.execution
        └── (Agent tool) subagent claude_code.llm_request / claude_code.tool spans

    Two more details matter at scale. First, subagents launched through the Agent tool nest their own llm_request and tool spans under the parent tool span, so a multi-agent run stays in one tree. Second, the interaction span records a parent.source attribute. It shows env when the span was parented under an inbound TRACEPARENT and none when it started its own trace. That lets the agent's spans join a trace your application already started. Hook spans are an exception: they need detailed beta tracing (ENABLE_BETA_TRACING_DETAILED=1 and BETA_TRACING_ENDPOINT), and that pair also changes where your logs and traces are sent.

    An enterprise is rolling out centrally managed OpenTelemetry export for Claude Code across 400 engineers on Kubernetes, and the observability team wants configuration that scales without per-developer setup and keeps backend storage costs predictable. Which combination of practices should they adopt? (Select all that apply)(Select 3)

    Sources12

    4.Tagging a fleet: service names and per-tenant attributes

    Once many agents report to the same backend, raw spans are only useful if you can filter them. The SDK supports the standard OpenTelemetry resource variables for this. OTEL_SERVICE_NAME names the agent, for example support-triage-agent. OTEL_RESOURCE_ATTRIBUTES adds labels such as service.version and deployment.environment. Because the per-call env option exists, you can also stamp each query with the end user and tenant it serves. The values are taken from the incoming request and URL-encoded:

    Per-request tenant and user attributes on one query() callpython
    from urllib.parse import quote
    
    options = ClaudeAgentOptions(
        env={
            # ... exporter configuration from the Enable telemetry export example ...
            # request is the incoming request object from your web framework.
            "OTEL_RESOURCE_ATTRIBUTES": f"enduser.id={quote(request.user_id)},tenant.id={quote(request.tenant_id)}",
        },
    )

    These labels also reach your metrics, not only your spans. OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTES defaults to true, so the keys you put in OTEL_RESOURCE_ATTRIBUTES appear on metric datapoints. That means a cost counter can be grouped by tenant.id directly. OTEL_METRICS_INCLUDE_SESSION_ID also defaults to true. OTEL_METRICS_INCLUDE_VERSION and OTEL_METRICS_INCLUDE_ENTRYPOINT are off unless you enable them.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting OTEL_TRACES_EXPORTER=otlp is enough to get agent spans, just as the metrics and logs exporters are enough for their signals.Why is that wrong?

      Traces are in beta and also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1. Metrics and log events work without that flag.

      Covered in Turning export on: process environment or per-call options

    2. 2.Passing telemetry variables through options.env adds them to the existing environment in both the Python and TypeScript SDKs.Why is that wrong?

      Only Python merges env on top of the inherited environment. TypeScript replaces it, so you must spread process.env into the object you pass.

      Covered in Turning export on: process environment or per-call options

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Structured records for each prompt, API request, API error, and tool result”
      ↩︎ Why agent runs are hard to observe, and the three signals that help
      “Counters for tokens, cost, sessions, lines of code, and tool decisions”
      ↩︎ Why agent runs are hard to observe, and the three signals that help
      “Every query() call picks them up automatically with no code change. This is the recommended approach for production deployments.”
      ↩︎ Turning export on: process environment or per-call options
      “claude_code.llm_request: wraps each call to the Claude API, with model name, latency, and token counts as attributes.”
      ↩︎ Reading the span tree: where did the time go?
      “wraps each tool invocation, with child spans for the permission wait (claude_code.tool.blocked_on_user) and the execution itself (claude_code.tool.execution).”
      ↩︎ Reading the span tree: where did the time go?
      “Spans for each interaction, model request, tool call, and hook (beta)”
      ↩︎ Key concept
      “Required for traces, which are in beta. Metrics and log events do not need this.”
      ↩︎ Exam trap 1
      “In TypeScript, env replaces the inherited environment entirely, so include ...process.env in the object you pass.”
      ↩︎ Exam trap 2
    2. 2.
      “Claude Code has no default protocol, so set this or the signal-specific protocol variable for each otlp exporter you enable”
      ↩︎ Turning export on: process environment or per-call options
      “env when it parented under an inbound TRACEPARENT, none when it started its own trace”
      ↩︎ Reading the span tree: where did the time go?
      “Include keys from OTEL_RESOURCE_ATTRIBUTES as attributes on metric datapoints”
      ↩︎ Tagging a fleet: service names and per-tenant attributes

    Continue to page 2 of 2

    Governing Agent Telemetry at Scale: Content Gates, Managed Collectors and Fleet Analytics