CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 25/38

    Claude Agent Observability with OpenTelemetry: Metrics, Logs and Traces

    Monitor system performance using logging and observability tools

    11 min read
    2.67% of exam
    3 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Name the three OpenTelemetry signals a Claude agent emits and say which question each one answers
    • Configure exporters, protocol and endpoint for Claude Code or the Agent SDK so that telemetry actually reaches a collector
    • Read the span hierarchy of an agent turn and tune which attributes go onto metrics
    • Decide which content-logging gates to enable, and know what each gate exposes

    Key concept

    OpenTelemetry signals: metrics, log events and traces — A Claude agent's runtime reports on itself through three separate OpenTelemetry streams. Metrics are aggregate counters, log events are one structured record per occurrence, and traces are timed spans nested along the agent loop. Each signal has its own exporter switch, and you choose each one based on the question you need answered.

    1.What an agent's telemetry has to tell you

    When an agent is running in production, the questions you ask it after the fact are nearly always the same. Which tools did it call? How long did each model request take? How many tokens did it spend? Where did it fail? The Agent SDK observability guide frames monitoring around exactly those four questions. Claude Code and the Agent SDK answer them by emitting standard OpenTelemetry (OTel) data. That data flows into whatever collector and backend you already run, and nothing about it is specific to Anthropic.

    The data is split into three signals, and much of this subdomain depends on keeping them apart. Metrics are counters: good for dashboards and alerts on spend or volume, useless for explaining one bad run. Log events are discrete structured records, so you can ask what happened on a particular API error or tool result. Traces are spans with durations and parent-child relationships, so you can see where the time in a single turn went.

    The three signals, what they carry, and the variable that turns each on
    SignalWhat it containsEnable with
    MetricsCounters for tokens, cost, sessions, lines of code, and tool decisionsOTEL_METRICS_EXPORTER
    Log eventsStructured records for each prompt, API request, API error, and tool resultOTEL_LOGS_EXPORTER
    TracesSpans for each interaction, model request, tool call, and hook (beta)OTEL_TRACES_EXPORTER plus CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1

    Sources1

    2.Turning telemetry on and pointing it at a collector

    Getting data out takes more than one switch. CLAUDE_CODE_ENABLE_TELEMETRY=1 is the required master switch, but it only turns collection on. Each signal also needs an exporter selected: OTEL_METRICS_EXPORTER accepts otlp, prometheus, console or none, and OTEL_LOGS_EXPORTER accepts otlp, console or none. The quick-start labels both exporter choices as optional, which is exactly why a runner can have telemetry enabled and still send nothing to your collector.

    For an otlp exporter you also need a transport. OTEL_EXPORTER_OTLP_ENDPOINT sets the collector for all signals, and OTEL_EXPORTER_OTLP_HEADERS carries authentication, such as a bearer token for the collector. The protocol is easy to overlook: Claude Code has no default protocol, so set OTEL_EXPORTER_OTLP_PROTOCOL (grpc, http/json or http/protobuf) or the per-signal protocol variable. Per-signal variants such as OTEL_EXPORTER_OTLP_METRICS_ENDPOINT override the general setting for one signal.

    Traces need one more switch. They are in beta, so the trace exporter only works when CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 is also set. Metrics and log events do not depend on that flag.

    Agent SDK telemetry configuration enabling all three signals over OTLPpython
    OTEL_ENV = {
        "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
        # Required for traces, which are in beta. Metrics and log events do not need this.
        "CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
        # Choose an exporter per signal. Use otlp for the SDK; see the Note below.
        "OTEL_TRACES_EXPORTER": "otlp",
        "OTEL_METRICS_EXPORTER": "otlp",
        "OTEL_LOGS_EXPORTER": "otlp",
        # Standard OTLP transport configuration.
        "OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
        "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4318",
        "OTEL_EXPORTER_OTLP_HEADERS": "Authorization=Bearer your-token",
    }

    In the Agent SDK there are two places to put these variables. The process environment (your shell, container or orchestrator) is the recommended approach for production, because every query() call picks the variables up without a code change. Per-call options (ClaudeAgentOptions.env in Python, options.env in TypeScript) are for giving different agents in the same process different telemetry settings. The two languages behave differently here. In Python, env is merged on top of the inherited environment. In TypeScript, env replaces the inherited environment entirely, so you spread process.env into the object you pass.

    Export timing is also adjustable. Metrics export every 60000 ms by default, and logs and trace spans every 5000 ms. The quick-start suggests shortening these intervals while debugging, then resetting them for production.

    A platform engineer sets CLAUDE_CODE_ENABLE_TELEMETRY=1 on every CI runner but the team's OTLP collector never receives any spans, metrics, or logs from Claude Code jobs. No other telemetry variables were configured. What is the most likely cause of the missing data?

    Sources21

    3.Reading traces and shaping metric attributes

    Traces are organised along the agent loop. The root span is one turn, and its children are the work done inside that turn. Once you know the tree, a slow run becomes a question of which child span is long. If a tool span's time sits mostly in the permission-wait child, the delay came from waiting for a user decision, not from the tool itself.

    Claude Code span names and what each one wraps
    SpanWraps
    claude_code.interactionA single turn of the agent loop, from receiving a prompt to producing a response
    claude_code.llm_requestEach call to the Claude API, with model name, latency, and token counts as attributes
    claude_code.toolEach tool invocation, with child spans claude_code.tool.blocked_on_user (permission wait) and claude_code.tool.execution
    claude_code.hookEach hook execution; requires detailed beta tracing (ENABLE_BETA_TRACING_DETAILED=1 and BETA_TRACING_ENDPOINT)

    When several agents report to one backend, you need to tell their telemetry apart. OTEL_SERVICE_NAME names the agent, and OTEL_RESOURCE_ATTRIBUTES attaches key-value pairs such as a version and a deployment environment. The same variable can carry per-request identity, such as enduser.id and tenant.id, taken from the incoming request and URL-encoded.

    Tagging an agent's telemetry with a service name and resource attributespython
    options = ClaudeAgentOptions(
        env={
            # ... exporter configuration from the Enable telemetry export example ...
            "OTEL_SERVICE_NAME": "support-triage-agent",
            "OTEL_RESOURCE_ATTRIBUTES": "service.version=1.4.0,deployment.environment=production",
        },
    )

    Every attribute on a metric becomes a dimension your backend has to store. The monitoring reference has a separate switch for each of the standard metric attributes. Some are on by default, and every one can be turned off without affecting the others.

    Metric attribute switches and their defaults
    VariableControlsDefault
    OTEL_METRICS_INCLUDE_SESSION_IDsession.id attribute in metricstrue
    OTEL_METRICS_INCLUDE_VERSIONapp.version attribute in metricsfalse
    OTEL_METRICS_INCLUDE_ACCOUNT_UUIDuser.account_uuid and user.account_id attributestrue
    OTEL_METRICS_INCLUDE_ENTRYPOINTapp.entrypoint attributefalse
    OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTESKeys from OTEL_RESOURCE_ATTRIBUTES on metric datapointstrue
    OTEL_METRICS_INCLUDE_REPOSITORYvcs.* repository identity attributes on metrics and eventsfalse

    A metrics backend's bill spikes because every exported time series carries a unique session.id, fragmenting cardinality across thousands of short-lived CLI sessions. The team still needs cost and token totals broken down by model. Which change addresses the cardinality problem without losing the model breakdown?

    Sources12

    4.What goes into the logs: content gates and locked destinations

    By default, Claude Code telemetry records that something happened, not what was said. User prompt text is disabled by default: on the interaction span, the user_prompt attribute reads <REDACTED> unless you set the gate. Each additional kind of content has its own opt-in variable, and each exposes more than the one before it.

    Content-logging gates in Claude Code and the Agent SDK
    VariableWhat it adds
    OTEL_LOG_USER_PROMPTS=1Prompt text on claude_code.user_prompt events and on the claude_code.interaction span
    OTEL_LOG_TOOL_DETAILS=1Tool input arguments (file paths, shell commands, search patterns) on claude_code.tool_result events
    OTEL_LOG_TOOL_CONTENT=1A tool.output span event on claude_code.tool with tool output; requires tracing
    OTEL_LOG_RAW_API_BODIESFull Messages API request and response JSON as claude_code.api_request_body and claude_code.api_response_body log events

    Raw API bodies are the most revealing gate. They contain the entire conversation history, with extended-thinking content redacted, and turning them on implies consent to everything the other three gates would reveal. Content-bearing attributes are truncated at 60 KB by default, a limit you set with CLAUDE_CODE_OTEL_CONTENT_MAX_LENGTH. The value 1 gives inline, truncated bodies. The value file:<dir> writes untruncated bodies to disk and puts a body_ref pointer in the event.

    These defaults belong to Claude Code, and other products differ. Cowork's OpenTelemetry export, which an admin configures under Organization settings, includes user prompt content in its events by default. Its guidance is to filter or redact in your collector before routing events downstream.

    Administrators can also pin where telemetry goes. When managed settings set OTEL_EXPORTER_OTLP_ENDPOINT, Claude Code removes every per-signal endpoint a developer has set. The exporter selectors, however, still follow normal per-key precedence, so a developer could switch a signal to console or none. If you need the selectors locked, set them in managed settings as well.

    Sources132

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting CLAUDE_CODE_ENABLE_TELEMETRY=1 on its own is enough to start sending telemetry to your collector.Why is that wrong?

      That variable only enables collection. Each signal still needs an exporter selected (for example OTEL_METRICS_EXPORTER=otlp), and an otlp exporter also needs a protocol, because Claude Code has none by default.

      Covered in Turning telemetry on and pointing it at a collector

    2. 2.Once exporters are configured, traces arrive the same way metrics and log events do.Why is that wrong?

      Traces are beta and also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1. Metrics and log events work without that flag, so traces can be missing while the other two signals arrive normally.

      Covered in Turning telemetry on and pointing it at a collector

    3. 3.Adding hook spans is a harmless extra switch that leaves the rest of your telemetry pipeline unchanged.Why is that wrong?

      Hook spans require detailed beta tracing (ENABLE_BETA_TRACING_DETAILED=1 plus BETA_TRACING_ENDPOINT). That pair sends logs and traces to BETA_TRACING_ENDPOINT instead of through your configured exporters.

      Covered in Reading traces and shaping metric attributes

    4. 4.Claude Code telemetry logs the text of user prompts by default, so you must opt out to keep prompts out of your backend.Why is that wrong?

      Claude Code works the other way round: prompt content is disabled by default and appears only when you set OTEL_LOG_USER_PROMPTS=1. Cowork's OpenTelemetry export is the product that includes prompt content by default.

      Covered in What goes into the logs: content gates and locked destinations

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Structured records for each prompt, API request, API error, and tool result”
      ↩︎ What an agent's telemetry has to tell you
      “Counters for tokens, cost, sessions, lines of code, and tool decisions”
      ↩︎ What an agent's telemetry has to tell you
      “This is the recommended approach for production deployments.”
      ↩︎ Turning telemetry on and pointing it at a collector
      “In TypeScript, env replaces the inherited environment entirely, so include ...process.env in the object you pass.”
      ↩︎ Turning telemetry on and pointing it at a collector
      “claude_code.llm_request: wraps each call to the Claude API, with model name, latency, and token counts as attributes.”
      ↩︎ Reading traces and shaping metric attributes
      “Bodies include the entire conversation history and have extended-thinking content redacted.”
      ↩︎ What goes into the logs: content gates and locked destinations
      “Spans for each interaction, model request, tool call, and hook (beta)”
      ↩︎ Key concept
      “Required for traces, which are in beta. Metrics and log events do not need this.”
      ↩︎ Exam trap 2
      “a pair that also changes where your logs and traces go.”
      ↩︎ Exam trap 3
    2. 2.
      “Claude Code has no default protocol, so set this or the signal-specific protocol variable for each otlp exporter you enable”
      ↩︎ Turning telemetry on and pointing it at a collector
      “For debugging: reduce export intervals, and reset them for production use”
      ↩︎ Turning telemetry on and pointing it at a collector
      “Include session.id attribute in metrics”
      ↩︎ Reading traces and shaping metric attributes
      “so set the selectors in managed settings too if you need them locked.”
      ↩︎ What goes into the logs: content gates and locked destinations
      “Choose exporters (both are optional - configure only what you need)”
      ↩︎ Exam trap 1
      “Enable logging of user prompt content (default: disabled)”
      ↩︎ Exam trap 4

    Continue to page 2 of 2

    Monitoring Claude Usage and Cost with the Admin APIs