What you will be able to do
- Query the Usage API at the granularity that fits the monitoring task, and filter or group the results
- Use the Cost API for spend breakdowns, and know its daily-only limit and currency format
- Choose between the Usage and Cost Admin API, the Claude Code Analytics API and the Enterprise Analytics API, and allow for each one's data freshness
- Use cache diagnostics to explain a drop in cached input tokens
1.The Usage API: token consumption across the organization
OpenTelemetry shows what individual agent runs are doing. For the organization as a whole, the Usage & Cost Admin API gives programmatic access to historical API usage and cost data, similar to the Usage and Cost pages of the Claude Console. The documentation lists what it is for: precise token counts instead of adding up response token counts yourself; reconciling cost with Anthropic billing; measuring whether a change to the system actually improved it, or setting up alerting; and tuning features such as prompt caching to make better use of rate limits.
It needs an Admin API credential, not an ordinary API key. Claude Enterprise organizations have no Admin API keys; they use an Analytics API key against a different API.
The usage endpoint is /v1/organizations/usage_report/messages. It reports uncached input, cached input, cache creation and output tokens in fixed time buckets. You can filter by API key, workspace, model, service tier, context window, data residency or speed, group results by the same dimensions, and track server-side tool usage such as web search.
curl "https://api.anthropic.com/v1/organizations/usage_report/messages?\
starting_at=2025-01-01T00:00:00Z&\
ending_at=2025-01-08T00:00:00Z&\
group_by[]=model&\
bucket_width=1d" \
-H "anthropic-version: 2023-06-01" \
-H "x-api-key: $ANTHROPIC_ADMIN_KEY"The bucket width you choose determines what the query is good for. Each width has a default number of buckets and a higher maximum you can request with limit.
| bucket_width | Default limit | Maximum limit | Use case |
|---|---|---|---|
| 1m | 60 buckets | 1,440 buckets | Real-time monitoring |
| 1h | 24 buckets | 168 buckets | Daily patterns |
| 1d | 7 buckets | 31 buckets | Weekly/monthly reports |
During an active incident, an on-call engineer needs a near real-time view of organization-wide token consumption spikes, checking for anomalies roughly once a minute. Which Usage API configuration fits this monitoring need within the documented granularity limits?
Correct answer: A — Request bucket_width=1m, raise the limit above the 60-bucket default up to the 1440-bucket maximum as needed, and poll about once per minute
- A. Correct. Minute-level buckets are intended for real-time monitoring, the limit can be raised well past its default to capture a longer recent window, and roughly once-per-minute polling matches the documented sustained polling guidance.
- B. Incorrect. Daily buckets aggregate an entire day into one data point, which cannot reveal minute-level spikes during an active incident regardless of how many buckets are requested.
- C. Incorrect. Hourly buckets aggregate a full hour of usage into a single value; they do not refresh at 60-second intervals, and they are too coarse to spot a spike within a given hour.
- D. Incorrect. Polling far more frequently than the recommended once-per-minute cadence does not produce fresher underlying data, since usage figures typically finalize within about five minutes of request completion, and it needlessly increases API load.
Sources1
2.The Cost API and paging through results
Spend comes from a separate endpoint, /v1/organizations/cost_report, which returns service-level cost breakdowns. It covers token usage, web search and code execution costs. You can group by workspace or by description; grouping by description adds parsed fields such as model and inference_geo to the response. Two details matter when you read the numbers. Costs are in USD as decimal strings in the lowest unit, cents. And unlike the usage endpoint, the cost endpoint has daily buckets only, so minute or hourly cost views are not available from it.
Both endpoints page the same way. Make the first request; if the response has has_more set to true, pass its next_page value as page in the next request; stop when has_more is false.
You do not have to build all of this yourself. The documentation notes that leading observability platforms offer ready-made Claude usage and cost integrations with dashboards, alerting and analytics, with no custom code required.
| Aspect | usage_report/messages | cost_report |
|---|---|---|
| Measures | Uncached input, cached input, cache creation, and output tokens | Token usage, web search, and code execution costs |
| Time buckets | 1m, 1h, or 1d | Daily granularity only (1d) |
| Breakdowns | Filter and group by API key, workspace, model, service tier, context window, data residency, or speed | Group by workspace or description |
| Units | Token counts | USD, decimal strings in cents |
Only a daily figure per workspace: the cost endpoint supports 1d buckets only. For an hourly picture, use the usage endpoint, which reports token counts in 1h buckets.
Sources1
3.Picking the right reporting source, and trusting its freshness
Each tool answers a different question, and they refresh on different schedules. For Claude Code specifically, the Claude Code Analytics Admin API sits between the basic Analytics dashboard and a full OpenTelemetry pipeline. Its endpoint, /v1/organizations/usage_report/claude_code, returns one record per user for the single day given by starting_at. Each record holds sessions, lines of code, commits, pull requests, tool accept and reject counts, and a model_breakdown of input, output, cache_read and cache_creation tokens with an estimated cost in cents. Metrics can arrive up to an hour late.
Claude Enterprise organizations use the Claude Enterprise Analytics API instead, and its freshness rules differ by endpoint. Engagement and adoption data normally arrives with a one-day lag; if you request a day that is not ready, you get a 400 error naming the most recent available day. Cost and usage data usually arrives within four hours but can take up to 24. Those values can also be revised for up to 30 days as late events arrive and reconciliation runs, so invoicing-grade totals should come from dates at least 30 days in the past.
| Source | Best for | Granularity and freshness |
|---|---|---|
| OpenTelemetry (metrics, log events, traces) | Per-run behaviour: tool calls, request latency, failures | Streamed at the export interval you configure |
| Usage & Cost Admin API | Organization-wide tokens and spend for Claude Platform API usage | Usage in 1m/1h/1d buckets; cost daily |
| Claude Code Analytics Admin API | Per-user daily Claude Code productivity and model cost | One day per request; up to 1-hour delay |
| Claude Enterprise Analytics API | Enterprise engagement, adoption, cost and usage | Engagement about 1 day behind; cost revisable for 30 days |
4.Diagnosing a cache-read drop from the usage signal
Cached input is one of the token categories the Usage API reports, and it is also returned on each response as usage.cache_read_input_tokens. If that number falls to zero, cost and latency rise, but the number alone does not tell you why. Cache diagnostics fills that gap. Add a diagnostics object to every request. On the first turn pass previous_message_id as null; on later turns pass the id of the previous response. The API then compares the two requests and reports the first place they diverged: the model, the system prompt, the tools or the message history.
The comparison looks at request structure, regardless of whether the cache actually hit, so read it alongside the usage figures. A null diagnostics object means no divergence was found. A null cache_miss_reason means the comparison is still pending. Only requests that include the diagnostics object are fingerprinted, and a fingerprint holds hashes and token-count estimates, never raw prompt content.
diagnostics = r2.diagnostics
if diagnostics is None:
print("No divergence detected.")
elif diagnostics.cache_miss_reason is None:
print("Comparison still pending.")
else:
print(f"cache_miss_reason: {diagnostics.cache_miss_reason.type}")Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The Cost API supports the same 1m and 1h buckets as the Usage API, so you can chart spend per hour.Why is that wrong?
The cost_report endpoint has daily buckets only. For finer-grained views, use token counts from the usage endpoint.
Covered in The Cost API and paging through results
2.Any API key that can call the Messages API can also pull the organization's usage and cost reports.Why is that wrong?
Claude Platform organizations need an Admin API credential for this API. Claude Enterprise organizations have no Admin API keys and use an Analytics API key with the Enterprise Analytics API.
Covered in The Usage API: token consumption across the organization
3.Enterprise Analytics cost figures are final once they appear, so yesterday's numbers are safe to invoice from.Why is that wrong?
Cost and usage values can be revised for up to 30 days while late events are reconciled. Invoicing-grade totals should use dates at least 30 days old.
Covered in Picking the right reporting source, and trusting its freshness
Practise it for real
Pull a week of organization token usage broken down by model, and page through a longer range, using the Usage API.
1.Export an Admin API key as ANTHROPIC_ADMIN_KEY, then call /v1/organizations/usage_report/messages with starting_at, ending_at and bucket_width=1d.
Why: The usage report requires an Admin API credential, and 1d buckets suit weekly reporting.
You should see: A response of daily buckets with uncached input, cached input, cache creation and output token counts.
2.Add group_by[]=model to the same request.
Why: Grouping splits each bucket by model, which shows which model is driving consumption.
You should see: Each daily bucket now contains one result per model.
3.Widen the range to a month and add limit=7.
Why: Seven buckets cannot cover a month, so the response has to be paged.
You should see: The response includes "has_more": true and a next_page value.
4.Repeat the request with page set to the next_page value, until has_more is false.
Why: This is the documented cursor pagination shared by the usage and cost endpoints.
You should see: The last page returns has_more false, and together the pages cover the whole month.
Stuck? Get a nudge
If you need minute resolution during an incident, switch to bucket_width=1m and raise limit toward its 1,440-bucket maximum.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Monitor product performance while measuring if changes to the system have improved it, or set up alerting”
↩︎ The Usage API: token consumption across the organization“Aggregate usage data in fixed intervals (1m, 1h, or 1d)”
↩︎ The Usage API: token consumption across the organization“All costs in USD, reported as decimal strings in lowest units (cents)”
↩︎ The Cost API and paging through results“Leading observability platforms offer ready-to-use integrations for monitoring your Claude API usage and cost, without writing custom code.”
↩︎ The Cost API and paging through results“Time buckets: Daily granularity only (1d)”
↩︎ Exam trap 1“Admin API key (sk-ant-admin01-...) or another Admin API credential”
↩︎ Exam trap 2 - 2.
“This API provides more detail than the basic Analytics dashboard without the complexity of the OpenTelemetry integration.”
↩︎ Picking the right reporting source, and trusting its freshness“Data freshness: Metrics are available with up to 1-hour delay for consistency”
↩︎ Picking the right reporting source, and trusting its freshness - 3.
“Values for a given date can be revised for up to 30 days as late events arrive and reconciliation runs.”
↩︎ Picking the right reporting source, and trusting its freshness“For invoicing-grade totals, query dates at least 30 days in the past.”
↩︎ Exam trap 3 - 4.
“Without cache diagnostics, the only signal is usage.cache_read_input_tokens dropping to zero, with no indication of what changed.”
↩︎ Diagnosing a cache-read drop from the usage signal“The comparison is about request structure, independent of whether the cache actually hit.”
↩︎ Diagnosing a cache-read drop from the usage signal