What you will be able to do
- Explain why production callers should pin an explicit prompt, agent or Skill version instead of taking the latest
- Use an evaluation suite as the gate for promoting a new version, and read a per-category regression
- Roll back to a known-good version, or canary a new one, without redeploying application code
- Keep a versioned artifact maintainable over time with a registry, Git history and a deprecation rule
1.Pin production to an explicit version
The artifacts that shape a Claude application's behaviour change more often than its code does: system prompts, agent definitions, Skills. When those live server-side and are versioned there, the question that matters in production is which version a request actually uses. Anthropic's guidance for Skills is direct. Pin production to a specific version, because leaving the version out means you get whatever was uploaded most recently.
Managed Agents work the same way for prompts. Creating an agent returns version 1, and every later update produces a new immutable version under the same agent ID. When you open a session, you choose. Pass the bare agent ID and you get the latest version. Pass an object with a type, an id and a version and the session is pinned to that exact version. The cookbook example uses the pinned form below so that its comparisons are controlled.
{"type": "agent", "id": ..., "version": ...}Pinning matters because an update comes with no approval step. Any key in the workspace can publish a new agent version, and every caller that passes the bare ID picks it up on its very next session. The fix is to make the pinned number the thing under change control. Creating v2, v3 or v10 is cheap and harmless, because those versions receive no traffic. Promotion means changing the configuration value that tells production callers which version to pass, and that change goes through your normal review process.
2.Promote only what passes evaluation
If promoting a version is the controlled step, you need evidence for the promotion decision. Anthropic's Skills guidance makes the evaluation suite that gate. Run the full suite before promoting, treat every update as a new deployment that needs a full security review, and use the latest versions only in development and testing to validate changes.
The prompt-versioning cookbook shows why the suite has to be broken down by category. A support-ticket triage agent was scored against 20 labelled tickets, five for each of four teams. Then a PM added a rule sending anything about API usage or rate limits to the platform team. On a usage-billed product, billing tickets mention API usage too.
| Team | v1 | v2 |
|---|---|---|
| api-platform | 5/5 | 5/5 |
| auth | 5/5 | 5/5 |
| billing | 4/5 | 2/5 |
| dashboard | 5/5 | 5/5 |
A fintech company runs automated compliance workflows against Claude Opus 4.5 and must guarantee that model behavior never changes between deployments without an explicit, reviewed migration step. When specifying the model in their Messages API calls, which practice should they adopt?
Correct answer: B — Specify the exact dated snapshot ID, and update it only during a deliberate, scheduled migration you control.
- A. Incorrect -- a convenience alias can silently resolve to a newer dated snapshot over time, which is exactly the unreviewed behavior change the compliance workflow needs to avoid.
- B. Correct -- pinning the exact dated snapshot ID keeps behavior fixed until the team deliberately migrates to a new snapshot on their own schedule.
- C. Incorrect -- there is no unpinned 'latest' pointer in the Messages API model field, and even if there were, it would introduce the same uncontrolled drift the team is trying to prevent.
- D. Incorrect -- the model parameter is required on Messages API requests; there is no default model selection to fall back on.
3.Rollback and canary without a deploy
Because older versions stay on the server, recovering from the billing regression was a configuration change, not a release. Callers went back to passing version 1, and billing routing returned to 4/5. Version 2 was also kept. The PM could keep iterating on it while production stayed on v1, then create v3 and send it through the same evaluate-then-promote process.
The Skills guidance asks for the same readiness in advance. Keep the previous version available as a fallback, and if a new version fails evaluations in production, revert immediately. For higher-stakes agents, the cookbook suggests a middle step before full promotion: send a fraction of traffic to the new version and compare results. Used this way, versioning works as a feature flag for prompt changes, and you can evaluate and promote prompts separately from application code.
4.Maintaining versioned artifacts over time
Maintenance is the long tail of the life cycle. For Skills, Anthropic recommends an internal registry entry for each Skill that records its purpose, owner, currently deployed version, dependencies such as MCP servers or packages, and evaluation status with the date of the last run. The source files belong in Git, which gives you history, review through pull requests, and rollback. Git is the single source of truth, including when the same Skills are deployed across several surfaces.
Evaluation also tells you when a Skill should be retired. Update a Skill when its workflow changes or its scores decline, and deprecate it when evaluations keep failing or the workflow no longer exists. Load count needs the same discipline: every Skill's name and description competes for attention in the system prompt, and an API request accepts at most 20 Skills.
The Managed Agents production cookbook covers the operational side. It lists resource lifecycle verbs (list, retrieve, update, archive, delete) for managing what a workspace builds up over time, and an inference geography pin, inference_geo, for cases where compliance requires model requests to run in a specific region.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Omitting the version number is the safe default, because production then keeps whatever it was using.Why is that wrong?
Omitting the version gives you the latest one. Anyone in the workspace who uploads a new version immediately changes what production runs, so production should pin an explicit version.
Covered in Pin production to an explicit version
2.Rolling back a regressed prompt requires reverting code and redeploying the application.Why is that wrong?
With server-side versions, rollback means pointing callers back at the last known-good version. Keep that version available as a fallback and revert to it straight away.
Covered in Rollback and canary without a deploy
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Production: Pin Skills to specific versions.”
↩︎ Pin production to an explicit version“Run the full evaluation suite before promoting a new version. Treat every update as a new deployment requiring full security review.”
↩︎ Promote only what passes evaluation“revert to the last known-good version immediately.”
↩︎ Rollback and canary without a deploy“Store Skill directories in Git for history tracking, code review through pull requests, and rollback capability.”
↩︎ Maintaining versioned artifacts over time“Deprecate Skills when evaluations consistently fail or the workflow is retired.”
↩︎ Maintaining versioned artifacts over time“a new version uploaded by anyone in the workspace immediately changes what production agents run”
↩︎ Exam trap 1“Rollback plan: Maintain the previous version as a fallback.”
↩︎ Exam trap 2 - 2.https://platform.claude.com/cookbook/managed-agents-cma-prompt-versioning-and-rollbackSecondary source
“Every agents.update produces a new immutable version, and sessions choose which version to use by ID.”
↩︎ Pin production to an explicit version“There's no built-in approval workflow on agents.update. Any key in the workspace can call it.”
↩︎ Pin production to an explicit version“The new rule is broad, and on a usage-billed API product the billing tickets talk about API usage too.”
↩︎ Promote only what passes evaluation“Rolling back isn't a deploy; callers just go back to passing version: 1.”
↩︎ Rollback and canary without a deploy“You can effectively use this versioning as a feature flag.”
↩︎ Rollback and canary without a deploy - 3.
“Inference geography pinning (inference_geo) when compliance requires model requests to run in a specific region.”
↩︎ Maintaining versioned artifacts over time