Model configuration
This is the canonical reference for which model the agent uses and where to change it. The model ID is configured across five independent layers in three languages, so read this section before changing a default — a mismatch between the layers fails every task on the stack at turn 0, not just an edge case.
The five layers
Section titled “The five layers”| # | Layer | What it controls | Where | ID form |
|---|---|---|---|---|
| 1 | IAM invoke allowlist | Which models the agent’s roles may invoke at all, and in which geography. The outer gate — everything below fails without it. | DEFAULT_BEDROCK_MODEL_IDS (cdk/src/constructs/bedrock-models.ts); override with CDK context bedrockModels. Geography via context bedrockGeoRegion (default us) | Bare (anthropic.claude-…) |
| 2 | Platform default model | The model used when nothing narrower is set. Injected by the stack as ANTHROPIC_MODEL, derived from bedrockGeoRegion; the Python literal is the fallback a run with no env reaches. | PLATFORM_DEFAULT_MODEL_ID (cdk/src/handlers/shared/bedrock-model-constants.ts) — the value actually injected, so the one to change; then the no-env fallbacks agent/src/config.py:563 and agent/src/models.py:157 (TaskConfig.anthropic_model) | Prefixed (<geo>.anthropic.…) |
| 3 | Auxiliary / fast model | The small model Claude Code uses for auxiliary work (WebFetch page summarization, the pre-flight safety check). | Stack env ANTHROPIC_DEFAULT_HAIKU_MODEL (cdk/src/stacks/agent.ts (the runtime environment block)); agent-side fallback at agent/src/config.py:569 | Prefixed (<geo>.anthropic.…) |
| 4 | Per-repo override | One repository’s model, with no agent redeploy. | Blueprint agent.modelId (cdk/src/constructs/blueprint.ts, BlueprintProps.agent.modelId) → RepoTable model_id (cdk/src/handlers/shared/repo-config.ts:37) → ECS injects ANTHROPIC_MODEL (cdk/src/handlers/shared/strategies/ecs-strategy.ts:217) | Prefixed (<geo>.anthropic.…) |
| 5 | Per-task / local | One task’s model. Payload model_id is aliased to anthropic_model (agent/src/pipeline.py, _PAYLOAD_KEY_ALIASES); local batch runs read ANTHROPIC_MODEL from the shell via agent/run.sh. | Task payload model_id; shell ANTHROPIC_MODEL | Prefixed (<geo>.anthropic.…) |
Environment variables
Section titled “Environment variables”| Variable | Who sets it | ID form | Purpose |
|---|---|---|---|
ANTHROPIC_MODEL | ECS strategy from the repo Blueprint (layer 4); you, in the shell, for local batch runs (layer 5) | Prefixed inference profile | The main coding model. Unset → the agent/src/config.py fallback. |
ANTHROPIC_DEFAULT_HAIKU_MODEL | The CDK stack, hardcoded at cdk/src/stacks/agent.ts (the runtime environment block) | Prefixed inference profile | The small/fast auxiliary model. Must be a granted profile, or the pre-flight check times out with “Pre-flight check is taking longer than expected”. |
CLAUDE_CODE_USE_BEDROCK | The CDK stack (='1') and agent/run.sh | — | Routes Claude Code to Bedrock instead of the Anthropic API. ABCA always runs on Bedrock. |
Precedence — narrowest wins
Section titled “Precedence — narrowest wins”per-task payload model_id (layer 5) > blueprint agent.modelId (layer 4, arrives as stack env ANTHROPIC_MODEL) > stack env ANTHROPIC_MODEL (layer 3-adjacent / local shell) > agent/src/config.py fallback (layer 2 — global.anthropic.claude-opus-5)Every one of those is gated by the IAM invoke allowlist (layer 1), which is itself gated by account-level Bedrock model access. Both gates are silent until invocation: a model that resolves fine through precedence still fails at turn 0 with AccessDenied if it is not in the grant list, and fails again if your account has not completed Bedrock model access for it.
Bare vs. prefixed IDs — the one rule that bites
Section titled “Bare vs. prefixed IDs — the one rule that bites”Layer 1 takes bare foundation-model IDs; every other layer takes the prefixed inference-profile ID. This asymmetry is deliberate: both grant sites derive the inference-profile ARN by adding the geo prefix themselves, so a prefixed entry in bedrockModels would produce an invalid us.us.anthropic.… ARN. The resolver rejects an entry carrying any modelled geo prefix so the typo fails at synth rather than at runtime.
In the other direction, a bare ID cannot be invoked on demand at all. Verified:
$ aws bedrock-runtime invoke-model --model-id anthropic.claude-opus-5 ...ValidationException: Invocation of model ID anthropic.claude-opus-5 with on-demandthroughput isn't supported. Retry your request with the ID or ARN of an inferenceprofile that contains this model.So: bedrockModels context → anthropic.claude-opus-5. Everywhere else → global.anthropic.claude-opus-5 (or whatever geo you have configured — see below).
Choosing the inference-profile geography
Section titled “Choosing the inference-profile geography”Which geography those prefixes name is itself configurable, via CDK context bedrockGeoRegion:
$ cdk deploy -c bedrockGeoRegion=globalor as a context entry in cdk/cdk.json. The shipped cdk.json sets global; the code default absent any context is us. The accepted values are whatever @aws-cdk/aws-bedrock-alpha models — currently global, us, us-gov, eu, apac, jp, au. An unrecognized value fails at synth rather than at deploy: an invented geography produces a well-formed ARN for a profile that does not exist, so the grant would authorize nothing and the agent would fail at turn 0 with AccessDenied and nothing to explain why.
One value drives everything that needs a prefix — every grant site (AgentCore runtime, ECS task role, Lambda MicroVM) and both injected env vars, ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL. That is deliberate: a deployment can never grant one geography’s profiles while telling the agent to call another’s.
Which to choose. A global. profile routes to any supported commercial Region, which gives better throughput and resilience under peak demand — worth having for tasks that run for hours and burst. A geo profile (us., eu., apac., …) keeps inference within that geography, which is what you need under a data-residency requirement. Pick the geo profile in that case; the throughput benefit is not worth a compliance breach.
Two things to check when changing it, both of which fail at runtime rather than synth:
- The models you use must have an active profile in the target geography. Not every model is published to every geo.
- Your account’s Bedrock model access must cover the deployment Region’s entitlements for that geography.
Changing the geography requires a redeploy, because the IAM grants are scoped to explicit profile ARNs resolved at synth. Switching among already-granted models does not — see layer 4.
Bumping the default model
Section titled “Bumping the default model”- Add the bare ID to
DEFAULT_BEDROCK_MODEL_IDS(cdk/src/handlers/shared/bedrock-model-constants.ts) and deploy, so the grant exists before anything tries to use it. - Confirm account-level Bedrock access for the model in the target Region.
- Change
PLATFORM_DEFAULT_MODEL_ID(cdk/src/handlers/shared/bedrock-model-constants.ts). This is the step that changes what deployed tasks run — the stack injects it asANTHROPIC_MODELon every substrate, so editing only the Python literals below has no effect on any deployed task. - Update the prefixed ID in
agent/src/config.pyandagent/src/models.pytoo. These are the fallback a run with no env reaches (a local batch run), and the doc-drift test compares against them. - Verify the SDK price table recognizes the model. The
max_budget_usdguardrail is computed from a price table bundled into the Claude Agent SDK at build time, so an unrecognized model silently degrades budget enforcement. Runagent/scripts/diagnostics/test_sdk_smoke.pywithANTHROPIC_MODELset to the new ID, divide the reported cost by the input-token count, and confirm the implied rate matches published Bedrock pricing. A$0.00or wildly-off result means the table does not know the model and budgets cannot be trusted. - The doc-drift test (
cdk/test/contracts/model-default-docs-parity.test.ts) fails until the documented defaults here and inagent/README.mdmatchconfig.py. That failure is the reminder, not a nuisance — update both.
Cost and model selection
Section titled “Cost and model selection”Model choice is a cost decision, which is why it is adjustable per repo and per task without a code change.
Per-token rate vs. token volume. Measured on the pinned toolchain, same one-turn prompt, same system prompt:
| Model | Input tokens | Reported cost_usd | Implied input rate |
|---|---|---|---|
us.anthropic.claude-opus-4-8 | 32,145 | $0.160850 | $5.00/MTok |
us.anthropic.claude-opus-5 | 37,584 | $0.188020 | $5.00/MTok |
Token ratio 1.169; cost ratio 1.169 — identical. The per-token rate is unchanged; the whole delta is token volume on an identical prompt. Read it that way: “Opus 5 costs ~17% more per task” invites the wrong remedy (switch models), while “same rate, more tokens” points at the real levers — prompt size, prompt caching, and max_turns.
Where can I set max_budget_usd?
| Surface | How | Status |
|---|---|---|
| Per task, CLI | bgagent submit --max-budget <dollars> (cli/src/commands/submit.ts:69), range 0.01–100 | Works |
| Per task, REST | max_budget_usd in the POST /v1/tasks body | Works |
| Local batch only | MAX_BUDGET_USD shell env, when running entrypoint.py directly | Works locally; ignored by the deployed AgentCore server mode, which reads the budget from the /invocations JSON body |
| Per repo, Blueprint | agent.maxBudgetUsd on the repo’s Blueprint construct | Works — persisted to RepoTable.max_budget_usd; same 0.01–100 range as the CLI, validated at CDK synth so an out-of-range value cannot deploy. See Per-repo overrides. |
| Platform default | — | None by design: unset means unlimited |
Per-task runtime budgets are unlimited by default. Because no platform-wide per-task runtime ceiling applies, the documented mitigation for cost is choosing a lighter-token model rather than relying only on a cap:
- Per repo: Blueprint
agent.modelId— no code change, no agent redeploy. Prefix it with the geography the stack grants, read from itsBedrockGeoRegionoutput; aus.prefix on aglobaldeployment is granted nothing and fails at turn 0. - Per task:
model_idin the task payload - Platform-wide:
PLATFORM_DEFAULT_MODEL_ID(what every substrate is told to invoke) plus thebedrockModelscontext (what it is granted) — both, or the deploy fails at synth for disagreeing
The model must be in the IAM grant list (layer 1) or the task fails at turn 0 with AccessDenied — the grant is the gate, so a lighter model is only reachable if it is granted.
Trust boundary on the number. cost_usd is the Claude Agent SDK’s client-side estimate from that bundled price table — not authoritative billing. It drifts when Bedrock pricing changes, when the SDK version does not recognize a model, or when discounts and commitments apply. See Cost attribution (the warning at line 6); authoritative cost comes from AWS Cost Explorer / CUR 2.0.
Fleet monthly budgets are a separate admission control. bgagent budget set --user|--team --monthly-usd <amount> [--hard-stop] stores a recurring user or Cognito-group limit. Terminal task costs roll up by UTC month from the TaskTable stream. The 80% and 100% crossings publish CloudWatch alarms; hard-stop scopes reject new task creation at 100% without terminating in-flight work. See Monthly user and team budgets.