Skip to content

Adr 021 lambda microvms compute backend

ADR-021: AWS Lambda MicroVMs as a third ComputeStrategy backend

Section titled “ADR-021: AWS Lambda MicroVMs as a third ComputeStrategy backend”

Number: candidate ADR-021 (ADR-020 is the highest accepted on main; ADR-018 is claimed by open PR #548, ADR-019 by open PR #663). Numbers are never reused. If a lower number frees before merge, renumber and coordinate with those PRs.

Status: proposed Date: 2026-07-29

ABCA selects a per-repo compute backend through the Blueprint’s compute_type field (cdk/src/handlers/shared/repo-config.ts). Two backends exist today, resolved by resolveComputeStrategy (cdk/src/handlers/shared/compute-strategy.ts) behind a uniform ComputeStrategy interface (startSession / pollSession / stopSession):

  • AgentCore Runtime (agentcore, default) — managed Firecracker MicroVM per session. Invoked via InvokeAgentRuntime; liveness is inferred from agent heartbeats in DynamoDB plus the FastAPI /ping endpoint (agent/src/server.py) — the strategy’s pollSession is a stub that always reports running. Constraints: 2 GB image limit, no substrate-level suspend API exposed to the orchestrator.
  • ECS on Fargate (ecs) — always-on Fargate task (16 vCPU / 120 GB, ARM64) for repos that exceed AgentCore’s limits (#596). Invoked via RunTask in batch mode (bypassing the HTTP server); liveness via DescribeTasks. No suspend — an idle task blocked on an approval wait burns full compute the whole time.

AWS Lambda MicroVMs (launched 2026-06-22) is a serverless Firecracker sandbox primitive that AWS positions explicitly for AI coding agents: VM-level isolation, snapshot-based near-instant launch, suspend/resume with full memory + disk state preserved (compute charges stop while suspended), a dedicated JWE-authenticated HTTPS endpoint per instance, lifecycle hooks (/run, /suspend, /resume, /terminate), and up to 8 hours per session. It is not classic Lambda: the 15-minute function cap does not apply, and COMPUTE.md’s “Lambda: poor fit” verdict refers to functions, not MicroVMs.

#645 proposes adding Lambda MicroVMs as a third ComputeStrategy. The fit is strong but not free — pre-implementation review of the service’s lifecycle model surfaced real design tensions this ADR resolves:

Capability comparison (delta rows only — full matrix in COMPUTE.md)

Section titled “Capability comparison (delta rows only — full matrix in COMPUTE.md)”
AgentCore RuntimeECS on FargateLambda MicroVMs
IsolationMicroVM (managed)Task-level (Firecracker)MicroVM (Firecracker)
Max duration8 hNo cap8 h (running + suspended) — verified (L-B430C318 = 8 hours)
Suspend/resumeNo orchestrator-visible APINoYes — explicit API + idle policy, state preserved, no compute charge while suspended. Verified: suspend reaches SUSPENDED in ~1 s with no idlePolicy; resume restores RUNNING in ~1 s with microvmId and endpoint byte-identical
ResourcesAgentCore-managed16 vCPU / 120 GB / 20–200 GB diskBaseline 8 GiB RAM / 4 vCPU, auto-scaling to a 32 GiB / 16 vCPU peak; 32 GB disk. minimumMemoryInMiB configures the BASELINE (max 8,192 MiB); the service scales vertically on demand — capacity is baseline-priced with 4× burst headroom
PackagingECR image ≤ 2 GBECR image, no hard capZip + Dockerfile in S3 → service-built snapshot image (versioned, storage billed)
InvocationInvokeAgentRuntime (SigV4)RunTask + container overridesRunMicrovm (image ARN required — a bare name is rejected) → dedicated HTTPS endpoint + JWE token (CreateMicrovmAuthToken, ≤ 60 min TTL)
LivenessAgent heartbeat + /pingDescribeTasksMicroVM state (RUNNING / SUSPENDED / TERMINATED) via control-plane API and agent heartbeat (see sub-decision 1)
Session storage/mnt/workspace FUSE (no flock())Ephemeral diskNative disk in snapshot — survives suspend/resume, flock() works
ArchitectureARM64ARM64ARM64 (Graviton)
Regions (launch)BroadBroad5 (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1)

Rows marked verified were discharged empirically on 2026-07-31 (us-east-1); see docs/verification/645-p1-lambda-microvm-runbook.md. Two originally documented constants were refuted by that run and are corrected throughout this ADR: the runHookPayload cap (16 KB → 4 KB) and the claim that RunMicrovm accepts a bare image name (it does not).

Source hierarchy for service facts. Where sources disagree, the higher tier wins, and the tier is recorded next to the claim:

  1. Live boundary probe — a request the service actually accepted or rejected in this account/Region. Strongest, and the only tier that can refute the others. (Both corrections above came from here.)
  2. Modeled constraints — the CLI/SDK request schema and enumerated allowed values (--generate-cli-skeleton, ARM_64, ENABLED|DISABLED). Authoritative about request shape; silent about runtime behaviour.
  3. Service developer guide — including its sizing/scaling tables. Authoritative about semantics a probe cannot see, which is exactly how the memory row was fixed: a probe can only show that 32,768 MiB is rejected as a minimumMemoryInMiB value; only the guide explains that the field is a BASELINE and that the service scales to a 32 GiB peak on its own. A boundary probe is the strongest evidence about a boundary and says nothing about what the boundary means.
  4. SDK docstrings — generated, and demonstrably stale here: runHookPayload documents “Maximum: 16,384 bytes” against an enforced 4,096.
  5. Launch blogs / skills / toolkit material — orientation only; never load-bearing on its own.

Omitted API fields mean “service default”, never “none”. Two live findings drove this rule: leaving ingressNetworkConnectors unset attaches a PUBLIC HTTP_INGRESS connector, and leaving /ready out of hooks makes the image un-creatable once any lifecycle hook is enabled. So for any security-relevant field, the desired posture must be requested explicitly and the test must assert the outcome (the ARN present in the request, the hook enabled) rather than the omission (expect(field).toBeUndefined()) — an omission assertion passes just as happily when the service is silently choosing something wider.

On the memory row specifically: CreateMicrovmImage enumerates the baseline sizes a base image supports — [512, 1024, 2048, 4096, 8192] MiB for al2023-1 — and rejects anything else, which is why the construct validates against that list at synth. The 32 GiB / 16 vCPU peak is reached by the service’s own vertical scaling, not by asking for it. The account memory quota (L-CD1C0CC4, 1024 GB, “burst up to 4×”) is an aggregate across MicroVMs, not a per-VM limit; note that concurrency arithmetic should be done against the PEAK, not the baseline, since that is what a busy fleet can actually consume.

  1. Idle detection is inbound-traffic-based; the ABCA agent is outbound-only. MicroVM idle policies suspend when no traffic arrives at the endpoint. A busy agent running a 40-minute build receives no inbound traffic and would be suspended mid-work by a naive idle policy. Conversely, “no inbound traffic” is the agent’s normal state.
  2. No self-suspend. The agent cannot suspend its own MicroVM from inside; only an external SuspendMicrovm call can. Suspend decisions must be owned by the orchestrator — which aligns with the unified liveness model proposed in #491.
  3. Snapshots bake state. The image snapshot is captured once at build time; every MicroVM resumes from it. Secrets, tokens, and per-task identity must arrive at run time (runHookPayload, ≤ 4 KB, or fetched in the /run hook), never at image build. CSPRNG reseeding is a consequence of the same property and is scoped to P3 — see the amended note under sub-decision 2 and the risk bullet, which record why the exposure is negligible today.
  4. Auth tokens are short-lived. JWE tokens max out at 60 minutes; any orchestrator→agent HTTP interaction over the endpoint needs token refresh, unlike AgentCore’s SigV4 invoke or ECS’s no-endpoint model.
  5. Identity delta — narrower than it looks. Most AgentCore services ABCA uses are standalone and substrate-portable: Memory is already consumed from ECS via an IAM grant plus MEMORY_ID (EcsAgentCluster), and Gateway (ADR-019) is portable by design (SigV4 inbound). The genuinely Runtime-coupled piece is the workload-access-token delivery mechanism (runtimeUserId → WorkloadAccessToken request header → BedrockAgentCoreContext, used by resolve_linear_api_token()), which has no MicroVM equivalent. The ECS backend already lives with this delta (env-var token delivery); MicroVMs inherit the same posture until the pluggable identity work (#249, ADR-016) redesigns the seam.
  6. The service’s defaults are not our posture. Two of them, both discovered live: RunMicrovm attaches a public HTTP_INGRESS connector (and mints a public *.lambda-microvm.<region>.on.aws endpoint) when ingressNetworkConnectors is omitted, and CreateMicrovmImage requires the /ready hook whenever any lifecycle hook is enabled. Neither posture can be reached by leaving a field out — each needs an explicit control (see sub-decisions 3 and 4).

Adopt AWS Lambda MicroVMs as a third, opt-in ComputeStrategy backend named lambda-microvm, selected per repo via Blueprint compute_type. AgentCore remains the default. Five sub-decisions:

1. Strategy shape: extend the interface with mandatory suspend/resume

Section titled “1. Strategy shape: extend the interface with mandatory suspend/resume”

ComputeType widens to 'agentcore' | 'ecs' | 'lambda-microvm' (mirrored in cli/src/types.ts and the CLI’s inline unions). SessionHandle gains a { strategyType: 'lambda-microvm', microvmId, endpoint } variant — microvmId because every lifecycle API (suspend-microvm, resume-microvm, terminate-microvm, get-microvm, and create-microvm-auth-token — the latter not called in P1–P3, see sub-decision 3) takes only the MicroVM identifier, and endpoint because it is per-session (minted by RunMicrovm) and required for any orchestrator→agent HTTP interaction. Note the naming seam: the handle field is microvmId (matching RunMicrovmResponse.microvmId), while the request key on every lifecycle command is microvmIdentifier — the strategy is the only place that translates between the two. The image ARN is deliberately not in the handle: like the ECS task definition ARN, it is deployment-time configuration consumed by startSession (from construct-injected environment) and recorded in the session-start log entry for diagnostics, not per-session lifecycle state.

A second, sharper naming seam: imageIdentifier must be an ARN. The name suggests a bare image name is acceptable — create-microvm-image --name takes one, and this ADR originally assumed run-microvm --image-identifier would too. It does not: a bare name is rejected with ValidationException: Malformed ARN - doesn't start with 'arn:', and so is list-microvm-image-builds --image-identifier <name> (Invalid ARN format). The construct therefore resolves an operator-supplied name to its exact arn:${Partition}:lambda:${Region}:${Account}:microvm-image:${Name} ARN once — the same value it scopes the lifecycle IAM grant to — and injects THAT as MICROVM_IMAGE_IDENTIFIER. One derivation, two consumers, so a request field and an IAM resource can never disagree. The strategy validates the invariant and fails fast with the remedy, because the service’s own error names neither the env var nor the fix.

The ComputeStrategy interface gains mandatory suspendSession(handle) / resumeSession(handle) methods returning a typed result ({ supported: false } | { supported: true }-shaped, exact type at implementation time); the widening lands in P3 (see sub-decision 5), in one commit across all three strategies. Mandatory-with-explicit-stub is the codebase idiom, not optional methods: no behavioral interface in the codebase has an optional method, AgentCoreComputeStrategy.pollSession is already a mandatory explicit stub rather than an optional member, and the exhaustive-never switch culture means a fourth backend must make a compile-checked decision about its suspend semantics instead of silently falling through a strategy.suspendSession?.() feature-detection. The agentcore and ecs strategies return unsupported (not a silent success — a suspend that silently no-ops would let the orchestrator believe compute billing stopped when it did not); the orchestrator gates its suspend policy on the typed response, consistent with how pollTaskStatus already branches explicitly on computeType.

Poll semantics — the strategy reports, the orchestrator interprets. pollSession(handle) receives only the session handle and cannot see task state, so the health rules must live where the DynamoDB status lives. SessionStatus gains a 'suspended' variant; the strategy maps GetMicrovm state mechanically and the orchestrator cross-references against the task row — the same division of labor finalPollState already uses for ECS (substrate stopped + non-terminal DynamoDB status → failed) and pollTaskStatus uses for agentcore heartbeats: substrate suspended + task AWAITING_APPROVAL is healthy (orchestrator-intended suspend); suspended with any other task status is an anomaly to surface, not fail-fast; substrate terminal + non-terminal task status → classify failed.

Liveness on this backend is substrate state AND agent heartbeat. The substrate cross-check above answers one question — “is the VM still there?” — and P2 established that it is not sufficient on its own. GetMicrovm catches a MicroVM that died; it cannot catch a MicroVM that is alive and reporting RUNNING while the pipeline inside the guest is hung, deadlocked, or was OOM-killed. That state is not hypothetical or self-correcting on this substrate, and the P2 live run narrowed why without weakening the conclusion. The service does reap a VM whose run hook FAILS (a 4xx makes it terminate within ~12 s — see sub-decision 2), so the P1 evidence for this paragraph (a hook-less image sitting in RUNNING indefinitely with no stateReason) no longer describes an ABCA image. What the service reaps is a hook result; it has no view into the guest afterwards. So the surviving — and more realistic — hang case is a /run hook that returned 200 and a pipeline that then hung, deadlocked or was OOM-killed behind it: the substrate stays RUNNING, the service is satisfied, and nothing else notices. Left to the substrate check alone, such a task would burn the orchestrator’s full ~8.5 h poll window — billing an 8-hour reservation — before the safety net fired.

The in-guest half is already being written: the agent updates agent_heartbeat_at on the task row unconditionally, with no backend awareness, so the timestamp exists on every substrate. Only the orchestrator’s reaction to it was backend-scoped — pollTaskStatus evaluated staleness for agentcore alone — which is the gap P2 closes by extending it to lambda-microvm. The grace and stale thresholds are the SAME on both: the timestamp is written by the same pipeline code at the same cadence, so a backend-specific window would encode a difference that does not exist. The two signals stay complementary rather than redundant — the substrate check is the crash detector, the heartbeat is the hang detector — and the check remains scoped to task status RUNNING, which is what keeps a deliberately suspended VM during an approval wait (sub-decision 2, P3) from being read as a dead one. ecs is deliberately left out: DescribeTasks reports a real container exit with an exit code (OOM-kill included) and the ECS poll block already interprets it with its own patience counters, so adding the heartbeat there would give one backend two independently-tuned kill paths for the same failure.

The service’s MicrovmState enum has six members, not three, so the mapping is stated exhaustively (one line of rationale each, mirrored in the strategy’s doc comment):

MicrovmStateSessionStatusWhy
PENDINGrunningStill booting; the same way ECS’s PENDING/PROVISIONING map to running.
RUNNINGrunning—
SUSPENDINGsuspendedAlready on its way to frozen; reporting running would tell the orchestrator compute is still progressing when it is not. Both suspend states land on a report the orchestrator treats as benign-or-anomalous depending on task status, never as failure. Never observable in practice — suspend reached SUSPENDED in under 1 s live — so it is mapped for completeness and nothing may wait for it.
SUSPENDEDsuspended—
TERMINATINGcompletedTerminal-bound and carries no exit code, so “the substrate is gone” is all the strategy can honestly say.
TERMINATEDcompletedSuccess vs failure is the orchestrator’s call — it cross-references the DynamoDB status. This is the load-bearing terminal signal, not NotFound (see below).
unrecognizedrunningA future service enum addition must never fail a healthy task; the strategy warns and keeps polling.

GetMicrovm ResourceNotFoundException → completed, but as a LATE fallback. This deliberately diverges from ecs-strategy, where DescribeTasks returning no task maps to failed. ECS keeps stopped tasks describable for roughly an hour, so a missing task there really is anomalous; a MicroVM is eventually reaped from the control plane by design, so failed would fail tasks that finished cleanly.

What the live run corrected is the timing, not the mapping. NotFound is not the near-term terminal signal: a terminated MicroVM reported TERMINATING at +1 s, TERMINATED at +3 s, and was still TERMINATED ~10 minutes later and at every subsequent checkpoint — ResourceNotFoundException was never observed in that window. So the branch that actually fires in practice is TERMINATED → completed in the table above; the NotFound rule covers only a VM reaped after a long gap (a poll resumed after a crash, a stranded-task reconciler sweep). Both are required and neither substitutes for the other: without the TERMINATED row the orchestrator would poll a finished VM until its safety net fired, and without the NotFound rule a late sweep would classify a cleanly-finished task as a poll error.

The divergence is safe because it does not weaken detection: the orchestrator still fails the task when a terminal report lands while the DynamoDB status is non-terminal, so a genuine mid-run disappearance is caught — it simply receives the substrate-failure classification instead of a misleading poll error. Because that cross-check acts on a status read earlier in the same poll cycle, the orchestrator re-reads the task row before failing (the normal shutdown order is “agent writes terminal status → agent exits → VM terminates”, which a stale read would otherwise turn into a spurious failure); ECS buys the same protection with a five-consecutive-poll patience counter instead.

Neither the mapping nor the NotFound rule is a health decision: both are mechanical restatements of substrate state, which is what keeps the “strategy reports, orchestrator interprets” split intact.

Normative requirements (EARS, per ADR-020):

  • When a task’s Blueprint sets compute_type: 'lambda-microvm', the orchestrator shall resolve the LambdaMicrovmComputeStrategy via resolveComputeStrategy.
  • When startSession is invoked, the strategy shall call RunMicrovm with maximumDurationInSeconds set to 28 800 (the service maximum, matching AgentCore’s 8-hour session cap and sitting inside the orchestrator’s ~8.5 h safety-net poll window).
  • When startSession is invoked, the strategy shall pass a fully-qualified MicroVM image ARN as imageIdentifier.
  • If the configured image identifier is not an ARN, then the strategy shall fail the session start with an error naming the environment variable and the redeploy remedy, before performing any AWS call.
  • When startSession returns, the orchestrator shall persist the MicroVM handle (microvmId, endpoint) in the task row’s compute_metadata (the field cancel-task.ts already reads ECS handles from).
  • The strategy shall omit idlePolicy on every RunMicrovm call, in every phase.
  • The orchestrator shall be the sole initiator of suspension, via suspendSession.
  • When pollSession observes MicroVM state SUSPENDED or SUSPENDING, the strategy shall report suspended without interpreting task state.
  • When pollSession observes MicroVM state TERMINATED or TERMINATING, the strategy shall report completed (the observable terminal state persists for at least ~10 minutes, so completed shall not depend on the MicroVM being reaped).
  • When pollSession observes a MicroVM state it does not recognize, the strategy shall report running.
  • If GetMicrovm reports that the MicroVM does not exist, then the strategy shall report completed.
  • If the strategy reports a terminal substrate state while the task’s DynamoDB status is non-terminal, then the orchestrator shall re-read the task row and, if it is still non-terminal, classify the task as failed with a substrate-failure remedy.
  • If the strategy reports suspended while the task’s DynamoDB status is not AWAITING_APPROVAL, then the orchestrator shall surface an anomaly event and shall not fail-fast the task.
  • While a lambda-microvm task’s DynamoDB status is RUNNING, if the task’s agent_heartbeat_at is stale (or absent past the grace window) by the same thresholds the orchestrator applies to agentcore, then the orchestrator shall treat the session as unhealthy and stop polling — the substrate GetMicrovm check shall remain the crash detector, and the heartbeat shall be the in-guest hang detector.
  • The task-detail API response shall include agent_heartbeat_at, and the CLI shall surface it while the task is non-terminal (P2r2-F11: the field drove the orchestrator’s hang detector but was never projected, so no operator could observe the signal — and its invisibility produced a wrong verification conclusion).
  • The task-summary API response (GET /v1/tasks) shall also include agent_heartbeat_at, and bgagent list shall render it as an age column. Extending the field to the list response is a deliberate widening of the requirement above rather than an incidental one: the detail-only projection makes liveness a per-task question, and an operator checking a fleet of tasks one bgagent status at a time is exactly how the hung task P2r2-F11 describes went unnoticed. Same suppression rule as the detail view (terminal tasks and never-beaten tasks render a placeholder) so the two views cannot disagree.
  • If suspendSession or resumeSession is invoked on a strategy that does not support suspension, then the strategy shall return an explicit unsupported result.
  • When the agent process reaches a terminal state, the agent shall exit.
  • When the orchestrator finalizes a lambda-microvm task, the orchestrator shall call terminate-microvm (termination shall not rely on any substrate timeout, and shall not rely on the MicroVM self-terminating — it does not).

On the omitted idlePolicy: if the block is present all three fields are required, so omission is the unambiguous disabled state the invariant test asserts. This deliberately forgoes suspendedDurationSeconds — it lives inside idlePolicy and cannot be set without re-enabling the traffic-idle machinery — so the suspended-state bound is maximumDurationInSeconds plus orchestrator termination and the stranded-approval reconciler (see sub-decision 2). A tighter substrate-level suspended-TTL remains available later as an additive idlePolicy change if operators want it. On the fixed maximumDurationInSeconds: no wall-clock task budget exists in the platform (budgets are max_turns / max_budget_usd), so the value is parity with AgentCore’s 8 h cap rather than derived policy; a Blueprint override can be added later if a real need appears.

2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll

Section titled “2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll”

The headline economic win is suspend during HITL approval waits (Cedar approval gates, CEDAR_HITL_GATES.md): while a task waits on a human decision, the MicroVM is suspended (compute charges stop; memory/disk state — cloned repo, warm build caches — is preserved) and resumed when the decision lands. Under Cedar decision #6 the approval window is bounded (default 300 s, ceiling 1 h, timeout → deny), so the saving per gate is bounded at ~1 h of compute — real at 16 vCPU, and it makes any future extension of gate ceilings (the off-hours posture §14.8 deliberately defers) cheap on this backend.

The handshake must respect the existing approval mechanics: the agent discovers decisions itself by polling DynamoDB (_poll_for_decision, monotonic timeout), the approve/deny Lambda writes only the decision rows, and AWAITING_APPROVAL holds the concurrency slot (Cedar decision #7). Nothing “delivers” an approval to the agent, and suspension freezes the agent’s monotonic clock — so the design is:

  • Suspend — orchestrator-owned. The orchestrator’s durable poll observes AWAITING_APPROVAL on a lambda-microvm task and calls suspendSession after a grace period, and only when the gate’s remaining window exceeds grace + resume overhead (suspending a 30 s gate is pure loss). Suspend is a policy decision on a poll observation, not a user action.

  • Resume — inline in the approve/deny Lambdas, orchestrator poll as backstop. After the transactional decision write commits, ApproveTaskFn/DenyTaskFn load the MicroVM handle from the task row’s compute_metadata (persisted at session start — the same field cancel-task.ts reads ECS handles from) via a post-commit strongly-consistent GetItem, then call resumeSession best-effort: on failure they log a warning and write a resume-orphan task event; the decision response never fails on a compute error (the decision row is already durable). The orchestrator poll reconciles: decision row present + MicroVM still SUSPENDED → retry resume (idempotent).

    Why inline rather than poll-only — codebase precedent: resume-on-approve is structurally identical to task cancellation — a user-initiated, latency-sensitive action whose purpose is an immediate compute-lifecycle side effect. cancel-task.ts already resolves this exact tension: the API-plane handler invokes ECS StopTask / AgentCore StopRuntimeSession inline, best-effort (a failed stop logs a warning and the state transition stands; a task_cancel_compute_orphan event is written when no stoppable compute handle exists, reason: missing_runtime_handle) — with the conditional IAM wired in task-api.ts. The resume path goes one step further than the precedent by also writing the orphan event on failed resume calls, because a failed resume strands a suspended VM awaiting a decision — a stronger liveness consequence than a failed stop of an already-cancelled task. The alternative (orchestrator-poll-only resume) preserves single-owner lifecycle purity but pays up to a full poll interval (~30 s) of latency on every approval, and the purity argument was already litigated and declined for cancel. approve-task.ts is deliberately minimal today (security-critical ownership comparison, Cedar finding #6); the resume call is therefore added after the transaction commits, cannot alter the decision outcome, and carries one conditional lambda:ResumeMicrovm grant — the same blast-radius trade the cancel handler accepted in review.

  • Timeout under freeze — the agent re-bases on the wall clock it already owns. The agent’s monotonic gate timer freezes while suspended, so resuming near the deadline is not enough: the frozen timer would still hold its remaining budget and fire the deny minutes after the user-visible window — colliding with the approval row’s TTL (created_at + timeout_s + 120s) and triggering the “row reaped → stranded” fallback on a healthy gate. Instead, the gate expires at min(monotonic budget, created_at + timeout_s), evaluated on each poll iteration and on /resume. This is not a new principle: Cedar decision #6 is already “min wins” for timeouts, the wall-clock deadline is already durable in the approval row the agent itself writes (created_at is in the agent’s own clock domain — no skew), and §13.12’s late-approval race fix already establishes that the durable row is authoritative over the agent’s local timer. Deny authority stays agent-side (the conditional TIMED_OUT write + ConsistentRead re-read race protection is untouched); the orchestrator’s resume at deadline − margin is purely the wake-up mechanism, with no correctness role.

  • Backstops, not mechanisms. maximumDurationInSeconds (mandatory on every RunMicrovm, pinned at 28 800 s — see sub-decision 1) is the substrate kill switch bounding running and suspended time; the orchestrator’s finalization terminate-microvm is the active cleanup path; the stranded-approval reconciler retains its role for orphaned waits. No idlePolicy-based bound is used in any phase — see sub-decision 1’s omit-idlePolicy invariant.

    The active terminate is still mandatory on the SUCCESS path, and P2 sharpened why. P1 concluded flatly that “nothing self-terminates”: a hook-less MicroVM reached RUNNING in 12 s and stayed there with no stateReason through every checkpoint. P2 refuted that for the failure path only — with run: ENABLED, a run hook that answers 4xx makes the service terminate the VM within ~12 s, stateReason: "Run lifecycle hook returned HTTP status 400. Please check your hook endpoint and application logs for more details.", after which suspend-microvm correctly refuses it. That is a real improvement in cost posture and a direct benefit of declaring hooks (see also the failure-path row in the phasing table, sub-decision 3).

    It does not relieve the orchestrator of anything, because the two cases are disjoint. The service reaps a hook result it did not like; it has no view of the guest once the hook returned 200. So a task that starts normally — the overwhelming majority — has no service-side reaper at all, and a VM whose pipeline finished, crashed after /run, or hung is reaped by nobody but TerminateMicrovm. A leaked handle therefore remains a cost incident that bills until the 8 h cap; only the “the guest rejected its own payload” corner now cleans itself up.

  • Concurrency slot stays held during suspend. Cedar decision #7’s rationale (“container alive, consuming memory”) weakens under suspend, and the harder replacement rationale — “AWS counts SUSPENDED MicroVMs toward the account memory quota, so releasing ABCA’s slot would not free real capacity” — is undischarged: the suspended VM stayed in list-microvms at every checkpoint, but that only proves listed. L-CD1C0CC4 (1024 GB, account-scoped) exposes no UsageMetric, AWS/Usage carries only CallCount per API, and no MicroVM memory metric exists in any namespace, so consumption is not observable safely — proving it would need a large concurrent fleet. The conclusion (hold the slot) stands as the conservative choice, not as a verified fact. Size the arithmetic against the 32 GiB peak rather than the 8 GiB baseline: a busy fleet scales up, so peak is what actually competes for the account quota.

The agent’s /suspend hook flushes progress events (durable writes before returning 200, within the 60 s hook budget); /resume reseeds CSPRNGs and refreshes cached credentials.

Amended (P2 review): CSPRNG reseeding is P3 scope, and the P2 exposure is negligible. The original EARS requirement below implied a /run-time reseed had to land with P2; it did not, and shipping P2 with the requirement unmet-but-asserted was itself the defect. Measured exposure, which is what changed the scoping:

  • The only consumer of a non-cryptographic PRNG in agent/src is progress_writer.py’s getrandbits(80), used for the random half of a ULID. os.urandom / secrets are not seeded from the snapshot at all, and nothing in the agent derives a key, token, or nonce from random.
  • That ULID is a DynamoDB sort key under a task_id partition. A collision therefore needs two events in the same task at the same millisecond with the same 80 random bits — and identical PRNG state across two MicroVMs restored from one snapshot does not produce that, because the events are in different partitions. The worst case is a duplicate progress event within one task, not a security boundary.
  • There is no credential exposure from the snapshot’s PRNG state: build-role credentials are kept out of the snapshot structurally (the build hooks make zero AWS calls, so boto3.DEFAULT_SESSION is never populated — see _aws_silent_log), and per-task credentials arrive via platform_config.agent_session_role_arn at /run.

So the reseed moves to P3 alongside /suspend + /resume, where a resumed VM — which really does continue with the exact PRNG state it was frozen with, repeatedly — makes it load-bearing rather than theoretical. P3 must reseed on both /run and /resume, and must not treat the ULID as the only consumer: any future use of random for anything security-relevant needs the reseed in place first.

Normative requirements (EARS):

  • While a lambda-microvm task is in AWAITING_APPROVAL and the gate’s remaining window exceeds the configured grace period plus resume overhead, the orchestrator shall call suspendSession after the grace period.
  • When the approve or deny Lambda commits a decision for a lambda-microvm task, the Lambda shall load the MicroVM handle from compute_metadata and call resumeSession best-effort.
  • If the inline resume fails, then the Lambda shall record a resume-orphan task event and shall still return the decision outcome.
  • While a decision row exists and the MicroVM remains SUSPENDED, the orchestrator shall retry resumeSession.
  • While a lambda-microvm task waits on an approval gate, the agent shall evaluate gate expiry as the earlier of its monotonic budget and the row’s wall-clock deadline (created_at + timeout_s), on each poll iteration and on /resume.
  • If gate expiry is reached without a decision, then the agent shall deny.
  • If no decision arrives by the gate’s wall-clock deadline minus the resume margin, then the orchestrator shall resume the MicroVM so the agent can evaluate expiry and fire the deny agent-side.

3. Packaging: same agent image source, new build path

Section titled “3. Packaging: same agent image source, new build path”

The existing agent container (agent/ Dockerfile, already ARM64) is repackaged as a zip + Dockerfile artifact in S3 and built into a versioned MicrovmImage via CreateMicrovmImage. The agent runs its existing FastAPI server (agent/src/server.py) — the MicroVM path uses the HTTP entrypoint like AgentCore, not ECS’s batch bypass — plus the runtime lifecycle hooks (/run, /suspend, /resume, /terminate) and the /ready + /validate build hooks, all on the same port the server already listens on (8080, declared as the image’s hooks.port). Runtime hooks are fast-notification only (1–60 s): /run validates the payload and starts the pipeline asynchronously, mirroring how the agent loop already runs in a background thread behind /ping on AgentCore.

Hook phasing — corrected: /ready + /run are both P1. The original plan split declaring a hook from serving it, putting /run’s declaration in P1 and its implementation in P2. Live verification proved that split is not a reachable service state, on two independent counts:

  • CreateMicrovmImage rejects an image that enables any lifecycle hook without /ready: “The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled.” So a P1 image declaring only /run is not creatable.
  • With /ready added but unserved, both chipset builds fail: “Ready hook check failed: the application returned a client error (HTTP 4xx) response.” So a declared hook must be served in the same phase.
  • And an image with no hooks at all — the only other creatable shape — cannot receive a payload: “The run hook must be enabled in the MicroVM image to pass the run hook payload.” So deferring hooks entirely also defers the whole payload-delivery channel.

The phasing is therefore:

HookDeclared byServed by the agentNotes
/readyP1 (construct enables hooks.microvmImageHooks.ready)P1MANDATORY, not a quality nicety — see above. A 200 proves uvicorn is bound and server imported cleanly (pulling in pipeline → runner → the policy engine), so a missing policy file fails the BUILD instead of the first task. Since P2-F5 it also WARMS the snapshot — the hook’s 200 is what the service waits for before capturing the snapshot, making this the only place a warm page can be created, and the 225 MiB claude binary was cold in it (see the P2-F5 correction below). A required warm-up failure answers 503, so a snapshot that cannot exec the agent’s own CLI fails the image build instead of every task. Still makes ZERO AWS calls, logging included (a --version exec is neither an AWS call nor a network call).
/runP1 (construct sets hooks.microvmHooks.run)P1The payload-delivery channel. Must be served in P1 because /ready forces hooks to exist at all, and a hook-less image cannot accept runHookPayload. Since P2 it is also the platform-configuration channel (see “Platform configuration delivery” below).
/validateP2 (construct sets hooks.microvmImageHooks.validate)P2An image (build-time) hook, and a shallow self-check only: server alive, hook routes registered, interpreter + contract sanity. It runs under the BUILD role, which deliberately holds no Bedrock / Secrets Manager / DynamoDB grants, so it must make zero AWS API calls and must not touch credential resolution — the “deeper warm-up assertions (Bedrock reachability, Memory access, tool availability)” this ADR originally assigned here are not implementable: every one of them would AccessDenied and fail every image build. They belong to the first task’s own error handling. 200 when the checks pass, 503 while still initialising.
/terminateP2 (construct sets hooks.microvmHooks.terminate)P2Best-effort final flush: a final structured log line, then 200 — always, inside the hook budget, even with nothing running. It must not write terminal task status (the orchestrator finalizes the task and then calls TerminateMicrovm, so a terminate hook that wrote a status would race that finalization and could clobber the real outcome) and must not join the pipeline thread. There is nothing buffered to flush: _ProgressWriter does a synchronous put_item per event, so durability is per-write. “Always 200” also covers the BODY: the handler reads the raw request rather than a typed model, because a typed body is validated before the handler runs and would answer 422 to malformed JSON — a reported hook failure on a successful teardown. Safe to declare because TerminateMicrovm removes the VM with or without in-guest cooperation. Correction (P2-F8): the service sends microvmId: "" on this hook, unlike /run where it is populated, so an empty id is expected-normal and this hook cannot join the guest record to the control-plane one — /run’s accepted line carries that correlation instead.
/suspend, /resumeP3P3Declaring a runtime hook the agent does not serve fails the corresponding lifecycle transition, so each is declared only in the phase that implements it. P1 termination is the orchestrator’s TerminateMicrovm, which needs no in-guest cooperation.

Consequence to state plainly, replacing the original “a P1-built MicroVM image is not runnable end to end”: a P1 image is creatable, launchable and payload-deliverable, but carries no smoke-parity guarantee. P1 delivers the strategy, the construct, the roles/buckets/connectors, the image resource, the packaging script, and the /ready + /run endpoints — so a lambda-microvm task can start a MicroVM and hand it a payload. What P1 has not established is anything P2 owns: AgentCore Memory grants and MEMORY_ID delivery, the agent’s non-secret env parity inside the snapshot, egress specifics from a running MicroVM, and heartbeat/progress behaviour end to end. No clone → change → PR run has happened on this substrate. P2 (“smoke parity”) is the phase that closes that gap. The construct and the packaging script both surface exactly this at synth/run time (abca:microvm-image-p1-smoke-unverified) so an operator cannot mistake a launchable substrate for a verified one.

The AWS::Lambda::MicrovmImage L1 enforces the API’s enums (P2-F2, live 2026-08-06). This closes the one item P1 left explicitly open, and it closes it against the construct’s own stated reasoning. CloudFormation’s generated types make cpuConfigurations[].architecture and all four hooks.* fields plain strings and document no allowed values, from which P1 concluded that the CloudFormation surface takes a hook path while the API takes an ENABLED/DISABLED flag, and that both were correct for their own surface. CloudFormation refused the change set at early validation — the stack was never touched, so there was no rollback and no runtime symptom to trace back — on five values:

/aws/lambda-microvms/runtime/v1/run is not a valid enum value. Supported values: [DISABLED, ENABLED]
(at /Resources/…/Properties/Hooks/MicrovmHooks/Run) … and the same for Terminate, Ready, Validate
arm64 is not a valid enum value. Supported values: [ARM_64]
(at /Resources/…/Properties/CpuConfigurations/0/Architecture)

Three consequences. First, the CloudFormation surface is identical to the API surface, and the packaging script (--cpu-configurations '[{"architecture":"ARM_64"}]', --hooks '{"microvmHooks":{"run":"ENABLED",…}}') had it right all along. Second, the “CDK-managed (recommended)” bootstrap path was non-functional for the whole of P1 and P2 — the out-of-band --create-image script was the only working path — and no unit test, cdk synth or cdk-nag rule could see it, because the types accept any string. Third, hook paths are not configurable on either surface: the service calls fixed well-known routes (proved by the build and run logs, which POST to exactly the /aws/lambda-microvms/runtime/v1/* paths the agent serves), so the route constants in the construct are an agent-side cross-package contract ONLY and must never be sent as property values again. Also discharged in passing: the microvmImageHooks property name and nesting are correct — CloudFormation resolved …/Hooks/MicrovmImageHooks/Ready and objected only to its value.

A snapshot is only as warm as the pages touched before it was captured (P2-F5, live 2026-08-07). This is the defect that stopped the P2 smoke run one step short of a pull request, and it is a property of the substrate rather than a bug in any one file. Every task failed at turn 0, reproducibly:

TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds

The binary was fine — in the identical image, locally, claude --version answers 2.1.191 (Claude Code) in under a second. It is a 225 MiB (236,305,136-byte) statically-linked ELF that nothing had exec’d before the snapshot was taken, so on a guest restored ~50 s earlier the first exec had to fault all of those pages in from lazily-restored storage, and 10 s was not enough. /ready existed precisely so “the snapshot is taken with a warm server”, and the snapshot was warm for uvicorn and stone cold for the binary that does all the work.

Both halves of the fix are kept, because they answer different questions. /ready now exec’s the heavyweight binaries before returning 200 (claude required, git/node best-effort), which is the only mechanism that can make the shipped snapshot warm — and its own budget rises to 300 s, well inside the 3600 s build-hook window, because it now does work whose duration is a cold exec. Two structural rules keep that honest, because per-command timeouts do not compose: the required command runs FIRST with its own budget so no best-effort warm-up can starve the one that decides whether the snapshot is usable, and the best-effort ones then SHARE the remainder of a total warm-up ceiling that sits inside the hook budget with margin (240 s against 300 s). Without them, three commands at 120 s each would be 360 s — a fix for a runtime failure that produces a build failure instead — and a single hung optional command could hold up a 200 that the required warm-up had already earned. Separately, the version probe’s timeout goes from 10 s to 60 s: a probe that exists to print a version string into a log line gains nothing from a tight bound and loses the whole task when it trips. The general rule this generalises to, and the reason it belongs in the ADR rather than only in a comment: on this backend, a first-touch cost that other substrates pay during container start is deferred to the first task instead, so anything large and lazily-loaded is a turn-0 hazard unless it is touched in /ready.

Payload delivery reuses the ECS strategy’s S3-pointer pattern, adapted to runHookPayload (≤ 4 KB — measured, see below): payloads that fit ride inline; the rest are uploaded by the strategy to a platform payload bucket (the ECS payload bucket pattern in ecs-agent-cluster.ts: orchestrator write access, compute-role read-only scoped to the bucket, lifecycle expiry on objects) with the S3 URI in runHookPayload in place of the payload itself — the MicroVM execution role holds the read grant, exactly as the ECS task role does today. Since P2 the hook body also carries platform_config in both branches, so runHookPayload is never only the URI (see the canonical shapes below).

The cap is 4 096 bytes, not the 16 384 the SDK documents. Measured exactly: 4 096 passes, 4 097 is rejected with “Value at ‘runHookPayload’ failed to satisfy constraint: Member must have length less than or equal to 4096”. Two consequences follow. First, the original threshold would have inlined every envelope between 4 097 and 16 384 bytes and had the service reject all of them. Second, and more structurally: the S3-pointer path is now the dominant one, and inline is the exception. A hydrated task payload (prompt + issue thread + repo context) essentially always exceeds 4 KB, so “small payloads ride inline” describes tiny repo-less prompts rather than the common case. The payload bucket is therefore not a rarely-exercised overflow valve but a required part of every normal task, which raises its lifecycle rule (MICROVM_PAYLOAD_TTL_DAYS) and the execution role’s read grant from edge-case plumbing to load-bearing.

Canonical wire shapes. Three, and the producer (lambda-microvm-strategy.ts) emits exactly these:

WhereExact shape
runHookPayload, inline branch{"agent_payload": {…}, "platform_config": {…}}
runHookPayload, pointer branch{"agent_payload_s3_uri": "s3://…", "platform_config": {…}}
the object at that S3 URI{…agent_payload fields…, "platform_config": {…}} — the payload’s own fields at the TOP level, with the config merged in beside them

Two asymmetries are deliberate and must not be “tidied” without changing both sides. First, the S3 object is not the envelope: the payload’s fields sit at the top level (that is what P1 uploaded, before platform_config existed) rather than nested under an agent_payload key. Second, platform_config is duplicated on the pointer path — once beside the pointer, once inside the uploaded object. It costs a few hundred bytes and buys the property that the config is reachable whichever end of the fetch a reader looks at, which matters because it is the agent’s only substitute for an env block.

The agent’s reader is deliberately more permissive than this contract: it also accepts an S3 object shaped like the envelope (agent_payload nested), and platform_config present in only one of the two places (the fetched object wins, the hook body is the fallback). Those are defensive compatibility for the independent deploy cadences of a snapshot image and the orchestrator Lambda — a tolerant reader, not an alternative contract. A producer must emit the three shapes above.

Platform configuration delivery (P2): payload-sourced, allowlisted, fail-closed. The other two backends hand the agent its non-secret platform env at launch — AgentCore Runtime env vars, ECS container overrides — and there is no equivalent on this substrate: a MicroVM starts from a snapshot, so its process environment is whatever was frozen at image build time and is then replayed by every MicroVM launched from that image version. Baking the deployment’s identifiers into the snapshot would make them version-frozen: a redeploy that renames a table, adds a bucket or rotates the session role would leave every existing image version describing a deployment that no longer exists, and the drift would surface as a task-time ResourceNotFound rather than a deploy-time error. So the values travel with the task instead: platform_config is a SIBLING of the payload — beside agent_payload in the inline branch, beside agent_payload_s3_uri in the pointer branch, and merged in beside the payload’s own fields inside the S3 object (the canonical shapes above give each one exactly) — whose snake_case keys the agent installs into os.environ as their UPPER_SNAKE equivalents. A payload value therefore wins over any pre-existing/image value — the orchestrator is describing the live deployment, the snapshot is describing a past one. platform_config carries non-secret identifiers only (table and bucket names, secret ARNs, the session-role ARN); secrets are still fetched at /run time from Secrets Manager using those ARNs, so the snapshot-must-stay-secret-free requirement above is untouched. Per-task fields — memory_id and friends — stay inside agent_payload: platform_config configures the process, agent_payload describes the task.

Two rules make it safe. First, the allowlist fails closed: the agent installs a fixed set of keys and rejects the entire run (HTTP 400, nothing spawned, not one key installed) if the block carries anything else. These values become environment variables of the process that spawns the agent’s tool subprocesses, so an unrecognised key is an attempt to set an arbitrary variable in the agent (AWS_ENDPOINT_URL, LD_PRELOAD, PATH, …) — an injection attempt, not a forward-compatibility gap, which is why unknown keys are refused rather than filtered out. Second, installation happens before any credential or pipeline initialisation on the hook path: the very next step reads GITHUB_TOKEN_SECRET_ARN to resolve the GitHub token and AGENT_SESSION_ROLE_ARN to scope the task’s credentials, so installing later would silently resolve the whole task against the snapshot’s frozen env. The one call that must precede installation is the S3 payload fetch (the config is inside the fetched object), which therefore runs on the ambient compute role via the attributed platform client — and it is the ONLY one: the same rule covers logging, so every /run log line before the install is stdout-only. The CloudWatch writer would otherwise resolve credentials and pin a boto3 default session (region included) off whatever a snapshot happened to bake, which is the build-hook defect one phase later. Nothing is lost — in the intended deployment there is no baked LOG_GROUP_NAME, so those lines would have gone to stdout anyway, and the reason for every pre-install rejection also travels in the structured 4xx/5xx body the service surfaces. A required subset (task table, task-events table, GitHub token secret ARN, session-role ARN) is rejected as …_INCOMPLETE when missing or blank — a distinct wire code from the …_INVALID allowlist rejection, because the remedies differ (deployment wiring vs. producer bug). A /run envelope with no platform_config at all is still accepted, loudly warned: the image snapshot and the orchestrator Lambda deploy on independent cadences, and a new image must not require a same-instant orchestrator. The key set is a cross-package contract in contracts/constants.json (microvm_platform_config), consumed by the agent’s /run hook and produced by the orchestrator, with shape and required-subset invariants enforced by scripts/check-constants-sync.ts.

No orchestrator→agent HTTP path exists in P1–P3: payload arrives through the /run hook, all agent work is outbound, and therefore no JWE auth tokens are minted at all — token minting (and its ≤ 60 min TTL refresh problem) is deferred until a real consumer exists (e.g. operator shell access, #391). The endpoint stays in the SessionHandle because it is genuinely per-session state that becomes load-bearing the day such a consumer appears. But note the service does not agree by default: omitting ingressNetworkConnectors on RunMicrovm attaches a public HTTP_INGRESS connector, so the strategy passes the Lambda-managed NO_INGRESS connector explicitly on every launch (see sub-decision 4’s security table).

Constraint accepted: the configured baseline is 8 GiB RAM / 4 vCPU and the service scales vertically to a 32 GiB / 16 vCPU peak on its own, with 32 GB of disk. So capacity is baseline-priced with 4× burst headroom — good for the bursty compile-and-test shape of an agent task — but the SUSTAINED ceiling is still 32 GiB, so repos that motivated the 120 GB ECS sizing stay on ecs. What the construct configures (and validates) is the baseline; the peak is not something a deployment asks for.

Normative requirements (EARS). Each requirement’s own (Pn) tag is authoritative; there is no blanket phase for the list. The tags are per-requirement because the original list was split P1/P2 on the assumption that a hook could be declared in one phase and served in a later one — which the service does not permit (see the phasing table above), so the hook-serving requirements collapsed into P1 while the P2 items below arrived with the P2 hooks and platform_config:

  • (P1) The image build shall not embed secrets, tokens, or per-task identity in the snapshot.
  • (P1) If the task payload exceeds the 4 KB runHookPayload limit, then the strategy shall upload the payload to the platform payload bucket and pass its S3 URI in runHookPayload in place of the payload (from P2, alongside platform_config).
  • (P1) The MicroVM execution role shall hold read-only access to the payload bucket, scoped to that bucket.
  • (P1) Where no ingress is configured for a deployment, the strategy shall pass the Lambda-managed NO_INGRESS network connector on every RunMicrovm call (the field shall not be omitted).
  • (P1) Where the image enables any MicroVM lifecycle hook, the image shall also enable the /ready hook and the agent shall serve it.
  • (P1) When the /run hook receives the task payload, the agent shall validate it, start the pipeline asynchronously, and return HTTP 200 within the hook budget.
  • (P1) If the /run hook payload cannot be resolved to a task payload, then the agent shall reject the hook with a client error and shall not start a pipeline.
  • (P1) The agent shall not execute the clone→verify→PR pipeline on the hook path.
  • (P1) The agent shall resolve credentials at /run time.
  • (P1) Where a deployment configures a MicroVM image before smoke parity is verified, the platform shall warn that the backend has no smoke-parity guarantee.
  • (P2) Where the image declares the /validate build hook, the agent shall serve it.
  • (P2) The /validate hook shall make no AWS API calls.
  • (P2) When the /run hook receives platform_config, the agent shall install only allowlisted keys into the environment before pipeline initialization.
  • (P2) If platform_config carries a key that is not on the allowlist, then the agent shall reject the run with a 400 and shall install none of the block’s keys.
  • (P2) If a required platform_config key is missing, then the agent shall reject the run with a 400.
  • (P2) Where a platform_config value and an image-baked environment value disagree, the agent shall use the platform_config value.
  • (P2) Until platform_config is installed, the /run hook shall make no AWS API call other than the payload fetch, and shall log to stdout only.
  • (P2) Where the image declares the /terminate hook, the agent shall return 200 within the hook budget for any request body — including a malformed, empty or absent one — and shall not write terminal task status.
  • (P2) When the /ready hook runs, the agent shall exec the agent CLI binary before returning 200, so that its pages are resident when the snapshot is captured.
  • (P2) If a required /ready warm-up does not complete successfully, then the agent shall report not-ready (HTTP 503) rather than allow the snapshot to be taken.
  • (P2) The /ready hook shall make no AWS API call, warm-up included.

4. Infra and IAM: conditional resources behind bootstrap ComputeTypes

Section titled “4. Infra and IAM: conditional resources behind bootstrap ComputeTypes”

Mirroring the ECS pattern: a compute-lambda-microvm bootstrap policy (cdk/src/bootstrap/policies/) gated on the ComputeTypes CFN parameter; a CDK construct provisioning the build role, execution role (admitted to the per-session role via AgentSessionRole.admitComputeRole, which was designed for exactly this), the S3 artifact bucket wiring, and image build automation. Egress uses the platform VPC via egress network connectors so the DNS Firewall / security-group / flow-log stack in COMPUTE.md applies unchanged; ingress is suppressed with the Lambda-managed NO_INGRESS connector (no SHELL_INGRESS — it is noted as a candidate for #391, operator session access, as a separate decision).

Two networking facts the construct has to encode, both established live:

  • A VPC_EGRESS connector requires an operator role. CloudFormation’s generated L1 types operatorRole as optional and this ADR originally assumed Lambda would manage the ENIs with its own service-linked role. It does not: the connector fails to create with “NetworkConnectorOperatorRole is required for VPC_EGRESS connector type”. The construct creates one role — trusting the bare lambda.amazonaws.com service principal (see the trust-policy fact below), carrying AWSLambdaVPCAccessExecutionRole plus the ENI / tag / private-IP actions that policy omits — and shares it across both connectors, since it manages interfaces rather than traffic.

  • The MicroVM-facing roles cannot carry a confused-deputy source condition. All three (build, execution, connector operator) trust the bare lambda.amazonaws.com service principal with no aws:SourceAccount / aws:SourceArn, and that is a forced choice, not an oversight: the Lambda MicroVMs service presents no source key when it assumes them, so a trust policy carrying one is unassumable. Two symptoms of the one cause, both live 2026-08-06/07 and both blocking:

    • Both AWS::Lambda::NetworkConnector resources CREATE_FAILED deterministically (on a freshly deleted stack, so not propagation lag — which matters, because “The service is unable to assume the provided NetworkConnectorOperatorRole. Please verify the trust policy on the role.” is also the classic propagation symptom and a re-run is the obvious wrong guess). Removing the condition → both created within a second.
    • RunMicrovm failed with a misleading iam:PassRole AccessDenied on the caller, with the orchestrator’s grant present, simulate-principal-policy returning allowed, no permissions boundary, and a temporary unconditioned iam:PassRole also denied. The real cause was the execution role’s trust; removing its conditions made the next submission reach RUNNING in 6 s. So the service reports a role it cannot pass-and-assume as an identity-policy denial on the principal passing it.

    Recorded plainly because the fix looks like a regression to anyone applying the standard service-principal pattern — and because it was a regression in the other direction: P1’s standalone-validated operator-role probe had no conditions and worked, and the P1 F2 fix then added them “to mirror the build/execution roles”. sts:TagSession stays: the service needs both actions and it was never implicated.

  • Neither can the iam:PassRole grants carry an iam:PassedToService condition — same root cause, identity side (P2r2-F9 + P2r2-F10, live 2026-08-07 run 2). An earlier revision of this ADR recorded the opposite, that the identity-side condition “was exonerated” by run 1’s elimination. That was a false negative, and its cause is worth recording because it is a general trap: run 1 tested the conditioned grant by adding a temporary unconditioned iam:PassRole and watching the task still fail — but the temporary grant remained attached through the later submissions that succeeded, so the conditioned grant was never once tested against a working trust policy. A contaminated control.

    Run 2 ran the clean experiment — same exact-ARN resource, same ~5-minute IAM settle, one variable. It removed the run-1 workaround first (submission 4: denied) and only then added the unconditioned grant back on the same resource (submission 5: RUNNING), which is the ordering run 1 got wrong:

    Orchestrator iam:PassRole on the execution roleResult
    exact ARN + iam:PassedToService: lambda.amazonaws.comDENIED (two independent submissions)
    exact ARN, no conditionRUNNING in 9 s

    The denial lands on the caller, which is what makes it so misleading — the statement names that exact ARN and simulate-principal-policy answers allowed:

    User: arn:aws:sts::<account>:assumed-role/backgroundagent-dev-TaskOrchestratorOrchestratorFn-…
    is not authorized to perform: iam:PassRole on resource:
    arn:aws:iam::<account>:role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-…
    because no identity-based policy allows the iam:PassRole action

    And the same key blocks the other PassRole path, which run 1 never reached because the enum defect (P2-F2) stopped it earlier: CloudFormation could not pass the build role at CreateMicrovmImage under the bootstrap infrastructure policy’s allowlisted IAMPassRole. Verbatim, so the diagnosis does not have to be taken on trust:

    LambdaMicrovmComputeImage… CREATE_FAILED
    User: arn:aws:sts::<account>:assumed-role/cdk-hnb659fds-cfn-exec-role-<account>-us-east-1/AWSCloudFormation
    is not authorized to perform: iam:PassRole on resource:
    arn:aws:iam::<account>:role/backgroundagent-dev-LambdaMicrovmComputeBuildRoleF0-…
    because no identity-based policy allows the iam:PassRole action
    (Service: LambdaMicrovms, Status Code: 403)

    Three pieces of evidence pin that to the condition rather than to a stale bootstrap or a wrong resource pattern:

    1. the live IaCRole-ABCA-Infrastructure policy was byte-identical to this branch’s cdk/bootstrap/policies/infrastructure.json, so cdk bootstrap --force would have changed nothing;
    2. aws iam simulate-principal-policy --policy-source-arn <cfn-exec-role> --action-names iam:PassRole --resource-arns <build-role-arn> returned allowed with --context-entries ContextKeyName=iam:PassedToService,ContextKeyValues=lambda.amazonaws.com,ContextKeyType=string and implicitDeny with no context entry — so the resource pattern matches and the condition key is the only remaining variable;
    3. the control: the out-of-band create-microvm-image call passed the same build role to the same service successfully, using operator credentials that carry no such condition. The role’s trust is therefore fine and the denial is genuinely caller-side.

    So: the Lambda MicroVMs service presents no usable value for iam:PassedToService on either PassRole path (CloudFormation → build role at CreateMicrovmImage; orchestrator → execution role at RunMicrovm), exactly as it presents no aws:SourceAccount on the assume-role path. One root cause, two more symptoms. Both statements therefore drop the condition, and the fix is deliberately asymmetric so it stays contained:

    • task-orchestrator.ts sid MicrovmPassExecutionRole — condition removed; the exact execution-role ARN is now the whole of the scoping, which is why that resource must never be relaxed to a prefix or *.
    • a new sid MicrovmPassRoles in the conditional compute-lambda-microvm bootstrap policy — unconditioned iam:PassRole on the build- and connector-operator role name prefixes only (not the execution role, which CloudFormation never passes). The shared infrastructure IAMPassRole keeps its allowlist, so no other role in the stack loses that constraint, and an agentcore-only bootstrap never gains an unconditioned pass at all. Operators must re-bootstrap (bundle ≥ 1.6.0) for the CDK-managed image path to work.

    If AWS documents the value the service does present, adding it to both statements restores the condition. microvms.lambda.amazonaws.com, lambda-microvms.amazonaws.com and microvms.amazonaws.com were all implicitDeny against the conditioned policy, so any one of them would serve as the allowlist entry if it turns out to be right. Note that CloudTrail carries no lambda-microvms management events at all today, so the value cannot be read out of a log — only confirmed by AWS or found by a bounded sweep.

    Compensating controls, enumerated per role. The deployment-role’s shared IAMPassRole grant is name-prefix-scoped, so it technically covers all three roles; in practice only the orchestrator actively invokes iam:PassRole on the execution role (CloudFormation never requests it). The other two roles are passed to the deployment role (not to themselves).

    RoleWho can pass it, and how that grant is scoped
    Execution roleThe orchestrator Lambda only, at RunMicrovm — iam:PassRole scoped to this role’s exact ARN, no condition (constructs/task-orchestrator.ts, sid MicrovmPassExecutionRole). The referenced comment contains the authoritative two-arm experiment evidence that the condition is the true blocker (not a permissions gap or stale bootstrap).
    Build roleThe CloudFormation deployment role, at CreateMicrovmImage (the L1’s buildRoleArn) — via the new MicrovmPassRoles statement, scoped to role/backgroundagent-dev-LambdaMicrovmComputeBuild*, no condition. Also whoever runs package-microvm-artifact.sh --create-image out of band, using their own credentials.
    Connector operator roleThe CloudFormation deployment role, at AWS::Lambda::NetworkConnector create/update (operatorRole) — the same statement, scoped to role/backgroundagent-dev-LambdaMicrovmComputeConnector*.

    The rest of the posture: every resource these roles can reach is account-scoped by ARN except two deliberate Resource: '*' statements — ec2:DescribeAvailabilityZones on the execution role (EC2 describe actions have no resource-level scoping; read-only, no mutation, no data access, needed so a CDK target repo’s cdk synth build gate can resolve AZ context on a fresh clone) and the connector operator role’s ENI/tag/private-IP statement (CreateNetworkInterface is authorized before the ENI exists and the Describe* calls take no resource, which is why the AWS-managed VPC-access policy uses * too). Both are justified in the construct’s cdk-nag AwsSolutions-IAM5 suppressions, which is where a reviewer should check them rather than here. The Logs grants are prefix-scoped (/aws/lambda-microvms/* plus one named log group), i.e. wildcards inside a namespace, not *. Separately, the orchestrator’s lambda:PassNetworkConnector is also Resource: '*' and unavoidably so — the AWS-managed connectors live in the aws account, outside any ARN we could enumerate (justified in task-orchestrator.ts, sid MicrovmPassNetworkConnector). Finally: none of the three roles holds iam:*, none has cross-account trust, and the only sts:AssumeRole any of them has is the execution role’s, scoped to the per-task SessionRole.

    If AWS later populates a source key on this path, adding it to the shared principal fixes all three roles and both sts actions at once.

  • Build-time egress needs port 80; runtime does not. agent/Dockerfile installs Debian packages and apt-get fetches over plain HTTP, so a 443-only egress path fails every snapshot build (Could not connect to deb.debian.org:80 … exit code: 100). Rather than widen the runtime posture, the construct provisions a second, build-only connector on the same private subnets with a 443 + 80 security group, referenced solely by the image resource and the packaging script. The agent at run time still has 443-only egress.

  • Where the bootstrap ComputeTypes parameter includes lambda-microvm, the generated template shall attach the IaCRole-ABCA-Compute-LambdaMicrovms policy to the CloudFormation execution role.

  • The orchestrator role shall receive only the MicroVM lifecycle actions it calls (lambda:RunMicrovm, lambda:SuspendMicrovm, lambda:ResumeMicrovm, lambda:TerminateMicrovm, lambda:GetMicrovm for pollSession, and lambda:PassNetworkConnector, which is required even for the default connectors), scoped to platform-created images.

  • Where the lambda-microvm backend is enabled, the approve and deny Lambdas shall receive lambda:ResumeMicrovm and lambda:GetMicrovm — conditionally, mirroring the cancel handler’s conditional RUNTIME_ARN wiring in task-api.ts.

  • The trust policy of every MicroVM-facing role shall name lambda.amazonaws.com and shall carry no source-condition key (the service presents none; see the trust-policy fact above).

  • The iam:PassRole grant the orchestrator uses for the MicroVM execution role shall carry no iam:PassedToService condition and shall be scoped to that role’s exact ARN.

  • Where the bootstrap ComputeTypes parameter includes lambda-microvm, the IaCRole-ABCA-Compute-LambdaMicrovms policy shall grant iam:PassRole without an iam:PassedToService condition, scoped to the MicroVM build- and connector-operator role name prefixes, and shall not extend that grant to the MicroVM execution role.

  • The shared IaCRole-ABCA-Infrastructure iam:PassRole statement shall retain its iam:PassedToService allowlist.

  • The MicroVM execution role shall hold logs:CreateLogStream and logs:PutLogEvents on the application log group whose name is delivered in platform_config, scoped to that log group.

lambda:CreateMicrovmAuthToken is granted to no role in P1–P3 (no JWE consumer exists; see sub-decision 3).

Cost attribution. cdk/src/main.ts currently tags the whole stack with a single compute_type context value (default agentcore) — already imprecise with two backends, wrong with three. P1 must add backend-identifying cost-allocation tags on the MicroVM-specific resources (images, payload/artifact bucket wiring, log groups) and revisit the stack-level tag semantics (e.g. a compute_types list), keeping attribution consistent with #645’s cost/attribution acceptance criterion.

  • Where a deployment enables the lambda-microvm backend, MicroVM-specific resources shall carry backend-identifying cost-allocation tags.
ControlAgentCoreECSLambda MicroVMsDelta
Egress, runtime (DNS Firewall, TCP 443 SG, flow logs)Platform VPCPlatform VPCPlatform VPC via egress network connectorNone
Egress, image buildECR build outside the platform VPCECR build outside the platform VPCPlatform VPC via a separate build-only connector, TCP 443 + 80 (apt-get is plain HTTP)New surface: build-time egress is wider than runtime egress by one port, on a connector no running MicroVM can use
Tenant-data scopingPer-session role (admitComputeRole)Per-session rolePer-session role, execution role admitted identicallyNone
Secrets deliveryRuntime env + Identity injectionTask env varsFetched at /run; never in snapshotNew surface: snapshot must stay secret-free (EARS req., sub-decision 3)
Non-secret platform config (table/bucket names, secret + role ARNs)Runtime env varsTask env varsplatform_config in the /run payload, installed into the process envNew surface: the values are attacker-relevant as env vars (LD_PRELOAD, AWS_ENDPOINT_URL), so the agent installs a fixed allowlist and rejects the whole run on any other key (EARS req., sub-decision 3)
Inbound exposureNone (SigV4 invoke only)None (no endpoint)None — but only because the strategy passes NO_INGRESS explicitly. The service default is a PUBLIC HTTP_INGRESS connector plus a public *.lambda-microvm.<region>.on.aws endpoint; no tokens are minted in P1–P3 either wayNew surface and a new failure mode: “no inbound” is an active control, not an absence. Drop the NO_INGRESS argument and every agent MicroVM gets a public endpoint (EARS req., sub-decision 3)
IAM condition keys on the compute-role trust and on the iam:PassRole grants that hand it overTrust pinned with aws:SourceAccount; PassRole under the allowlisted bootstrap statementTrust pinned per-service; PassRole under the allowlisted bootstrap statementNeither is possible. All three MicroVM-facing roles trust the bare lambda.amazonaws.com with no aws:SourceAccount/aws:SourceArn, and both iam:PassRole grants (orchestrator → execution role at RunMicrovm; CloudFormation → build role at CreateMicrovmImage) carry no iam:PassedToService — the service presents no usable value for any of those keys, and each condition is a hard blocker while present (live-verified, blocking, four times across two runs)Real, evidenced gap that does not close from our side, and it is wider than the trust policy alone. lambda.amazonaws.com is shared with every other Lambda feature, so neither the account pin nor the passed-to-service pin is available on this path. Compensated per role (table in sub-decision 4): the execution role is passable by the orchestrator only (at RunMicrovm), restricted to its exact ARN; the build and connector-operator roles are passable by the CloudFormation deployment role under a new conditional, per-backend, name-prefix-scoped statement (MicrovmPassRoles, bootstrap ≥ 1.6.0) that deliberately excludes the execution role. The shared allowlisted IAMPassRole (role/backgroundagent-dev-*) is left intact to avoid widening the grant for ~30 other roles, so while it technically matches the execution role, only the orchestrator actively reaches for it. Resources are account-scoped by ARN apart from two justified Resource: \'*\' statements (ec2:DescribeAvailabilityZones; the operator role’s ENI management — both carry cdk-nag IAM5 suppressions). No iam:*, no cross-account trust. Revisit if AWS ever documents the values the service presents; CloudTrail records no lambda-microvms events, so they cannot be read from logs
Per-task observability writesRuntime writes to the vended APPLICATION_LOGS groupTask role writes to the task log groupExecution role writes to the SAME APPLICATION_LOGS group, granted against the group platform_config names (P2-F4)None — but only after P2-F4: the name was delivered a phase before the grant, so the agent attempted the write and every per-task line (and METRICS_REPORT) was AccessDenied, degrading silently to guest stdout
Session isolationMicroVMTask-levelMicroVM (Firecracker)None (≥ ECS)
State reuseNoneNoneSnapshot shared across MicroVMsNew surface: CSPRNG reseed + credential refresh on /run//resume — P3 scope (P2 exposure measured as negligible: sole random consumer is a ULID sort key under a task_id partition; os.urandom/secrets unaffected; no credential derives from random). Credential refresh IS in P2: per-task credentials arrive via platform_config at /run
Workload-token injectionYes (Runtime-coupled)No (env-var posture)No (env-var posture)Shared with ECS; deferred to #249/ADR-016
Operator shell accessNoNoNot enabled (SHELL_INGRESS omitted; candidate for #391)None by default
Auth-token mintingn/an/aCreateMicrovmAuthToken granted to no role in any phaseVerified static-only: any principal holding the action can mint a working JWE, including against a SUSPENDED MicroVM, so the posture rests entirely on the grant being absent

Lambda MicroVMs launched in 5 regions (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1) and will expand. ABCA is a single-region deployment, so the constraint is binary per stack: either the stack region supports the backend or the backend does not exist there. Enforcement is layered — one static check where offline determinism is required, live probes everywhere else so the platform self-heals as AWS adds regions (list-managed-microvm-images is the documented read-only availability probe):

StageMechanismCheck
CDK synth/deployStatic region constant (single exported list, documented update path)Synth fails when ComputeTypes includes lambda-microvm in an unlisted region; context-flag escape hatch for newly launched regions ahead of the constant update
Repo onboardingLive probe from the CLIbgagent repo onboard --compute-type lambda-microvm calls list-managed-microvm-images in the stack region and rejects with a remedy (supported-region list + suggest agentcore/ecs)
bgagent platform doctorLive probe (precedent: checkBedrockModel)Reports backend availability for the stack region whenever any active blueprint selects lambda-microvm
Orchestration (defense in depth)Error classificationstartSession failures from a missing regional endpoint classify to a typed remedy in error-classifier.ts, never a cryptic SDK error on the task
  • If an operator onboards a repo with compute_type: 'lambda-microvm' and the availability probe fails for the stack region, then the CLI shall reject the onboarding with the supported-region list and alternative backends as the remedy.
  • If startSession fails because the MicroVM service is unavailable in the stack region, then the orchestrator shall classify the failure with a configuration remedy and shall not retry.
  • When the platform doctor runs in a deployment where any active blueprint selects lambda-microvm, the doctor shall probe MicroVM availability in the stack region and report the result.
  • P1 — strategy + infra + minimal hook serving: LambdaMicrovmComputeStrategy (start/poll/stop), CDK construct, bootstrap policy, types sync, unit + CDK assertion tests, and the agent’s /ready + /run endpoints. No suspend yet. The image IS creatable and launchable and the payload DOES reach the agent — but there is no smoke-parity guarantee (sub-decision 3’s phasing table).
  • P2 — smoke parity: the agent serves /terminate + /validate and installs its platform env from the /run payload (see sub-decision 3’s “Platform configuration delivery”); agent completes clone → change → PR on the backend with progress visible to bgagent watch; failure classification entries in error-classifier.ts; AgentCore Memory parity (IAM grant + MEMORY_ID delivery, following the EcsAgentCluster pattern — Memory is a standalone service already consumed cross-substrate, and omitting the grant silently no-ops cross-session learning); the agent’s remaining non-secret env parity inside the snapshot.
  • P3 — suspend/resume: the interface widening from sub-decision 1 (mandatory methods, all three strategies in one commit), HITL-wait suspend policy, inline resume in the approve/deny Lambda with orchestrator-poll reconciliation (sub-decision 2), timeout-under-freeze wall-clock handling; coordinate with #491’s unified liveness model and update Cedar decision #7’s rationale note.
  • Out of scope: replacing AgentCore as default; classic Lambda functions as a runtime; GPU; the Runtime-coupled workload-access-token injection path (delivery mechanism exists only on AgentCore Runtime; MicroVMs adopt the ECS env-var posture until #249/ADR-016 redesign the seam). Gateway integration is orthogonal: ADR-019/#641 is substrate-portable by design and applies to this backend when it lands.
  • (+) Suspend/resume economics. Tasks idling on approval waits stop billing compute while preserving full state — bounded at ~1 h per gate under the current Cedar ceiling (decision #6), and the enabler for cheap off-hours gate-ceiling extensions later (§14.8). Verified end to end at the substrate level: suspend and resume each complete in ~1 s, and microvmId and endpoint survive the cycle byte-identical, so a stored SessionHandle remains valid.
  • (+) VM-level isolation without cluster ops. Firecracker isolation with no ECS cluster, task definition, or capacity management; one-session-per-MicroVM maps 1:1 onto ABCA’s task model.
  • (+) Escapes AgentCore’s 2 GB image limit and FUSE flock() workaround — native disk in the snapshot supports uv/mise without the split-storage scheme. Note the size comparison must say WHICH measure it means: the same agent tree is 1.799 GB as an OCI image (629.7 MB compressed, i.e. under AgentCore’s limit) but reports codeInstallSizeInBytes of 2.17 GiB as a MicroVM snapshot (i.e. over it). The two straddle the limit and are not interchangeable; memory/disk snapshot sizes are a third thing again and must not be summed into the comparison.
  • (+) Liveness becomes explicit. Unlike AgentCore’s stub pollSession, the strategy can report real substrate state, strengthening the #491 unification.
  • (−) 32 GiB sustained ceiling (8 GiB baseline + automatic 4× vertical scaling), 32 GB disk. Not a successor to the ECS backend for heavy CI-parity builds; the platform now maintains three backends.
  • (−) Capacity is baseline-priced with burst headroom, which is a narrower promise than “32 GiB”. The deployment configures an 8 GiB / 4 vCPU baseline and the service scales to 32 GiB / 16 vCPU on demand — well matched to an agent task, which is idle-ish while waiting on the model and spiky during builds, and cheaper than reserving the peak. But it is burst, not a reservation: a workload that needs 32 GiB sustained is relying on scaling behaviour this ADR has not measured, and against ECS’s 120 GB the gap for sustained-memory workloads is unchanged. So the value proposition remains the suspend economics, the observable control-plane state machine, and the absence of cluster ops — with capacity now a fair-to-good fit rather than a hard blocker. Repos with genuinely sustained heavy builds still belong on ecs.
  • (−) New packaging pipeline. Zip + Dockerfile + service-side image builds with versioned snapshots (storage billed per version) alongside the existing ECR flow; image versions need lifecycle cleanup — including versions left behind by FAILED builds, and noting the last version of an image cannot be deleted individually (delete the image, which reaps it).
  • (−) The payload bucket is on the hot path, not the overflow path. With a 4 KB runHookPayload cap, virtually every real task delivers its payload via S3, so the bucket, its TTL rule and the execution role’s read grant are load-bearing for normal operation rather than an edge case (sub-decision 3).
  • (−) 8-hour hard cap includes suspended time, and with idlePolicy omitted there is no tighter substrate-level suspended-TTL — the suspended-state bound is maximumDurationInSeconds plus orchestrator termination and the stranded reconciler. A manually suspended VM was observed alive at 1 h with no TTL in sight (observation truncated there), so nothing contradicts this bound, but nothing narrows it either. Under today’s 1 h gate ceiling it is comfortably sufficient; any future extension of gate ceilings must revisit the bound (an additive idlePolicy change) and give the orchestrator a checkpoint-and-restart path (push branch, new session) beyond the cap.
  • (!) Idle-policy foot-gun. Traffic-based auto-suspend would freeze a busy outbound-only agent; the decision to disable auto-suspend must be enforced in code and covered by tests, not left to configuration discipline.
  • (!) Service defaults are not the desired posture. Two live-caught cases (public HTTP_INGRESS by default; /ready mandatory) mean an omitted field on this backend does not mean “off” — it can mean “the service picks, and it picks wider than we want”. Every new RunMicrovm / CreateMicrovmImage field should be assumed to have an opinionated default until checked.
  • (!) Nothing self-terminates on the paths that matter — superseding P1’s unqualified version of this bullet. With run: ENABLED the service DOES reap a VM whose run hook returns 4xx (~12 s, stateReason: "Run lifecycle hook returned HTTP status 400.", live-verified), so a guest that rejects its own payload cleans itself up. That is the only self-cleaning case: the service reaps a hook result, and once /run has answered 200 it has no view of the guest. A VM whose task finished, crashed after /run, or hung stays RUNNING and billing until the 8 h cap, so the orchestrator’s TerminateMicrovm on finalize remains the only cleanup for normal operation and a leaked handle is still a cost incident.
  • (!) Snapshot uniqueness. Shared memory snapshots mean every MicroVM restored from one image starts with identical PRNG state. Re-scoped to P3 in the P2 review, with the exposure measured rather than assumed (see the amendment under sub-decision 2): the sole random consumer in agent/src is a ULID sort key under a task_id partition, os.urandom/secrets are unaffected, and no credential or token derives from random — so the P2 exposure is a possible duplicate progress event, not a security boundary. It becomes load-bearing at P3, where a resumed VM continues from frozen state repeatedly; the reseed must land on /run and /resume together with those hooks. Asserting the requirement while leaving it unimplemented was the real defect, and this amendment is the fix.
  • (−) The agent stack template is at 98.6 % of CloudFormation’s 1 MB limit (985,886 bytes) and 486 of 500 resources with a MicroVM image configured — ~14 KB of headroom, i.e. roughly one more construct, and down from 98.4 % / ~16 KB one run earlier. Not caused by this backend (the MicroVM construct is ~6 KB of it) but reached by it, and it will block deploys for reasons that have nothing to do with MicroVMs. Tracked in #735; the candidate remedies are suppressTemplateIndentation and a stack split.
  • (!) Snapshot WARMTH is a first-class property, not an optimisation. A snapshot inherits only the pages something touched before it was captured, so a large lazily-loaded artifact — the 225 MiB claude binary, and anything similar added later — pays its first-touch cost on the first task instead of at container start. That cost failed every task at turn 0 in the P2 smoke run (P2-F5). Anything heavyweight added to the image must be exec’d in /ready, and any timeout guarding a first touch must be sized for a cold page fault rather than for the work itself.
  • (!) Regional availability (5 regions at launch, expanding) — enforced in layers (synth-time static check, onboarding + doctor live probes, orchestration-time classification; see sub-decision 4). The static CDK constant is the one piece that rots as AWS expands; its update path and context-flag escape hatch are deliberate.
  • (!) Workload-token injection delta persists (shared with the ECS backend) until #249/ADR-016 land; document it in the security bar comparison rather than blocking on it. Memory and Gateway are explicitly not deltas — both are standalone services consumed via IAM from any substrate.

P1 (start/poll/stop — no suspend):

  • Unit tests for the strategy: start/poll/stop mapping (including SessionStatus 'suspended' reported mechanically, without task-state interpretation), payload-size branching (inline vs S3 pointer, at the exact 4 096/4 097-byte boundary), the image-identifier-must-be-an-ARN guard, the explicit NO_INGRESS argument (including the blank-env-var fallback, which must never omit the field), error classification (ServiceQuotaExceededException, ThrottlingException, ResourceNotFoundException, regional-unavailability), the omit-idlePolicy invariant, and maximumDurationInSeconds fixed at 28 800.
  • Agent tests: /ready returns 200 once the server is up and starts nothing; /run accepts both envelope shapes (inline and S3 pointer), starts the pipeline asynchronously through the same mapper /invocations uses, returns before the pipeline finishes, and rejects every unusable envelope with a named code before spawning; /suspend and /resume are NOT served (/validate and /terminate joined the served set in P2).
  • Orchestrator tests: substrate-terminal + non-terminal task status → failed classification; suspended + non-AWAITING_APPROVAL status → anomaly event, no fail-fast; compute_metadata persisted with microvmId/endpoint after startSession.
  • CDK assertions: MicroVM resources present only when ComputeTypes includes the backend; synth failure for unsupported regions (plus the context-flag escape hatch); memory size validated against the accepted list at synth; the connector operator role and its trust; two connectors with the build-only one carrying port 80 and the runtime one not; the /ready + /run hook declaration and the absence of the others; IAM actions scoped as specified (orchestrator lifecycle set; no CreateMicrovmAuthToken anywhere); payload-bucket grants (execution role read-only); backend cost-allocation tags; types-sync check covers the widened ComputeType.
  • CLI tests: onboarding rejection with remedy when the availability probe fails; doctor check present when a blueprint selects the backend.
  • P1 verification items (external service facts) — executed 2026-07-31, us-east-1; see docs/verification/645-p1-lambda-microvm-runbook.md for the full evidence. Discharged: runHookPayload limit (4 096, not 16 KB), the accepted baseline memory sizes ([512…8192] MiB — note the developer guide, not the probe, is what establishes that this is a BASELINE with a 32 GiB peak), image-identifier ARN requirement, IAM action names and the observed image-ARN shape, region probe behaviour, manual suspend/resume without idlePolicy, terminate timing and the TERMINATED-persists-≥10-min finding, the default public HTTP_INGRESS, and the /ready requirement. Not discharged: account-quota treatment of SUSPENDED MicroVMs (not observable safely), suspended TTL beyond 1 h (truncated), the vertical-scaling behaviour itself (no workload here approached the baseline, so the 4× peak is documented rather than observed), and the AWS::Lambda::MicrovmImage CloudFormation value shapes (never exercised — the run used the out-of-band script path; discharged, and REFUTED, by the P2 run — see P2-F2 in sub-decision 3). Record the closed answers in COMPUTE.md.

P2 (smoke parity):

  • Agent tests: /validate returns 200 with its individual check results, 503 while initialising, reports a missing hook route / unsupported interpreter, starts nothing, and makes zero AWS calls even with LOG_GROUP_NAME set (asserted by poisoning the boto3 and CloudWatch-writer seams — the same assertion covers /ready); /terminate returns 200 with no body at all, with a malformed / non-object / wrong-content-type / whitespace-only body, when the body read itself fails, with a pipeline still running (without joining it), and when its own best-effort step raises — and never calls task_state.write_terminal; a structural assertion that the route carries no typed body param keeps the 422 from being reintroduced.
  • platform_config tests: the allowlist and required subset are read from contracts/constants.json (the wire key set is additionally asserted literally, as the agent-side tripwire on a contract edit); an unknown key rejects the whole block with nothing installed; a non-object block and a non-string value are rejected; a blank/null optional value is skipped without clobbering an image value while a blank required value is rejected; a payload value beats a pre-existing env value; installation is observed to happen before the GitHub-token resolver runs and before any pipeline thread exists; the config is picked up from the inline envelope, from beside the S3 pointer, and from inside the fetched object (inner wins); an envelope with no platform_config is still accepted with a warning.
  • Snapshot credential hygiene: a subprocess probe asserts that importing server and serving /ready + /validate imports neither boto3 nor botocore, caches no aws_session session, and spawns no CloudWatch writer thread — the property that keeps a build-role credential chain and the build-time region out of the snapshot.
  • /run pre-install silence: with a baked LOG_GROUP_NAME (the hostile case — without it the assertions pass vacuously) every AWS/credential seam (boto3.client/Session, the aws_session factories, _debug_cw/_warn_cw) is armed to raise until the install succeeds. Asserted on the accepted path, on all three rejection paths (bad envelope, platform_config invalid, platform_config incomplete) and on the failed-fetch 500 — where the seams stay armed for the whole request, because a rejected run installed nothing and so earns no AWS call. The permitted exception is asserted POSITIVELY: exactly one client is built pre-install, for s3, through the attributed factory.
  • /ready warm-up tests (P2-F5): the hook exec’s each configured binary exactly once with a generous timeout; claude is the only REQUIRED entry; a timeout, a missing binary, a non-zero exit and an unexpected OSError each produce 503 with the reason logged to stdout rather than a 200 or a 500; a best-effort failure still reports ready; the warm-up makes zero AWS calls with LOG_GROUP_NAME baked. Plus the backstop half: the claude --version probe’s bound is asserted to be ≥ 60 s and to be applied to the exec rather than to the PATH lookup, and a missing CLI warns instead of raising.
  • CDK assertions (P2-F1/F2/F4): no source-condition key on any of the three MicroVM-facing role trusts, and no aws:SourceAccount/aws:SourceArn string anywhere in them; hook properties are ENABLED and the architecture is ARM_64, with a negative assertion that no hook route string appears anywhere in the rendered image resource; the agent hook routes are asserted against their own dedicated constant (the template no longer carries a path to compare); the execution role holds logs:CreateLogStream/PutLogEvents on the application log group and the two logs grants stay separate; the stack wires the SAME log group it delivers as platform_config.log_group_name.
  • Smoke (gated like the ECS backend): clone → change → PR with bgagent watch progress; Memory write parity (no AccessDenied no-op). Run 1 (2026-08-06) FAILED at implement, turn 0 — no PR. Run 2 (2026-08-07) PASSED: two tasks clone → change → commit → push → PR, COMPLETED, 12 turns / $0.279 / 153 s (docs/verification/645-p2-smoke-runbook.md), which also discharged P2-F1, P2-F2, P2-F4, P2-F5 and the dual-signal-liveness item (45 s heartbeat cadence observed across a 181 s RUNNING window). The row is not yet fully closed: run 2 needed one live IAM workaround, and establishing why produced P2r2-F10 (the identity-side iam:PassedToService) and P2r2-F9 (its CloudFormation twin). Both are fixed in source above and neither has been re-exercised live, so what remains is a re-run on a re-bootstrapped account with no workarounds.

P3 (suspend/resume):

  • Unit tests: suspend/resume mapping; agentcore/ecs unsupported stubs; approve/deny inline resume with handle loaded from compute_metadata.
  • HITL lifecycle tests: inline resume failure leaves the decision outcome intact and records the orphan event; orchestrator backstop retries resume; gate expiry fires at min(monotonic budget, created_at + timeout_s) — including the suspend/resume case where the monotonic budget exceeds the wall-clock remainder — without disturbing the §13.12 late-approval race protection.
  • Smoke: suspend/resume across a simulated approval wait preserving workspace state.

All phases: docs sync for COMPUTE.md (new column distinguishing MicroVMs from classic Lambda) and ORCHESTRATOR.md (liveness + suspend lifecycle).

  • Issue #645 — originating RFC proposal
  • Issue #491 — unified liveness decision model (soft dependency, P3)
  • Issue #641 / ADR-019 (PR #663) — substrate-portable tool plane
  • PR #596 — ECS Fargate backend (pattern source for conditional wiring)
  • AWS Lambda MicroVMs — developer guide; Running and using MicroVMs — lifecycle APIs and hooks
  • Agent Toolkit for AWS — aws-lambda-microvms skill — operational constraints (no self-suspend, idle-policy semantics, snapshot uniqueness, size limits)
  • ADR-020 — EARS syntax used for the normative requirements above
  • CEDAR_HITL_GATES.md — approval-gate mechanics (decisions #6, #7) the suspend/resume handshake preserves; cancel-task.ts / task-api.ts — the inline best-effort + reconciler-backstop pattern the resume path mirrors
  • COMPUTE.md, ORCHESTRATOR.md — design docs to be updated by the implementing PRs