AgentCore Runtime V2 GA: A UK SRE Playbook for Elastic Memory and Cold Starts
AgentCore Runtime V2 GA (18 Sep 2026) reclaims session memory and holds P75 cold starts near 2s from 200 MB–2 GB images, including eu-west-1. A UK SRE canary and CAB checklist before flipping platformVersion to V2.

AgentCore Runtime V2 GA (18 Sep 2026) reclaims session memory and holds P75 cold starts near 2s from 200 MB–2 GB images, including eu-west-1. A UK SRE canary and CAB checklist before flipping platformVersion to V2.
- AgentCore Runtime V2 GA 18 Sep 2026 — set platformVersion to V2 on create/update.
- Elastic memory: small start, on-demand page-in, reclaim during session (not peak-until-end).
- AWS test P75 cold start ~1.9–2.0s for 200 MB–2 GB images vs ~5.4–30s on V1.
- GA Regions include eu-west-1 (Ireland) plus us-east-1/2, us-west-2, ap-northeast-1.
- Snapshot-after-healthy means one-time init should finish before the platform captures state.
- Billing: higher rate on fewer GB-hours; most agents should see lower cost if footprint drops.
- Hide interactive cold start by opening the session when chat loads, before submit.
- Not a Gateway/Identity/Policy replacement; not every Region on day one — verify before multi-Region promises.
What shipped on 18 September 2026
AWS announced general availability of the next-generation Amazon Bedrock AgentCore Runtime — the serverless microVM compute layer for production agents. Set platformVersion to V2 when you create or update a runtime.
Two headline fixes for UK SREs who already run agents in production:
- Elastic memory — sessions start small; memory pages in on demand; unused memory is reclaimed during the session, so you stop paying the peak watermark until session end.
- Consistent cold starts — the platform snapshots a healthy environment once; new instances restore that snapshot. In AWS testing, P75 cold start stayed ~1.9–2.0 seconds for images from 200 MB to 2 GB, versus roughly 5.4–30 seconds on V1 as image size grew.
Regions at GA include eu-west-1 (Ireland) alongside us-east-1, us-east-2, us-west-2, and ap-northeast-1 — so UK latency-sensitive estates can stay in-region without waiting for a later EU wave.
Why UK production agents felt the V1 pain
Agents moved past chat. Coding agents run for minutes; ambient agents stay event-driven for hours. Under V1:
- Memory tracked the high watermark for the whole session — expensive for bursty UK batch and overnight reconciliation agents.
- Cold starts grew with image size and concurrency — worst under the exact burst traffic that opens Monday morning queues.
- Teams compensated with warm pools and home-grown snapshot tricks — undifferentiated heavy lifting.
V2 targets that gap without abandoning serverless, session isolation, scale-to-zero, or consumption billing.
How V2 works (SRE-relevant)
| Mechanism | Operator takeaway |
|---|---|
| Small resident footprint + on-demand page-in | Design agents to release per-request buffers; do not assume peak RAM is "free once allocated" |
| Reclaim when memory goes cold | Long sessions no longer bill like a permanent reservation at the spike |
| Snapshot after healthy | One-time init (model artifacts, static config) should complete before the platform snapshots |
| Compact snapshot | Restore latency stays flat as container image grows — stop treating image diet as the only cold-start lever |
| Billing | Higher rate on fewer GB-hours; for most agents footprint drop outweighs rate rise |
AWS measured cold starts with an empty echo agent (no model, no tools) across Regions over the public internet — so published P75 numbers include cross-Region RTT. Treat them as platform start path, not end-to-end user latency. In that test the agent body itself was ~34 ms at P75; production wall-clock is still dominated by the agent loop and model calls.
UK production gate before flipping platformVersion: V2
- Region check — confirm your AgentCore Runtime Region is on the GA list (
eu-west-1for most UK primary estates). - Baseline V1 metrics — capture P50/P75/P99 time-to-first-token or time-to-ready, session GB-hours, and error rate for 7 days.
- Non-prod dual-run — deploy the same agent image on V1 and V2 under a canary alias; compare cold-start and cost, not just happy-path answers.
- Init hygiene — move heavy one-time loads into startup that completes before healthy; anything lazy that must run every cold start still costs users.
- Interactive UX tip — open the session when the user enters chat (before they submit), so restore finishes while they type.
- Observability — keep your existing AgentCore / CloudWatch dashboards; add a panel for cold-start vs in-session latency so regressions are obvious.
- Rollback — document how to pin
platformVersionback to V1 (or previous) if cost or latency regresses for a specific agent class. - Security unchanged expectation — hardware-enforced session isolation remains; do not weaken Identity / Policy / Gateway gates because compute got cheaper.
What V2 is not
- Not a replacement for AgentCore Gateway, Identity, or Policy — those are still your tool and auth planes (see prior AIATS Gateway and Policy playbooks).
- Not a guarantee that a 30-second agent loop becomes 2 seconds — only the platform start portion is being flattened.
- Not available in every commercial Region on day one — check the GA Region list before promising UK multi-Region active-active.
- Not a license to ship multi-GB images without review — smaller images still help pull times elsewhere in the pipeline.
Coming soon (plan, do not block on)
AWS called out roadmap items: committed baseline discounts for steady sessions, larger compute/storage, x86 microVMs, suspend/resume with memory snapshotting, and scoped identity for unattended agents. Useful for architecture reviews — not a reason to delay a V2 canary if you already pay V1 peak-memory tax in eu-west-1.
14-day UK cutover sketch
Days 1–3: Inventory runtimes; tag interactive vs long-running; pull cost and latency baselines.
Days 4–7: V2 canary on one interactive agent in eu-west-1; compare P75 ready-time and GB-hours.
Days 8–11: Expand to one long-running / ambient agent; validate reclaim behaviour under idle gaps.
Days 12–14: CAB note with before/after; promote remaining agents by class; keep one V1 control agent for a further week.
Ali's take
V2 is the AgentCore runtime UK SREs asked for after the first production winter: pay for the work, not the watermark, and stop gambling cold starts on image size. Flip the version flag behind a canary, keep Gateway and Policy as hard gates, and measure platform start separately from model time — or you will "prove" V2 failed when the LLM was the slow part all along.
V2 fixes the peak-memory bill and image-sized cold-start lottery. Canary in eu-west-1, measure platform ready-time separately from model latency, and keep Gateway/Identity/Policy gates unchanged when you flip platformVersion.


