AWS Lambda 90-Minute Timeout on Managed Instances: A UK Playbook for Long-Running Workloads
From 9 Sep 2026, Lambda Managed Instances support up to 90-minute async/ESM timeouts (6x prior 15-minute limit). UK playbook: eu-west-1/2, clean stop, idempotency and when to keep sync at 15 minutes.

From 9 Sep 2026, Lambda Managed Instances support up to 90-minute async/ESM timeouts (6x prior 15-minute limit). UK playbook: eu-west-1/2, clean stop, idempotency and when to keep sync at 15 minutes.
- Async and ESM invocations on Lambda Managed Instances can run up to 90 minutes — 6x the prior 15-minute limit (announced ~9 Sep 2026).
- Synchronous invocations remain capped at 15 minutes.
- LMI: multi-concurrent requests per instance, specialised compute, EC2 pricing advantages, no self-managed infra.
- Also applies inside Lambda durable functions; async multi-step durable execution can run up to 1 year overall.
- On timeout Lambda marks failure but does not forcibly kill code in the shared environment — check remaining time and stop cleanly.
- Configure via Console, CLI, APIs, IaC or Agent Toolkit; plan UK workloads for eu-west-1 / eu-west-2.
90 minutes on Managed Instances — a UK playbook, not a free pass
On 9 September 2026, AWS announced that AWS Lambda now supports a 90-minute function timeout for asynchronous and event source mapping (ESM) invocations on Lambda Managed Instances (LMI) — a 6× increase from the previous 15-minute limit. For UK platforms running data processing, media transcoding, Monte Carlo-style financial calculations, AI inference and batch jobs, that removes a class of awkward workarounds. It does not remove the need for clean shutdown, idempotency and regional planning in eu-west-1 / eu-west-2.
This is a UK architecture playbook: what changed, what stayed at 15 minutes, how LMI differs from default Lambda, the timeout nuance that can cause duplicate side effects, and a practical day plan before you raise timeouts in production.
Cloud architecture
What AWS announced
| Fact | Detail |
|---|---|
| New timeout | Up to 90 minutes for async and ESM invocations on Lambda Managed Instances |
| Previous limit | 15 minutes (still the ceiling for many patterns — see below) |
| Use cases called out | Data processing, media transcoding, financial calculations (e.g. Monte Carlo), AI inference, batch |
| Synchronous invocations | Remain at the existing 15-minute maximum |
| Configuration | Console, CLI, Lambda APIs, Infrastructure as Code, or the Agent Toolkit for AWS |
| Regional availability | All Regions where Lambda Managed Instances is available — plan UK workloads for eu-west-1 and eu-west-2 |
| Durable functions | The increased timeout also applies to invocations within Lambda durable functions (checkpoint / replay); when invoked asynchronously, a multi-step durable execution can run for up to 1 year overall |
Primary source: AWS What's New — 90-minute function timeout on Managed Instances.
Why the old 15-minute ceiling forced bad shapes
Before this launch, teams often split continuous jobs into:
- Step Functions with many short Lambdas (orchestration tax, more failure modes).
- ECS / Batch for anything "a bit long" (operational surface some serverless programmes were trying to avoid).
- Fragile self-chaining Lambdas that re-invoke themselves with pagination tokens (easy to get wrong under retry).
Those patterns remain valid. The point of the 90-minute ceiling is that some workloads were only fragmented because of the timeout, not because the business process naturally chunked. Media pipelines, Monte Carlo sweeps and certain inference batches are continuous compute with a clear success criterion — they benefit from fewer seams if you invest in clean stop and idempotency.
Lambda Managed Instances in one table
LMI is not "longer Lambda with the same execution model". It is Lambda on managed EC2 capacity in your account:
| Dimension | Lambda (default) | Lambda Managed Instances |
|---|---|---|
| Concurrency | One invocation per execution environment | Multi-concurrent requests per instance / environment |
| Compute | Shared Lambda fleets | Specialised EC2 instance types (e.g. Graviton4, network-optimised) via capacity providers |
| Pricing shape | Per-request duration | EC2-based pricing advantages (Savings Plans / RIs apply to underlying compute; management fee sits on top) |
| Ops | No instances to manage | Still no self-managed infra — Lambda provisions, patches, routes and scales; you define capacity providers / VPC |
| Scaling | Scale on cold start / free environments; can scale to zero | Scales on utilisation patterns; designed for steadier high-volume traffic |
| Best fit | Bursty / scale-to-zero | High-volume predictable, performance-critical, or longer continuous jobs |
You attach functions to a capacity provider, publish a version, and Lambda launches managed instances (commonly three for AZ resiliency before ACTIVE). Multi-concurrency means thread safety and shared-state discipline matter more than on classic single-concurrency Lambda. A 90-minute timeout on a multi-concurrent environment is a different risk profile than a 15-minute single-concurrency function.
Capacity providers — what UK platform teams own
Treat capacity providers as a platform primitive:
- VPC, subnets, security groups and placement decisions live with the cloud platform team.
- Application teams request attachment + timeout raises via change, not ad-hoc console edits.
- Tag capacity providers with cost centre, data class and environment so FinOps and security reviews stay tractable.
- Deleting a capacity provider is how you tear down managed instances — document that blast radius before anyone "cleans up" in prod.
The timeout nuance UK teams miss
From Lambda Managed Instances behaviour: when an invocation times out, Lambda marks the invocation failed but does not forcibly kill your code in the shared execution environment. On LMI, that environment may still be serving other concurrent work.
Implication: developers must check remaining time via the invocation context and stop cleanly. If you ignore the deadline:
- Downstream writes may continue after the caller already saw failure.
- Retries (async / ESM) can create duplicate side effects.
- Sibling concurrent invocations on the same environment can observe inconsistent in-memory state.
- Observability lies to you: the platform says failed; your side effects say succeeded.
Clean-stop checklist
- Poll
context.getRemainingTimeInMillis()(or language equivalent) on a tight loop for long steps. - Checkpoint durable / idempotent progress before the soft deadline (leave headroom — e.g. stop at 60–90 seconds remaining).
- Make every external write idempotent (client tokens, conditional writes, exactly-once outbox).
- On soft timeout: flush metrics, release locks, return a controlled failure — do not "just finish the last batch".
- Load-test timeout + retry paths, not only the happy 90-minute path.
- Never store critical mutable state only in environment memory under multi-concurrency.
Idempotency patterns that survive a 90-minute retry
| Pattern | Use when |
|---|---|
| Idempotency key on every write | API calls, queue publishes, ledger postings |
| Conditional put / version attribute | DynamoDB or similar — retry becomes a no-op |
| Outbox + poller | You need exactly-once effect with at-least-once invoke |
| Durable function checkpoints | Multi-step work that should resume, not restart blindly |
| DLQ + replay tooling | Poison messages after repeated timeout failures |
If your job cannot be made idempotent, do not raise the timeout and hope — redesign the boundary first.
When to use 90 minutes vs keep 15 (or step up to durable)
| Pattern | Guidance |
|---|---|
| API / synchronous user path | Still 15 minutes max — do not design UX around LMI async timeouts |
| ESM / async batch on LMI | Candidate for up to 90 minutes if the job is continuous and hard to chunk |
| Multi-step saga over hours/days | Prefer Lambda durable functions (checkpoint/replay); async durable execution can span up to 1 year |
| Always-on high RPS | LMI capacity providers + multi-concurrency; timeout is secondary to scaling and tenancy design |
| Strict cost isolation | Model EC2 + management fee vs default Lambda duration; use Savings Plans where steady |
| Regulated batch with human approval mid-flight | Durable steps with explicit wait / approval — not one 90-minute opaque invoke |
UK architecture controls
Regions and residency
- Default new LMI workloads to eu-west-2 (London) or eu-west-1 (Ireland) per your data-residency and DR policy.
- Confirm LMI availability in the chosen Region before promising 90-minute SLAs to the business.
- Document cross-Region retry behaviour — a failed 80-minute job that retries in another Region is a GDPR design issue if personal data moves.
- Keep CloudWatch logs and artefacts in-Region unless a DPIA already covers replication.
Security and tenancy
- Capacity providers are the security boundary for LMI; functions run in containers on EC2 Nitro instances in your account.
- Restrict who can raise timeouts and attach capacity providers (IAM separation: platform vs app teams).
- VPC placement, egress and Secrets access still need the same CAB packet as any long-running data job.
- Remember managed instances may be hidden from default EC2 console views — adjust visibility settings so FinOps and security tooling still see billable resources.
FinOps acceptance
Longer timeouts increase the blast radius of a stuck loop. Require:
- Duration and error-rate alarms before production timeout raises.
- Cost anomaly detection on the capacity provider's linked instances.
- A documented kill switch (throttle ESM, disable trigger, or detach provider) owned by on-call.
Observability acceptance tests
- Async invocation exceeding configured timeout is marked failed in CloudWatch / metrics within expected latency.
- Application logs show a clean stop (checkpoint written) before hard failure.
- Retry after timeout does not duplicate side effects (verified with idempotency keys).
- Multi-concurrent load on one environment does not corrupt shared mutable state.
- Synchronous invoke still rejects timeouts > 15 minutes at configure or invoke time.
- Durable function path checkpoints and can replay after a mid-flight timeout.
- Kill switch tested: stopping the ESM / trigger halts new work within the agreed SLO.
10-day UK day plan
| Day | Outcome |
|---|---|
| 1 | Inventory functions still fragmented solely to dodge the old 15-minute async limit |
| 2 | Confirm LMI + target Region (eu-west-1 / eu-west-2); capacity provider + VPC design |
| 3 | Pick one pilot: media, batch, Monte Carlo or inference — define success metrics and cost envelope |
| 4 | Implement context-based soft stop + idempotent writes; add timeout chaos tests |
| 5 | Configure timeout via IaC (Console/CLI/API/Agent Toolkit as needed — prefer IaC); peer review IAM |
| 6 | Dual-run: chunked 15-minute design vs single longer LMI invoke on non-prod data |
| 7 | Cost model: EC2 + management fee vs default Lambda; Savings Plan impact |
| 8 | CAB / change: timeout raise, retry policy, DLQ, data-residency note, kill switch owner |
| 9 | Limited production with alarms on duration, concurrent errors and duplicate-detection metrics |
| 10 | Decide: standardise LMI 90-minute pattern, or move multi-hour work to durable functions |
CAB packet template (copy into your change)
Use this as the minimum evidence pack when raising an LMI timeout in a UK production account:
- Function ARN / alias, capacity provider ID, Region (eu-west-1 or eu-west-2).
- Invocation mode (async / ESM only — confirm sync paths remain ≤15 minutes).
- Business justification for continuous execution vs chunking.
- Soft-stop design (context polling, checkpoint store, headroom seconds).
- Idempotency proof (keys, conditional writes, dual-run results).
- Retry / DLQ policy and expected duplicate-window behaviour.
- Data classes processed and residency statement.
- Kill switch owner and tested procedure.
- Cost envelope (EC2 + management fee vs prior design) and alarm thresholds.
- Rollback: previous timeout value in IaC and how to redeploy within one change window.
Without items 4–6, a 90-minute ceiling is an incident waiting for a retry storm.
Strategic takeaway
The 90-minute timeout on Lambda Managed Instances is a real unlock for UK long-running serverless jobs — provided you treat LMI as a different execution model (multi-concurrency, EC2 pricing, clean stop on timeout) and keep synchronous paths on the 15-minute ceiling. Extend duration where the job earns it; invest in idempotency where retries will punish you.
If you want a structured review of Lambda Managed Instances, durable functions and a UK regional playbook for long-running workloads, AIATS offers a Free Evaluation for AWS architecture on UK estates.
Ninety minutes without context-based soft stop and idempotency is a retry storm. Treat LMI as a different execution model — multi-concurrency and EC2 pricing — not just a longer classic Lambda.

