AWSArchitecture

AWS Lambda 90-Minute Timeout on Managed Instances: A UK Playbook for Long-Running Workloads

From 9 Sep 2026, Lambda Managed Instances support up to 90-minute async/ESM timeouts (6x prior 15-minute limit). UK playbook: eu-west-1/2, clean stop, idempotency and when to keep sync at 15 minutes.

AQ
Ali Qaiser
AWS Certified | ServiceNow Architect | Enterprise AI Consultant
16 September 2026
8 min read
AWS Lambda 90-Minute Timeout on Managed Instances: A UK Playbook for Long-Running Workloads
In brief

From 9 Sep 2026, Lambda Managed Instances support up to 90-minute async/ESM timeouts (6x prior 15-minute limit). UK playbook: eu-west-1/2, clean stop, idempotency and when to keep sync at 15 minutes.

Key Takeaways
  • Async and ESM invocations on Lambda Managed Instances can run up to 90 minutes — 6x the prior 15-minute limit (announced ~9 Sep 2026).
  • Synchronous invocations remain capped at 15 minutes.
  • LMI: multi-concurrent requests per instance, specialised compute, EC2 pricing advantages, no self-managed infra.
  • Also applies inside Lambda durable functions; async multi-step durable execution can run up to 1 year overall.
  • On timeout Lambda marks failure but does not forcibly kill code in the shared environment — check remaining time and stop cleanly.
  • Configure via Console, CLI, APIs, IaC or Agent Toolkit; plan UK workloads for eu-west-1 / eu-west-2.

90 minutes on Managed Instances — a UK playbook, not a free pass

On 9 September 2026, AWS announced that AWS Lambda now supports a 90-minute function timeout for asynchronous and event source mapping (ESM) invocations on Lambda Managed Instances (LMI) — a 6× increase from the previous 15-minute limit. For UK platforms running data processing, media transcoding, Monte Carlo-style financial calculations, AI inference and batch jobs, that removes a class of awkward workarounds. It does not remove the need for clean shutdown, idempotency and regional planning in eu-west-1 / eu-west-2.

This is a UK architecture playbook: what changed, what stayed at 15 minutes, how LMI differs from default Lambda, the timeout nuance that can cause duplicate side effects, and a practical day plan before you raise timeouts in production.

Cloud architectureCloud architecture

What AWS announced

FactDetail
New timeoutUp to 90 minutes for async and ESM invocations on Lambda Managed Instances
Previous limit15 minutes (still the ceiling for many patterns — see below)
Use cases called outData processing, media transcoding, financial calculations (e.g. Monte Carlo), AI inference, batch
Synchronous invocationsRemain at the existing 15-minute maximum
ConfigurationConsole, CLI, Lambda APIs, Infrastructure as Code, or the Agent Toolkit for AWS
Regional availabilityAll Regions where Lambda Managed Instances is available — plan UK workloads for eu-west-1 and eu-west-2
Durable functionsThe increased timeout also applies to invocations within Lambda durable functions (checkpoint / replay); when invoked asynchronously, a multi-step durable execution can run for up to 1 year overall

Primary source: AWS What's New — 90-minute function timeout on Managed Instances.

Why the old 15-minute ceiling forced bad shapes

Before this launch, teams often split continuous jobs into:

  • Step Functions with many short Lambdas (orchestration tax, more failure modes).
  • ECS / Batch for anything "a bit long" (operational surface some serverless programmes were trying to avoid).
  • Fragile self-chaining Lambdas that re-invoke themselves with pagination tokens (easy to get wrong under retry).

Those patterns remain valid. The point of the 90-minute ceiling is that some workloads were only fragmented because of the timeout, not because the business process naturally chunked. Media pipelines, Monte Carlo sweeps and certain inference batches are continuous compute with a clear success criterion — they benefit from fewer seams if you invest in clean stop and idempotency.

Lambda Managed Instances in one table

LMI is not "longer Lambda with the same execution model". It is Lambda on managed EC2 capacity in your account:

DimensionLambda (default)Lambda Managed Instances
ConcurrencyOne invocation per execution environmentMulti-concurrent requests per instance / environment
ComputeShared Lambda fleetsSpecialised EC2 instance types (e.g. Graviton4, network-optimised) via capacity providers
Pricing shapePer-request durationEC2-based pricing advantages (Savings Plans / RIs apply to underlying compute; management fee sits on top)
OpsNo instances to manageStill no self-managed infra — Lambda provisions, patches, routes and scales; you define capacity providers / VPC
ScalingScale on cold start / free environments; can scale to zeroScales on utilisation patterns; designed for steadier high-volume traffic
Best fitBursty / scale-to-zeroHigh-volume predictable, performance-critical, or longer continuous jobs

You attach functions to a capacity provider, publish a version, and Lambda launches managed instances (commonly three for AZ resiliency before ACTIVE). Multi-concurrency means thread safety and shared-state discipline matter more than on classic single-concurrency Lambda. A 90-minute timeout on a multi-concurrent environment is a different risk profile than a 15-minute single-concurrency function.

Capacity providers — what UK platform teams own

Treat capacity providers as a platform primitive:

  • VPC, subnets, security groups and placement decisions live with the cloud platform team.
  • Application teams request attachment + timeout raises via change, not ad-hoc console edits.
  • Tag capacity providers with cost centre, data class and environment so FinOps and security reviews stay tractable.
  • Deleting a capacity provider is how you tear down managed instances — document that blast radius before anyone "cleans up" in prod.

The timeout nuance UK teams miss

From Lambda Managed Instances behaviour: when an invocation times out, Lambda marks the invocation failed but does not forcibly kill your code in the shared execution environment. On LMI, that environment may still be serving other concurrent work.

Implication: developers must check remaining time via the invocation context and stop cleanly. If you ignore the deadline:

  • Downstream writes may continue after the caller already saw failure.
  • Retries (async / ESM) can create duplicate side effects.
  • Sibling concurrent invocations on the same environment can observe inconsistent in-memory state.
  • Observability lies to you: the platform says failed; your side effects say succeeded.

Clean-stop checklist

  1. Poll context.getRemainingTimeInMillis() (or language equivalent) on a tight loop for long steps.
  2. Checkpoint durable / idempotent progress before the soft deadline (leave headroom — e.g. stop at 60–90 seconds remaining).
  3. Make every external write idempotent (client tokens, conditional writes, exactly-once outbox).
  4. On soft timeout: flush metrics, release locks, return a controlled failure — do not "just finish the last batch".
  5. Load-test timeout + retry paths, not only the happy 90-minute path.
  6. Never store critical mutable state only in environment memory under multi-concurrency.

Idempotency patterns that survive a 90-minute retry

PatternUse when
Idempotency key on every writeAPI calls, queue publishes, ledger postings
Conditional put / version attributeDynamoDB or similar — retry becomes a no-op
Outbox + pollerYou need exactly-once effect with at-least-once invoke
Durable function checkpointsMulti-step work that should resume, not restart blindly
DLQ + replay toolingPoison messages after repeated timeout failures

If your job cannot be made idempotent, do not raise the timeout and hope — redesign the boundary first.

When to use 90 minutes vs keep 15 (or step up to durable)

PatternGuidance
API / synchronous user pathStill 15 minutes max — do not design UX around LMI async timeouts
ESM / async batch on LMICandidate for up to 90 minutes if the job is continuous and hard to chunk
Multi-step saga over hours/daysPrefer Lambda durable functions (checkpoint/replay); async durable execution can span up to 1 year
Always-on high RPSLMI capacity providers + multi-concurrency; timeout is secondary to scaling and tenancy design
Strict cost isolationModel EC2 + management fee vs default Lambda duration; use Savings Plans where steady
Regulated batch with human approval mid-flightDurable steps with explicit wait / approval — not one 90-minute opaque invoke

UK architecture controls

Regions and residency

  • Default new LMI workloads to eu-west-2 (London) or eu-west-1 (Ireland) per your data-residency and DR policy.
  • Confirm LMI availability in the chosen Region before promising 90-minute SLAs to the business.
  • Document cross-Region retry behaviour — a failed 80-minute job that retries in another Region is a GDPR design issue if personal data moves.
  • Keep CloudWatch logs and artefacts in-Region unless a DPIA already covers replication.

Security and tenancy

  • Capacity providers are the security boundary for LMI; functions run in containers on EC2 Nitro instances in your account.
  • Restrict who can raise timeouts and attach capacity providers (IAM separation: platform vs app teams).
  • VPC placement, egress and Secrets access still need the same CAB packet as any long-running data job.
  • Remember managed instances may be hidden from default EC2 console views — adjust visibility settings so FinOps and security tooling still see billable resources.

FinOps acceptance

Longer timeouts increase the blast radius of a stuck loop. Require:

  • Duration and error-rate alarms before production timeout raises.
  • Cost anomaly detection on the capacity provider's linked instances.
  • A documented kill switch (throttle ESM, disable trigger, or detach provider) owned by on-call.

Observability acceptance tests

  1. Async invocation exceeding configured timeout is marked failed in CloudWatch / metrics within expected latency.
  2. Application logs show a clean stop (checkpoint written) before hard failure.
  3. Retry after timeout does not duplicate side effects (verified with idempotency keys).
  4. Multi-concurrent load on one environment does not corrupt shared mutable state.
  5. Synchronous invoke still rejects timeouts > 15 minutes at configure or invoke time.
  6. Durable function path checkpoints and can replay after a mid-flight timeout.
  7. Kill switch tested: stopping the ESM / trigger halts new work within the agreed SLO.

10-day UK day plan

DayOutcome
1Inventory functions still fragmented solely to dodge the old 15-minute async limit
2Confirm LMI + target Region (eu-west-1 / eu-west-2); capacity provider + VPC design
3Pick one pilot: media, batch, Monte Carlo or inference — define success metrics and cost envelope
4Implement context-based soft stop + idempotent writes; add timeout chaos tests
5Configure timeout via IaC (Console/CLI/API/Agent Toolkit as needed — prefer IaC); peer review IAM
6Dual-run: chunked 15-minute design vs single longer LMI invoke on non-prod data
7Cost model: EC2 + management fee vs default Lambda; Savings Plan impact
8CAB / change: timeout raise, retry policy, DLQ, data-residency note, kill switch owner
9Limited production with alarms on duration, concurrent errors and duplicate-detection metrics
10Decide: standardise LMI 90-minute pattern, or move multi-hour work to durable functions

CAB packet template (copy into your change)

Use this as the minimum evidence pack when raising an LMI timeout in a UK production account:

  1. Function ARN / alias, capacity provider ID, Region (eu-west-1 or eu-west-2).
  2. Invocation mode (async / ESM only — confirm sync paths remain ≤15 minutes).
  3. Business justification for continuous execution vs chunking.
  4. Soft-stop design (context polling, checkpoint store, headroom seconds).
  5. Idempotency proof (keys, conditional writes, dual-run results).
  6. Retry / DLQ policy and expected duplicate-window behaviour.
  7. Data classes processed and residency statement.
  8. Kill switch owner and tested procedure.
  9. Cost envelope (EC2 + management fee vs prior design) and alarm thresholds.
  10. Rollback: previous timeout value in IaC and how to redeploy within one change window.

Without items 4–6, a 90-minute ceiling is an incident waiting for a retry storm.

Strategic takeaway

The 90-minute timeout on Lambda Managed Instances is a real unlock for UK long-running serverless jobs — provided you treat LMI as a different execution model (multi-concurrency, EC2 pricing, clean stop on timeout) and keep synchronous paths on the 15-minute ceiling. Extend duration where the job earns it; invest in idempotency where retries will punish you.

If you want a structured review of Lambda Managed Instances, durable functions and a UK regional playbook for long-running workloads, AIATS offers a Free Evaluation for AWS architecture on UK estates.

Expert Commentary

Ninety minutes without context-based soft stop and idempotency is a retry storm. Treat LMI as a different execution model — multi-concurrency and EC2 pricing — not just a longer classic Lambda.

Topics
AWSAWS LambdaLambda Managed InstancesLMITimeoutEvent Source MappingDurable Functionseu-west-1eu-west-2UKServerlessIdempotency
All insights

Need Help With Your Implementation?

Get expert guidance from our certified ServiceNow and AWS architects.

Schedule a Consultation