AWS MCP Server Serverless Lambda Diagnostics: A UK SRE Playbook
AWS MCP Server (4 Sep 2026) lets coding agents diagnose Lambda plus API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines and fewer tokens. UK playbook: eu-central-1 MCP preference, IAM least privilege, PII in logs and SRE runbook gates.

AWS MCP Server (4 Sep 2026) lets coding agents diagnose Lambda plus API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines and fewer tokens. UK playbook: eu-central-1 MCP preference, IAM least privilege, PII in logs and SRE runbook gates.
- On 4 Sep 2026, AWS MCP Server added a serverless capability for coding agents (Claude Code, Kiro, etc.) to diagnose Lambda and connected resources.
- Inspects API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions alongside the function.
- Correlates errors vs a 7-day baseline; surfaces recurring errors, deployed config, recent change timelines and cross-resource latency — richer data in a single call (fewer tokens).
- Start with `aws configure agent-toolkit` or enable AWS MCP Server directly; available via Agent Toolkit or standalone.
- MCP Server runs in us-east-1 and eu-central-1 (Frankfurt) but can access services in all commercial Regions — prefer Frankfurt for EU/UK estates.
- No additional cost for the serverless diagnostic capabilities.
- UK gates: IAM least privilege, change windows, PII in logs, SRE runbook diagnose-only — keep distinct from Lambda Managed Instances 90-minute timeout work.
Lambda diagnostics for coding agents — a UK SRE playbook
On 4 September 2026, AWS announced that the AWS Model Context Protocol (MCP) Server added a serverless capability so coding agents such as Claude Code and Kiro can diagnose issues with AWS Lambda functions and their connected resources. For UK platform and SRE teams, this is not "another AI demo". It is a way to compress incident triage — if IAM, change windows, log PII and EU data paths are designed first.
The AWS MCP Server is available through the Agent Toolkit for AWS or as a standalone installation. It is a managed service that gives AI coding agents secure access to AWS services. The new serverless capability focuses the agent on running Lambda functions and the services they talk to, returning richer correlated data in fewer tokens than a hand-rolled multi-API investigation.
Treat this as adjacent to — but distinct from — the 9 September 2026 Lambda Managed Instances 90-minute timeout announcement. Longer async/ESM timeouts change how long work can run; MCP serverless diagnostics change how agents inspect failures. Do not conflate them in the same CAB ticket without naming both.
Cloud infrastructure and observability
What AWS shipped (facts for the architecture board)
| Fact | Detail |
|---|---|
| Announcement | 4 Sep 2026 — serverless capability on AWS MCP Server |
| Agents called out | Claude Code, Kiro (and other MCP-capable coding agents) |
| Distribution | Agent Toolkit for AWS or standalone AWS MCP Server |
| Primary target | Diagnose issues with Lambda and connected resources |
| Connected resources inspected | API Gateway, EventBridge, S3, DynamoDB, SNS, SQS, Step Functions |
| Analytics behaviours | Correlate errors vs 7-day baseline; recurring errors/trends; deployed config; timeline of recent changes; latency across connected resources |
| Token efficiency | Comprehensive data in a single call → fewer tokens than orchestrating many APIs |
| Start command | aws configure agent-toolkit or enable AWS MCP Server directly |
| MCP Server regions | Runs in us-east-1 (N. Virginia) and eu-central-1 (Frankfurt) |
| Service reach | Can access services in all commercial AWS Regions |
| Price | Serverless diagnostic capabilities at no additional cost |
Primary source: AWS What's New — AWS MCP Server serverless capability.
Why UK estates should care
Lambda incidents rarely live in the function alone. A 5xx from API Gateway, a poison SQS message, a DynamoDB throttle, a Step Functions state failure and a recent IAM policy change often arrive as one pager. Human SRE already knows how to fan out across CloudWatch, X-Ray, Config and deployment history. Coding agents with MCP access can do the same fan-out faster — which raises the governance bar, not lowers it.
| Stakeholder | Why it lands |
|---|---|
| SRE / Platform | Faster correlation of Lambda + event/API/data plane symptoms |
| Security / IAM | Agent Toolkit principals need least privilege and break-glass |
| DPO / Privacy | Logs and payloads may contain PII — agent reads amplify exposure |
| FinOps | Diagnostics themselves are no extra MCP fee; agent token spend and engineer time still matter |
| CAB / Change | Agent-driven inspection during incidents must sit inside approved change/incident process |
Capability deep dive — what the agent actually gets
Use this table in runbooks so on-call knows what "ask the agent" can mean:
| Capability | Operational meaning | UK gate |
|---|---|---|
| Lambda + connected resource inspect | Pull config and status across API Gateway, EventBridge, S3, DynamoDB, SNS, SQS, Step Functions | Scope which accounts/environments the toolkit may see |
| 7-day error baseline correlation | Spot "what changed vs last week" without manual dashboard archaeology | Baseline windows must match your release cadence |
| Recurring errors / trends | Surface systematic failures, not one-off noise | Feed into problem management, not only chat |
| Deployed configuration retrieval | Compare live config to IaC expectations | Prefer read-only roles in prod |
| Recent change timeline | Link symptoms to deploys / config edits | Align with CodePipeline / Terraform / CDK change tickets |
| Cross-resource latency | Find slow hops between Lambda and neighbours | Pair with existing SLOs; do not invent new SLOs from a chat session |
| Single-call richness | Fewer tokens / fewer brittle multi-tool loops | Still log what the agent queried for audit |
Getting started — controlled, not cowboy
Enablement path
- In a non-production account: run
aws configure agent-toolkitor enable the AWS MCP Server directly per AWS user guide. - Wire your coding agent (Claude Code, Kiro, etc.) to the MCP endpoint with explicit server allowlists.
- Prove diagnostics against a known broken Lambda (injected fault) and capture the agent's evidence pack.
- Only then propose production read-only access with a named IAM role and CloudTrail review.
Regional design for UK / EU estates
| Choice | Recommendation |
|---|---|
| Where the MCP Server runs | Prefer eu-central-1 (Frankfurt) for EU/UK estates so the managed MCP control plane sits in the EU |
| Where your workloads run | UK production often eu-west-2 / eu-west-1; MCP can still access commercial Regions |
| Data path narrative | Document that diagnostic queries may be orchestrated via the MCP region while targeting workload Regions |
| Sovereignty programmes | If your policy forbids us-east-1 dependencies for tooling, do not default the toolkit to N. Virginia |
The MCP Server itself runs only in us-east-1 and eu-central-1. That is a residency conversation for the DPO even when Lambda is in London — be precise in the DPIA: workload data plane vs MCP orchestration plane.
IAM least privilege — non-negotiable
Coding agents with broad * on Lambda, logs and data stores are a credentialed insider. Design the Agent Toolkit / MCP principal like a break-glass SRE role:
| Principle | Practice |
|---|---|
| Separate roles | mcp-diag-readonly-nonprod vs mcp-diag-readonly-prod |
| Read-first | Prefer Get/List/Describe and CloudWatch Logs filter; deny UpdateFunctionConfiguration in prod by default |
| Resource tags | Constrain by Environment=nonprod / Team=... where IAM conditions allow |
| Time-bound elevation | Production write or broader log access via short-lived role assumption with ticket ID |
| CloudTrail | Alert on MCP/toolkit role usage outside incident windows |
| Secrets | Never grant the agent general Secrets Manager / SSM Parameter wildcards "for convenience" |
PII, logs and GDPR
Lambda logs and event payloads frequently contain personal data (emails, account IDs, free-text case notes). An agent that "reads everything to help" is still a processing activity.
UK checklist:
- Confirm log redaction / data protection policies on the functions in scope.
- Exclude known PII-heavy log groups from agent scope until redaction is proven.
- Record the MCP diagnostic use case in the AI / processing inventory.
- Prefer aggregated metrics and error codes over raw payload dumps in agent prompts.
- Retain agent session transcripts per your existing AI tooling retention standard.
Change windows and SRE runbook integration
Fold MCP diagnostics into the existing incident process — do not create a parallel "AI on-call":
| Phase | Human SRE | Agent + MCP |
|---|---|---|
| Detect | Alert / SLO burn | Optional assist to summarise first signals |
| Triage | Assign severity, bridge | Correlate Lambda vs 7-day baseline + connected resources |
| Mitigate | Rollback / throttle / feature flag | Suggest candidates; human executes change |
| Recover | Validate SLOs | Re-check latency across neighbours |
| Learn | Problem record | Attach agent timeline + recurring-error notes |
CAB language: "Read-only MCP diagnostics via Agent Toolkit in eu-central-1; no automatic remediations; changes remain under existing Lambda change model."
Adjacent context — Lambda Managed Instances 90-minute timeout
On 9 Sep 2026, AWS raised async/ESM timeouts on Lambda Managed Instances to 90 minutes (from 15). That matters for long batch and inference jobs. It does not mean every Lambda can run for 90 minutes, and it does not replace diagnostics.
| Topic | MCP serverless diagnostics | LMI 90-min timeout |
|---|---|---|
| Date | 4 Sep 2026 | 9 Sep 2026 |
| Job | Inspect / correlate failures | Allow longer async/ESM execution on Managed Instances |
| UK playbook | IAM + EU MCP region + PII + runbooks | Idempotency, shutdown, regional LMI availability |
| Same CAB? | Only if both are in scope — list them as two decisions |
If yesterday's architecture review already approved longer LMI timeouts, today's review should still ask: who may read those longer-running functions' logs via an agent?
10-day UK adoption plan
Days 1–2 — Policy
- Name owners: platform SRE, IAM, DPO.
- Decide MCP region preference: eu-central-1 for EU estates.
- Draft allowed accounts/OU for Agent Toolkit.
Days 3–5 — Non-prod proof
- Enable toolkit / MCP; connect one coding agent.
- Break a sandbox Lambda on purpose; capture diagnostic quality and token use.
- Review CloudTrail for the toolkit role.
Days 6–8 — Runbook
- Add MCP triage steps to the Lambda incident runbook.
- Define "human executes change" rule; ban autonomous remediations in v1.
- Align with existing observability (CloudWatch, X-Ray, dashboards).
Days 9–10 — Prod read-only proposal
- CAB paper: scope, IAM, region, PII, cost (no MCP fee; token/ops cost noted).
- Production role with read-only + tag conditions.
- Success metric: mean time to first useful correlation on Sev-2 Lambda incidents.
Risks
| Risk | Mitigation |
|---|---|
| Over-privileged toolkit role | Separate nonprod/prod; deny config writes by default |
| us-east-1 default for EU estate | Explicitly select eu-central-1 MCP region |
| PII in agent context | Scope log groups; redaction first |
| Agent "fixes" production | Runbook: diagnose only; humans change |
| Confused with LMI timeout work | Separate tickets and acceptance criteria |
| Shadow IT agents | Inventory MCP clients; SSO-backed principals only |
Closing
AWS MCP Server's serverless capability turns coding agents into faster Lambda diagnosticians across API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines, change timelines and cross-resource latency in fewer tokens, at no additional cost for the diagnostic capability. For UK production, the win is real only when Frankfurt MCP preference, least-privilege IAM, PII-aware log scope and SRE runbook discipline ship in the same change.
If you want a structured Agent Toolkit / MCP diagnostics gate for a UK Lambda estate — IAM, eu-central-1 pathing and incident runbook integration — AIATS offers a Free Evaluation: practical, no theatre.
Questions for the architecture / SRE review
- Will production MCP diagnostics use eu-central-1 or us-east-1 — and is that written in the DPIA?
- What is the exact IAM permission set for the toolkit role in prod?
- Which log groups are out of scope for PII reasons?
- Who is allowed to connect Claude Code / Kiro to the AWS MCP Server?
- Are remediations forbidden until a later CAB, or never via agent?
- How do we keep MCP diagnostics separate from the LMI 90-minute timeout programme in reporting?
MCP serverless diagnostics compress Lambda triage for Claude Code and Kiro — but only after least-privilege IAM, Frankfurt MCP preference for EU estates, and a diagnose-only runbook. Do not conflate this with the separate LMI 90-minute timeout change.


