AWSArchitecture

AWS MCP Server Serverless Lambda Diagnostics: A UK SRE Playbook

AWS MCP Server (4 Sep 2026) lets coding agents diagnose Lambda plus API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines and fewer tokens. UK playbook: eu-central-1 MCP preference, IAM least privilege, PII in logs and SRE runbook gates.

AQ
Ali Qaiser
AWS Certified | ServiceNow Architect | Enterprise AI Consultant
17 September 2026
10 min read
AWS MCP Server Serverless Lambda Diagnostics: A UK SRE Playbook
In brief

AWS MCP Server (4 Sep 2026) lets coding agents diagnose Lambda plus API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines and fewer tokens. UK playbook: eu-central-1 MCP preference, IAM least privilege, PII in logs and SRE runbook gates.

Key Takeaways
  • On 4 Sep 2026, AWS MCP Server added a serverless capability for coding agents (Claude Code, Kiro, etc.) to diagnose Lambda and connected resources.
  • Inspects API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions alongside the function.
  • Correlates errors vs a 7-day baseline; surfaces recurring errors, deployed config, recent change timelines and cross-resource latency — richer data in a single call (fewer tokens).
  • Start with `aws configure agent-toolkit` or enable AWS MCP Server directly; available via Agent Toolkit or standalone.
  • MCP Server runs in us-east-1 and eu-central-1 (Frankfurt) but can access services in all commercial Regions — prefer Frankfurt for EU/UK estates.
  • No additional cost for the serverless diagnostic capabilities.
  • UK gates: IAM least privilege, change windows, PII in logs, SRE runbook diagnose-only — keep distinct from Lambda Managed Instances 90-minute timeout work.

Lambda diagnostics for coding agents — a UK SRE playbook

On 4 September 2026, AWS announced that the AWS Model Context Protocol (MCP) Server added a serverless capability so coding agents such as Claude Code and Kiro can diagnose issues with AWS Lambda functions and their connected resources. For UK platform and SRE teams, this is not "another AI demo". It is a way to compress incident triage — if IAM, change windows, log PII and EU data paths are designed first.

The AWS MCP Server is available through the Agent Toolkit for AWS or as a standalone installation. It is a managed service that gives AI coding agents secure access to AWS services. The new serverless capability focuses the agent on running Lambda functions and the services they talk to, returning richer correlated data in fewer tokens than a hand-rolled multi-API investigation.

Treat this as adjacent to — but distinct from — the 9 September 2026 Lambda Managed Instances 90-minute timeout announcement. Longer async/ESM timeouts change how long work can run; MCP serverless diagnostics change how agents inspect failures. Do not conflate them in the same CAB ticket without naming both.

Cloud infrastructure and observabilityCloud infrastructure and observability

What AWS shipped (facts for the architecture board)

FactDetail
Announcement4 Sep 2026 — serverless capability on AWS MCP Server
Agents called outClaude Code, Kiro (and other MCP-capable coding agents)
DistributionAgent Toolkit for AWS or standalone AWS MCP Server
Primary targetDiagnose issues with Lambda and connected resources
Connected resources inspectedAPI Gateway, EventBridge, S3, DynamoDB, SNS, SQS, Step Functions
Analytics behavioursCorrelate errors vs 7-day baseline; recurring errors/trends; deployed config; timeline of recent changes; latency across connected resources
Token efficiencyComprehensive data in a single call → fewer tokens than orchestrating many APIs
Start commandaws configure agent-toolkit or enable AWS MCP Server directly
MCP Server regionsRuns in us-east-1 (N. Virginia) and eu-central-1 (Frankfurt)
Service reachCan access services in all commercial AWS Regions
PriceServerless diagnostic capabilities at no additional cost

Primary source: AWS What's New — AWS MCP Server serverless capability.

Why UK estates should care

Lambda incidents rarely live in the function alone. A 5xx from API Gateway, a poison SQS message, a DynamoDB throttle, a Step Functions state failure and a recent IAM policy change often arrive as one pager. Human SRE already knows how to fan out across CloudWatch, X-Ray, Config and deployment history. Coding agents with MCP access can do the same fan-out faster — which raises the governance bar, not lowers it.

StakeholderWhy it lands
SRE / PlatformFaster correlation of Lambda + event/API/data plane symptoms
Security / IAMAgent Toolkit principals need least privilege and break-glass
DPO / PrivacyLogs and payloads may contain PII — agent reads amplify exposure
FinOpsDiagnostics themselves are no extra MCP fee; agent token spend and engineer time still matter
CAB / ChangeAgent-driven inspection during incidents must sit inside approved change/incident process

Capability deep dive — what the agent actually gets

Use this table in runbooks so on-call knows what "ask the agent" can mean:

CapabilityOperational meaningUK gate
Lambda + connected resource inspectPull config and status across API Gateway, EventBridge, S3, DynamoDB, SNS, SQS, Step FunctionsScope which accounts/environments the toolkit may see
7-day error baseline correlationSpot "what changed vs last week" without manual dashboard archaeologyBaseline windows must match your release cadence
Recurring errors / trendsSurface systematic failures, not one-off noiseFeed into problem management, not only chat
Deployed configuration retrievalCompare live config to IaC expectationsPrefer read-only roles in prod
Recent change timelineLink symptoms to deploys / config editsAlign with CodePipeline / Terraform / CDK change tickets
Cross-resource latencyFind slow hops between Lambda and neighboursPair with existing SLOs; do not invent new SLOs from a chat session
Single-call richnessFewer tokens / fewer brittle multi-tool loopsStill log what the agent queried for audit

Getting started — controlled, not cowboy

Enablement path

  1. In a non-production account: run aws configure agent-toolkit or enable the AWS MCP Server directly per AWS user guide.
  2. Wire your coding agent (Claude Code, Kiro, etc.) to the MCP endpoint with explicit server allowlists.
  3. Prove diagnostics against a known broken Lambda (injected fault) and capture the agent's evidence pack.
  4. Only then propose production read-only access with a named IAM role and CloudTrail review.

Regional design for UK / EU estates

ChoiceRecommendation
Where the MCP Server runsPrefer eu-central-1 (Frankfurt) for EU/UK estates so the managed MCP control plane sits in the EU
Where your workloads runUK production often eu-west-2 / eu-west-1; MCP can still access commercial Regions
Data path narrativeDocument that diagnostic queries may be orchestrated via the MCP region while targeting workload Regions
Sovereignty programmesIf your policy forbids us-east-1 dependencies for tooling, do not default the toolkit to N. Virginia

The MCP Server itself runs only in us-east-1 and eu-central-1. That is a residency conversation for the DPO even when Lambda is in London — be precise in the DPIA: workload data plane vs MCP orchestration plane.

IAM least privilege — non-negotiable

Coding agents with broad * on Lambda, logs and data stores are a credentialed insider. Design the Agent Toolkit / MCP principal like a break-glass SRE role:

PrinciplePractice
Separate rolesmcp-diag-readonly-nonprod vs mcp-diag-readonly-prod
Read-firstPrefer Get/List/Describe and CloudWatch Logs filter; deny UpdateFunctionConfiguration in prod by default
Resource tagsConstrain by Environment=nonprod / Team=... where IAM conditions allow
Time-bound elevationProduction write or broader log access via short-lived role assumption with ticket ID
CloudTrailAlert on MCP/toolkit role usage outside incident windows
SecretsNever grant the agent general Secrets Manager / SSM Parameter wildcards "for convenience"

PII, logs and GDPR

Lambda logs and event payloads frequently contain personal data (emails, account IDs, free-text case notes). An agent that "reads everything to help" is still a processing activity.

UK checklist:

  1. Confirm log redaction / data protection policies on the functions in scope.
  2. Exclude known PII-heavy log groups from agent scope until redaction is proven.
  3. Record the MCP diagnostic use case in the AI / processing inventory.
  4. Prefer aggregated metrics and error codes over raw payload dumps in agent prompts.
  5. Retain agent session transcripts per your existing AI tooling retention standard.

Change windows and SRE runbook integration

Fold MCP diagnostics into the existing incident process — do not create a parallel "AI on-call":

PhaseHuman SREAgent + MCP
DetectAlert / SLO burnOptional assist to summarise first signals
TriageAssign severity, bridgeCorrelate Lambda vs 7-day baseline + connected resources
MitigateRollback / throttle / feature flagSuggest candidates; human executes change
RecoverValidate SLOsRe-check latency across neighbours
LearnProblem recordAttach agent timeline + recurring-error notes

CAB language: "Read-only MCP diagnostics via Agent Toolkit in eu-central-1; no automatic remediations; changes remain under existing Lambda change model."

Adjacent context — Lambda Managed Instances 90-minute timeout

On 9 Sep 2026, AWS raised async/ESM timeouts on Lambda Managed Instances to 90 minutes (from 15). That matters for long batch and inference jobs. It does not mean every Lambda can run for 90 minutes, and it does not replace diagnostics.

TopicMCP serverless diagnosticsLMI 90-min timeout
Date4 Sep 20269 Sep 2026
JobInspect / correlate failuresAllow longer async/ESM execution on Managed Instances
UK playbookIAM + EU MCP region + PII + runbooksIdempotency, shutdown, regional LMI availability
Same CAB?Only if both are in scope — list them as two decisions

If yesterday's architecture review already approved longer LMI timeouts, today's review should still ask: who may read those longer-running functions' logs via an agent?

10-day UK adoption plan

Days 1–2 — Policy

  • Name owners: platform SRE, IAM, DPO.
  • Decide MCP region preference: eu-central-1 for EU estates.
  • Draft allowed accounts/OU for Agent Toolkit.

Days 3–5 — Non-prod proof

  • Enable toolkit / MCP; connect one coding agent.
  • Break a sandbox Lambda on purpose; capture diagnostic quality and token use.
  • Review CloudTrail for the toolkit role.

Days 6–8 — Runbook

  • Add MCP triage steps to the Lambda incident runbook.
  • Define "human executes change" rule; ban autonomous remediations in v1.
  • Align with existing observability (CloudWatch, X-Ray, dashboards).

Days 9–10 — Prod read-only proposal

  • CAB paper: scope, IAM, region, PII, cost (no MCP fee; token/ops cost noted).
  • Production role with read-only + tag conditions.
  • Success metric: mean time to first useful correlation on Sev-2 Lambda incidents.

Risks

RiskMitigation
Over-privileged toolkit roleSeparate nonprod/prod; deny config writes by default
us-east-1 default for EU estateExplicitly select eu-central-1 MCP region
PII in agent contextScope log groups; redaction first
Agent "fixes" productionRunbook: diagnose only; humans change
Confused with LMI timeout workSeparate tickets and acceptance criteria
Shadow IT agentsInventory MCP clients; SSO-backed principals only

Closing

AWS MCP Server's serverless capability turns coding agents into faster Lambda diagnosticians across API Gateway, EventBridge, S3, DynamoDB, SNS, SQS and Step Functions — with 7-day baselines, change timelines and cross-resource latency in fewer tokens, at no additional cost for the diagnostic capability. For UK production, the win is real only when Frankfurt MCP preference, least-privilege IAM, PII-aware log scope and SRE runbook discipline ship in the same change.

If you want a structured Agent Toolkit / MCP diagnostics gate for a UK Lambda estate — IAM, eu-central-1 pathing and incident runbook integration — AIATS offers a Free Evaluation: practical, no theatre.

Questions for the architecture / SRE review

  1. Will production MCP diagnostics use eu-central-1 or us-east-1 — and is that written in the DPIA?
  2. What is the exact IAM permission set for the toolkit role in prod?
  3. Which log groups are out of scope for PII reasons?
  4. Who is allowed to connect Claude Code / Kiro to the AWS MCP Server?
  5. Are remediations forbidden until a later CAB, or never via agent?
  6. How do we keep MCP diagnostics separate from the LMI 90-minute timeout programme in reporting?
Expert Commentary

MCP serverless diagnostics compress Lambda triage for Claude Code and Kiro — but only after least-privilege IAM, Frankfurt MCP preference for EU estates, and a diagnose-only runbook. Do not conflate this with the separate LMI 90-minute timeout change.

Topics
AWSAWS MCP ServerLambdaServerlessAgent ToolkitClaude CodeKiroDiagnosticseu-central-1UKSREObservability
All insights

Need Help With Your Implementation?

Get expert guidance from our certified ServiceNow and AWS architects.

Schedule a Consultation