AgentCore Evaluations + GitHub Actions: A UK CI Quality Gate for Bedrock Agents
AWS's September 2026 pattern wires Amazon Bedrock AgentCore Evaluations into GitHub Actions so PRs fail when agent quality regresses. Here is the UK-ready auth, threshold and evaluator checklist.

AWS's September 2026 pattern wires Amazon Bedrock AgentCore Evaluations into GitHub Actions so PRs fail when agent quality regresses. Here is the UK-ready auth, threshold and evaluator checklist.
- AgentCore Evaluations scores OpenTelemetry traces (on-demand, online, batch) with built-in and custom judges.
- GitHub Actions + OIDC can deploy, invoke and gate PRs when scores fall below a threshold.
- Solve CI OAuth with stored traces, a service-account user, or M2M client_credentials — pick based on role-testing needs.
- Budget judge calls, use ARM64 builds, wait for runtime READY, and destroy ephemeral stacks every run.
From "it worked in the demo" to a PR quality gate
You can host a solid agent on Amazon Bedrock AgentCore Runtime with MCP tools and OAuth. The harder question for UK delivery teams is: how do you know a prompt, model or tool change made it worse before it reaches production?
AWS's September 2026 pattern — automated agent evaluation with AgentCore Evaluations and GitHub Actions — turns that into a CI quality gate: deploy, invoke, score with LLM-as-judge (and optional code evaluators), block the PR when scores drop.
Cloud architecture
This article is a UK-oriented playbook for platform and MLOps leads adopting AgentCore — not a copy of the AWS sample repo.
What AgentCore Evaluations actually is
AgentCore Evaluations scores agent behaviour from OpenTelemetry traces (the same stream AgentCore Observability already emits to CloudWatch). Modes:
- On-demand — score specific sessions; powers CI gates
- Online — sample production traffic continuously
- Batch — aggregate scoring for baselines and pre/post tests
Built-in evaluators cover goal success, correctness, helpfulness, tool selection/parameter accuracy and trajectory checks. You can add custom LLM judges, Lambda code evaluators, or managed third-party evaluators.
The Evaluate API accepts sessionSpans for a single session (mixing sessions fails validation). Ground truth (expectedResponse, assertions, expectedTrajectory) is optional but powerful for tool-heavy agents.
Why CI needs a deliberate auth story
Production MCP servers often expect user JWTs with roles. GitHub Actions has no interactive user. Three approaches:
| Approach | Idea | Best for |
|---|---|---|
| A — Stored traces | Evaluate committed OpenTelemetry fixtures | Fast first gate; no live OAuth |
| B — Service account | Cached refresh token as a test user | Need to test role enforcement |
| C — M2M | client_credentials scopes; role checks bypassed for CI | Full E2E on PR code (AWS sample default) |
Start with A to prove the gate. Graduate to C when you need the PR's actual runtime behaviour. Use B when role denial paths are in scope.
Reference architecture (UK production-minded)
GitHub Actions (OIDC → IAM role)
│ CDK deploy (dev stack)
▼
Cognito (or Entra ID)
• M2M client for CI
• User client for interactive
│
├── AgentCore Runtime (Strands / LangGraph / …)
└── AgentCore Runtime (MCP server)
│
▼
CloudWatch traces → AgentCore Evaluate → threshold gate
Three-layer MCP auth (Approach C):
- Platform JWT validation (AgentCore custom JWT authorizer)
- Authorization header passthrough into the container
- Middleware that enforces
custom:rolesfor user tokens; M2M (scopes, no roles) gets full tool access for CI only
Keep M2M client secrets out of the browser and out of feature branches. Prefer GitHub OIDC over long-lived AWS keys.
Practical UK CI checklist
- OIDC federation from GitHub to a least-privilege IAM role (CDK, AgentCore, Cognito/Entra, ECR, Bedrock).
- Threshold with margin — LLM-as-judge varies; e.g. target 0.85 reliability, gate at 0.8.
- Start with 4–5 evaluators —
GoalSuccessRate,Correctness,ToolSelectionAccuracy,ToolParameterAccuracy; add trajectory and safety evaluators for customer-facing agents. - ARM64 images — AgentCore expects ARM; GitHub x86 runners need QEMU/Buildx.
- Wait for READY — invoking before runtime readiness yields 424; poll then warm MCP.
- Trace lag — allow 30–90s (retries up to minutes) before evaluation.
- Always destroy ephemeral CDK stacks (
if: always()) to control cost. - Entra ID path — same pattern; swap token/discovery URLs and audience; keep JWT authorizer shape.
Deliberate fail → fix → pass
Break the system prompt so the agent always refuses. The gate should fail goal/correctness even if tool selection sometimes still "passes". Restore a real prompt and confirm green. That rehearsal builds trust with UK change boards faster than a slide deck.
Cost and governance notes for regulated estates
- Rough order: evaluators × prompts per PR (e.g. 4 × 5 = 20 judge calls) — budget Bedrock and set sampling for online mode.
- Log who can raise/lower
EVAL_THRESHOLDper environment (dev 0.7 / staging 0.8 / prod promotion 0.9). - Treat evaluation datasets as controlled assets — no live PII in CI prompts.
- Pair CI gates with online evaluation in production for drift after merge.
Bottom line
AgentCore gives UK teams a managed agent runtime; Evaluations + GitHub Actions give them a merge-time proof that agent quality did not regress. Start with stored-trace gates, then wire M2M end-to-end when the estate is ready.
AIATS helps UK organisations land AgentCore Runtime, MCP gateways and CI quality gates alongside existing AWS and ServiceNow programmes.
A demo agent without a merge-time quality gate is a liability. Start with stored-trace evaluation, then M2M end-to-end — and keep thresholds with margin for LLM-as-judge variance.

