AWSAI Agents

AgentCore Evaluations + GitHub Actions: A UK CI Quality Gate for Bedrock Agents

AWS's September 2026 pattern wires Amazon Bedrock AgentCore Evaluations into GitHub Actions so PRs fail when agent quality regresses. Here is the UK-ready auth, threshold and evaluator checklist.

AQ
Ali Qaiser
AWS Certified | ServiceNow Architect | Enterprise AI Consultant
10 September 2026
9 min read
AgentCore Evaluations + GitHub Actions: A UK CI Quality Gate for Bedrock Agents
In brief

AWS's September 2026 pattern wires Amazon Bedrock AgentCore Evaluations into GitHub Actions so PRs fail when agent quality regresses. Here is the UK-ready auth, threshold and evaluator checklist.

Key Takeaways
  • AgentCore Evaluations scores OpenTelemetry traces (on-demand, online, batch) with built-in and custom judges.
  • GitHub Actions + OIDC can deploy, invoke and gate PRs when scores fall below a threshold.
  • Solve CI OAuth with stored traces, a service-account user, or M2M client_credentials — pick based on role-testing needs.
  • Budget judge calls, use ARM64 builds, wait for runtime READY, and destroy ephemeral stacks every run.

From "it worked in the demo" to a PR quality gate

You can host a solid agent on Amazon Bedrock AgentCore Runtime with MCP tools and OAuth. The harder question for UK delivery teams is: how do you know a prompt, model or tool change made it worse before it reaches production?

AWS's September 2026 pattern — automated agent evaluation with AgentCore Evaluations and GitHub Actions — turns that into a CI quality gate: deploy, invoke, score with LLM-as-judge (and optional code evaluators), block the PR when scores drop.

Cloud architectureCloud architecture

This article is a UK-oriented playbook for platform and MLOps leads adopting AgentCore — not a copy of the AWS sample repo.

What AgentCore Evaluations actually is

AgentCore Evaluations scores agent behaviour from OpenTelemetry traces (the same stream AgentCore Observability already emits to CloudWatch). Modes:

  • On-demand — score specific sessions; powers CI gates
  • Online — sample production traffic continuously
  • Batch — aggregate scoring for baselines and pre/post tests

Built-in evaluators cover goal success, correctness, helpfulness, tool selection/parameter accuracy and trajectory checks. You can add custom LLM judges, Lambda code evaluators, or managed third-party evaluators.

The Evaluate API accepts sessionSpans for a single session (mixing sessions fails validation). Ground truth (expectedResponse, assertions, expectedTrajectory) is optional but powerful for tool-heavy agents.

Why CI needs a deliberate auth story

Production MCP servers often expect user JWTs with roles. GitHub Actions has no interactive user. Three approaches:

ApproachIdeaBest for
A — Stored tracesEvaluate committed OpenTelemetry fixturesFast first gate; no live OAuth
B — Service accountCached refresh token as a test userNeed to test role enforcement
C — M2Mclient_credentials scopes; role checks bypassed for CIFull E2E on PR code (AWS sample default)

Start with A to prove the gate. Graduate to C when you need the PR's actual runtime behaviour. Use B when role denial paths are in scope.

Reference architecture (UK production-minded)

GitHub Actions (OIDC → IAM role)
        │  CDK deploy (dev stack)
        ▼
  Cognito (or Entra ID)
   • M2M client for CI
   • User client for interactive
        │
        ├── AgentCore Runtime (Strands / LangGraph / …)
        └── AgentCore Runtime (MCP server)
                │
                ▼
        CloudWatch traces → AgentCore Evaluate → threshold gate

Three-layer MCP auth (Approach C):

  1. Platform JWT validation (AgentCore custom JWT authorizer)
  2. Authorization header passthrough into the container
  3. Middleware that enforces custom:roles for user tokens; M2M (scopes, no roles) gets full tool access for CI only

Keep M2M client secrets out of the browser and out of feature branches. Prefer GitHub OIDC over long-lived AWS keys.

Practical UK CI checklist

  1. OIDC federation from GitHub to a least-privilege IAM role (CDK, AgentCore, Cognito/Entra, ECR, Bedrock).
  2. Threshold with margin — LLM-as-judge varies; e.g. target 0.85 reliability, gate at 0.8.
  3. Start with 4–5 evaluators — GoalSuccessRate, Correctness, ToolSelectionAccuracy, ToolParameterAccuracy; add trajectory and safety evaluators for customer-facing agents.
  4. ARM64 images — AgentCore expects ARM; GitHub x86 runners need QEMU/Buildx.
  5. Wait for READY — invoking before runtime readiness yields 424; poll then warm MCP.
  6. Trace lag — allow 30–90s (retries up to minutes) before evaluation.
  7. Always destroy ephemeral CDK stacks (if: always()) to control cost.
  8. Entra ID path — same pattern; swap token/discovery URLs and audience; keep JWT authorizer shape.

Deliberate fail → fix → pass

Break the system prompt so the agent always refuses. The gate should fail goal/correctness even if tool selection sometimes still "passes". Restore a real prompt and confirm green. That rehearsal builds trust with UK change boards faster than a slide deck.

Cost and governance notes for regulated estates

  • Rough order: evaluators × prompts per PR (e.g. 4 × 5 = 20 judge calls) — budget Bedrock and set sampling for online mode.
  • Log who can raise/lower EVAL_THRESHOLD per environment (dev 0.7 / staging 0.8 / prod promotion 0.9).
  • Treat evaluation datasets as controlled assets — no live PII in CI prompts.
  • Pair CI gates with online evaluation in production for drift after merge.

Bottom line

AgentCore gives UK teams a managed agent runtime; Evaluations + GitHub Actions give them a merge-time proof that agent quality did not regress. Start with stored-trace gates, then wire M2M end-to-end when the estate is ready.

AIATS helps UK organisations land AgentCore Runtime, MCP gateways and CI quality gates alongside existing AWS and ServiceNow programmes.

Expert Commentary

A demo agent without a merge-time quality gate is a liability. Start with stored-trace evaluation, then M2M end-to-end — and keep thresholds with margin for LLM-as-judge variance.

Topics
AWSAmazon BedrockAgentCoreEvaluationsGitHub ActionsMCPCI/CDStrandsUKMLOps
All insights

Need Help With Your Implementation?

Get expert guidance from our certified ServiceNow and AWS architects.

Schedule a Consultation