ServiceNow AI Control Tower Runtime Observability: A UK AI Steward Playbook
AI Control Tower Aug/Sep 2026 adds runtime observability and LLM-as-a-judge evaluations (including Traceloop for external agents). UK Steward playbook: Zurich P11 / Australia P4+, managed inventory, Trace Collectors vs SDK, Monitor dashboard — not Kill Switch.

AI Control Tower Aug/Sep 2026 adds runtime observability and LLM-as-a-judge evaluations (including Traceloop for external agents). UK Steward playbook: Zurich P11 / Australia P4+, managed inventory, Trace Collectors vs SDK, Monitor dashboard — not Kill Switch.
- AICT Aug/Sep 2026 ships runtime monitoring with LLM-as-a-judge quality and safety evaluations.
- Prerequisites: Zurich Patch 11 or Australia Patch 4+, AI Native Experience, role sn_ai_observe.ai_observability_admin.
- Mark agents managed in inventory before evaluations run; results land on the Monitor dashboard.
- Traceloop (acquired 2026) covers external agents; Trace Collectors poll AWS/Azure/GCP; SDK pushes OTLP for other frameworks.
- Sixteen OOTB metrics; configurable sampling (default ~5%); typically ≤5 minutes to dashboard visibility.
- Do not conflate with Kill Switch containment — observability detects drift; Kill Switch contains compromise.
- Budget assists: vendor FAQ cites 500 assists/month per managed AI asset for evaluations.
- UK gate: patch/plugin/role → managed inventory → ingest path → metrics/weights → scored Monitor proof + DPIA.
Runtime observability is the Steward control — not another dashboard tile
ServiceNow's August & September 2026 AI Control Tower (AICT) wave ships runtime monitoring and evaluations: continuous telemetry from ServiceNow and external agents, scored by LLM-as-a-judge against quality and safety metrics, with results on the Monitor dashboard. Through the Traceloop acquisition (integrated into AICT), Stewards can observe external agents — not only Now-platform agents — via Trace Collectors or the Traceloop SDK.
This playbook is the production gate for observability. It deliberately does not rehash Kill Switch (credential revoke / runtime stop across Okta, GCP and Bedrock) — that is a separate IR control covered in our Kill Switch UK CISO playbook. Observability tells you what is degrading; Kill Switch contains what is compromised.
Enterprise security and observability
What shipped (facts for CAB / AI Steward packs)
| Topic | Fact |
|---|---|
| Product | AI Control Tower runtime observability & evaluations (Aug/Sep 2026 wave) |
| Prerequisites | Instance on Zurich Patch 11 or Australia Patch 4+; AICT licence; AI Native Experience plugin |
| Role | sn_ai_observe.ai_observability_admin (included in sn_ai_governance.ai_steward) |
| Inventory gate | At least one AI agent marked managed in AICT inventory |
| Evaluation | LLM-as-a-judge against 16 OOTB quality/safety metrics (plus performance: latency, tokens) |
| Providers | ServiceNow Autoeval (ServiceNow agents) and Traceloop (external agents) — no separate Traceloop licence |
| Ingest paths | Trace Collectors (MID polls AWS CloudWatch, Azure/Microsoft Foundry, Application Insights, GCP Vertex AI) or Traceloop SDK (push / OTLP for unsupported frameworks) |
| Sampling | Configurable per metric; default ~5% |
| Visibility | Typically ≤5 minutes from trace ingestion to Monitor dashboard |
| UI note | Majority of new AICT features require the new AICT UI |
Why UK estates should care
UK AI Stewards and platform owners are already inventorying agents. Without runtime evaluations:
- Quality drift and tool-path failures show up as user complaints, not Steward tickets.
- External agents (CrewAI, LangChain, hyperscaler runtimes) remain blind spots if you only watch Now Assist.
- DPIA / audit packs need trace-level evidence and judge reasoning — not screenshots of a happy path.
Observability closes that gap: sessions → traces → spans, scored and rolled up to system and portfolio quality/safety scores on Monitor.
UK production gate — numbered steps
1. Patch, plugin and role before any metric toggle
- Confirm instance is on Zurich Patch 11 or Australia Patch 4+.
- Install / verify AI Native Experience for AI Control Tower.
- Assign
sn_ai_observe.ai_observability_admin(or fullsn_ai_governance.ai_steward) to named Stewards — not a shared admin account. - Plan the new AICT UI cutover; most observability tiles live there.
2. Inventory first — mark agents managed
- Populate AI inventory (manual entry, Service Graph connectors, or Shadow AI detection).
- Mark pilot agents managed — evaluations only run on managed assets.
- Prefer connectors that bring the metadata Trace Collectors need (provider, region, resource IDs).
- If you use domain separation, verify Steward visibility and MID scope per domain before go-live.
3. Choose ingest path: Trace Collector vs Traceloop SDK
| Path | Use when | Notes |
|---|---|---|
| Trace Collector | AWS / Azure / GCP / Microsoft Foundry agents | MID Server polls; set collection frequency; unique connection name per account |
| Traceloop SDK | Frameworks not covered by collectors (e.g. CrewAI, LangChain, custom) | Push-based; OpenTelemetry; only when you control the agent codebase |
| ServiceNow agents | Now-platform agents | Traces gathered automatically once managed |
For collectors you need: validated MID Server, cloud credentials, region (AWS), and managed AI asset records. For SDK: valid x-snc-observability-token and a user with sn_ai_observe.ai_data_sender.
4. Select metrics, sampling and weighting
- Open Rules and Templates → Evaluations.
- Enable quality and safety metrics for ServiceNow and external agents separately (lists differ).
- Set sampling above 0% (default 5%); raise for pilots, then tune for cost/assists.
- Optionally edit metric templates so quality/safety weights sum to 100%.
- Remember: assists for evaluation are billed at 500 assists/month per managed AI asset (vendor FAQ) — model that into CAB cost notes.
5. Prove Monitor before you claim production readiness
- Generate real traffic on the pilot agent.
- Confirm sessions/traces/spans land (
sn_ai_observe_ai_session/_trace/_span). - Open Monitor: portfolio scores, ranked systems, top action items, session drill-down with LLM judge reasoning.
- Build a Workflow Studio path from low scores on
sn_ai_observe_ai_trace→ AI task / incident (in-product threshold triggers are limited today). - Document DPIA delta: what prompt/tool content enters evaluation, retention, and who can see traces (Steward role sees broadly — plan ACLs if needed).
What good looks like in CAB evidence
- Patch level + AI Native Experience confirmed.
- Named Steward with observability admin role.
- Managed inventory for pilot agents (internal + at least one external path if in scope).
- Working Trace Collector or Traceloop SDK with a successful scored session on Monitor.
- Sampling and assist-cost note signed by product owner.
- Explicit boundary: observability ≠ Kill Switch; both owners and runbooks linked.
Adjacent AICT items (do not conflate)
- Kill Switch — contain by revoking credentials / stopping runtime (IR).
- AI Gateway MCP pause — pause MCP servers at the gateway boundary.
- Post-runtime OWASP security evaluations — complementary security scoring, not a substitute for quality Monitor.
- Otto / Guided Setup — accelerate Day-0 wiring; they do not replace metric ownership.
Closing
Treat AICT runtime observability as the UK Steward operating system for production agents: inventory → managed → traces → evaluated → Monitor → remediate. Wire Traceloop or Trace Collectors once, rehearse a scored session in non-prod, then promote with assist-cost and DPIA notes in the CAB pack — and keep Kill Switch as a separate, rehearsed containment control.
Observability is the Steward operating loop — inventory, managed, traces, evaluated, remediate. Keep Kill Switch as a separate IR control; do not treat Monitor scores as containment.
