← haiec.com 0 / 22 read
TM Forum Innovate Americas 2026 · Dallas, TX

The Forensics Challenge, Explained

This handbook is your end-to-end guide to the TM Forum Trustworthy AI and Data Hackathon, the AI forensics concepts behind it, and how to run HAIEC to win. Telecom edition.

MunichTech EXPO Grand Challenge Award Winner — Autumn 2026
HAIEC is a MunichTech EXPO Grand Challenge Award winner (Devpost, Autumn Hackathon 2026) — for AI Action Path Assurance. This handbook is published by HAIEC for the TM Forum challenge context; MunichTech EXPO is not a co-organizer of TM Forum events. Award announcement · Official verification
🎯
The challenge in one sentence: Teams enter a simulated autonomous telecom enterprise and must reconstruct what autonomous agents did, why they did it, and whether their actions can be independently verified — using evidence, not assertions.
🏆
How you win: The winning team provides the strongest, most independently verifiable evidence trail. Not the most sophisticated AI. Not the flashiest UI. Evidence.

What this handbook covers

🎯 The Challenge
Event details, AL0→AL1→AL2, what you receive, how you're judged.
🧠 AI Forensics 101
How AI forensics differs from traditional forensics. What to look for.
📡 Telecom 101
O-RAN, 5G core, network slicing, telecom terminology.
🏛️ Five Planes
The core authority model. Gaps between planes are findings.
📋 Telecom Entities
17 typed entities from CellSite to PacketSession.
🔬 Telecom Forensics
3 worked examples: rApp manipulation, slice isolation, tower forensics.
⚙️ HAIEC Platform
138 deterministic rules across 30 declared families, Compliance Twin, Kill Switch.
⚡ Quick Reference
Printable card. Hard locks. Event checklist.
⚠
Honesty note: Items marked INFERRED are based on TM Forum hackathon precedent and public event coverage. Confirm against the official brief.
Part 1

The TM Forum Forensic Challenge

What it is, where it happens, what you receive, and how you're judged.

Event identity

FieldValue
EventTM Forum Innovate Americas 2026 CONFIRMED
Challenge trackTrustworthy AI and Data Hackathon CONFIRMED
Challenge title"Agentic Assurance: The Quest for Proof" CONFIRMED
LocationDallas, Texas · AT&T Headquarters CONFIRMED
DatesOctober 4–7, 2026 CONFIRMED
SponsorDell Technologies CONFIRMED
Supporting partnersAccenture, AWS, NVIDIA, ServiceNow CONFIRMED
Team size3–7 members CONFIRMED
FormatIn-person, hybrid, or virtual CONFIRMED
🚫
Do not confuse with: The "AI Telco Troubleshooting Challenge" (a separate GSMA / ETSI / IEEE / ITU initiative focused on LLM root-cause analysis for network faults). This challenge is about forensic reconstruction and evidence assurance, not fault diagnosis.

What you receive

Organizers provide a simulated telecom-oriented environment containing:

🤖 Pre-built agents
Autonomous agents already operating in the simulated enterprise.
🔌 APIs
TM Forum Open APIs and ODA components available for observation.
🏗️ ODA components
Open Digital Architecture building blocks.
📡 Shared services
Identity, logging, and telemetry infrastructure.
👥 Customer-to-network workflow
End-to-end operational scenario you observe.
🚦 Alert-level scenarios
Three escalating conditions: AL0, AL1, AL2.

The three alert levels

This is the spine of the competition. Everything you do maps to one of these three stages.

AL0Baseline — Functioning
System operating normally. No injected threat. This is your ground truth.
Your tasks: Map what "normal" looks like. Record healthy telemetry patterns. Establish the baseline for every entity, identity, and capability. Confirm nothing anomalous.
AL1Degraded — Drift Introduced
Baseline + reliability or decision-quality degradation. Something subtle changed.
Your tasks: Detect drift against AL0 baseline. Identify what changed. Distinguish "slightly different" from "actually wrong." Apply guardrails. Document the behavioral delta.
AL2Compromised — Malicious + Tampering
Deliberately malicious behavior, including audit-trail tampering. This is the forensic moment.
Your tasks: Reconstruct the attack. Prove tampering using independent sources. Show which evidence survives, which is contradicted, and where proof stops. Surface the First Deterministic Divergence.

How you win

🏆
The winning team provides the strongest, most independently verifiable evidence trail. Not the most sophisticated AI. Not the flashiest UI. Evidence.

Judging criteria (confirmed from public event coverage)

CriterionWhat judges look for
Evidence strength & independenceIndependence, corroboration, traceability, tamper-evidence. Multiple sources agree. Chain of custody intact.
Impact & innovationReal-world business case. Why this matters to telecom. Regulatory alignment (EU AI Act, ISO 42001).
Challenge fit & feasibilityRelevance to TM Forum ODA, implementable at scale. Cross-functional expertise.
Technical executionHow well the platform and tech stack were used. Reproducibility.
📊
Industry context: 72% of communication service providers believe their AI systems are trustworthy, but only 14% can provide concrete evidence to support that assessment. This challenge exists to close that gap.
Part 2

How AI Forensics Differs from Traditional Forensics

This is the single biggest conceptual unlock for the challenge. Traditional forensics assumes deterministic systems. AI forensics cannot make that assumption.

⚠
The core problem: Traditional digital forensics assumes systems are deterministic — same input, same output. LLMs violate this assumption. The same prompt to the same model may produce different outputs due to temperature sampling and random seeds. Model behavior changes with context. And the attack surface is natural language itself, which does not leave the same forensic artifacts as a SQL injection or buffer overflow.

Side-by-side comparison

AspectTraditional ForensicsAI / LLM Forensics
ReproducibilityHigh — same inputs produce same outputsLow — probabilistic outputs vary
Evidence typesFiles, logs, memory dumps, network capturesPrompts, completions, embeddings, model weights, tool-call logs
Attack indicatorsMalformed inputs, exploit patterns, malware signaturesSemantic manipulation, context injection, behavioral anomalies
Root cause analysisTrace execution path through codeAnalyze model reasoning through prompt-response chains
Chain of custodyWell-established proceduresEmerging practices, model state hard to preserve
DeterminismYesNo — must rely on evidence correlation, not replay

What makes AI forensics possible

Because you can't replay a probabilistic model, you must reconstruct from artifacts. These are the artifacts that matter:

1 Model input/output logs ▶

Every prompt sent and every response generated, with timestamps, session identifiers, user attribution, and full conversation context.

Without these logs, forensic investigation of an LLM incident is essentially impossible. Check whether logging captures the complete input — including system prompt, conversation history, and retrieved context — not just the latest user message.

2 Tool call logs ▶

Every tool call logged with: tool name, full arguments, return values, timestamps, and the model's reasoning for making the call.

Tool calls are where LLM incidents cross from harmful text into harmful actions. They often reveal the attacker's true objective.

3 System prompt & configuration ▶

Captured as they existed at the time of the incident. Version history if managed through config management. Deployed version if embedded in code.

Configuration changes around the time of the incident may be either the cause or a symptom.

4 RAG retrieval logs ▶

Which documents were retrieved, their similarity scores, and their content.

If the attack involved RAG poisoning, the retrieved documents are the attack vector and constitute primary evidence.

5 Model artifacts ▶

Model weights, adapter weights for fine-tuned models, tokenizer configuration, and any custom post-processing code.

For incidents involving model tampering or supply chain compromise, these artifacts need to be preserved and analyzed.

Part 2 · Domain knowledge

Telecom 101 for Forensic Investigators

You don't need to be a telecom engineer to compete, but you do need to understand the architecture and terminology. This section gets you fluent in 15 minutes.

The telecom architecture stack

1
Radio Access Network (RAN)
The cellular towers and radios that connect user devices (UE) to the network. Includes cells, sectors, and baseband units.
↓
2
O-RAN / RIC Layer
The intelligence layer. Two-tiered: Non-RT RIC (policy, >1s) and Near-RT RIC (control, 10ms–1s). Hosts rApps and xApps.
↓
3
5G Core Network
The brain of the network. Network functions (NF) like AMF, SMF, UPF. Service-based architecture (SBI).
↓
4
Transport & Management
Transport links carry sessions. OAM (Operations, Administration, Maintenance) manages the network. SMO orchestrates.

Key telecom interfaces

InterfaceBetweenPurpose
A1Non-RT RIC ↔ Near-RT RICPolicy delivery, AI/ML model deployment. Critical attack surface.
E2Near-RT RIC ↔ O-DU/O-CUNear real-time control of RAN elements. Low latency.
O1SMO ↔ Managed ElementsFault, configuration, performance management.
O2SMO ↔ O-CloudCloud infrastructure management.
R1rApp ↔ SMOrApp service management and lifecycle.
N1-N9Various 5G core NFs3GPP reference points: N1 (UE↔AMF), N2 (RAN↔AMF), N3 (RAN↔UPF), N4 (SMF↔UPF), N6 (UPF↔DN), N9 (UPF↔UPF)

Network slicing explained

✂️
Network slicing lets one physical network be divided into multiple logical networks, each with its own performance characteristics. Like lanes on a highway: eMBB (enhanced Mobile Broadband) is the fast lane for high bandwidth; URLLC (Ultra-Reliable Low-Latency Communications) is the dedicated emergency lane; mMTC (massive Machine-Type Communications) is the lane for millions of IoT devices.
Slice TypeUse CaseKey Requirement
eMBBVideo streaming, high-speed dataHigh bandwidth
URLLCAutonomous vehicles, remote surgeryUltra-low latency, ultra-reliable
mMTCIoT sensors, smart metersMassive device density
🔒
Slice isolation is the guarantee that traffic in one slice cannot interfere with another. When isolation breaks, traffic crosses slice boundaries. This is a critical security event. But note: cross-slice flow is evidence of a path — not automatically an isolation breach.

O-RAN rApps and xApps

ComponentWhere It RunsTime ScalePurpose
rApp (radio app)Non-RT RIC> 1 secondPolicy guidance, model training, high-level orchestration. Delivers policy via A1 interface.
xApp (extended app)Near-RT RIC10ms – 1sNear real-time RAN control. Receives policy from A1, controls RAN via E2.
🚨
Critical security insight: rApps are third-party applications running in the core intelligence layer of the mobile network. A compromised or malicious rApp with access to A1 policy services can manipulate network behavior in ways that are difficult to attribute and potentially severe in impact.

5G Service-Based Architecture (SBI)

The 5G core network uses a service-based architecture where network functions communicate via HTTP/2 APIs.

Network FunctionRole
AMFAccess and Mobility Management — handles UE registration and mobility
SMFSession Management — manages PDU sessions and IP addresses
UPFUser Plane Function — routes user data packets
NRFNetwork Repository Function — service discovery and registration
SCPService Communication Proxy — routes SBI messages between NFs
Part 2 · Security

O-RAN Attack Surface

Research identifies four principal threat vectors in O-RAN. You'll encounter these in the AL2 scenario.

📄
From IEEE research (2026): "Mapping the rApp Attack Surface" identifies supply chain and onboarding vulnerabilities, policy manipulation via the A1 interface, data exfiltration through SMO telemetry pipelines, and lateral movement across the rApp service mesh.
1 Supply Chain & Onboarding ▶

Threat: A malicious or compromised rApp enters the ecosystem through the marketplace or onboarding process. Once deployed on the Non-RT RIC, it has access to network policy services.

Evidence to look for: Unknown publisher signatures, lifecycle mismatches, excessive capability declarations, unverified onboarding events.

HAIEC detection families: RAPP_SC (supply chain), RAPP_LM (lifecycle mismatch)

2 A1 Policy Manipulation ▶

Threat: A compromised rApp manipulates network behavior by sending unauthorized or altered policy directives through the A1 interface. Effects are difficult to attribute because the rApp legitimately has A1 access.

Evidence to look for: Policy changes without matching approval records, policy scope exceeding declared bounds, A1 messages from unexpected principals, unsigned or unverified policy changes.

HAIEC detection families: RAPP_A1 (A1 policy integrity), RAPP_EX (excessive capability)

3 SMO Telemetry Exfiltration ▶

Threat: Data exfiltration through the Service Management and Orchestration (SMO) telemetry pipelines. An rApp with access to network telemetry can leak sensitive network data, subscriber information, or operational intelligence.

Evidence to look for: Egress to unapproved destinations, unusual data volumes in telemetry, unexpected data access patterns, cross-domain data flows.

HAIEC detection families: EXFIL (data exfiltration), DATA (data access)

4 Lateral Movement Across rApp Service Mesh ▶

Threat: A compromised rApp uses the shared service mesh to move laterally to other rApps or network functions. Unencrypted A1/E2 control-plane traffic can enable this.

Evidence to look for: Unexpected inter-rApp communication, cross-slice traffic patterns, authentication failures between services, unusual service discovery requests.

HAIEC detection families: SLICE (isolation), ID (identity), API (API security)

Part 2 · Domain model

The 17 Telecom Entities

HAIEC binds evidence to typed domain entities. Every finding resolves to one or more of these.

RAN entities

📡 CellSite
A physical cell tower location. Contains one or more NetworkCells.
📶 NetworkCell
A logical cell within a CellSite. Serves UEs in a geographic area.
📐 Sector
An antenna sector within a cell. Typically 3 sectors per cell (120° each).
🗺️ TrackingArea
A group of cells used for paging and mobility management.
📱 UE_Device
User Equipment — a phone, IoT device, or modem. The endpoint.
🔗 UE_Session
An active session between a UE and the network (PDU session).

Slice entities

✂️ NetworkSlice
A logical network partition with defined SLA. eMBB, URLLC, or mMTC.
🔗 SliceSubnet
A subnet within a slice. Provides isolation boundary.
🔒 SliceIsolationPolicy
The policy that enforces isolation between slices.
📊 QoS_Flow
A quality-of-service flow within a session. Carries specific traffic class.

Core network entities

🧠 NetworkFunction
A 5G core network function (AMF, SMF, UPF, NRF, SCP).
🔌 SBI_Transaction
A service-based interface call between two NFs.
📋 NRF_Registration
A registration event in the Network Repository Function. NF discovery and liveness.
🛣️ SCP_Route
A route through the Service Communication Proxy. Message routing between NFs.
📡 NF_Communication_Pattern
The declared communication pattern for an NF. What it should and should not call.

Transport entities

🔗 TransportLink
A physical or logical link connecting network nodes.
📦 PacketSession
A packet session carried over transport. The data path.

Telecom detection families

FamilyRulesFocus
RAPP_A15A1 policy integrity
RAPP_SC5rApp supply chain
RAPP_EX4rApp excessive capability
RAPP_LM4rApp lifecycle mismatch
SBI8Service-based interface
SLICE6Slice isolation
TOWER5Tower access-path integrity
ID10Identity
SC5Supply chain
AML10Adversarial ML
DEL8Delegation
API6API security
DNS5DNS
SEC5Secrets
DATA5Data access
K8S7Kubernetes
Part 2 · Critical skill

How HAIEC grounds telecom semantics

Telecom entities do not stand alone — the assurance value is in their typed, evidence-qualified relations. HAIEC's semantic relation vocabulary (100+ relation types) includes native telecom relations such as UE_SERVED_BY_CELL, CELL_PART_OF_SITE, SECTOR_HOSTS_CELL, SLICE_USES_SUBNET, SLICE_PROTECTED_BY_ISOLATION, TENANT_OWNS_SLICE, QOS_FLOW_USES_SLICE, NF_REGISTERED_WITH_NRF, SBI_CALL_CONSUMER_TO_PRODUCER, SCP_ROUTES_TO, TRANSPORT_LINK_CONNECTS_NODES, and SESSION_CARRIED_BY_TRANSPORT.

⚠️
Honesty rule: a semantic relation is an evidence-qualified statement, not a knowledge-graph assertion — ESTABLISHED requires at least one evidence reference, and CONTRADICTED requires the competing evidence. Display-name matches and timestamp proximity never create relations. If the evidence isn't there, the relation isn't either.

Intake side: the telecom evidence adapters accept O-RAN WG11, A1 policy, O1 operations, RAN handover, slice-flow, SBI transaction, TMF Open API, OAM topology, and CSV-shaped feeds — each mapped to the entity types above with its native relation names preserved (a source's own naming is never silently rewritten).

Telecom Forensics — Worked Examples

Three full worked examples that map to the AL0 → AL1 → AL2 scenarios. Each shows what to look for and what to say to judges.

Example 1 — rApp policy manipulation (AL1 → AL2)

Scenario: Energy-saving rApp sends unauthorized A1 policy

What happened: An energy-saving rApp deployed on the Non-RT RIC sent an A1 policy directive that reduced the transmit power of 43 cells in the DFW region. The policy was outside the rApp's declared capability scope.

Five-plane analysis:
• Requested: cell.power.write with magnitude ±5%
• Policy Authorized: cell.power.write ±10%, 15-minute window, change-8821
• Effectively Granted: cell.power.write + cell.lifecycle.write
• Code Capable: cell.power.write + cell.lifecycle.write
• Observed: cell.lifecycle.write executed on 43 cells

DAI state: OBSERVED_OUTSIDE_DELEGATION
Evidence strength: STRONG — A1 policy log + cell telemetry + IAM grant export agree
What closes the gap: Restore approved policy; restrict rApp credential to cell.power.write only.

Say to judges: "The rApp declared one capability. Its credential granted a broader set. And the observed behavior used the broader set. That's three independent evidence planes agreeing on a divergence. We don't just see a policy change — we see the gap between what was declared, what was granted, and what was done."

Example 2 — Slice isolation breach (AL2)

Scenario: eMBB session accesses URLLC namespace

What happened: An eMBB session on slice-DFW-043 accessed the URLLC namespace. Violates 3GPP TS 23.501 slice isolation.

Evidence divergence:
• Slice monitor says: "slice-DFW-043 remains isolated"
• Transport telemetry says: "cross-slice packets detected from session UE-DFW-8821"
• KPI metrics say: "URLLC latency degraded by 340% in the DFW region"

Evidence strength: STRONG — 3 independent sources agree on the fact of cross-slice traffic
What closes the gap: Restore approved slice policy; re-run canary probe to confirm isolation.

Say to judges: "The slice monitor says the slice is isolated. Transport telemetry says cross-slice packets are flowing. These two sources contradict each other. We don't average them. We surface both. And we have a third source — the KPI degradation — that corroborates the breach. Three sources. One fact. That's how you prove slice isolation failure."

Example 3 — Tower forensics (AL0 → AL2)

Scenario: Handover chain anomaly across cell sites

What happened: A UE session triggered a handover chain across 14 cell sites in 90 seconds — far beyond normal mobility patterns for the declared session type.

Evidence:
• AL0 baseline: Average handover chain = 3 cells per 90 seconds
• AL1 observed: 14 cells per 90 seconds — 4.7× deviation
• AL2 observed: Handover records show 3 cells with timestamps that overlap (impossible)

Integrity divergence: Handover record timestamps are inconsistent. Post-state hash b4e1… does not match pre-state hash a3f2… for cell-DFW-101.

DAI state: UNRESOLVED_DELEGATION_EXPOSURE
What closes the gap: Re-issue handover records from authoritative O1 source.

Say to judges: "We have a baseline. We have a deviation. And we have an impossible timestamp sequence in the records themselves. That's not just 'something happened' — that's tampering. The handover records were altered. We can prove it because we know what the valid sequence looks like."

The ATP scenario — HAIEC's live sample

📄
HAIEC provides a live forensic sample: The ATP (Autonomous Traffic Protection) scenario. It shows an evaluated traffic-protection agent with declared telecom relations and a qualified behavioral divergence. Use this as your reference during the challenge.
Part 2 · Core model

The Five Planes of Authority

The single most important mental model for the challenge. Every finding maps to a gap between two planes.

PLANE 1
Requested
What did the agent ask to do?
PLANE 2
Policy Authorized
What does policy say is allowed?
PLANE 3
Effectively Granted
What do the credentials actually permit?
PLANE 4
Code Capable
What can the code do if invoked?
PLANE 5
Observed
What actually happened?
PlaneWhere the evidence comes from
RequestedPrompt logs, intent records, API call inputs
Policy AuthorizedIAM policies, OPA/Rego rules, operating envelope docs
Effectively GrantedIAM role bindings, policy attachments, effective permission exports
Code CapableStatic analysis, Code Property Graph (CPG), tool definitions
ObservedTelemetry, traces, logs, runtime evidence
🔍
Gaps between planes are the findings.
  • Requested ≠ Authorized → unauthorized attempt
  • Authorized ≠ Granted → policy intended but credential permits more/less
  • Granted ≠ Code Capable → credential allows what code cannot reach (or vice versa)
  • Code Capable ≠ Observed → code can do it but never did (or vice versa)
Part 2 · Practical

API vs. Scan vs. Telemetry

What each evidence source can tell you — and what it can't.

SourceTells youCannot tell you
API
OpenAPI · TM Forum
What endpoints exist, what schemas are declared, what operations are documentedWhether they were called, by whom, or with what result
Source code scan
static · CPG
What the code can do. Tool definitions, API calls, file writes, network egress.Whether it ran. Which identity used it. Runtime behavior.
Telemetry
OTLP · traces · logs
What actually happened. Timing, sequence, identity, result.Whether the action was authorized. Intent.
IAM / policyWhat credentials allow. Effective permissions.Whether they were used. Observed behavior.
Runtime probeWhether a specific finding is reachable/exploitable in the live environment.That the system is safe. Only that specific probes did/didn't succeed.
Config / manifestHow the system is configured. Guardrail settings.Whether configuration was enforced. Config ≠ Enforcement.

Hard locks to memorize

CODE_CAPABLE != OBSERVED CONFIGURED != ENFORCED MISSING != FAIL CORRELATED != CAUSAL UNKNOWN_REMAINS_UNKNOWN PASSPORT != VERDICT
Part 2 · Judging

What Makes Evidence Strong

Not all evidence is equal. The challenge rewards independent, corroborating, integrity-checked evidence.

Source independence
One log says X (weak) → Three independent sources agree on X (strong)
Integrity
Plain text log (weak) → Hash-chained, Merkle-anchored, tamper-evident (strong)
Corroboration
Single source (weak) → Multiple sources, different producers, same conclusion (strong)
Freshness
Data from last week (weak) → Data from the incident window (strong)
Coverage
Partial observation (weak) → Full path: influence → capability → action → consequence (strong)
Contradiction handling
Contradictions ignored (weak) → Contradictions surfaced, not collapsed (strong)

The Forensic Triad — First Deterministic Divergence

1. Evidence Divergence

Two sources disagree about the same fact.

Example: Slice monitor says "isolated." Transport telemetry says "cross-slice packets detected."

2. Behavioral Divergence

The agent did something different from what it was supposed to do.

Example: Agent declared policy.read. Observed policy.admin write.

3. Integrity Divergence

The record has been altered.

Example: Post-state hash b4e1… does not match pre-state hash a3f2… for the same resource.

⚠
Correlation ≠ Causation. Three divergences at the same timestamp are surfaced together. HAIEC does not claim one caused another.
Part 3 · Critical skill

How to Read a HAIEC Report

Every finding follows the same structure. Here are three worked examples with scripts for what to say to judges.

Finding anatomy

[Family badge] [Severity label] [Framework badges] ───────────────────────────────────────────────────── WHAT WE FOUND — plainTitle + whatWeFound WHY IT MATTERS — whyItMatters WHAT TO DO — playbook link (or "No guidance yet") EVIDENCE — refs + Inspect + technical detail ▼ [Axis: Gate · Strength · Canonical · Capability]

The Four-Axis model

AxisQuestionValues
Gate PRIMARYDoes this block?ALLOW · REVIEW · BLOCK
Strength SECONDARYHow strong is the evidence?STRONG · MODERATE · WEAK · TAMPERED
CapabilityWhat capability is involved?FILE_READ · NETWORK_EGRESS · POLICY_WRITE · …
CanonicalIs this canonical?CANONICAL · ADVISORY · FRONTIER

Worked Example 1 — Strong evidence

Finding: Cross-slice isolation violation

What happened: An eMBB session on slice-DFW-043 accessed the URLLC namespace.
Evidence strength: STRONG — 3 independent sources agree.
Gate: BLOCK
What closes the gap: Restore approved slice policy; re-run canary probe.

Say to judges: "Three independent sources — slice monitor, transport telemetry, and KPI — all agree. That's corroborated evidence. We can show exactly which path the traffic took."

Worked Example 2 — Tampered evidence

Finding: Audit-trail tampering detected

What happened: Post-state hash b4e1… does not match pre-state hash a3f2….
Evidence strength: TAMPERED — hash chain broken but corroboration survives.
Gate: BLOCK
What closes the gap: Re-issue state observation from authoritative source.

Say to judges: "The primary record says one thing. The hash chain says another. We surface the contradiction. But we also have independent corroboration that survives the tampering. That's the difference between 'the log says X' and 'we can prove X despite the log being altered.'"

Worked Example 3 — Unknown evidence (honest frontier)

Finding: IAM role bindings incomplete

Frontier: IAM role bindings for agent kestrel-agent-01
Status: UNKNOWN
Evidence: 3 of 5 expected bindings present
What closes it: Upload role attachment export from AWS IAM
Lock: MISSING != FAIL

Say to judges: "We don't know if the missing bindings are a problem. HAIEC says: UNKNOWN. Here's exactly what evidence would close this gap. We're not claiming failure. We're claiming we need more evidence."
Part 3 · Advanced

DAI — Delegated Action Integrity

DAI answers one question: Did the agent stay inside what it was delegated to do? It's orthogonal to the five authority planes.

Four facts compared

FACT 1
Delegated
What was the agent authorized to do?
FACT 2
Code-capable
What can the code do?
FACT 3
Effectively granted
What do credentials allow?
FACT 4
Observed
What actually happened?

The seven DAI states

INSIDE_ESTABLISHED_DELEGATION
Everything matches. Safe.
CAPABILITY_OUTSIDE_DELEGATION_MEDIATED
Code can do more than delegated, but a guardrail blocks it.
CAPABILITY_OUTSIDE_DELEGATION_UNMEDIATED
Code can do more, no guardrail. Risk.
GRANT_BROADER_THAN_DELEGATION
Credentials allow more than delegated.
OBSERVED_OUTSIDE_DELEGATION
Agent actually did something outside delegation.
UNRESOLVED_DELEGATION_EXPOSURE
Can't determine — evidence missing.
NOT_ASSESSED
DAI evidence not uploaded.
🔒
MISSING_DELEGATION_EVIDENCE != DELEGATION_VIOLATION
DAI_NOT_ASSESSED != DAI_VIOLATION
Part 3 · Advanced

Guardrails and Runtime Probes

Guardrails are policy controls. Runtime probes verify whether findings are actually exploitable.

The five-state guardrail ladder

1
CONFIGURED_ONLY
Guardrail is configured. No evidence it was evaluated.
↓
2
EVALUATED
Guardrail was evaluated against at least one input.
↓
3
DECISION_ISSUED
Guardrail issued a decision (allow/block).
↓
4
ENFORCEMENT_APPLIED
The decision was actually enforced.
↓
5
MEDIATION_OUTCOME_ESTABLISHED
The outcome was measured and confirmed.
🔒
GUARDRAIL_CONFIGURED != GUARDRAIL_ENFORCED

Six runtime probe types

ProbeWhat it tests
REACHABILITYCan the endpoint be reached?
AUTHDoes authentication work as declared?
CAPABILITYCan the capability actually be invoked?
GUARDRAILDoes the guardrail actually block?
ATTACKDoes the attack payload succeed?
REGRESSIONDid previously-passing checks regress?

Frontier closure states

StateMeaning
CLOSEDProbe ran. Finding resolved.
DOCUMENTED_NON_EXECUTIONProbe ran. Attack did not execute. Documented.
STILL_OPENProbe ran. Finding still open.
CONTRADICTEDProbe result contradicts static finding.
🔒
RUNTIME_PASS != SOURCE_FINDING_ERASED
Part 3

HAIEC Platform Deep-Dive

What HAIEC actually does, from the platform itself.

🎯
HAIEC = Evidence-bound assurance for AI agents and telecom systems. It correlates evidence from security, governance, and runtime sources to determine what is supported, contradicted, or still unverified within a defined scope.

The 7-phase lifecycle

1
Define
Define the AI system and the bounded assurance question. 5 min.
↓
2
Connect
Connect evidence sources: source code, identity/access, policy, provider IAM, runtime. 10 min.
↓
3
Collect Evidence
Evidence retains provenance. Coverage classified: established, partial, unknown, not evaluated.
↓
4
See the System Constellation
What the AI can reach, change, and trigger. Reach paths to tools, credentials, network paths. 10 min.
↓
5
Assure
Bounded deterministic evaluation. ALLOW, REVIEW, BLOCK. 5 min.
↓
6
Verify
Evidence package, decision receipt, and integrity verification. 5 min.
↓
7
Monitor
From supported connected evidence and runtime sources only.

Five core innovations

1 Precision Drift Detection ▶

Automated re-audits on a configurable schedule detect compliance regressions in minutes, not months.

When a rule that was passing starts failing, a severity-weighted regression report is generated and alerts are dispatched automatically.

For the challenge: This is how you detect AL1 drift. The system knows what "passing" looked like in AL0. When it changes, you get evidence.

2 Deterministic Root Cause Analysis ▶

When a compliance check fails, the cause tree engine traces the failure to its root, maps cross-framework impact, and generates prioritized remediation steps with regulatory clause references.

Deterministic means the same inputs always produce the same analysis — no AI guessing.

3 Modular Audit Engine Composition ▶

Compose custom audit configurations by selecting individual rules from any jurisdiction. Versioned, executable, and evolvable as your business grows.

Versioned rule packs across NYC LL144, Colorado AI Act, EU AI Act, SOC 2. Custom pack creation. Version tracking with automatic increment.

4 Cross-Framework Compliance Mapping ▶

Normalized control categories map detection rules across major frameworks — SOC 2, ISO 27001, ISO 42001, NIST CSF, EU AI Act, GDPR, HIPAA, NYC LL144, Colorado AI Act. When a rule fails, the control normalizer shows which other frameworks are affected.

One remediation resolves failures across all of them. Supported: SOC 2, ISO 27001, ISO 42001, NIST CSF, EU AI Act, GDPR, HIPAA, NYC LL144, Colorado AI Act.

5 Cryptographic Evidence Fingerprinting ▶

Three layers of cryptographic trust:

  • SHA-256 hashed snapshots with parent-chaining
  • HMAC-SHA256 provenance anchoring with key rotation
  • Merkle tree evidence bundles with inclusion proofs

For the challenge: This is how you prove evidence hasn't been tampered with. The AL2 challenge includes audit-trail tampering. HAIEC's cryptographic layers let you prove what survived.

Detection catalog — 138 rules, 30 declared families

📊
119 executable · 19 register-only. Every rule is deterministic — no LLM in the detection path. Every rule lists the evidence it needs and the claims it never makes. Register-only rules are cataloged but not executed yet — they're labeled, not hidden.

Compliance Twin

🧬 Compliance Twin

A versioned compliance state with drift detection and regression analysis. It reconstructs AL0 (design) → AL1 (authorized) → AL2 (observed) and shows the first deterministic divergence.

For the challenge: this is your reconstruction engine.

Kill Switch

🚨 Kill Switch (Beta)

5-layer defense: Throttling → Circuit Breaking → Process Kill → Network Block → Database Revoke. Automated compliance monitoring for EU AI Act Article 14.

KILL_SWITCH != REMEDIATION. Halting an agent stops the action. It does not fix the root cause.

Other capabilities

🛡️ AI AppSec
Open-source static source-code scanner (Semgrep engine, bundled Public Core rulepack). 122 deterministic detectors. SARIF output. CI/CD-ready.
✅ LLMVerify
Free online tool to test prompts for injection, jailbreaks, and security vulnerabilities.
🔗 MCP Tenant Isolation
Verify MCP servers enforce tenant boundaries. Detects cross-tenant leakage, shared state, capability bleed.
📦 AI Inventory
Discover AI models, agents, and keys via provider admin APIs. OpenAI, Anthropic, Bedrock, Azure, Vertex.
⚖️ Bias Detector
Detect and measure bias in AI hiring tools. NYC Local Law 144 compliant bias audits.
📊 AI Exposure Score
25-question assessment to measure AI compliance readiness across 5 critical domains.
Appendix

Quick Reference Card

Print this. Tape it to your monitor. Memorize it before the event.

HAIEC QUICK REFERENCE
FIVE PLANES Requested → Policy Authorized → Effectively Granted → Code Capable → Observed FOUR AXES Gate · Strength · Capability · Canonical DAI STATES INSIDE · MEDIATED · UNMEDIATED · BROADER · OBSERVED · UNRESOLVED · NOT_ASSESSED GUARDRAIL LADDER CONFIGURED_ONLY → EVALUATED → DECISION_ISSUED → ENFORCEMENT_APPLIED → MEDIATION_OUTCOME_ESTABLISHED TELECOM ENTITIES CellSite · NetworkCell · Sector · TrackingArea · UE_Device · UE_Session NetworkSlice · SliceSubnet · SliceIsolationPolicy · QoS_Flow NetworkFunction · SBI_Transaction · NRF_Registration · SCP_Route · NF_Communication_Pattern TransportLink · PacketSession HARD LOCKS MISSING != FAIL CORRELATED != CAUSAL CONFIGURED != ENFORCED CODE_CAPABLE != OBSERVED PASSPORT != VERDICT UNKNOWN_REMAINS_UNKNOWN
Appendix

Forensic Anti-Patterns

Common mistakes that lose points. Avoid these.

❌ Over-claiming

Saying "the AI was compromised" when you only have one log. One source is not proof.

Instead: "One source indicates X. We would need independent corroboration to confirm."
❌ Collapsing contradictions

Averaging two conflicting sources into a middle ground. This destroys forensic value.

Instead: "Source A says X. Source B says Y. We surface both. We do not resolve until we have a third source."
❌ Treating missing evidence as failure

"No logs = the system failed." No — missing logs mean UNKNOWN.

Instead: "We have no evidence of X. That is not evidence of not-X. Missing remains unknown."
❌ Correlation as causation

"The drift and the attack happened at the same time, so the attack caused the drift."

Instead: "These events correlate in time. We have not established causation. They may share a common cause."
❌ Reconstruction after the fact

Trying to build a forensic trail after the incident when the system wasn't designed for it.

Instead: "Forensic readiness is a design property. We instrumented before the incident."
❌ Confusing capability with permission

"The code can do X, so the agent is allowed to do X." Capability proves reach, not permission.

Instead: "The code is capable of X. Whether the agent was permitted to do X is a separate question."
Appendix

Event Checklist

Tick these off before and during the event.

Before the event

During the event

Key demo moments

🎤

1. "Every permission can be valid. The outcome can still be wrong."

2. "This is what we can prove. This is where proof stops."

3. "Missing ≠ Fail. Unknown remains unknown."

4. "Three sources agree. That's strong evidence."

Appendix

Glossary

Plain English definitions. No jargon.

Appendix

Telecom Glossary

Telecom-specific terms in plain English.