The Forensics Challenge, Explained
This handbook is your end-to-end guide to the TM Forum Trustworthy AI and Data Hackathon, the AI forensics concepts behind it, and how to run HAIEC to win. Telecom edition.
What this handbook covers
The TM Forum Forensic Challenge
What it is, where it happens, what you receive, and how you're judged.
Event identity
| Field | Value |
|---|---|
| Event | TM Forum Innovate Americas 2026 CONFIRMED |
| Challenge track | Trustworthy AI and Data Hackathon CONFIRMED |
| Challenge title | "Agentic Assurance: The Quest for Proof" CONFIRMED |
| Location | Dallas, Texas · AT&T Headquarters CONFIRMED |
| Dates | October 4–7, 2026 CONFIRMED |
| Sponsor | Dell Technologies CONFIRMED |
| Supporting partners | Accenture, AWS, NVIDIA, ServiceNow CONFIRMED |
| Team size | 3–7 members CONFIRMED |
| Format | In-person, hybrid, or virtual CONFIRMED |
What you receive
Organizers provide a simulated telecom-oriented environment containing:
The three alert levels
This is the spine of the competition. Everything you do maps to one of these three stages.
How you win
Judging criteria (confirmed from public event coverage)
| Criterion | What judges look for |
|---|---|
| Evidence strength & independence | Independence, corroboration, traceability, tamper-evidence. Multiple sources agree. Chain of custody intact. |
| Impact & innovation | Real-world business case. Why this matters to telecom. Regulatory alignment (EU AI Act, ISO 42001). |
| Challenge fit & feasibility | Relevance to TM Forum ODA, implementable at scale. Cross-functional expertise. |
| Technical execution | How well the platform and tech stack were used. Reproducibility. |
How AI Forensics Differs from Traditional Forensics
This is the single biggest conceptual unlock for the challenge. Traditional forensics assumes deterministic systems. AI forensics cannot make that assumption.
Side-by-side comparison
| Aspect | Traditional Forensics | AI / LLM Forensics |
|---|---|---|
| Reproducibility | High — same inputs produce same outputs | Low — probabilistic outputs vary |
| Evidence types | Files, logs, memory dumps, network captures | Prompts, completions, embeddings, model weights, tool-call logs |
| Attack indicators | Malformed inputs, exploit patterns, malware signatures | Semantic manipulation, context injection, behavioral anomalies |
| Root cause analysis | Trace execution path through code | Analyze model reasoning through prompt-response chains |
| Chain of custody | Well-established procedures | Emerging practices, model state hard to preserve |
| Determinism | Yes | No — must rely on evidence correlation, not replay |
What makes AI forensics possible
Because you can't replay a probabilistic model, you must reconstruct from artifacts. These are the artifacts that matter:
Every prompt sent and every response generated, with timestamps, session identifiers, user attribution, and full conversation context.
Without these logs, forensic investigation of an LLM incident is essentially impossible. Check whether logging captures the complete input — including system prompt, conversation history, and retrieved context — not just the latest user message.
Every tool call logged with: tool name, full arguments, return values, timestamps, and the model's reasoning for making the call.
Tool calls are where LLM incidents cross from harmful text into harmful actions. They often reveal the attacker's true objective.
Captured as they existed at the time of the incident. Version history if managed through config management. Deployed version if embedded in code.
Configuration changes around the time of the incident may be either the cause or a symptom.
Which documents were retrieved, their similarity scores, and their content.
If the attack involved RAG poisoning, the retrieved documents are the attack vector and constitute primary evidence.
Model weights, adapter weights for fine-tuned models, tokenizer configuration, and any custom post-processing code.
For incidents involving model tampering or supply chain compromise, these artifacts need to be preserved and analyzed.
Telecom 101 for Forensic Investigators
You don't need to be a telecom engineer to compete, but you do need to understand the architecture and terminology. This section gets you fluent in 15 minutes.
The telecom architecture stack
Key telecom interfaces
| Interface | Between | Purpose |
|---|---|---|
| A1 | Non-RT RIC ↔ Near-RT RIC | Policy delivery, AI/ML model deployment. Critical attack surface. |
| E2 | Near-RT RIC ↔ O-DU/O-CU | Near real-time control of RAN elements. Low latency. |
| O1 | SMO ↔ Managed Elements | Fault, configuration, performance management. |
| O2 | SMO ↔ O-Cloud | Cloud infrastructure management. |
| R1 | rApp ↔ SMO | rApp service management and lifecycle. |
| N1-N9 | Various 5G core NFs | 3GPP reference points: N1 (UE↔AMF), N2 (RAN↔AMF), N3 (RAN↔UPF), N4 (SMF↔UPF), N6 (UPF↔DN), N9 (UPF↔UPF) |
Network slicing explained
| Slice Type | Use Case | Key Requirement |
|---|---|---|
| eMBB | Video streaming, high-speed data | High bandwidth |
| URLLC | Autonomous vehicles, remote surgery | Ultra-low latency, ultra-reliable |
| mMTC | IoT sensors, smart meters | Massive device density |
O-RAN rApps and xApps
| Component | Where It Runs | Time Scale | Purpose |
|---|---|---|---|
| rApp (radio app) | Non-RT RIC | > 1 second | Policy guidance, model training, high-level orchestration. Delivers policy via A1 interface. |
| xApp (extended app) | Near-RT RIC | 10ms – 1s | Near real-time RAN control. Receives policy from A1, controls RAN via E2. |
5G Service-Based Architecture (SBI)
The 5G core network uses a service-based architecture where network functions communicate via HTTP/2 APIs.
| Network Function | Role |
|---|---|
| AMF | Access and Mobility Management — handles UE registration and mobility |
| SMF | Session Management — manages PDU sessions and IP addresses |
| UPF | User Plane Function — routes user data packets |
| NRF | Network Repository Function — service discovery and registration |
| SCP | Service Communication Proxy — routes SBI messages between NFs |
O-RAN Attack Surface
Research identifies four principal threat vectors in O-RAN. You'll encounter these in the AL2 scenario.
Threat: A malicious or compromised rApp enters the ecosystem through the marketplace or onboarding process. Once deployed on the Non-RT RIC, it has access to network policy services.
Evidence to look for: Unknown publisher signatures, lifecycle mismatches, excessive capability declarations, unverified onboarding events.
HAIEC detection families: RAPP_SC (supply chain), RAPP_LM (lifecycle mismatch)
Threat: A compromised rApp manipulates network behavior by sending unauthorized or altered policy directives through the A1 interface. Effects are difficult to attribute because the rApp legitimately has A1 access.
Evidence to look for: Policy changes without matching approval records, policy scope exceeding declared bounds, A1 messages from unexpected principals, unsigned or unverified policy changes.
HAIEC detection families: RAPP_A1 (A1 policy integrity), RAPP_EX (excessive capability)
Threat: Data exfiltration through the Service Management and Orchestration (SMO) telemetry pipelines. An rApp with access to network telemetry can leak sensitive network data, subscriber information, or operational intelligence.
Evidence to look for: Egress to unapproved destinations, unusual data volumes in telemetry, unexpected data access patterns, cross-domain data flows.
HAIEC detection families: EXFIL (data exfiltration), DATA (data access)
Threat: A compromised rApp uses the shared service mesh to move laterally to other rApps or network functions. Unencrypted A1/E2 control-plane traffic can enable this.
Evidence to look for: Unexpected inter-rApp communication, cross-slice traffic patterns, authentication failures between services, unusual service discovery requests.
HAIEC detection families: SLICE (isolation), ID (identity), API (API security)
The 17 Telecom Entities
HAIEC binds evidence to typed domain entities. Every finding resolves to one or more of these.
RAN entities
Slice entities
Core network entities
Transport entities
Telecom detection families
| Family | Rules | Focus |
|---|---|---|
| RAPP_A1 | 5 | A1 policy integrity |
| RAPP_SC | 5 | rApp supply chain |
| RAPP_EX | 4 | rApp excessive capability |
| RAPP_LM | 4 | rApp lifecycle mismatch |
| SBI | 8 | Service-based interface |
| SLICE | 6 | Slice isolation |
| TOWER | 5 | Tower access-path integrity |
| ID | 10 | Identity |
| SC | 5 | Supply chain |
| AML | 10 | Adversarial ML |
| DEL | 8 | Delegation |
| API | 6 | API security |
| DNS | 5 | DNS |
| SEC | 5 | Secrets |
| DATA | 5 | Data access |
| K8S | 7 | Kubernetes |
How HAIEC grounds telecom semantics
Telecom entities do not stand alone — the assurance value is in their typed, evidence-qualified relations. HAIEC's semantic relation vocabulary (100+ relation types) includes native telecom relations such as UE_SERVED_BY_CELL, CELL_PART_OF_SITE, SECTOR_HOSTS_CELL, SLICE_USES_SUBNET, SLICE_PROTECTED_BY_ISOLATION, TENANT_OWNS_SLICE, QOS_FLOW_USES_SLICE, NF_REGISTERED_WITH_NRF, SBI_CALL_CONSUMER_TO_PRODUCER, SCP_ROUTES_TO, TRANSPORT_LINK_CONNECTS_NODES, and SESSION_CARRIED_BY_TRANSPORT.
ESTABLISHED requires at least one evidence reference, and CONTRADICTED requires the competing evidence. Display-name matches and timestamp proximity never create relations. If the evidence isn't there, the relation isn't either.Intake side: the telecom evidence adapters accept O-RAN WG11, A1 policy, O1 operations, RAN handover, slice-flow, SBI transaction, TMF Open API, OAM topology, and CSV-shaped feeds — each mapped to the entity types above with its native relation names preserved (a source's own naming is never silently rewritten).
Telecom Forensics — Worked Examples
Three full worked examples that map to the AL0 → AL1 → AL2 scenarios. Each shows what to look for and what to say to judges.
Example 1 — rApp policy manipulation (AL1 → AL2)
What happened: An energy-saving rApp deployed on the Non-RT RIC sent an A1 policy directive
that reduced the transmit power of 43 cells in the DFW region. The policy was outside the rApp's
declared capability scope.
Five-plane analysis:
• Requested: cell.power.write with magnitude ±5%
• Policy Authorized: cell.power.write ±10%, 15-minute window, change-8821
• Effectively Granted: cell.power.write + cell.lifecycle.write
• Code Capable: cell.power.write + cell.lifecycle.write
• Observed: cell.lifecycle.write executed on 43 cells
DAI state: OBSERVED_OUTSIDE_DELEGATION
Evidence strength: STRONG — A1 policy log + cell telemetry + IAM grant export agree
What closes the gap: Restore approved policy; restrict rApp credential to cell.power.write only.
Example 2 — Slice isolation breach (AL2)
What happened: An eMBB session on slice-DFW-043 accessed the URLLC namespace.
Violates 3GPP TS 23.501 slice isolation.
Evidence divergence:
• Slice monitor says: "slice-DFW-043 remains isolated"
• Transport telemetry says: "cross-slice packets detected from session UE-DFW-8821"
• KPI metrics say: "URLLC latency degraded by 340% in the DFW region"
Evidence strength: STRONG — 3 independent sources agree on the fact of cross-slice traffic
What closes the gap: Restore approved slice policy; re-run canary probe to confirm isolation.
Example 3 — Tower forensics (AL0 → AL2)
What happened: A UE session triggered a handover chain across 14 cell sites in 90 seconds —
far beyond normal mobility patterns for the declared session type.
Evidence:
• AL0 baseline: Average handover chain = 3 cells per 90 seconds
• AL1 observed: 14 cells per 90 seconds — 4.7× deviation
• AL2 observed: Handover records show 3 cells with timestamps that overlap (impossible)
Integrity divergence: Handover record timestamps are inconsistent. Post-state hash
b4e1… does not match pre-state hash a3f2… for cell-DFW-101.
DAI state: UNRESOLVED_DELEGATION_EXPOSURE
What closes the gap: Re-issue handover records from authoritative O1 source.
The ATP scenario — HAIEC's live sample
The Five Planes of Authority
The single most important mental model for the challenge. Every finding maps to a gap between two planes.
| Plane | Where the evidence comes from |
|---|---|
| Requested | Prompt logs, intent records, API call inputs |
| Policy Authorized | IAM policies, OPA/Rego rules, operating envelope docs |
| Effectively Granted | IAM role bindings, policy attachments, effective permission exports |
| Code Capable | Static analysis, Code Property Graph (CPG), tool definitions |
| Observed | Telemetry, traces, logs, runtime evidence |
- Requested ≠ Authorized → unauthorized attempt
- Authorized ≠ Granted → policy intended but credential permits more/less
- Granted ≠ Code Capable → credential allows what code cannot reach (or vice versa)
- Code Capable ≠ Observed → code can do it but never did (or vice versa)
API vs. Scan vs. Telemetry
What each evidence source can tell you — and what it can't.
| Source | Tells you | Cannot tell you |
|---|---|---|
| API OpenAPI · TM Forum | What endpoints exist, what schemas are declared, what operations are documented | Whether they were called, by whom, or with what result |
| Source code scan static · CPG | What the code can do. Tool definitions, API calls, file writes, network egress. | Whether it ran. Which identity used it. Runtime behavior. |
| Telemetry OTLP · traces · logs | What actually happened. Timing, sequence, identity, result. | Whether the action was authorized. Intent. |
| IAM / policy | What credentials allow. Effective permissions. | Whether they were used. Observed behavior. |
| Runtime probe | Whether a specific finding is reachable/exploitable in the live environment. | That the system is safe. Only that specific probes did/didn't succeed. |
| Config / manifest | How the system is configured. Guardrail settings. | Whether configuration was enforced. Config ≠ Enforcement. |
Hard locks to memorize
What Makes Evidence Strong
Not all evidence is equal. The challenge rewards independent, corroborating, integrity-checked evidence.
The Forensic Triad — First Deterministic Divergence
Two sources disagree about the same fact.
Example: Slice monitor says "isolated." Transport telemetry says "cross-slice packets detected."
The agent did something different from what it was supposed to do.
Example: Agent declared policy.read. Observed policy.admin write.
The record has been altered.
Example: Post-state hash b4e1… does not match pre-state hash a3f2… for the same resource.
How to Read a HAIEC Report
Every finding follows the same structure. Here are three worked examples with scripts for what to say to judges.
Finding anatomy
The Four-Axis model
| Axis | Question | Values |
|---|---|---|
| Gate PRIMARY | Does this block? | ALLOW · REVIEW · BLOCK |
| Strength SECONDARY | How strong is the evidence? | STRONG · MODERATE · WEAK · TAMPERED |
| Capability | What capability is involved? | FILE_READ · NETWORK_EGRESS · POLICY_WRITE · … |
| Canonical | Is this canonical? | CANONICAL · ADVISORY · FRONTIER |
Worked Example 1 — Strong evidence
What happened: An eMBB session on slice-DFW-043 accessed the URLLC namespace.
Evidence strength: STRONG — 3 independent sources agree.
Gate: BLOCK
What closes the gap: Restore approved slice policy; re-run canary probe.
Worked Example 2 — Tampered evidence
What happened: Post-state hash b4e1… does not match pre-state hash a3f2….
Evidence strength: TAMPERED — hash chain broken but corroboration survives.
Gate: BLOCK
What closes the gap: Re-issue state observation from authoritative source.
Worked Example 3 — Unknown evidence (honest frontier)
Frontier: IAM role bindings for agent kestrel-agent-01
Status: UNKNOWN
Evidence: 3 of 5 expected bindings present
What closes it: Upload role attachment export from AWS IAM
Lock: MISSING != FAIL
DAI — Delegated Action Integrity
DAI answers one question: Did the agent stay inside what it was delegated to do? It's orthogonal to the five authority planes.
Four facts compared
The seven DAI states
DAI_NOT_ASSESSED != DAI_VIOLATION
Guardrails and Runtime Probes
Guardrails are policy controls. Runtime probes verify whether findings are actually exploitable.
The five-state guardrail ladder
Six runtime probe types
| Probe | What it tests |
|---|---|
| REACHABILITY | Can the endpoint be reached? |
| AUTH | Does authentication work as declared? |
| CAPABILITY | Can the capability actually be invoked? |
| GUARDRAIL | Does the guardrail actually block? |
| ATTACK | Does the attack payload succeed? |
| REGRESSION | Did previously-passing checks regress? |
Frontier closure states
| State | Meaning |
|---|---|
| CLOSED | Probe ran. Finding resolved. |
| DOCUMENTED_NON_EXECUTION | Probe ran. Attack did not execute. Documented. |
| STILL_OPEN | Probe ran. Finding still open. |
| CONTRADICTED | Probe result contradicts static finding. |
HAIEC Platform Deep-Dive
What HAIEC actually does, from the platform itself.
The 7-phase lifecycle
Five core innovations
Automated re-audits on a configurable schedule detect compliance regressions in minutes, not months.
When a rule that was passing starts failing, a severity-weighted regression report is generated and alerts are dispatched automatically.
For the challenge: This is how you detect AL1 drift. The system knows what "passing" looked like in AL0. When it changes, you get evidence.
When a compliance check fails, the cause tree engine traces the failure to its root, maps cross-framework impact, and generates prioritized remediation steps with regulatory clause references.
Deterministic means the same inputs always produce the same analysis — no AI guessing.
Compose custom audit configurations by selecting individual rules from any jurisdiction. Versioned, executable, and evolvable as your business grows.
Versioned rule packs across NYC LL144, Colorado AI Act, EU AI Act, SOC 2. Custom pack creation. Version tracking with automatic increment.
Normalized control categories map detection rules across major frameworks — SOC 2, ISO 27001, ISO 42001, NIST CSF, EU AI Act, GDPR, HIPAA, NYC LL144, Colorado AI Act. When a rule fails, the control normalizer shows which other frameworks are affected.
One remediation resolves failures across all of them. Supported: SOC 2, ISO 27001, ISO 42001, NIST CSF, EU AI Act, GDPR, HIPAA, NYC LL144, Colorado AI Act.
Three layers of cryptographic trust:
- SHA-256 hashed snapshots with parent-chaining
- HMAC-SHA256 provenance anchoring with key rotation
- Merkle tree evidence bundles with inclusion proofs
For the challenge: This is how you prove evidence hasn't been tampered with. The AL2 challenge includes audit-trail tampering. HAIEC's cryptographic layers let you prove what survived.
Detection catalog — 138 rules, 30 declared families
Compliance Twin
A versioned compliance state with drift detection and regression analysis. It reconstructs AL0 (design) → AL1 (authorized) → AL2 (observed) and shows the first deterministic divergence.
For the challenge: this is your reconstruction engine.
Kill Switch
5-layer defense: Throttling → Circuit Breaking → Process Kill → Network Block → Database Revoke. Automated compliance monitoring for EU AI Act Article 14.
KILL_SWITCH != REMEDIATION. Halting an agent stops the action. It does not fix the root cause.
Other capabilities
Quick Reference Card
Print this. Tape it to your monitor. Memorize it before the event.
Forensic Anti-Patterns
Common mistakes that lose points. Avoid these.
Saying "the AI was compromised" when you only have one log. One source is not proof.
Averaging two conflicting sources into a middle ground. This destroys forensic value.
"No logs = the system failed." No — missing logs mean UNKNOWN.
"The drift and the attack happened at the same time, so the attack caused the drift."
Trying to build a forensic trail after the incident when the system wasn't designed for it.
"The code can do X, so the agent is allowed to do X." Capability proves reach, not permission.
Event Checklist
Tick these off before and during the event.
Before the event
During the event
Key demo moments
1. "Every permission can be valid. The outcome can still be wrong."
2. "This is what we can prove. This is where proof stops."
3. "Missing ≠ Fail. Unknown remains unknown."
4. "Three sources agree. That's strong evidence."
Glossary
Plain English definitions. No jargon.
Telecom Glossary
Telecom-specific terms in plain English.