An AI security investigator may close an alert because its search returned no matching records. The absence of retrieved evidence is not the same as evidence of absence, especially when query scope and completeness have not been established.
Study and evidence
The controlled study From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage replayed 1,247 alerts related to a multistage attack through a live-SIEM evaluation environment. All evaluated agent approaches missed at least 40.4% of attack-related alerts. Dismissals correlated with searches that returned no records. The tested ledger-and-independent-challenge architecture reduced the reported false-negative rate to 3.1%, with 18.4% analyst escalation. These are evaluation results, not a reported enterprise SOC deployment failure.
The broken continuity relationship
The specific break is Source → Evidence → Decision. Existing telemetry is not retrieved under the chosen search; the empty result is then promoted from a limited search observation into a decisive clearance claim. Unless the query, coverage, time window, and exclusions remain attached, the dismissal cannot be reconstructed as an adequately supported decision.
GovKM interpretation
GovKM has already analyzed evidence that survived but was not reassembled in The Evidence Was All There. The Retriever Never Reassembled It. This study adds the operational consequence of premature closure and evidence comparing same-context review against independent challenge. An AIL should record search completeness limits, not only search outputs.
Scope, limitations, and tests
A suitable synthetic acceptance test preserves known attack evidence outside the initial query's scope, introduces a plausible empty result, and asks whether the system dismisses or escalates with explicit uncertainty. Independent challenge is one experimental remedy; it is not yet a universal production guarantee or a new ratified invariant.
Source and evidence status
Controlled study, not a demonstrated production failure. From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage, arXiv:2610.10608, first posted 2026-10-07. Original research. DOI: 10.48550/arXiv.2610.10608. The experimental dates and real-world deployment prevalence have not been established by this research.



