A correct conclusion is not the same as an evidenced conclusion. In institutional forensic reconstruction, the difference is material: a record may support the right answer in retrospect while the evidence actually available to the reviewer cannot establish which record was cited.
The controlled study Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs tested 64 mechanically checkable cases from saved AgentDojo Banking executions. When the mapping from execution-local citation identifiers to source records was withheld, one reader made unsupported assertions in 26 of 28 binding-dependent cases, although 22 happened to agree with the complete reference. Explicit identifier mappings improved grounded reconstruction; a deterministic comparator could resolve or abstain.
What failed across the continuity topology
Source → Evidence → Authority became disconnected: citation identifiers survived in the evidence packet, but not their binding to particular records. An authentic-looking citation cannot justify attribution if the reader cannot determine which saved artifact it denotes. The model's confidence or a matching value elsewhere in the record is not a substitute for that missing relationship.
Why it advances GovKM
This finding extends The Evidence Was All There. The Retriever Never Reassembled It. The prior case concerns evidence retrieval and assembly. The new study isolates a more exacting distinction: retrieval may succeed, and the answer may be right, yet the binding between a conclusion and its cited record can remain unproven.
For an Agent Interaction Ledger (AIL), the proposed research implication is to preserve an execution-scoped relationship among claim, citation identifier, original record, and derivation. If the required binding is absent, the system should record uncertainty or abstain rather than manufacture evidentiary authority. This is a GovKM interpretation, not a newly ratified invariant.
Evidence limits and source
This is a controlled forensic benchmark using simulated banking-agent traces, not a reported real-bank production incident. The study evaluates recorded citations, not private model thought processes. See Taehyeon Yun et al., Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs, arXiv:2610.09581, October 7, 2026; full paper.



