The Tool Said Success. The Evidence Was Missing.
A successful tool call can still return incomplete evidence.
Researchers auditing 15 scientific tools and wrappers in the ToolUniverse environment manually validated 91 silent failures: cases where an invocation appeared successful while information or functionality was missing and neither the user nor the agent was told. Most failures originated in API or wrapper layers and could propagate into downstream scientific outputs that appeared valid.
The continuity failure
The result preserved success status while losing completeness, filtering semantics, and missingness state.
The GovKM interpretation
Retrieval evidence needs contextual reliability metadata. A downstream model cannot reason correctly about evidence gaps if the gap itself does not travel with the returned object.
Continuity path: Source / scientific database → Evidence / API result → Context / completeness and filtering state → Decision → Record / scientific output.
Source
Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan, “Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse,” arXiv, September 21, 2026.



