A skill may appear benign when inspected yet execute different behavior through a cached Python bytecode artifact. The distinction is not an ordinary scanner miss: the scanner and interpreter can operate on different representations of the supposed same skill.
Study and evidence
In the controlled study PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners, researchers paired benign Python source with substituted cache-resident bytecode across 100 constructed skills and seven scanning systems. The reported 94–100% attack success concerns that experimental setup, not observed exploitation of named production organizations. Execution-aware validation identified the 100 tested substitutions.
The broken continuity relationship
The topology failure arises at Source → Evidence → Authority → Action. A scan report authenticates the source text it inspected; the runtime may select a different bytecode payload. Admission based on source inspection then grants action authority to an artifact outside the inspected evidence. This is an inference from the tested mechanism, not proof of a specific operational compromise.
GovKM interpretation
GovKM has examined the difference between a source artifact and its executed effect in The User Could Author a Workflow. The Template Engine Gave Them the Host. The new study adds a precise inspection/execution identity check. An AIL record should distinguish inspected artifact, resolved runtime artifact, loader behavior, and the authority decision binding them. It should not claim that a scan certified runtime behavior when it certified a different object.
Scope, limitations, and tests
The evidence supports a synthetic regression challenge: present readable source and cached executable content with divergent behavior, and require the verifier either to bind the runtime-selected artifact to the admission decision or refuse to treat the scan as sufficient. No new GovKM invariant is ratified. Generalization is limited by Python cache-loading conditions and scanner configurations tested.
Source and evidence status
Controlled study, not a demonstrated production failure. PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners, arXiv:2610.10612, first posted 2026-10-07. Original research. DOI: 10.48550/arXiv.2610.10612. The experimental dates and real-world deployment prevalence have not been established by this research.



