GovKM
Controlled Study — Agent Skill Supply Chain

The Scanner Approved the Source. Python Ran the Cache.

A controlled agent-skill study shows why inspecting benign Python source cannot certify different cached bytecode selected at execution.
GovKM reconstruction illustration representing the gap between inspected source files and executed Python bytecode.
Expand image

A skill may appear benign when inspected yet execute different behavior through a cached Python bytecode artifact. The distinction is not an ordinary scanner miss: the scanner and interpreter can operate on different representations of the supposed same skill.

Study and evidence

In the controlled study PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners, researchers paired benign Python source with substituted cache-resident bytecode across 100 constructed skills and seven scanning systems. The reported 94–100% attack success concerns that experimental setup, not observed exploitation of named production organizations. Execution-aware validation identified the 100 tested substitutions.

The broken continuity relationship

The topology failure arises at Source → Evidence → Authority → Action. A scan report authenticates the source text it inspected; the runtime may select a different bytecode payload. Admission based on source inspection then grants action authority to an artifact outside the inspected evidence. This is an inference from the tested mechanism, not proof of a specific operational compromise.

GovKM interpretation

GovKM has examined the difference between a source artifact and its executed effect in The User Could Author a Workflow. The Template Engine Gave Them the Host. The new study adds a precise inspection/execution identity check. An AIL record should distinguish inspected artifact, resolved runtime artifact, loader behavior, and the authority decision binding them. It should not claim that a scan certified runtime behavior when it certified a different object.

Scope, limitations, and tests

The evidence supports a synthetic regression challenge: present readable source and cached executable content with divergent behavior, and require the verifier either to bind the runtime-selected artifact to the admission decision or refuse to treat the scan as sufficient. No new GovKM invariant is ratified. Generalization is limited by Python cache-loading conditions and scanner configurations tested.

Source and evidence status

Controlled study, not a demonstrated production failure. PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners, arXiv:2610.10612, first posted 2026-10-07. Original research. DOI: 10.48550/arXiv.2610.10612. The experimental dates and real-world deployment prevalence have not been established by this research.

RELATED KNOWLEDGE

Continue Exploring

Explore related research, framework domains, and continuity concepts.
CONTINUE WITH THE FRAMEWORK

Explore the continuity relationships that support trustworthy organizational intelligence.

Continue through the GovKM Framework to examine the doctrine, knowledge, and implementation guidance behind Organizational Continuity.