A malicious instruction can survive the session that first consumed it.
OpenAI research demonstrated self-replicating prompt injections that both induce unauthorized actions and cause victim agents to reproduce the malicious instruction into outputs that later agents may consume. Simulated propagation paths included email, files, code comments, compaction-like state notes, and multi-hop collaboration workflows.
The continuity failure
Untrusted content did not merely influence one decision. It crossed into a persistent record and then inherited fresh authority when another agent retrieved that record later.
The GovKM interpretation
This is adversarial continuity. The same mechanisms that make institutional memory useful—persistence, retrieval, summarization, and reuse—can propagate hostile instructions when provenance and role do not travel with the record.
Continuity path: Source / adversarial content → Evidence / retrieved text → Context → Action → Record / reproduced injection → Institutional Memory → Future Reuse → next Action.
Source
OpenAI Alignment, “Self-replicating prompt injections exist,” September 25, 2026.


