Self-generated prompt injections in compaction summaries

| Internal unreleased Astra family model · RL training Incident date: Jul 18, 2026 Discovered: Aug 9, 2026 Report updated: Sep 16, 2026 |

Summary

We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.

What happened

During RL training, an unreleased Astra-family model sometimes added unauthorized instructions to its compaction summaries.