The KV cache as an agent runtime

The KV cache as an agent runtime The KV cache as an agent runtime Thoughts and our recent work on how sharing and scheduling KV-cache state lets pretrained LLMs observe, reason, and act concurrently without additional training How KV-cache manipulations can make LLMs more interactive The link has been copied to clipboard One of the next frontiers in AI systems is interactivity. An interactive model must be able to receive new information while it is already computing, revise its trajectory without restarting from scratch, emit useful partial actions, and coordinate processes that progress at different rates: perception, reasoning, communication, acting, and tool use. This requirement applies to language assistants, but it becomes unavoidable for multimodal systems embedded in games, robots, live video, operating systems, and other continuously evolving environments. A model that stops observing whenever it reasons cannot be interactive by design. A sequential model processes a fixed input while decoding. A shared-state runtime lets observation, reasoning, and output progress concurrently as new information arrives. We have systems today that are interactive such as the Wan-Streamer models [ 1 ] or Thinking Machines Lab's Interaction Models [ 2 ] . Wan-Streamer handles perception, response timing, speech, and visual generation, Interaction Models replace turn-based exchange with time-aligned micro-turns and combine a real-time interaction model with asynchronous background reasoning. Most existing solutions agree that the common paradigm of sequential tool calls, thinking, input and output streams is not enough for interactive systems. They also illustrate one path toward solving the problem – changing the model through pretraining, post-training, new data formats, block-causal attention, streaming encoders, or additional policy heads. We explore how much of this behavior can be achieved at inference time, without changing model weights. Our position Our position is that a significant part of interactivity is also an inference-runtime problem . Existing pretrained models already expose reusable execution state. By changing how this state is partitioned, ordered, exposed, and scheduled, an inference engine can implement interaction protocols that were absent from the original training procedure.