UNITED STATES, May 27 — Language Models Need Sleep Report GitHub Issue × Title: Content selection saved. Describe the issue below: Description: Submit without GitHub Submit in GitHub Back to arXiv Why HTML? Report Issue Back to Abstract Download PDF Abstract 1 Introduction 2 Related Work 3 Preliminaries 3.
1 Sequence mixers 3. 2 Synthetic reasoning tasks 4 Motivating example: Can attention-SSM hybrid models reason about context they can no longer attend to? 5 LLM Sleep: Offline Recursive Memory Consolidation 6 Experiments 6.
1 Task: Cellular automaton 6. 2 Task: Depo 6. 3 Task: GSM-Infinite 6.
4 Sliding-window eviction 6. 5 Training throughput 7 Discussion and Limitations 8 Conclusion References License: CC BY 4. 0 arXiv:2605.
26099v1 [cs. CL] 25 May 2026 Language Models Need Sleep Sangyun Lee Carnegie Mellon University &Sean McLeish University of Maryland &Tom Goldstein University of Maryland &Giulia Fanti Carnegie Mellon University Correspondence to: sangyunl@andrew. cmu.
edu Abstract Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache. During the sleep, the model performs N N offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule.
During inference, this shifts extra computation to the sleep while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration N N for our models improves performance, with the largest gains on examples that require deeper reasoning.
1 Introduction Large Language Models (LLMs) are commonly based on the transformer architecture [ 51 ] , which stores context in an attention cache and retrieves past tokens as needed. This memory mechanism is central to their performance, but it scales poorly: total attention compute grows quadratically with context length, while cache memory grows linearly. Recent efficient sequence models [ 42 , 18 , 16 , 2 ] mitigate this cost by introducing fixed-size fast weight memories [ 53 , 14 , 43 ] interleaved with full self-attention.
This hybrid design brings together two complementary forms of memory: attention for high-fidelity access to recent tokens, and weight-based memory for compressed information beyond the active context window. Hybrid models are now common among large scale frontier models [ 49 ] . However, scalable memory is not the same as scalable reasoning.
A fast weight memory may support long-range recall [ 42 ] , but it is unclear whether it can support deep computation over tokens that are no longer present in the KV cache. We find that the performance of vanilla SSM-attention hybrid models degrades (under the same token budget) as the required reasoning depth increases even when the amount of information to store is held fixed. This suggests that the bottleneck is not merely memory capacity as suggested by prior work [ 27 , 2 ] , but the amount of computation available for transforming evicted context into a useful internal state.
Sleep. In animals, the transfer from short-term memory to long-term memory is thought to be supported by hippocampal replay [ 33 ] , especially during sleep [ 41 ] ; in this phase, short-term hippocampal memories are reactivated and consolidated into cortical synaptic weights. Sleep makes animals unable to respond to external stimuli, suggesting that it must provide enough cognitive benefit to justify this cost [ 41 ] .
Inspired by these biological processes, we propose a method for transferring context-window memory into persistent weights. When the model’s context window becomes full during inference, the model enters a “sleep” in which it performs multiple forward passes over the accumulated context and recursively updates its fast weights via a learned local rule. As in animal sleep, the model receives no external input tokens during this phase.
After consolidation, the context window is cleared, and the model resumes operation with updated fast weights. During training, the model is optimized end-to-end by backpropagating through the entire process to maximize task performance after sleep. Our architecture is also motivated by results on depth-recurrent or looped neural networks [ 23 , 17 , 4 ] .
Prior work shows that dynamic-depth models can outperform fixed-depth counterparts on sequential reasoning tasks and solve hard problem instances that fixed-depth models cannot by scaling amount of compute spent at prediction. Our key insight is that recurrence can be used not only for prediction but also for memory consolidation. Converting observed tokens into useful weight memory is itself a nontrivial computation, and need not be achievable in a single pass.
Indeed, many learning algorithms, such as gradient descent, improve through iterative weight updates. Thus, allocating more recurrent computation during fast weight formation gives the model more steps to transform context into representations that support later prediction. We find that increasing the depth of recurrence, or sleep duration , improves reasoning after sleep.
Unlike previous looped models, our model does not need to loop at prediction time: the additional computation has already been spent on forming fast weights that support later single-pass prediction. We introduce and evaluate LLM sleep on carefully designed synthetic tasks where a model must answer questions about context that has already been evicted, using only a single forward pass. These synthetic tasks allow us to vary reasoning depth while holding memory load fixed, providing a clean stress test of whether sleep-time computation can convert transient context into fast weights that support later inference.
We summarize our contributions as follows: • In a controlled setting, we show that as the reasoning depth of a problem increases, vanilla State-Space Models (SSMs) such as Gated Delta Nets (GDNs) fail despite having enough fast weight capacity. • We propose an architecture that combines recurrent computation with fast weight memory blocks, and show that increasing the number of recursions for our architecture improves performance over GDNs. We observe the largest gains on problem instances that require the deepest reasoning.
• We further validate the efficacy of our architecture on GSM-Infinite, a natural language math-reasoning dataset, using pre-trained LLM initializations. Overall, these results support the central claim that a sleep-like offline recurrence can organize evicted context into weights to support later reasoning. 2 Related Work Fast weights and linear recurrent neural networks.
Linear recurrent neural networks or SSMs can be viewed as maintaining an online fast weight memory rather than a KV cache which grows quadratically with sequence length. In this view, linear attention corresponds to a recurrent update over a fixed-size, matrix-valued, state, where key-value mappings are written and queried [ 29 , 43 ] . Recent variants improve this memory with delta-rule updates and gates, enabling more selective writing, overwriting, and forgetting [ 54 , 53 , 55 , 14 ] .
These mechanisms underlie recent efficient hybrid language models [ 24 , 39 ] and help explain why linear networks can offer a favorable recall, throughput, and memory tradeoffs. They still struggle with exact copying and retrieval relative to full attention in some cases due to a fixed memory size, as pointed out by prior work [ 2 , 27 ] . Contrary to these works, we show that such models can fail as the required reasoning depth to solve a task increases, even when the amount of information to store is held fixed .
Context compression. There are several methods for processing long contexts at test time by condensing contextual information. Ge et al.
[ 21 ] propose using a language model to compress long contexts into a shorter sequence of hidden states, which are then passed to the language model in place of the original long context. Eyuboglu et al. [ 20 ] use offline self-study to learn a small KV cache that can substitute for the full-context cache.
This line of work shares our goal of spending offline computation once to turn a long context into a compact state that can be reused later. These methods shorten what remains in the attention context, whereas our method transfers evicted context into weight-based memory. Context distillation.
Context distillation [ 46 , 3 ] aims to distill active context into model weights by training a model without it to imitate a contextful teacher [ 46 , 3 , 8 ] , reconstruct it [ 11 ] , predict its continuation [ 8 , 11 ] , or answer questions about it [ 47 , 9 , 8 ] . Instead of doing gradient descent on predefined losses, our method uses a learned recurrent forward pass to transfer context to weights. Test-time training.
Tandon et al. [ 48 ] replace full attention with sliding-window attention and perform test-time gradient updates on a subset of MLP layers. At inference time, their method optimizes a standard cross-entropy loss on the observed context, storing long-range information in temporary parameter updates rather than in a full KV cache.
They perform only one gradient step for distilling each context chunk. By contrast, our method uses a learned recurrent forward pass as the memory-update rule, allowing more flexible forms of consolidation that need not correspond to a one-step gradient descent on a fixed scalar objective. They primarily evaluate perplexity on general web-text data, where retrieval and reasoning demands are entangled; we instead use synthetic tasks that independently control reasoning depth and problem length, showing that additional sleep-time computation is most beneficial when reasoning depth increases.
Zhang et al. [ 56 ] attach a LoRA adapter that updates model weights from the current context chunk and evaluate this approach in a reinforcement-learning setting. Unlike ours, their method updates the weights only once per chunk.
Depth-recurrent models. Increasing the depth of language models is known to increase their expressivity [ 35 ] . Depth-recurrence, is one way to increase depth in transformer models and is one method to make them Turing complete [ 17 ] .
Moreover, the depth of these models can be adaptive [ 23 , 19 , 44 , 5 ] . Recent work has scaled these depth-adaptive language models to large scales, both training from scratch [ 22 , 58 ] and as a post-training objective [ 34 ] . Detailed analyses of how best to train depth recurrent models suggest the recurrent depth should be scaled with training compute [ 40 , 45 ] .
Offline planning. Successful planning in structured environments often requires combining newly-observed information with memories of earlier states. A longstanding view is that animals perform this integration online at choice time [ 50 , 36 ] .
However, integrating distant memories at choice time can be time-consuming, and offline planning during off-task rest can amortize such cost [ 36 ] . Consistent with this view, Momennejad et al. [ 36 ] show that neural evidence of offline replay during rest predicts improved planning performance for human subjects.
Recent work from the machine learning community studies related mechanisms with artificial neural networks. Lin et al. [ 30 ] propose scaling offline compute by letting LLMs generate expected questions from users and precompute quantities needed to solve them.
Chalvidal et al. [ 10 ] train a single-layer network on reinforcement-learning environments and show that recursive Hebbian-like weight updates support fast adaptation. In this paper, we show that recursively updating fast weights during a sleep-like offline phase improves reasoning over evicted context while preserving a strict prediction-phase latency constraint.
3 Preliminaries 3. 1 Sequence mixers Attention. Softmax attention [ 51 ] is a sequence-mixing operation in which each token retrieves information from previous tokens according to query-key similarity.
Source Affidavit & Verification Notice
Transcribed from a verified dispatch by arxiv.org (US). Entered into The Oddyist permanent ledger.
The Readers' Verdict
Mark one. The count is taken at press time.
Circulate this dispatch · · Telegraph · Broadsheet · Post