One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction
Quick summary
arXiv:2609.13500v1 Announce Type: cross Abstract: How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing $\mu$ for the resulting energy spectrum, we prove the minimax law $\mathfrak R^*_{\rm value}(n,B)\asymp_R \Phi_\mu(n^{-1})+\Phi_\mu(\tau_B), \Phi_\mu(t)=\int\min\{x,t\}\,\mu(\mathrm dx),$ for $n$ prediction blocks and a learned state with at most $2^B
Key takeaways
- arXiv:2609.13500v1 Announce Type: cross Abstract: How much learned memory is needed to benefit from more data?
- We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source.
- Each coordinate contributes its query probability times the squared radius of its unknown logit.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction” may reshape data collection, model training, output accountability and market access.

Member comments