What Pretraining and Midtraining Make Learnable from Rewards?
Quick summary
arXiv:2609.38446v1 Announce Type: cross Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and r
Key takeaways
- arXiv:2609.38446v1 Announce Type: cross Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined.
- We study how pretraining and midtraining supply the information and computation that make reward adaptation effective.
- In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers.
Why it matters
“What Pretraining and Midtraining Make Learnable from Rewards?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments