Carryover Drafting: Recycling Rejected States for Speculative Decoding
Quick summary
arXiv:2609.14717v1 Announce Type: cross Abstract: Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We find that these discarded hidden states generated during target forward retain useful information about future tokens that
Key takeaways
- arXiv:2609.14717v1 Announce Type: cross Abstract: Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens.
- By construction, verification computes representations for both accepted and rejected tokens.
- Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted.
Why it matters
The importance of “Carryover Drafting: Recycling Rejected States for Speculative Decoding” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments