AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
Quick summary
arXiv:2601.06022v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors. Inference-time ensembling provides a practical way to combine these capabilities without retraining. However, existing ensemble approaches suffer from fundamental limitations. Most rely on fixed fusion granularity, which lacks the flexibility required for mid-generation adaptation and fails to adapt to different generation characteristics across tasks. To address these challenges, we pro
Key takeaways
- arXiv:2601.06022v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors.
- Inference-time ensembling provides a practical way to combine these capabilities without retraining.
- However, existing ensemble approaches suffer from fundamental limitations.
Why it matters
“AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments