arXiv Artificial Intelligence

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

Quick summary

arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, pruning protects weights by statistical salience rather than by their contribution to correct reasoning,

Key takeaways

  • arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead.
  • Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance.
  • Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning.

Why it matters

“OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗