arXiv Artificial Intelligence

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

Quick summary

arXiv:2605.29843v2 Announce Type: replace-cross Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains highly sensitive to activation outliers and anisotropic weight curvature. Existing incoherence-based PTQ methods mitigate this issue with fixed randomized Hadamard transforms (RHTs), which improve quantization robustness but cannot adapt the rotated basis to the layer, calibration distribution, or quantizer. We introduce HARP (Hadamard-preconditioned Adaptive Rotation Processor), a learna

Key takeaways

  • arXiv:2605.29843v2 Announce Type: replace-cross Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints.
  • However, extreme low-bit quantization remains highly sensitive to activation outliers and anisotropic weight curvature.
  • Existing incoherence-based PTQ methods mitigate this issue with fixed randomized Hadamard transforms (RHTs), which improve quantization robustness but cannot adapt the rotated basis to the layer, calibration distribution, or quantizer.

Why it matters

“HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗