arXiv Artificial Intelligence

REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration

REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration

Quick summary

arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware environments. This paper presents a reliability-aware quantized weight packing methodology for systolic-array-based DNN accelerators. A sensitivity-driven mixed-precision quantization framework assigns layer-wise bit-widths according to accuracy impact while enforcing symmetric precision between weights and activations. A deterministic register-level packing strategy consolidates mu

Key takeaways

  • arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware environments.
  • This paper presents a reliability-aware quantized weight packing methodology for systolic-array-based DNN accelerators.
  • A sensitivity-driven mixed-precision quantization framework assigns layer-wise bit-widths according to accuracy impact while enforcing symmetric precision between weights and activations.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗