Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
Quick summary
arXiv:2605.20641v2 Announce Type: replace-cross Abstract: Inference optimization aims to minimize the latency and resource consumption of LLM inference while preserving output quality, making large-scale deployment practical and cost-effective. However, optimized execution can introduce small numerical inconsistencies from the original model. We reveal that this inconsistency not only causes the model's outputs to diverge, but more critically can introduce hidden backdoors. The backdoor remains dormant under standard unoptimized execution and is activated only when inference optimization is en
Key takeaways
- arXiv:2605.20641v2 Announce Type: replace-cross Abstract: Inference optimization aims to minimize the latency and resource consumption of LLM inference while preserving output quality, making large-scale deployment practical and cost-effective.
- However, optimized execution can introduce small numerical inconsistencies from the original model.
- We reveal that this inconsistency not only causes the model's outputs to diverge, but more critically can introduce hidden backdoors.
Why it matters
“Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments