arXiv Artificial Intelligence

SkillPoison: Progressive Skill Poisoning via Successful Experiences

SkillPoison: Progressive Skill Poisoning via Successful Experiences

Quick summary

arXiv:2610.07645v1 Announce Type: cross Abstract: Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. B

Key takeaways

  • arXiv:2610.07645v1 Announce Type: cross Abstract: Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills.
  • Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills.
  • However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗