arXiv Artificial Intelligence

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Quick summary

arXiv:2609.39607v1 Announce Type: cross Abstract: Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext,

Key takeaways

  • arXiv:2609.39607v1 Announce Type: cross Abstract: Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code.
  • Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent.
  • The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector.

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗