arXiv Artificial Intelligence

Render Before Reading: Visual Rendering as a Prompt Injection Defense

Render Before Reading: Visual Rendering as a Prompt Injection Defense

Quick summary

arXiv:2609.36121v1 Announce Type: cross Abstract: Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to o

Key takeaways

  • arXiv:2609.36121v1 Announce Type: cross Abstract: Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior.
  • In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image).
  • We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to o

Why it matters

“Render Before Reading: Visual Rendering as a Prompt Injection Defense” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗