Render Before Reading: Visual Rendering as a Prompt Injection Defense
Quick summary
arXiv:2609.36121v1 Announce Type: cross Abstract: Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to o
Key takeaways
- arXiv:2609.36121v1 Announce Type: cross Abstract: Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior.
- In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image).
- We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to o
Why it matters
“Render Before Reading: Visual Rendering as a Prompt Injection Defense” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments