ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation
Quick summary
arXiv:2609.29948v1 Announce Type: new Abstract: Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, dis
Key takeaways
- arXiv:2609.29948v1 Announce Type: new Abstract: Prompt injection can degrade benign task performance without eliciting harmful content.
- Yet many attack objectives depend on task labels or predefined target responses.
- We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions.
Why it matters
“ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments