arXiv Artificial Intelligence

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

Quick summary

arXiv:2609.29948v1 Announce Type: new Abstract: Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, dis

Key takeaways

  • arXiv:2609.29948v1 Announce Type: new Abstract: Prompt injection can degrade benign task performance without eliciting harmful content.
  • Yet many attack objectives depend on task labels or predefined target responses.
  • We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions.

Why it matters

“ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗