arXiv Artificial Intelligence

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

Quick summary

arXiv:2608.15594v1 Announce Type: new Abstract: Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, ev

Key takeaways

  • arXiv:2608.15594v1 Announce Type: new Abstract: Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails.
  • Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics.
  • We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning.

Why it matters

“TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗