arXiv Artificial Intelligence

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

Quick summary

arXiv:2609.27336v1 Announce Type: new Abstract: Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judg

Key takeaways

  • arXiv:2609.27336v1 Announce Type: new Abstract: Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing.
  • We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next.
  • CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding.

Why it matters

“CART: Closed-Loop Adaptive Red Teaming for Large Language Models” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗