arXiv Artificial Intelligence

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Quick summary

arXiv:2603.10044v3 Announce Type: replace Abstract: Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds. How much do these scaffolds affect model safety as measured by benchmarks? We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce. We conducted 62,808 scored evaluations. How safety is measured matters more than scaffolding does: we find that using a multiple choice vs. open-ended format for otherwise-i

Key takeaways

  • arXiv:2603.10044v3 Announce Type: replace Abstract: Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds.
  • How much do these scaffolds affect model safety as measured by benchmarks?
  • We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce.

Why it matters

“Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗