arXiv Artificial Intelligence

Risky Business: Measuring The Faithfulness-Safety Tension

Risky Business: Measuring The Faithfulness-Safety Tension

Quick summary

arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning. We demonstrate that this counterbalance exists in current Large Reasoning Models (LRMs), and show ways in which it can be addressed. We introduce HazMart, a human-written dataset set in an autonomous AI shopkeeper scenario. Un

Key takeaways

  • arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring.
  • However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace.
  • We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning.

Why it matters

“Risky Business: Measuring The Faithfulness-Safety Tension” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗