Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure
Quick summary
arXiv:2609.27763v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) rely on explicit chain-of-thought (CoT) reasoning and large context windows to achieve strong performance on complex tasks, but these features also introduce new attack surfaces. We show that LRMs' reasoning processes can be systematically steered by prepending counter-aligned few-shot conversations containing explicit CoT traces, leading to unsafe generations on harmful queries and unwarranted refusals on benign ones. We formalize this attack as SRCF (Steering Reasoning via Counter-Aligned Few-shot Conversations) th
Key takeaways
- arXiv:2609.27763v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) rely on explicit chain-of-thought (CoT) reasoning and large context windows to achieve strong performance on complex tasks, but these features also introduce new attack surfaces.
- We show that LRMs' reasoning processes can be systematically steered by prepending counter-aligned few-shot conversations containing explicit CoT traces, leading to unsafe generations on harmful queries and unwarranted refusals on benign ones.
- We formalize this attack as SRCF (Steering Reasoning via Counter-Aligned Few-shot Conversations) th
Why it matters
“Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments