FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Quick summary
arXiv:2609.39863v1 Announce Type: new Abstract: Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current evaluations, however, rely on rigid, single-turn tests or fixed scripts that fail to capture these natural dynamics. Furthermore, thes
Key takeaways
- arXiv:2609.39863v1 Announce Type: new Abstract: Large language models frequently fail to balance staying truthful with being supportive.
- They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views.
- In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time.
Why it matters
“FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments