Thought without systematicity? Evaluating reasoning models on rule induction tasks
Quick summary
arXiv:2609.13948v1 Announce Type: cross Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equiv
Key takeaways
- arXiv:2609.13948v1 Announce Type: cross Abstract: A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept.
- Do reasoning models robustly exhibit such systematicity?
- If so, we would expect consistent performance on structurally equivalent variants of the same task.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments