Available but Unclaimed: An Empirical Study of Human-AI Synergy
Quick summary
arXiv:2609.16793v1 Announce Type: cross Abstract: People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference incre
Key takeaways
- arXiv:2609.16793v1 Announce Type: cross Abstract: People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components.
- In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3.
- Each assisted trial required consultation with the model.
Why it matters
“Available but Unclaimed: An Empirical Study of Human-AI Synergy” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments