Heads, Tails, and AI Fails: LLMs, Randomness, and Human Judgments
Quick summary
arXiv:2406.00092v2 Announce Type: replace Abstract: Randomness is central to human cognition and to many applications in which large language models are deployed, yet probabilistic token generation does not imply that LLMs can produce unbiased random sequences. We study how contemporary LLMs generate binary random sequences using the classic behavioral-science paradigm of simulated coin flips. Across single flips, 20-flip sequences, n-gram statistics, run lengths, alternation rates, and next-flip predictability, we compare model outputs to both true Bernoulli baselines and human data from prio
Key takeaways
- arXiv:2406.00092v2 Announce Type: replace Abstract: Randomness is central to human cognition and to many applications in which large language models are deployed, yet probabilistic token generation does not imply that LLMs can produce unbiased random sequences.
- We study how contemporary LLMs generate binary random sequences using the classic behavioral-science paradigm of simulated coin flips.
- Across single flips, 20-flip sequences, n-gram statistics, run lengths, alternation rates, and next-flip predictability, we compare model outputs to both true Bernoulli baselines and human data from prio
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments