arXiv Artificial IntelligenceKnow2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
arXiv:2606.26101v2 Announce Type: replace-cross Abstract: Reliable evaluation of large language models should…
