arXiv Artificial Intelligence

WorldBench: Evaluating LLMs on Three.js Voxel World Generation

WorldBench: Evaluating LLMs on Three.js Voxel World Generation

Quick summary

arXiv:2610.10622v1 Announce Type: cross Abstract: Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents. We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.

Key takeaways

  • arXiv:2610.10622v1 Announce Type: cross Abstract: Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable.
  • Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source.
  • On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents.

Why it matters

“WorldBench: Evaluating LLMs on Three.js Voxel World Generation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗