arXiv Artificial Intelligence

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

Quick summary

arXiv:2610.03631v1 Announce Type: new Abstract: Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with h

Key takeaways

  • arXiv:2610.03631v1 Announce Type: new Abstract: Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with.
  • We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge.
  • Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with h

Why it matters

“NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗