Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning
Quick summary
arXiv:2606.07602v2 Announce Type: replace-cross Abstract: LLM-based LEGO assembly requires both semantic grounding and physical feasibility. In this paper, we identify a data-induced failure mode, physhack, in which generated assemblies satisfy physical-validity constraints while remaining geometrically misaligned, semantically inconsistent, or poorly calibrated. To address this challenge, we propose a model-based data selection approach that uses only 5% of the original training data while improving semantic and structural alignment across multiple independent evaluators. We further introduce
Key takeaways
- arXiv:2606.07602v2 Announce Type: replace-cross Abstract: LLM-based LEGO assembly requires both semantic grounding and physical feasibility.
- In this paper, we identify a data-induced failure mode, physhack, in which generated assemblies satisfy physical-validity constraints while remaining geometrically misaligned, semantically inconsistent, or poorly calibrated.
- To address this challenge, we propose a model-based data selection approach that uses only 5% of the original training data while improving semantic and structural alignment across multiple independent evaluators.
Why it matters
“Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments