arXiv Artificial Intelligence

Calibration-risk routing for controlled world-model adaptation

Calibration-risk routing for controlled world-model adaptation

Quick summary

arXiv:2610.01001v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy u

Key takeaways

  • arXiv:2610.01001v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data.
  • We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk.
  • A learned confidence signal and deterministic validity predicates weight one-step imagined policy u

Why it matters

“Calibration-risk routing for controlled world-model adaptation” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗