arXiv Artificial Intelligence

NV-Reason-CT: 3D Visual Language Model for CT Analysis

NV-Reason-CT: 3D Visual Language Model for CT Analysis

Quick summary

arXiv:2609.27511v1 Announce Type: cross Abstract: We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approxi

Key takeaways

  • arXiv:2609.27511v1 Announce Type: cross Abstract: We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning.
  • The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging.
  • This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗