Uncertainty-Aware Calibrated Clinical Text Classification with Large Language Models
Quick summary
arXiv:2509.19375v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect patient care. Existing black-box uncertainty methods attach a confidence score to a fixed LLM prediction using softmax probabilities, verbalised confidence, prompt agreement, or generation consistency. These signals are often poorly calibrated and offer no mechanism for combining model evidence with prior clinical belief. We instead formulate closed-set clinical classification as likelihood-free poster
Key takeaways
- arXiv:2509.19375v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect patient care.
- Existing black-box uncertainty methods attach a confidence score to a fixed LLM prediction using softmax probabilities, verbalised confidence, prompt agreement, or generation consistency.
- These signals are often poorly calibrated and offer no mechanism for combining model evidence with prior clinical belief.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments