arXiv Artificial Intelligence

Authority-Preserving Evaluation of Medical Vision-Language Assistants

Authority-Preserving Evaluation of Medical Vision-Language Assistants

Quick summary

arXiv:2609.22302v1 Announce Type: cross Abstract: Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context. Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score. We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declar

Key takeaways

  • arXiv:2609.22302v1 Announce Type: cross Abstract: Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context.
  • Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score.
  • We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declar

Why it matters

“Authority-Preserving Evaluation of Medical Vision-Language Assistants” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗