arXiv Artificial Intelligence

MIRA: A Bilingual Benchmark for Medical Information Response Audit

MIRA: A Bilingual Benchmark for Medical Information Response Audit

Quick summary

arXiv:2605.28025v2 Announce Type: replace Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answe

Key takeaways

  • arXiv:2605.28025v2 Announce Type: replace Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question.
  • To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals.
  • MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions.

Why it matters

“MIRA: A Bilingual Benchmark for Medical Information Response Audit” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗