arXiv Artificial Intelligence

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

Quick summary

arXiv:2609.10016v1 Announce Type: cross Abstract: We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-

Key takeaways

  • arXiv:2609.10016v1 Announce Type: cross Abstract: We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk.
  • It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input.
  • In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗