arXiv Artificial Intelligence

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

Quick summary

arXiv:2608.01033v2 Announce Type: replace-cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters

Key takeaways

  • arXiv:2608.01033v2 Announce Type: replace-cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible.
  • One such task is answering the phone.
  • A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle.

Why it matters

The importance of “CallScreenBench: Benchmarking Small Language Models as Phone Secretaries” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗