CallScreenBench: Benchmarking Small Language Models as Phone Secretaries
Quick summary
arXiv:2608.01033v2 Announce Type: replace-cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters
Key takeaways
- arXiv:2608.01033v2 Announce Type: replace-cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible.
- One such task is answering the phone.
- A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle.
Why it matters
The importance of “CallScreenBench: Benchmarking Small Language Models as Phone Secretaries” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments