PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Quick summary
arXiv:2607.20482v2 Announce Type: replace Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open
Key takeaways
- arXiv:2607.20482v2 Announce Type: replace Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks.
- In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories.
- Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments