Same Text, Different Numbers: The Divergence of LLM-Based Measures
Quick summary
arXiv:2609.31013v1 Announce Type: new Abstract: Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across provid
Key takeaways
- arXiv:2609.31013v1 Announce Type: new Abstract: Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables.
- We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk.
- Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs.
Why it matters
“Same Text, Different Numbers: The Divergence of LLM-Based Measures” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments