CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
Quick summary
arXiv:2509.22360v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with
Key takeaways
- arXiv:2509.22360v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web.
- Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation.
- To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with
Why it matters
The importance of “CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments