arXiv Artificial Intelligence

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

Quick summary

arXiv:2509.22360v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with

Key takeaways

  • arXiv:2509.22360v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web.
  • Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation.
  • To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with

Why it matters

The importance of “CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗