arXiv Artificial Intelligence

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

Quick summary

arXiv:2608.20365v1 Announce Type: cross Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology. We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling. A hybrid semantic-lexical extension, BiTopic, is further explored to improv

Key takeaways

  • arXiv:2608.20365v1 Announce Type: cross Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology.
  • We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling.
  • A hybrid semantic-lexical extension, BiTopic, is further explored to improv

Why it matters

The importance of “Trilingual Topic Modeling of Sri Lankan Parliamentary Debates” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗