AI & ML interests
Natural Language Processing, Multilingual NLP, Low-resource languages, Sinhala NLP, Tamil NLP, English NLP, Topic Modeling, Parliamentary Discourse Analysis, Political NLP, Information Extraction, Text Mining, Computational Social Science
Recent Activity
Sri Lanka Parliamentary NLP
Sri Lanka Parliamentary NLP is a research organization for multilingual NLP resources built around Sri Lankan parliamentary discourse. We host datasets, code, and analysis artifacts for Sinhala, Tamil, English, and code-mixed parliamentary text.
Our current work focuses on the Sri Lankan Hansard corpus: official parliamentary debate records published by the Parliament of Sri Lanka. The project supports research in low-resource NLP, trilingual topic modeling, parliamentary discourse analysis, and computational social science.
This is a research organization. It is not an official account, publication channel, or endorsement of the Parliament of Sri Lanka.
Featured Resources
- Dataset: https://huggingface.co/datasets/sl-parliamentary-nlp/sl-parliamentary-hansard-17-26
- Dashboard: https://hansards.vercel.app/
- Code repository: https://github.com/HimathX/lk-hansard-topic-modeling
- Publication venue: https://mercon.uom.lk/
- Primary Hansard source: https://www.parliament.lk/en/business-of-parliament/hansards
Current Dataset
Sri Lanka Parliamentary Hansard (2017-2026)
A trilingual speech-level corpus of Sri Lankan parliamentary debates containing Sinhala, Tamil, English, and code-mixed text.
The dataset includes:
- 19,699 rows in the Hugging Face dataset viewer
- 19,553 topic-labeled speeches
- Parliamentary sessions from 2017 to 2026
- Speaker names, session dates, speech text, and year metadata
- HDBSCAN micro-topic labels
- 30 macro-topic labels derived through hierarchical topic aggregation
The topic-modeling pipeline uses multilingual embeddings, dimensionality reduction, density-based clustering, and macro-topic aggregation to recover interpretable parliamentary themes without supervised topic labels.
Research Context
This organization supports the research paper:
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
Associated with MERCon 2026, this work presents an end-to-end pipeline for extracting and analyzing Sri Lankan parliamentary debate text across Sinhala, Tamil, English, and code-mixed speech.
Intended Uses
Resources hosted by this organization are intended for:
- Multilingual NLP research
- Sinhala and Tamil low-resource language research
- Parliamentary discourse analysis
- Topic modeling and clustering research
- Political agenda and temporal attention analysis
- Computational social science and public-interest research
Responsible Use
Parliamentary text can contain sensitive political claims, allegations, partisan language, and socially biased statements. Topic labels and derived metadata are analytical outputs, not official parliamentary categories.
Please verify against the original Hansard records before using these resources for legal, journalistic, electoral, policy, or other high-stakes decisions.
Citation
If you use the Sri Lanka Parliamentary Hansard dataset or related artifacts, please cite:
@inproceedings{dhanapala2026trilingual,
title={Trilingual Topic Modeling of Sri Lankan Parliamentary Debates},
author={Dhanapala, Himath and Daishika, Haren and Kuruppu, Himandhi and Seneviratne, Sithija and Kavindya, Ashini and Narasinghe, Patalee and Weerasekara, Sandeepa and de Silva, Nisansa and Wickramanayake, Sandareka},
booktitle={Proceedings of the 12th Moratuwa Engineering Research Conference (MERCon 2026)},
year={2026},
publisher={University of Moratuwa, Sri Lanka},
url={https://mercon.uom.lk/}
}
Project Team
- Himath Dhanapala
- Haren Daishika
- Himandhi Kuruppu
- Sithija Seneviratne
- Ashini Kavindya
- Patalee Narasinghe
- Sandeepa Weerasekara
- Nisansa de Silva
- Sandareka Wickramanayake
Contact
For code, reproducibility notes, or issue reports, use the companion repository: