LKValues / README.md
Nethmi14's picture
Update README.md
4e93857 verified
|
Raw History Blame Contribute Delete
3.41 kB
---
license: cc-by-4.0
task_categories:
- text-generation
- question-answering
language:
- si
- en
tags:
- societalValues
- cultural-alignment
- multilingual
- low-resource
- instruction-tuning
- benchmark
- llm-evaluation
pretty_name: LKValues
size_categories:
- 100K<n<1M
---
# LKValues
LKValues is a survey-grounded Sinhala–English resource suite for studying the alignment of large language models with Sri Lankan societal values.
It contains two complementary resources:
- **LKvaluesIT** — an instruction-tuning dataset for training models to generate value-grounded explanations.
- **LKvaluesBench** — an evaluation benchmark for testing controlled value-sensitive judgment.
LKValues is based on 40 societal values retained through a trilingual Sinhala–Tamil–English survey of 205 Sri Lankan respondents.
These values should not be interpreted as an official, exhaustive, or universally applicable definition of Sri Lankan values.
## LKvaluesIT — Instruction-Tuning Dataset
LKvaluesIT is the training resource of LKValues.
It is designed for **value-grounded generation and supervised instruction tuning**.
Given a situation and a target Sri Lankan societal value, the model is expected to generate a short explanation describing how the situation supports that value.
The public release contains the full bilingual instruction datasets:
| Language | Instances |
| --- | ---: |
| English | ~150,000 |
| Sinhala | ~150,000 |
| Total | ~300,000 |
Each instance contains:
```json
{
"instruction": "A situation or scenario grounded in a Sri Lankan context",
"input": "Target Sri Lankan societal value",
"output": "A short value-grounded explanation"
}
```
## LKvaluesBench — Benchmark Dataset
LKvaluesBench is the evaluation resource of LKValues.
It is designed to evaluate **controlled value-sensitive judgment**.
Each benchmark item contains:
```json
{
"ID": 2001,
"question": "Question or value-sensitive scenario",
"Statement_A": "First candidate statement",
"Statement_B": "Second candidate statement",
"CorrectChoice": "A",
"primary_value": "Health",
"Source": "SinhalaMMLU"
}
```
LKvaluesBench contains **1,000 aligned benchmark instances** in English and Sinhala.
Models should return one of the following labels:
| Label | Meaning |
| --- | --- |
| `A` | Only Statement A is justifiable |
| `B` | Only Statement B is justifiable |
| `BOTH` | Both statements are justifiable |
| `0` | Neither statement is justifiable |
## Dataset Files
```text
LKvaluesIT/
├── train_en.json
└── train_si.json
LKvaluesBench/
├── test_en.json
└── test_si.json
```
## Languages
- English
- Sinhala
## Citation
If you use LKValues, LKvaluesIT, or LKvaluesBench in your research, please cite:
```bibtex
@misc{muthugala2026lkvaluesaligninglargelanguage,
title={LKValues: Aligning Large Language Models with Sri Lankan Societal Values},
author={Nethmi Muthugala and Supryadi and Surangika Ranathunga and Nisansa de Silva and Ruijie Tao and Ovindu Gunatunga and Pengyun Zhu and Shaowei Zhang and Jingting Zheng and Deyi Xiong},
year={2026},
eprint={2607.20410},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.20410},
}
```
Paper: https://arxiv.org/abs/2607.20410
## Contact
**Nethmi Muthugala**
neth.muthugala1@gmail.com
**Deyi Xiong**
dyxiong@tju.edu.cn