Building a Functional Machine Translation Corpus for Kpelle
Original authors: Kweku Andoh Yamoah, Jackson Weako, Emmanuel J. Dorley
Original authors: Kweku Andoh Yamoah, Jackson Weako, Emmanuel J. Dorley
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Building a Functional Machine Translation Corpus for Kpelle
Problem Statement
Despite having over one million speakers across Liberia and Guinea, the Kpelle language remains severely underrepresented in Natural Language Processing (NLP) research. As a low-resource language, Kpelle suffers from data scarcity, a lack of standardized orthography, and dialectical variations (specifically between Liberian and Guinean Kpelle). While initiatives like Masakhane, the Lacuna Fund, and Meta's "No Language Left Behind" (NLLB) project have advanced resources for other African languages, Kpelle has been largely excluded, leaving speakers without access to digital tools such as machine translation (MT) or speech recognition.
Methodology
To address this gap, the authors constructed the first publicly available bilingual English-Kpelle corpus. The methodology involved a multi-stage process:
- Data Collection: The dataset was compiled from three primary sources:
- Travel and Tourism: Common phrases sourced from established language learning websites.
- Religious Texts: Excerpts from publicly available translated religious literature.
- Educational Materials: Content from textbooks and dictionaries, including A Learner Directed Approach to Kpelle and the English-Kpelle Dictionary.
- Processing and Alignment:
- Translation: Native Kpelle speakers with linguistic expertise translated English paragraphs lacking corresponding Kpelle text.
- Segmentation: Paragraphs were segmented into sentence pairs to increase granularity for MT tasks.
- Cleaning and Normalization: The raw data underwent spell-checking, duplicate removal, and standardization of the Latin-based orthography. Special attention was paid to diacritical marks to accurately represent Kpelle's tonal system (High, Mid, Low, and contour tones).
- Verification: All translations were reviewed by language experts to ensure grammatical correctness and contextual appropriateness.
- Experimental Setup:
- The authors utilized Meta's NLLB-200 model as a baseline.
- Two versions of the dataset were created: Version 1 (V1) with 1,518 translation pairs (1,667 Kpelle/1,638 English sentences), and Version 2 (V2) with 2,005 translation pairs (2,202 Kpelle/2,167 English sentences), achieved through data augmentation.
- The datasets were split 9:1 for training and testing.
- Models were fine-tuned for 10k, 30k, and 60k steps using the Adafactor optimizer on a Quadro RTX 6000 GPU.
- A custom Kpelle-specific SentencePiece tokenizer was trained to handle out-of-vocabulary tokens.
- Evaluation was performed using sacreBLEU, chrF2++, and precision metrics.
Key Contributions
The paper outlines three primary contributions:
- Dataset Creation: The release of a bilingual English-Kpelle corpus containing 3,234 unique translation pairs in total (across the dataset versions), comprising over 30,000 words. This is currently the largest publicly available resource for Kpelle.
- Methodological Framework: The authors provide a replicable framework for data collection, cleaning, and alignment specifically designed for low-resource languages with limited written traditions and orthographic inconsistencies.
- Benchmarking: The dataset was benchmarked against the NLLB model, establishing performance baselines for Kpelle machine translation.
Results
The fine-tuning experiments yielded the following results:
- Kpelle-to-English (kpe_Latn → eng_Latn): The best performance was achieved with Version 2 of the dataset at 60k training steps, reaching a BLEU score of 30.28 (and a chrF2++ of 44.28).
- English-to-Kpelle (eng_Latn → kpe_Latn): The best result was a BLEU score of 24.46, achieved with Version 1 at 30k steps. While Version 2 showed improvements at higher step counts (reaching 20.79 at 60k steps), it did not surpass the peak performance of Version 1 in this direction.
- Data Augmentation Impact: Expanding the corpus from V1 to V2 generally improved performance, particularly in the Kpelle-to-English direction at higher training steps. However, for English-to-Kpelle translation, the paper notes that gains were modest, with V1 outperforming V2 at the 30k step mark.
- Comparative Analysis: The achieved BLEU scores (ranging roughly 20–30) are consistent with NLLB-200's performance on other low-resource African languages (e.g., Wolof, Luo, Yoruba), though they remain below high-performing languages like Swahili.
Significance and Claims
The authors claim that this work lays the foundation for intensive research into Kpelle and other low-resource Liberian languages. By demonstrating that competitive MT performance is achievable despite the language's low-resource status, the paper underscores the potential for inclusive language technology development.
The significance of the work is framed around:
- Enabling NLP Applications: The dataset serves as a necessary resource for developing not only machine translation but also speech recognition and language modeling tools.
- Community and Inclusion: The project contributes to the broader goal of promoting language diversity and inclusion for African languages, specifically addressing the marginalization of Mande languages in NLP.
- Future Roadmap: The authors emphasize that while the current results are promising, future work must focus on orthographic consistency, expanding domain coverage (specifically in health, education, and religion), and community-driven validation to further improve model performance.
The paper concludes with a call to action for researchers and linguists to collaborate in refining the dataset and developing novel NLP techniques to bridge the digital divide for Kpelle speakers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.
Get the best NLP papers every week.
Trusted by researchers at Stanford, Cambridge, and the French Academy of Sciences.
Check your inbox to confirm your subscription.
Something went wrong. Try again?
No spam, unsubscribe anytime.