Papers
arxiv:2610.01634

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Published on Oct 1
Authors:
,

Abstract

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.01634
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.01634 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.