Antoine Nzeyimana

dblp:278/2652 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-4567-3471ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 44% Machine translation · 26% Deep learning architectures and training · 26%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 77% Health and well-being technologies · 23%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Energy-efficient computing · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
1.012026
Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026
Machine learning › Deep learning architectures and training
positional encoding
1.012026
Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026
Energy-efficient computing
energy harvesting
1.012026
[Emerging Ideas] Towards Practical Metabolic Sensing with Wearable TEGs · MobiSys 2026
Natural language and speech › Language models and text generation › language modeling
low-resource language modeling
0.612022
KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022
Natural language and speech › Language models and text generation › language modeling
morphological language modeling
0.612022
KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022
Natural language and speech › Language models and text generation
multilingual language models
0.612022
KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212022
KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

thermoelectric energy harvesting · 2.0thermal context sensing · 2.0two-tier architecture · 0.6morphological analysis · 0.6BERT · 0.6
YearPublicationVenuePosition
2026 Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation
abstract
Multilingual neural machine translation (MNMT) models degrade in performance as input context length increases, causing positional encoding schemes to misinterpret token distances.Existing absolute and relative positional encodings rely on fixed token indices and implicitly assume uniform semantic density, which breaks down for long-context inputs.We introduce DCARPE, a tokenization-aware adaptive positional encoding that conditions relative positional bias on inputlevel sequence length and fragmentation statistics, allowing the model to reinterpret positional distance when tokenization-induced inflation arises rather than semantic factors.Evaluations on JW300 and out-of-distribution FLORES-200 demonstrate consistent improvements in long-context robustness, achieving gains of up to +10.81 ChrF++ and +8.00 BLEU over baselines.
Khotso Selialia, Antoine Nzeyimana, Fatima M. Anwar 0001
ACL (1)2
2026 [Emerging Ideas] Towards Practical Metabolic Sensing with Wearable TEGs
abstract
This paper explores an emerging sensing modality for wearable metabolic-rate estimation: using a harvesting-optimized thermoelectric generator (TEG) as a metabolic sensor. Rather than optimizing harvesting hardware or studying intermittent runtimes, we treat the electrical energy produced by a commercial wrist-worn TEG as a physiological signal reflecting on-body heat transfer. Because harvested energy is influenced by both metabolic heat and transport effects (e.g., convection, contact, and microclimate), we pair it with lightweight thermal and motion context to assess feasibility.
Jean Bosco Nkurunziza, Antoine Nzeyimana, Luke Arieta, Michael Busa, Jeremy Gummeson
MobiSys2
2026 RangeTag: Adaptive Ultra-Wideband Ranging for Energy Cost-Accuracy Trade-Off in Wearable Systems
Antoine Nzeyimana, Jean Bosco Nkurunziza, Jeremy Gummeson
SECON1
2024 Improving Kinyarwanda Speech Recognition Via Semi-Supervised Learning
abstract
Achieving robust speech recognition for Kinyarwanda remains a challenging task. In this work, we empirically show that using self-supervised pretraining, following a curriculum schedule during fine-tuning and using semi-supervised learning improve speech recognition for Kinyarwanda. Our approach focuses on using public domain data only. A new studio-quality speech dataset is collected from a public website, aligned to text and then used to formulate a simple curriculum learning schedule for training on a larger, noisier dataset. After four generations of semi-supervised learning, our final model achieves 3.2% word error rate (WER) on the new dataset and 15.9% WER on Mozilla Common Voice benchmark. These results improve upon off-the-shelf models that use English language self-supervised representations. Our experiments also indicate that using syllabic rather than character-based tokenization results in better speech recognition for Kinyarwanda.
Antoine Nzeyimana
ICASSP1
2022 KinyaBERT: a Morphology-aware Kinyarwanda Language Model
abstract
Pre-trained language models such as BERT have been successful at tackling many natural language processing tasks. However, the unsupervised sub-word tokenization methods commonly used in these models (e.g., byte-pair encoding - BPE) are sub-optimal at handling morphologically rich languages. Even given a morphological analyzer, naive sequencing of morphemes into a standard BERT architecture is inefficient at capturing morphological compositionality and expressing word-relative syntactic regularities. We address these challenges by proposing a simple yet effective two-tier BERT architecture that leverages a morphological analyzer and explicitly represents morphological compositionality. Despite the success of BERT, most of its evaluations have been conducted on high-resource languages, obscuring its applicability on low-resource languages. We evaluate our proposed method on the low-resource morphologically rich Kinyarwanda language, naming the proposed model architecture KinyaBERT. A robust set of experimental results reveal that KinyaBERT outperforms solid baselines by 2% in F1 score on a named entity recognition task and by 4.3% in average score of a machine-translated GLUE benchmark. KinyaBERT fine-tuning has better convergence and achieves more robust results on multiple tasks even in the presence of translation noise.
Antoine Nzeyimana, Rubungo Andre Niyongabo
ACL (1)1
2020 Morphological disambiguation from stemming data
abstract
Morphological analysis and disambiguation is an important task and a crucial preprocessing step in natural language processing of morphologically rich languages.Kinyarwanda, a morphologically rich language, currently lacks tools for automated morphological analysis.While linguistically curated finite state tools can be easily developed for morphological analysis, the morphological richness of the language allows many ambiguous analyses to be produced, requiring effective disambiguation.In this paper, we propose learning to morphologically disambiguate Kinyarwanda verbal forms from a new stemming dataset collected through crowd-sourcing.Using feature engineering and a feed-forward neural network based classifier, we achieve about 89% non-contextualized disambiguation accuracy.Our experiments reveal that inflectional properties of stems and morpheme association rules are the most discriminative features for disambiguation.
Antoine Nzeyimana
COLING1