VLDB 2026 Research / reviewers in the wild / expert
Antoine Nzeyimana
dblp:278/2652
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-4567-3471ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 44% Machine translation · 26% Deep learning architectures and training · 26% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 77% Health and well-being technologies · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Energy-efficient computing · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
1.0 | 1 | 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
positional encoding |
1.0 | 1 | 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026 |
Energy-efficient computing
energy harvesting |
1.0 | 1 | 2026 | [Emerging Ideas] Towards Practical Metabolic Sensing with Wearable TEGs · MobiSys 2026 |
Natural language and speech › Language models and text generation › language modeling
low-resource language modeling |
0.6 | 1 | 2022 | KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022 |
Natural language and speech › Language models and text generation › language modeling
morphological language modeling |
0.6 | 1 | 2022 | KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022 |
Natural language and speech › Language models and text generation
multilingual language models |
0.6 | 1 | 2022 | KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.2 | 1 | 2022 | KinyaBERT: a Morphology-aware Kinyarwanda Language Model · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
thermoelectric energy harvesting · 2.0thermal context sensing · 2.0two-tier architecture · 0.6morphological analysis · 0.6BERT · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine TranslationabstractMultilingual neural machine translation (MNMT) models degrade in performance as input context length increases, causing positional encoding schemes to misinterpret token distances.Existing absolute and relative positional encodings rely on fixed token indices and implicitly assume uniform semantic density, which breaks down for long-context inputs.We introduce DCARPE, a tokenization-aware adaptive positional encoding that conditions relative positional bias on inputlevel sequence length and fragmentation statistics, allowing the model to reinterpret positional distance when tokenization-induced inflation arises rather than semantic factors.Evaluations on JW300 and out-of-distribution FLORES-200 demonstrate consistent improvements in long-context robustness, achieving gains of up to +10.81 ChrF++ and +8.00 BLEU over baselines. Khotso Selialia, Antoine Nzeyimana, Fatima M. Anwar 0001 |
ACL (1) | 2 |
| 2026 | [Emerging Ideas] Towards Practical Metabolic Sensing with Wearable TEGsabstractThis paper explores an emerging sensing modality for wearable metabolic-rate estimation: using a harvesting-optimized thermoelectric generator (TEG) as a metabolic sensor. Rather than optimizing harvesting hardware or studying intermittent runtimes, we treat the electrical energy produced by a commercial wrist-worn TEG as a physiological signal reflecting on-body heat transfer. Because harvested energy is influenced by both metabolic heat and transport effects (e.g., convection, contact, and microclimate), we pair it with lightweight thermal and motion context to assess feasibility. Jean Bosco Nkurunziza, Antoine Nzeyimana, Luke Arieta, Michael Busa, Jeremy Gummeson |
MobiSys | 2 |
| 2026 | RangeTag: Adaptive Ultra-Wideband Ranging for Energy Cost-Accuracy Trade-Off in Wearable Systems
Antoine Nzeyimana, Jean Bosco Nkurunziza, Jeremy Gummeson |
SECON | 1 |
| 2024 | Improving Kinyarwanda Speech Recognition Via Semi-Supervised LearningabstractAchieving robust speech recognition for Kinyarwanda remains a challenging task. In this work, we empirically show that using self-supervised pretraining, following a curriculum schedule during fine-tuning and using semi-supervised learning improve speech recognition for Kinyarwanda. Our approach focuses on using public domain data only. A new studio-quality speech dataset is collected from a public website, aligned to text and then used to formulate a simple curriculum learning schedule for training on a larger, noisier dataset. After four generations of semi-supervised learning, our final model achieves 3.2% word error rate (WER) on the new dataset and 15.9% WER on Mozilla Common Voice benchmark. These results improve upon off-the-shelf models that use English language self-supervised representations. Our experiments also indicate that using syllabic rather than character-based tokenization results in better speech recognition for Kinyarwanda. Antoine Nzeyimana |
ICASSP | 1 |
| 2022 | KinyaBERT: a Morphology-aware Kinyarwanda Language ModelabstractPre-trained language models such as BERT have been successful at tackling many natural language processing tasks. However, the unsupervised sub-word tokenization methods commonly used in these models (e.g., byte-pair encoding - BPE) are sub-optimal at handling morphologically rich languages. Even given a morphological analyzer, naive sequencing of morphemes into a standard BERT architecture is inefficient at capturing morphological compositionality and expressing word-relative syntactic regularities. We address these challenges by proposing a simple yet effective two-tier BERT architecture that leverages a morphological analyzer and explicitly represents morphological compositionality. Despite the success of BERT, most of its evaluations have been conducted on high-resource languages, obscuring its applicability on low-resource languages. We evaluate our proposed method on the low-resource morphologically rich Kinyarwanda language, naming the proposed model architecture KinyaBERT. A robust set of experimental results reveal that KinyaBERT outperforms solid baselines by 2% in F1 score on a named entity recognition task and by 4.3% in average score of a machine-translated GLUE benchmark. KinyaBERT fine-tuning has better convergence and achieves more robust results on multiple tasks even in the presence of translation noise. Antoine Nzeyimana, Rubungo Andre Niyongabo |
ACL (1) | 1 |
| 2020 | Morphological disambiguation from stemming dataabstractMorphological analysis and disambiguation is an important task and a crucial preprocessing step in natural language processing of morphologically rich languages.Kinyarwanda, a morphologically rich language, currently lacks tools for automated morphological analysis.While linguistically curated finite state tools can be easily developed for morphological analysis, the morphological richness of the language allows many ambiguous analyses to be produced, requiring effective disambiguation.In this paper, we propose learning to morphologically disambiguate Kinyarwanda verbal forms from a new stemming dataset collected through crowd-sourcing.Using feature engineering and a feed-forward neural network based classifier, we achieve about 89% non-contextualized disambiguation accuracy.Our experiments reveal that inflectional properties of stems and morpheme association rules are the most discriminative features for disambiguation. Antoine Nzeyimana |
COLING | 1 |