VLDB 2026 Research / reviewers in the wild / expert
John Bosco Mugeni
dblp:319/9422
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-2464-413XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AssistEM: Domain Instruction Tuning for Enhanced Entity Matching
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono |
PAKDD (5) | 1 |
| 2024 | MultiMatch: Low-Resource Generalized Entity Matching Using Task-Conditioned Hyperadapters in Multitask Learning
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono |
DaWaK | 1 |
| 2023 | AdapterEM: Pre-trained Language Model Adaptation for Generalized Entity Matching using Adapter-tuningabstractEntity Matching (EM) involves identifying different data representations referring to the same entity from multiple data sources and is typically formulated as a binary classification problem. It is a challenging problem in data integration due to the heterogeneity of data representations. State-of-the-art solutions have adopted NLP techniques based on pre-trained language models (PrLMs) via the fine-tuning paradigm, however, sequential fine-tuning of overparameterized PrLMs can lead to catastrophic forgetting, especially in low-resource scenarios. In this study, we propose a parameter-efficient paradigm for fine-tuning PrLMs based on adapters, small neural networks encapsulated between layers of a PrLM, by optimizing only the adapter and classifier weights while the PrLMs parameters are frozen. Adapter-based methods have been successfully applied to multilingual speech problems achieving promising results, however, the effectiveness of these methods when applied to EM is not yet well understood, particularly for generalized EM with heterogeneous data. Furthermore, we explore using (i) pre-trained adapters and (ii) invertible adapters to capture token-level language representations and demonstrate their benefits for transfer learning on the generalized EM benchmark. Our results show that our solution achieves comparable or superior performance to full-scale PrLM fine-tuning and prompt-tuning baselines while utilizing a significantly smaller computational footprint of the PrLM parameters. John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono |
IDEAS | 1 |