VLDB 2026 Research / reviewers in the wild / expert
Baraka William Nyamtiga
dblp:329/3698
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0002-3291-067XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
1.0 | 1 | 2026 | Building Foundations for Information Retrieval in Extremely Low-Resource Ethnic Languages: Evidence from Tanzania · SIGIR 2026 |
Information retrieval › cross-language information retrieval
low-resource language retrieval |
1.0 | 1 | 2026 | Building Foundations for Information Retrieval in Extremely Low-Resource Ethnic Languages: Evidence from Tanzania · SIGIR 2026 |
Information retrieval › document retrieval
spoken document retrieval |
0.3 | 1 | 2026 | Building Foundations for Information Retrieval in Extremely Low-Resource Ethnic Languages: Evidence from Tanzania · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
speech data collection · 1.0parallel corpus construction · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building Foundations for Information Retrieval in Extremely Low-Resource Ethnic Languages: Evidence from TanzaniaabstractDespite major advances in information retrieval (IR) and language technologies, speakers of extremely low-resource ethnic community languages (ECLs) remain largely excluded from digital information access. With Tanzania having over 120 languages, yet most IR systems and digital assistants primarily support English and, to a limited extent, Kiswahili. This disproportionately affects rural populations (over 65%), where ECLs are used alongside Kiswahili in daily life. This paper presents the ongoing Ethnic Voices initiative, which addresses this gap by building foundational text and speech resources explicitly relevant to IR research in extremely low-resource multilingual settings. The project targets six ECLs in Tanzania, with current data collection completed for Hehe and Luguru. Data collection from the two ECLs has produced over 800 Kiswahili-ECL parallel sentences per language and over 95 hours of speech. Preliminary survey results show strong demand for multilingual and voice-based access to information. It further identifies language barriers as a primary limitation of existing English-centric IR systems. This also paper discusses the IR implications of these findings and outlines directions toward inclusive retrieval and evaluation in extremely low-resource languages. Beyond resource creation, we position the collected dataset as IR-enabling infrastructure that supports the formulation of retrieval tasks, the development of baseline systems, and the design of evaluation protocols in benchmark-scarce, extremely low-resource multilingual settings. Joseph P. Telemala, Onesmo S. Nyinondi, Neema Nicodemus Lyimo, Baraka William Nyamtiga, Farian Severine Ishengoma, Wema L. Msigwa |
SIGIR | 4 |