VLDB 2026 Research / reviewers in the wild / expert
Khuyagbaatar Batsuren
dblp:157/7623
· DBLP profile ↗
7ranked-venue papers
3as first author
3since 2021 · last 2022
0000-0002-6819-5444ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 54% Knowledge representation and reasoning · 46% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › lexical resources
lexical resource construction |
0.4 | 1 | 2019 | CogNet: A Large-Scale Cognate Database · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
lexical knowledge base |
0.3 | 1 | 2017 | Understanding and Exploiting Language Diversity · IJCAI 2017 |
Natural language and speech › Information extraction and text analysis
lexical semantics |
0.3 | 1 | 2017 | Understanding and Exploiting Language Diversity · IJCAI 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology |
0.3 | 1 | 2017 | Understanding and Exploiting Language Diversity · IJCAI 2017 |
Computational social science and digital humanities › computational linguistics
cognate identification |
0.1 | 1 | 2019 | CogNet: A Large-Scale Cognate Database · ACL (1) 2019 |
Computational social science and digital humanities
historical linguistics |
0.1 | 1 | 2019 | CogNet: A Large-Scale Cognate Database · ACL (1) 2019 |
Methods — techniques the papers use, named apart from their topics
wordnet-based computation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Metonymy as a Universal Cognitive Phenomenon: Evidence from Multilingual Lexicons
Temuulen Khishigsuren, Gábor Bella, Thomas Brochhagen, Daariimaa Marav, Fausto Giunchiglia, Khuyagbaatar Batsuren |
CogSci | 6 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 1 |
| 2022 | Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in KinshipabstractThis paper describes a method to enrich lexical resources with content relating to linguistic diversity, based on knowledge from the field of lexical typology. We capture the phenomenon of diversity through the notion of lexical gap and use a systematic method to infer gaps semi-automatically on a large scale, which we demonstrate on the kinship domain. The resulting free diversity-aware terminological resource consists of 198 concepts, 1,911 words, and 37,370 gaps in 699 languages. We see great potential in the use of resources such as ours for the improvement of a variety of cross-lingual NLP tasks, which we illustrate through an application in the evaluation of machine translation systems. Temuulen Khishigsuren, Gábor Bella, Khuyagbaatar Batsuren, Abed Alhakim Freihat, Nandu Chandran Nair, Amarsanaa Ganbold, Hadi Khalilia, Yamini Chandrashekar, Fausto Giunchiglia |
LREC | 3 |
| 2019 | CogNet: A Large-Scale Cognate DatabaseabstractThis paper introduces CogNet, a new, large-scale lexical database that provides cognates-words of common origin and meaning-across languages.The database currently contains 3.1 million cognate pairs across 338 languages using 35 writing systems.The paper also describes the automated method by which cognates were computed from publicly available wordnets, with an accuracy evaluated to 94%.Finally, statistics and early insights about the cognate data are presented, hinting at a possible future exploitation of the resource 1 by various fields of lingustics. Khuyagbaatar Batsuren, Gábor Bella, Fausto Giunchiglia |
ACL (1) | 1 |
| 2019 | Building the Mongolian WordNetabstractThis paper presents the Mongolian Wordnet (MOW), and a general methodology of how to construct it from various sources e.g.lexical resources and expert translations.As of today, the MOW contains 23,665 synsets, 26,875 words, 2,979 glosses, and 213 examples.The manual evaluation of the resource 1 estimated its quality at 96.4%. Khuyagbaatar Batsuren, Amarsanaa Ganbold, Altangerel Chagnaa, Fausto Giunchiglia |
GWC | 1 |
| 2018 | One World - Seven Thousand Languages (Best Paper Award, Third Place)
Fausto Giunchiglia, Khuyagbaatar Batsuren, Abed Alhakim Freihat |
CICLing (1) | 2 |
| 2017 | Understanding and Exploiting Language DiversityabstractThe main goal of this paper is to describe a general approach to the problem of understanding linguistic phenomena, as they appear in lexical semantics, through the analysis of large scale resources, while exploiting these results to improve the quality of the resources themselves. The main contributions are: the approach itself, a formal quantitative measure of language diversity; a set of formal quantitative measures of resource incompleteness and a large scale resource, called the Universal Knowledge Core (UKC) built following the methodology proposed. As a concrete example of an application, we provide an algorithm for distinguishing polysemes from homonyms, as stored in the UKC. Fausto Giunchiglia, Khuyagbaatar Batsuren, Gábor Bella |
IJCAI | 2 |