VLDB 2026 Research / reviewers in the wild / expert
Bradley Hauer
dblp:127/6967 · also Bradley M. Hauer
· DBLP profile ↗
15ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-2715-0223ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semi-Automated Construction of Sense-Annotated Datasets for Practically Any LanguageabstractHigh-quality sense-annotated datasets are vital for evaluating and comparing WSD systems. We present a novel approach to creating parallel sense-annotated datasets, which can be applied to any language that English can be translated into. The method incorporates machine translation, word alignment, sense projection, and sense filtering to produce silver annotations, which can then be revised manually to obtain gold datasets. By applying our method to Farsi, Chinese, and Bengali, we produce new parallel benchmark datasets, which are vetted by native speakers of each language. Our automatically-generated silver datasets are of higher quality than the annotations obtained with recent multilingual WSD systems, particularly on non-European languages. Jai Riley, Bradley Hauer, Nafisa Sadaf Hriti, Guoqing Luo, Amirreza Mirzaei, Ali Rafiei, Hadi Sheikhi, Mahvash Siavashpour, Mohammad Tavakoli, Ning Shi, Grzegorz Kondrak |
COLING | 2 |
| 2024 | Translation-based Lexicalization Generation and Lexical Gap Detection: Application to Kinship TermsabstractConstructing lexicons with explicitly identified lexical gaps is a vital part of building multilingual lexical resources.Prior work has leveraged bilingual dictionaries and linguistic typologies for semi-automatic identification of lexical gaps.Instead, we propose a generallyapplicable algorithmic method to automatically generate concept lexicalizations, which is based on machine translation and hypernymy relations between concepts.The absence of a lexicalization implies a lexical gap.We apply our method to kinship terms, which make a suitable case study because of their explicit definitions and regular structure.Empirical evaluations demonstrate that our approach yields higher accuracy than BabelNet and ChatGPT.Our error analysis indicates that enhancing the quality of translations can further improve the accuracy of our method. Senyu Li, Bradley Hauer, Ning Shi, Grzegorz Kondrak |
ACL (1) | 2 |
| 2023 | Bridging the Gap Between BabelNet and HowNet: Unsupervised Sense Alignment and Sememe PredictionabstractAs the minimum semantic units of natural languages, sememes can provide interpretable representations of concepts.Despite the widespread utilization of lexical resources for semantic tasks, the use of sememes is limited by a lack of available sememe knowledge bases.Recent efforts have been made to connect Ba-belNet with HowNet by automating sememe prediction.However, these methods depend on large manually annotated datasets.Instead, we propose to use sense alignment via a novel unsupervised and explainable method.Our method consists of four stages, each relaxing predefined constraints until a complete alignment of BabelNet synsets to HowNet senses is achieved.Experimental results demonstrate the superiority of our unsupervised method over previous supervised ones by an improvement of 12% overall F1 score, setting a new state of the art.Our work is grounded in an interpretable propagation of sememe information between lexical resources, and may benefit downstream applications which can incorporate sememe information. Xiang Zhang 0011, Ning Shi, Bradley Hauer, Grzegorz Kondrak |
EACL | 3 |
| 2023 | Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsabstractLarge language models (LLMs) have demonstrated exceptional natural language understanding abilities, and have excelled in a variety of natural language processing (NLP) tasks.Despite the fact that most LLMs are trained predominantly on English, multiple studies have demonstrated their capabilities in a variety of languages.However, fundamental questions persist regarding how LLMs acquire their multilingual abilities and how performance varies across different languages.These inquiries are crucial for the study of LLMs since users and researchers often come from diverse language backgrounds, potentially influencing how they use LLMs and interpret their output.In this work, we propose a systematic way of qualitatively and quantitatively evaluating the multilingual capabilities of LLMs.We investigate the phenomenon of cross-language generalization in LLMs, wherein limited multilingual training data leads to advanced multilingual capabilities.To accomplish this, we employ a novel prompt back-translation method.The results demonstrate that LLMs, such as GPT, can effectively transfer learned knowledge across different languages, yielding relatively consistent results in translation-equivariant tasks, in which the correct output does not depend on the language of the input.However, LLMs struggle to provide accurate results in translation-variant tasks, which lack this property, requiring careful user judgment to evaluate the answers. Xiang Zhang 0011, Senyu Li, Bradley Hauer, Ning Shi, Grzegorz Kondrak |
EMNLP | 3 |
| 2023 | One Sense per TranslationabstractBradley Hauer, Grzegorz Kondrak. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bradley Hauer, Grzegorz Kondrak |
IJCNLP (1) | 1 |
| 2022 | Lexical Resource Mapping via TranslationsabstractAligning lexical resources that associate words with concepts in multiple languages increases the total amount of semantic information that can be leveraged for various NLP tasks. We present a translation-based approach to mapping concepts across diverse resources. Our methods depend only on multilingual lexicalization information. When applied to align WordNet/BabelNet to CLICS and OmegaWiki, our methods achieve state-of-the-art accuracy, without any dependence on other sources of semantic knowledge. Since each word-concept pair corresponds to a unique sense of the word, we also demonstrate that the mapping task can be framed as word sense disambiguation. To facilitate future work, we release a set of high-precision WordNet-CLICS alignments, produced by combining three different mapping methods. Hongchang Bao, Bradley Hauer, Grzegorz Kondrak |
LREC | 2 |
| 2022 | WiC = TSV = WSD: On the Equivalence of Three Semantic TasksabstractThe Word-in-Context (WiC) task has attracted considerable attention in the NLP community, as demonstrated by the popularity of the recent MCL-WiC SemEval shared task.Systems and lexical resources from word sense disambiguation (WSD) are often used for the WiC task and WiC dataset construction.In this paper, we establish the exact relationship between WiC and WSD, as well as the related task of target sense verification (TSV).Building upon a novel hypothesis on the equivalence of sense and meaning distinctions, we demonstrate through the application of tools from theoretical computer science that these three semantic classification problems can be pairwise reduced to each other, and therefore are equivalent.The results of experiments that involve systems and datasets for both WiC and WSD provide strong empirical evidence that our problem reductions work in practice. Bradley Hauer, Grzegorz Kondrak |
NAACL-HLT | 1 |
| 2021 | On Universal ColexificationsabstractColexification occurs when two distinct concepts are lexified by the same word.The term covers both polysemy and homonymy.We posit and investigate the hypothesis that no pair of concepts are colexified in every language.We test our hypothesis by analyzing colexification data from BabelNet, Open Multilingual WordNet, and CLICS.The results show that our hypothesis is supported by over 99.9% of colexified concept pairs in these three lexical resources. Hongchang Bao, Bradley Hauer, Grzegorz Kondrak |
GWC | 2 |
| 2021 | Homonymy and Polysemy Detection with Multilingual InformationabstractDeciding whether a semantically ambiguous word is homonymous or polysemous is equivalent to establishing whether it has any pair of senses that are semantically unrelated.We present novel methods for this task that leverage information from multilingual lexical resources.We formally prove the theoretical properties that provide the foundation for our methods.In particular, we show how the One Homonym Per Translation hypothesis of Hauer and Kondrak (2020a) follows from the synset properties formulated by Hauer and Kondrak (2020b).Experimental evaluation shows that our approach sets a new state of the art for homonymy detection. Amir Ahmad Habibi, Bradley Hauer, Grzegorz Kondrak |
GWC | 2 |
| 2020 | One Homonym per TranslationabstractThe study of homonymy is vital to resolving fundamental problems in lexical semantics. In this paper, we propose four hypotheses that characterize the unique behavior of homonyms in the context of translations, discourses, collocations, and sense clusters. We present a new annotated homonym resource that allows us to test our hypotheses on existing WSD resources. The results of the experiments provide strong empirical evidence for the hypotheses. This study represents a step towards a computational method for distinguishing between homonymy and polysemy, and constructing a definitive inventory of coarse-grained senses. Bradley Hauer, Grzegorz Kondrak |
AAAI | 1 |
| 2020 | Improving Word Sense Disambiguation with TranslationsabstractIt has been conjectured that multilingual information can help monolingual word sense disambiguation (WSD).However, existing WSD systems rarely consider multilingual information, and no effective method has been proposed for improving WSD by generating translations.In this paper, we present a novel approach that improves the performance of a base WSD system using machine translation.Since our approach is language independent, we perform WSD experiments on several languages.The results demonstrate that our methods can consistently improve the performance of WSD systems, and obtain state-ofthe-art results in both English and multilingual WSD.To facilitate the use of lexical translation information, we also propose BABALIGN, an precise bitext alignment algorithm which is guided by multilingual lexical correspondences from BabelNet. Yixing Luan, Bradley Hauer, Lili Mou, Grzegorz Kondrak |
EMNLP (1) | 2 |
| 2016 | Decoding Anagrammed Texts Written in an Unknown Language and ScriptabstractAlgorithmic decipherment is a prime example of a truly unsupervised problem. The first step in the decipherment process is the identification of the encrypted language. We propose three methods for determining the source language of a document enciphered with a monoalphabetic substitution cipher. The best method achieves 97% accuracy on 380 languages. We then present an approach to decoding anagrammed substitution ciphers, in which the letters within words have been arbitrarily transposed. It obtains the average decryption word accuracy of 93% on a set of 50 ciphertexts in 5 languages. Finally, we report the results on the Voynich manuscript, an unsolved fifteenth century cipher, which suggest Hebrew as the language of the document. Bradley Hauer, Grzegorz Kondrak |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Solving Substitution Ciphers with Combined Language Models
Bradley Hauer, Ryan B. Hayward, Grzegorz Kondrak |
COLING | 1 |
| 2013 | Automatic Generation of English Respellings
Bradley Hauer, Grzegorz Kondrak |
HLT-NAACL | 1 |
| 2011 | Clustering Semantically Equivalent Words into Cognate Sets in Multilingual Lists
Bradley Hauer, Grzegorz Kondrak |
IJCNLP | 1 |