Grzegorz Kondrak

dblp:40/3774 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2025 Semi-Automated Construction of Sense-Annotated Datasets for Practically Any Language
abstract
High-quality sense-annotated datasets are vital for evaluating and comparing WSD systems. We present a novel approach to creating parallel sense-annotated datasets, which can be applied to any language that English can be translated into. The method incorporates machine translation, word alignment, sense projection, and sense filtering to produce silver annotations, which can then be revised manually to obtain gold datasets. By applying our method to Farsi, Chinese, and Bengali, we produce new parallel benchmark datasets, which are vetted by native speakers of each language. Our automatically-generated silver datasets are of higher quality than the annotations obtained with recent multilingual WSD systems, particularly on non-European languages.
Jai Riley, Bradley Hauer, Nafisa Sadaf Hriti, Guoqing Luo, Amirreza Mirzaei, Ali Rafiei, Hadi Sheikhi, Mahvash Siavashpour, Mohammad Tavakoli, Ning Shi, Grzegorz Kondrak
COLING11
2024 Translation-based Lexicalization Generation and Lexical Gap Detection: Application to Kinship Terms
abstract
Constructing lexicons with explicitly identified lexical gaps is a vital part of building multilingual lexical resources.Prior work has leveraged bilingual dictionaries and linguistic typologies for semi-automatic identification of lexical gaps.Instead, we propose a generallyapplicable algorithmic method to automatically generate concept lexicalizations, which is based on machine translation and hypernymy relations between concepts.The absence of a lexicalization implies a lexical gap.We apply our method to kinship terms, which make a suitable case study because of their explicit definitions and regular structure.Empirical evaluations demonstrate that our approach yields higher accuracy than BabelNet and ChatGPT.Our error analysis indicates that enhancing the quality of translations can further improve the accuracy of our method.
Senyu Li, Bradley Hauer, Ning Shi, Grzegorz Kondrak
ACL (1)4
2023 Bridging the Gap Between BabelNet and HowNet: Unsupervised Sense Alignment and Sememe Prediction
abstract
As the minimum semantic units of natural languages, sememes can provide interpretable representations of concepts.Despite the widespread utilization of lexical resources for semantic tasks, the use of sememes is limited by a lack of available sememe knowledge bases.Recent efforts have been made to connect Ba-belNet with HowNet by automating sememe prediction.However, these methods depend on large manually annotated datasets.Instead, we propose to use sense alignment via a novel unsupervised and explainable method.Our method consists of four stages, each relaxing predefined constraints until a complete alignment of BabelNet synsets to HowNet senses is achieved.Experimental results demonstrate the superiority of our unsupervised method over previous supervised ones by an improvement of 12% overall F1 score, setting a new state of the art.Our work is grounded in an interpretable propagation of sememe information between lexical resources, and may benefit downstream applications which can incorporate sememe information.
Xiang Zhang 0011, Ning Shi, Bradley Hauer, Grzegorz Kondrak
EACL4
2023 Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMs
abstract
Large language models (LLMs) have demonstrated exceptional natural language understanding abilities, and have excelled in a variety of natural language processing (NLP) tasks.Despite the fact that most LLMs are trained predominantly on English, multiple studies have demonstrated their capabilities in a variety of languages.However, fundamental questions persist regarding how LLMs acquire their multilingual abilities and how performance varies across different languages.These inquiries are crucial for the study of LLMs since users and researchers often come from diverse language backgrounds, potentially influencing how they use LLMs and interpret their output.In this work, we propose a systematic way of qualitatively and quantitatively evaluating the multilingual capabilities of LLMs.We investigate the phenomenon of cross-language generalization in LLMs, wherein limited multilingual training data leads to advanced multilingual capabilities.To accomplish this, we employ a novel prompt back-translation method.The results demonstrate that LLMs, such as GPT, can effectively transfer learned knowledge across different languages, yielding relatively consistent results in translation-equivariant tasks, in which the correct output does not depend on the language of the input.However, LLMs struggle to provide accurate results in translation-variant tasks, which lack this property, requiring careful user judgment to evaluate the answers.
Xiang Zhang 0011, Senyu Li, Bradley Hauer, Ning Shi, Grzegorz Kondrak
EMNLP5
2023 One Sense per Translation
abstract
Bradley Hauer, Grzegorz Kondrak. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Bradley Hauer, Grzegorz Kondrak
IJCNLP (1)2
2023 Correcting Sense Annotations Using Wordnets and Translations
abstract
Acquiring large amounts of high-quality annotated data is an open issue in word sense disambiguation.This problem has become more critical recently with the advent of supervised models based on neural networks, which require large amounts of annotated data.We propose two algorithms for making selective corrections on a sense-annotated parallel corpus, based on cross-lingual synset mappings.We show that, when applied to bilingual parallel corpora, these algorithms can rectify noisy sense annotations, and thereby produce multilingual sense-annotated data of improved quality.
Arnob Mallik, Grzegorz Kondrak
GWC2
2022 Lexical Resource Mapping via Translations
abstract
Aligning lexical resources that associate words with concepts in multiple languages increases the total amount of semantic information that can be leveraged for various NLP tasks. We present a translation-based approach to mapping concepts across diverse resources. Our methods depend only on multilingual lexicalization information. When applied to align WordNet/BabelNet to CLICS and OmegaWiki, our methods achieve state-of-the-art accuracy, without any dependence on other sources of semantic knowledge. Since each word-concept pair corresponds to a unique sense of the word, we also demonstrate that the mapping task can be framed as word sense disambiguation. To facilitate future work, we release a set of high-precision WordNet-CLICS alignments, produced by combining three different mapping methods.
Hongchang Bao, Bradley Hauer, Grzegorz Kondrak
LREC3
2022 WiC = TSV = WSD: On the Equivalence of Three Semantic Tasks
abstract
The Word-in-Context (WiC) task has attracted considerable attention in the NLP community, as demonstrated by the popularity of the recent MCL-WiC SemEval shared task.Systems and lexical resources from word sense disambiguation (WSD) are often used for the WiC task and WiC dataset construction.In this paper, we establish the exact relationship between WiC and WSD, as well as the related task of target sense verification (TSV).Building upon a novel hypothesis on the equivalence of sense and meaning distinctions, we demonstrate through the application of tools from theoretical computer science that these three semantic classification problems can be pairwise reduced to each other, and therefore are equivalent.The results of experiments that involve systems and datasets for both WiC and WSD provide strong empirical evidence that our problem reductions work in practice.
Bradley Hauer, Grzegorz Kondrak
NAACL-HLT2
2021 On Universal Colexifications
abstract
Colexification occurs when two distinct concepts are lexified by the same word.The term covers both polysemy and homonymy.We posit and investigate the hypothesis that no pair of concepts are colexified in every language.We test our hypothesis by analyzing colexification data from BabelNet, Open Multilingual WordNet, and CLICS.The results show that our hypothesis is supported by over 99.9% of colexified concept pairs in these three lexical resources.
Hongchang Bao, Bradley Hauer, Grzegorz Kondrak
GWC3
2021 Homonymy and Polysemy Detection with Multilingual Information
abstract
Deciding whether a semantically ambiguous word is homonymous or polysemous is equivalent to establishing whether it has any pair of senses that are semantically unrelated.We present novel methods for this task that leverage information from multilingual lexical resources.We formally prove the theoretical properties that provide the foundation for our methods.In particular, we show how the One Homonym Per Translation hypothesis of Hauer and Kondrak (2020a) follows from the synset properties formulated by Hauer and Kondrak (2020b).Experimental evaluation shows that our approach sets a new state of the art for homonymy detection.
Amir Ahmad Habibi, Bradley Hauer, Grzegorz Kondrak
GWC3
2020 One Homonym per Translation
abstract
The study of homonymy is vital to resolving fundamental problems in lexical semantics. In this paper, we propose four hypotheses that characterize the unique behavior of homonyms in the context of translations, discourses, collocations, and sense clusters. We present a new annotated homonym resource that allows us to test our hypotheses on existing WSD resources. The results of the experiments provide strong empirical evidence for the hypotheses. This study represents a step towards a computational method for distinguishing between homonymy and polysemy, and constructing a definitive inventory of coarse-grained senses.
Bradley Hauer, Grzegorz Kondrak
AAAI2
2020 Improving Word Sense Disambiguation with Translations
abstract
It has been conjectured that multilingual information can help monolingual word sense disambiguation (WSD).However, existing WSD systems rarely consider multilingual information, and no effective method has been proposed for improving WSD by generating translations.In this paper, we present a novel approach that improves the performance of a base WSD system using machine translation.Since our approach is language independent, we perform WSD experiments on several languages.The results demonstrate that our methods can consistently improve the performance of WSD systems, and obtain state-ofthe-art results in both English and multilingual WSD.To facilitate the use of lexical translation information, we also propose BABALIGN, an precise bitext alignment algorithm which is guided by multilingual lexical correspondences from BabelNet.
Yixing Luan, Bradley Hauer, Lili Mou, Grzegorz Kondrak
EMNLP (1)4
2017 Identifying Cognate Sets Across Dictionaries of Related Languages
abstract
We present a system for identifying cognate sets across dictionaries of related languages.The likelihood of a cognate relationship is calculated on the basis of a rich set of features that capture both phonetic and semantic similarity, as well as the presence of regular sound correspondences.The similarity scores are used to cluster words from different languages that may originate from a common protoword.When tested on the Algonquian language family, our system detects 63% of cognate sets while maintaining cluster purity of 70%.
Adam St. Arnaud, David Beck, Grzegorz Kondrak
EMNLP3
2016 Leveraging Inflection Tables for Stemming and Lemmatization
abstract
We present several methods for stemming and lemmatization based on discriminative string transduction. We exploit the paradigmatic regularity of semi-structured inflection tables to identify stems in an unsupervised manner with over 85% accuracy. Experiments on English, Dutch and German show that our stemmers substantially outperform Snowball and Morfessor, and approach the accuracy of a supervised model. Furthermore, the generated stems are more consistent than those annotated by experts. Our direct lemmatization model is more accurate than Morfette and Lemming on most datasets. Finally, we test our methods on the data from the shared task on morphological reinflection.
Garrett Nicolai, Grzegorz Kondrak
ACL (1)2
2016 Integrating Morphological Desegmentation into Phrase-based Decoding
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
HLT-NAACL3
2016 Decoding Anagrammed Texts Written in an Unknown Language and Script
abstract
Algorithmic decipherment is a prime example of a truly unsupervised problem. The first step in the decipherment process is the identification of the encrypted language. We propose three methods for determining the source language of a document enciphered with a monoalphabetic substitution cipher. The best method achieves 97% accuracy on 380 languages. We then present an approach to decoding anagrammed substitution ciphers, in which the letters within words have been arbitrarily transposed. It obtains the average decryption word accuracy of 93% on a set of 50 ciphertexts in 5 languages. Finally, we report the results on the Voynich manuscript, an unsolved fifteenth century cipher, which suggest Hebrew as the language of the document.
Bradley Hauer, Grzegorz Kondrak
Trans. Assoc. Comput. Linguistics2
2015 Inflection Generation as Discriminative String Transduction
abstract
We approach the task of morphological inflection generation as discriminative string transduction.Our supervised system learns to generate word-forms from lemmas accompanied by morphological tags, and refines them by referring to the other forms within a paradigm.Results of experiments on six diverse languages with varying amounts of training data demonstrate that our approach improves the state of the art in terms of predicting inflected word-forms.
Garrett Nicolai, Colin Cherry, Grzegorz Kondrak
HLT-NAACL3
2015 English orthography is not "close to optimal"
abstract
In spite of the apparent irregularity of the English spelling system, Chomsky and Halle (1968) characterize it as "near optimal".We investigate this assertion using computational techniques and resources.We design an algorithm to generate word spellings that maximize both phonemic transparency and morphological consistency.Experimental results demonstrate that the constructed system is much closer to optimality than the traditional English orthography.
Garrett Nicolai, Grzegorz Kondrak
HLT-NAACL2
2015 Joint Generation of Transliterations from Multiple Representations
abstract
Machine transliteration is often referred to as phonetic translation. We show that transliterations incorporate information from both spelling and pronunciation, and propose an effective model for joint transliteration generation from both representations. We further generalize this model to include transliterations from other languages, and enhance it with reranking and lexicon features. We demonstrate significant improvements in transliteration accuracy on several datasets.
Grzegorz Kondrak
HLT-NAACL2
2014 Lattice Desegmentation for Statistical Machine Translation
abstract
Morphological segmentation is an effec-tive sparsity reduction strategy for statis-tical machine translation (SMT) involv-ing morphologically complex languages. When translating into a segmented lan-guage, an extra step is required to deseg-ment the output; previous studies have de-segmented the 1-best output from the de-coder. In this paper, we expand our trans-lation options by desegmenting n-best lists or lattices. Our novel lattice desegmenta-tion algorithm effectively combines both segmented and desegmented views of the target language for a large subspace of possible translation outputs, which allows for inclusion of features related to the de-segmentation process, as well as an un-segmented language model (LM). We in-vestigate this technique in the context of English-to-Arabic and English-to-Finnish translation, showing significant improve-ments in translation quality over deseg-mentation of 1-best decoder outputs. 1
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
ACL (1)3
2014 Solving Substitution Ciphers with Combined Language Models
Bradley Hauer, Ryan B. Hayward, Grzegorz Kondrak
COLING3
2013 Identification of Speakers in Novels
Denilson Barbosa 0001, Grzegorz Kondrak
ACL (1)3
2013 Automatic Generation of English Respellings
Bradley Hauer, Grzegorz Kondrak
HLT-NAACL2
2013 Reversing Morphological Tokenization in English-to-Arabic SMT
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
HLT-NAACL3
2012 Leveraging supplemental representations for sequential transduction
Aditya Bhargava, Grzegorz Kondrak
HLT-NAACL2
2011 How do you pronounce your name? Improving G2P with transliterations
Aditya Bhargava, Grzegorz Kondrak
ACL2
2011 The application of chordal graphs to inferring phylogenetic trees of languages
Jessica A. Enright, Grzegorz Kondrak
IJCNLP2
2011 Clustering Semantically Equivalent Words into Cognate Sets in Multilingual Lists
Bradley Hauer, Grzegorz Kondrak
IJCNLP2
2010 Letter-Phoneme Alignment: An Exploration
Sittichai Jiampojamarn, Grzegorz Kondrak
ACL2
2010 Predicting the Semantic Compositionality of Prefix Verbs
Shane Bergsma, Aditya Bhargava, Grzegorz Kondrak
EMNLP4
2010 Language identification of names with SVMs
Aditya Bhargava, Grzegorz Kondrak
HLT-NAACL2
2010 Integrating Joint n-gram Features into a Discriminative Training Framework
Sittichai Jiampojamarn, Colin Cherry, Grzegorz Kondrak
HLT-NAACL3
2009 A Ranking Approach to Stress Prediction for Letter-to-Phoneme Conversion
Qing Dou, Shane Bergsma, Sittichai Jiampojamarn, Grzegorz Kondrak
ACL/IJCNLP4
2009 Reducing the Annotation Effort for Letter-to-Phoneme Conversion
Kenneth Dwyer, Grzegorz Kondrak
ACL/IJCNLP2
2009 Online discriminative training for grapheme-to-phoneme conversion
abstract
We present an online discriminative training approach to grapheme-to-phoneme (g2p) conversion. We employ a manyto-many alignment between graphemes and phonemes, which overcomes the limitations of widely used one-to-one alignments. The discriminative structure-prediction model incorporates input segmentation, phoneme prediction, and sequence modeling in a unified dynamic programming framework. The learning model is able to capture both local context features in inputs, as well as non-local dependency features in sequence outputs. Experimental results show that our system surpasses the state-of-the-art on several data sets. Index Terms: grapheme-to-phoneme conversion, speech synthesis, discriminative training
Sittichai Jiampojamarn, Grzegorz Kondrak
INTERSPEECH2
2009 On the Syllabification of Phonemes
Susan Bartlett, Grzegorz Kondrak, Colin Cherry
HLT-NAACL2
2008 Automatic Syllabification with Structured SVMs for Letter-to-Phoneme Conversion
Susan Bartlett, Grzegorz Kondrak, Colin Cherry
ACL2
2008 Joint Processing and Discriminative Training for Letter-to-Phoneme Conversion
Sittichai Jiampojamarn, Colin Cherry, Grzegorz Kondrak
ACL3
2007 Alignment-Based Discriminative String Similarity
Shane Bergsma, Grzegorz Kondrak
ACL2
2007 Bootstrapping a Stochastic Transducer for Arabic-English Transliteration Extraction
Tarek Sherif, Grzegorz Kondrak
ACL2
2007 Substring-Based Transliteration
Tarek Sherif, Grzegorz Kondrak
ACL2
2007 Applying Many-to-Many Alignments and Hidden Markov Models to Letter-to-Phoneme Conversion
Sittichai Jiampojamarn, Grzegorz Kondrak, Tarek Sherif
HLT-NAACL2
2006 Automatic identification of confusable drug names
Grzegorz Kondrak, Bonnie J. Dorr
Artif. Intell. Medicine1
2005 Computing Word Similarity and Identifying Cognates with Pair Hidden Markov Models
Wesley Mackay, Grzegorz Kondrak
CoNLL2
2005 Cognates and Word Alignment in Bitexts
abstract
We evaluate several orthographic word similarity measures in the context of bitext word alignment. We investigate the relationship between the length of the words and the length of their longest common subsequence. We present an alternative to the longest common subsequence ratio (LCSR), a widely-used orthographic word similarity measure. Experiments involving identification of cognates in bitexts suggest that the alternative method outperforms LCSR. Our results also indicate that alignment links can be used as a substitute for cognates for the purpose of evaluating word similarity measures.
Grzegorz Kondrak
MTSummit1
2005 N-Gram Similarity and Distance
Grzegorz Kondrak
SPIRE1
2004 Identification of Confusable Drug Names: A New Approach and Evaluation Methodology
Grzegorz Kondrak, Bonnie J. Dorr
COLING1
2003 Identifying Complex Sound Correspondences in Bilingual Wordlists
Grzegorz Kondrak
CICLing1
2003 Cognates Can Improve Statistical Translation Models
Grzegorz Kondrak, Daniel Marcu, Kevin Knight
HLT-NAACL1
2002 Determining Recurrent Sound Correspondences by Inducing Translation Models
Grzegorz Kondrak
COLING1
2001 Identifying Cognates by Phonetic and Semantic Similarity
Grzegorz Kondrak
NAACL1
1997 A Theoretical Evaluation of Selected Backtracking Algorithms
Grzegorz Kondrak, Peter van Beek
Artif. Intell.1
1995 A Theoretical Evaluation of Selected Backtracking Algorithms
Grzegorz Kondrak, Peter van Beek
IJCAI1