EDBT 2026 Demo / reviewers in the wild / expert
Kevin Duh
dblp:58/3217
· DBLP profile ↗
93ranked-venue papers
12as first author
28since 2021 · last 2026
0000-0001-8107-4383ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
Debashish Chakraborty, Eugene Yang 0001, Daniel Khashabi, Dawn J. Lawrie, Kevin Duh |
ECIR (2) | 5 |
| 2026 | FACTUM: Mechanistic Detection of Citation Hallucination in Long-Form RAG
Maxime Dassen, Rebecca Kotula, Kenton Murray, Andrew Yates, Dawn J. Lawrie, Efsun Selin Kayi, James Mayfield, Kevin Duh |
ECIR (1) | 8 |
| 2025 | GenVC: Self-Supervised Zero-Shot Voice ConversionabstractMost current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.11Audio samples, code, and model checkpoints are available at https://caizexin.github.io/GenVC/index.html Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews |
ASRU | 5 |
| 2025 | Scalable Controllable Accented TTSabstractWe propose a method to scale accented TTS training to large, accent-diverse datasets that often lack consistent, high-quality accent labels. Our approach relies on a speech geolocation model to infer accent labels directly from audio. To improve speaker generalization and encourage disentangling speaker from accent we explore timbre augmentation through kNN voice conversion. We validate our approach on CommonVoice by fine-tuning XTTS-v2 with accent labels inferred or improved via geolocation. According to various automated metrics based on embeddings extracted from an accent identification model, the resulting accented TTS model produces speech with better accent fidelity compared to XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, or other existing accented TTS models. According to human evaluation, it was clear that the geolocation model based data discovery and enhancement improved the naturalness and accent fidelity of generated speech. However, the effect of different data augmentation strategies was less clear. Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
ASRU | 4 |
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 9 |
| 2025 | Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationabstractAs people increasingly use AI systems in work and daily life, mechanisms that help them use AI responsibly are urgently needed, especially when they are not equipped to verify AI predictions themselves.We study a realistic Machine Translation (MT) scenario where monolingual users decide whether to share an MT output, first without and then with quality feedback.We compare four types of quality feedback: explicit feedback that directly give users an assessment of translation quality using (1) error highlights and (2) LLM explanations, and implicit feedback that helps users compare MT inputs and outputs through (3) backtranslation and ( 4) question-answer (QA) tables.We find that all feedback types, except error highlights, significantly improve both decision accuracy and appropriate reliance.Notably, implicit feedback, especially QA tables, yields significantly greater gains than explicit feedback in terms of decision accuracy, appropriate reliance, and user perceptions -receiving the highest ratings for helpfulness and trust, and the lowest for mental burden. Dayeon Ki, Kevin Duh, Marine Carpuat |
EMNLP | 2 |
| 2025 | Whisper-UT: A Unified Translation Framework for Speech and TextabstractCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Cihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur |
EMNLP | 7 |
| 2025 | HLTCOE Submission to the VoicePrivacy Attacker ChallengeabstractWe describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By simultaneously employing one or more of these methods, we were able to achieve a significant reduction in EER against all of the submitted anonymization systems in the VoicePrivacy Challenge. Henry Li Xinyuan, Ashi Garg, Zexin Cai, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
ICASSP | 4 |
| 2025 | OJ4OCRMT: A Large Multilingual Dataset for OCR-MT EvaluationabstractWe introduce OJ4OCRMT, an Optical Character Recognition (OCR) dataset for Machine Translation (MT). The dataset supports research on automatic extraction, recognition, and translation of text from document images. The Official Journal of the European Union (OJEU), is the official gazette for the EU. Tens of thousands of pages of legislative acts and regulatory notices are published annually, and parallel translations are available in each of the official languages. Due to its large size, high degree of multilinguality, and carefully produced human translations, the OJEU is a singular resource for language processing research. We have assembled a large collection of parallel pages from the OJEU and have created a dataset to support translation of document images. In this work we introduce the dataset, describe the design decisions which we undertook, and report baseline performance figures for the translation task. It is our hope that this dataset will significantly add to the comparatively few resources presently available for evaluating OCR-MT systems. Paul McNamee, Kevin Duh, Cameron Carpenter, Ronald Colaianni, Nolan King, Kenton Murray |
MTSummit (1) | 2 |
| 2024 | Exploring Geometric Representational Disparities between Multilingual and Bilingual Translation ModelsabstractMultilingual machine translation has proven immensely useful for both parameter efficiency and overall performance across many language pairs via complete multilingual parameter sharing. However, some language pairs in multilingual models can see worse performance than in bilingual models, especially in the one-to-many translation setting. Motivated by their empirical differences, we examine the geometric differences in representations from bilingual models versus those from one-to-many multilingual models. Specifically, we compute the isotropy of these representations using intrinsic dimensionality and IsoScore, in order to measure how the representations utilize the dimensions in their underlying vector space. Using the same evaluation data in both models, we find that for a given language pair, its multilingual model decoder representations are consistently less isotropic and occupy fewer dimensions than comparable bilingual model decoder representations. Additionally, we show that much of the anisotropy in multilingual decoder representations can be attributed to modeling language-specific information, therefore limiting remaining representational capacity. Neha Verma 0001, Kenton Murray, Kevin Duh |
LREC/COLING | 3 |
| 2024 | Large-Scale Bitext Corpora Provide New Evidence for Cognitive Representations of Spatial TermsabstractPeter Viechnicki, Kevin Duh, Anthony Kostacos, Barbara Landau. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peter Viechnicki, Kevin Duh, Anthony Kostacos, Barbara Landau |
EACL (1) | 2 |
| 2024 | SpeechQE: Estimating the Quality of Direct Speech TranslationabstractRecent advances in automatic quality estimation for machine translation have exclusively focused on written language, leaving the speech modality underexplored.In this work, we formulate the task of quality estimation for speech translation, construct a benchmark, and evaluate a family of systems based on cascaded and end-to-end architectures.In this process, we introduce a novel end-to-end system leveraging pre-trained text LLM.Results suggest that end-to-end approaches are better suited to estimating the quality of direct speech translation than using quality estimation systems designed for text in cascaded systems.More broadly, we argue that quality estimation of speech translation needs to be studied as a separate problem from that of text, and release our data and models to guide further research in this space.1 HyoJung Han 0001, Kevin Duh, Marine Carpuat |
EMNLP | 2 |
| 2024 | Where does In-context Learning Happen in Large Language Models?abstractSelf-supervised large language models have demonstrated the ability to perform various tasks via in-context learning, but little is known about where the model locates the task with respect to prompt instructions and demonstration examples. In this work, we attempt to characterize the region where large language models transition from recognizing the task to performing the task. Through a series of layer-wise context-masking experiments on GPTNeo2.7B, Bloom3B, Starcoder2-7B, Llama3.1-8B, Llama3.1-8B-Instruct, on Machine Translation and Code generation, we demonstrate evidence of a "task recognition" point where the task is encoded into the input representations and attention to context is no longer necessary. Taking advantage of this redundancy results in 45% computational savings when prompting with 5 examples, and task recognition achieved at layer 14 / 32 using an example with Machine Translation. Our findings also have implications for resource and parameter efficient fine-tuning; we observe a correspondence between strong fine-tuning performance of individual LoRA layers and the task recognition layers. Suzanna Sia, David Mueller, Kevin Duh |
NeurIPS | 3 |
| 2024 | Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker AnonymizationabstractAdvances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its utility, including linguistic and paralinguistic aspects. However, anonymizing speech while maintaining emotional state remains challenging. We explore this problem in the context of the VoicePrivacy 2024 challenge. Specifically, we developed various speaker anonymization pipelines and find that approaches either excel at anonymization or preserving emotion state, but not both simultaneously. Achieving both would require an in-domain emotion recognizer. Additionally, we found that it is feasible to train a semi-effective speaker verification system using only emotion representations, demonstrating the challenge of separating these two modalities. Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner |
SLT | 5 |
| 2023 | HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation
Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur |
INTERSPEECH | 6 |
| 2023 | In-context Learning as Maintaining Coherency: A Study of On-the-fly Machine Translation Using Large Language ModelsabstractThe phenomena of in-context learning has typically been thought of as “learning from examples”. In this work which focuses on Machine Translation, we present a perspective of in-context learning as the desired generation task maintaining coherency with its context, i.e., the prompt examples. We first investigate randomly sampled prompts across 4 domains, and find that translation performance improves when shown in-domain prompts. Next, we investigate coherency for the in-domain setting, which uses prompt examples from a moving window. We study this with respect to other factors that have previously been identified in the literature such as length, surface similarity and sentence embedding similarity. Our results across 3 models (GPTNeo2.7B, Bloom3B, XGLM2.9B), and three translation directions (en\rightarrow{pt, de, fr}) suggest that the long-term coherency of the prompts and the test sentence is a good indicator of downstream translation performance. In doing so, we demonstrate the efficacy of in-context Machine Translation for on-the-fly adaptation. Suzanna Sia, Kevin Duh |
MTSummit (1) | 2 |
| 2022 | Modeling Constraints Can Identify Winning Arguments in Multi-Party Interactions (Student Abstract)abstractIn contexts where debate and deliberation is the norm, participants are regularly presented with new information that conflicts with their original beliefs. When required to update their beliefs (belief alignment), they may choose arguments that align with their worldview (confirmation bias). We test this and competing hypotheses in a constraint-based modeling approach to predict the winning arguments in multi-party interactions in the Reddit ChangeMyView dataset. We impose structural constraints that reflect competing hypotheses on a hierarchical generative Variational Auto-encoder. Our findings suggest that when arguments are further from the initial belief state of the target, they are more likely to succeed. Suzanna Sia, Kokil Jaidka, Niyati Chayya, Kevin Duh |
AAAI | 4 |
| 2022 | Transfer Learning Approaches for Building Cross-Language Dense Retrieval Models
Suraj Nair 0001, Eugene Yang 0001, Dawn J. Lawrie, Kevin Duh, Paul McNamee, Kenton Murray, James Mayfield, Douglas W. Oard |
ECIR (1) | 4 |
| 2022 | Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal TransportabstractBilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval.We improve bilingual lexicon induction performance across 40 language pairs with a graph-matching method based on optimal transport.The method is especially strong with low amounts of supervision. Kelly Marchisio, Ali Saad-Eldin, Kevin Duh, Carey E. Priebe, Philipp Koehn |
EMNLP | 3 |
| 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding SpacesabstractThe ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces-their degree of "isomorphism."We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the underlying spaces being non-isomorphic.We incorporate global measures of isomorphism directly into the Skip-gram loss function, successfully increasing the relative isomorphism of trained word embedding spaces and improving their ability to be mapped to a shared crosslingual space.The result is improved bilingual lexicon induction in general data conditions, under domain mismatch, and with training algorithm dissimilarities.We release IsoVec at https://github.com/ kellymarchisio/isovec. Kelly Marchisio, Neha Verma 0001, Kevin Duh, Philipp Koehn |
EMNLP | 3 |
| 2022 | AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African LanguagesabstractLanguage diversity in NLP is critical in enabling the development of tools for a wide range of users.However, there are limited resources for building such tools for many languages, particularly those spoken in Africa.For search, most existing datasets feature few or no African languages, directly impacting researchers' ability to build and improve information access capabilities in those languages.Motivated by this, we created AfriCLIRMatrix, a test collection for cross-lingual information retrieval research in 15 diverse African languages.In total, our dataset contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection.In addition, we release BM25, dense retrieval, and sparsedense hybrid baselines to provide a starting point for the development of future systems.We hope that these efforts can spur additional work in search for African languages.Afri-CLIRMatrix can be downloaded at https:// github.com/castorini/africlirmatrix. Odunayo Ogundepo, Xinyu Zhang 0018, Kevin Duh, Jimmy Lin |
EMNLP | 4 |
| 2022 | Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party DebatesabstractIn contexts where debate and deliberation are the norm, the participants are regularly presented with new information that conflicts with their original beliefs.When required to update their beliefs (belief alignment), they may choose arguments that align with their worldview (confirmation bias).We test this and competing hypotheses in a constraintbased modeling approach to predict the winning arguments in multi-party interactions in the Reddit Change My View and Intelligence Squared debates datasets.We adopt a hierarchical generative Variational Autoencoder as our model and impose structural constraints that reflect competing hypotheses about the nature of argumentation.Our findings suggest that in most settings, predictive models that anticipate winning arguments to be further from the initial argument of the opinion holder are more likely to succeed. Suzanna Sia, Kokil Jaidka, Hansin Ahuja, Niyati Chhaya, Kevin Duh |
EMNLP | 5 |
| 2022 | The Multilingual Microblog Translation Corpus: Improving and Evaluating Translation of User-Generated TextabstractTranslation of the noisy, informal language found in social media has been an understudied problem, with a principal factor being the limited availability of translation corpora in many languages. To address this need we have developed a new corpus containing over 200,000 translations of microblog posts that supports translation of thirteen languages into English. The languages are: Arabic, Chinese, Farsi, French, German, Hindi, Korean, Pashto, Portuguese, Russian, Spanish, Tagalog, and Urdu. We are releasing these data as the Multilingual Microblog Translation Corpus to support futher research in translation of informal language. We establish baselines using this new resource, and we further demonstrate the utility of the corpus by conducting experiments with fine-tuning to improve translation quality from a high performing neural machine translation (NMT) system. Fine-tuning provided substantial gains, ranging from +3.4 to +11.1 BLEU. On average, a relative gain of 21% was observed, demonstrating the utility of the corpus. Paul McNamee, Kevin Duh |
LREC | 2 |
| 2021 | Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yolóxochitl MixtecabstractJiatong Shi, Jonathan D. Amith, Rey Castillo García, Esteban Guadalupe Sierra, Kevin Duh, Shinji Watanabe. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jiatong Shi, Jonathan D. Amith, Rey Castillo García, Esteban Guadalupe Sierra, Kevin Duh, Shinji Watanabe 0001 |
EACL | 5 |
| 2021 | Adaptive Mixed Component LDA for Low Resource Topic ModelingabstractProbabilistic topic models in low data resource scenarios are faced with less reliable estimates due to sparsity of discrete word cooccurrence counts, and do not have the luxury of retraining word or topic embeddings using neural methods.In this challenging resource constrained setting, we introduce an automatic trade-off between the discrete and continuous representations via an adaptive mixture coefficient, which places greater weight on the discrete representation when the corpus statistics are more reliable.The adaptive mixture coefficient takes into account global corpus statistics, and the uncertainty in each topic's continuous distribution.Our approach outperforms the fully discrete, fully continuous, and static mixture model on topic coherence in low resource monolingual and multilingual settings. Suzanna Sia, Kevin Duh |
EACL | 2 |
| 2021 | Data and Parameter Scaling Laws for Neural Machine TranslationabstractWe observe that the development crossentropy loss of supervised neural machine translation models scales like a power law with the amount of training data and the number of non-embedding parameters in the model.We discuss some practical implications of these results, such as predicting BLEU achieved by large scale models and predicting the ROI of labeling data in low-resource language pairs. Mitchell A. Gordon, Kevin Duh, Jared Kaplan |
EMNLP (1) | 2 |
| 2021 | ORTHROS: non-autoregressive end-to-end speech translation With dual-decoderabstractFast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not been explored so far. Inspired by recent progress in non-autoregressive (NAR) methods in text-based translation, which generates target tokens in parallel by eliminating conditional dependencies, we study the problem of NAR decoding for E2E-ST. We propose a novel NAR E2E-ST framework, Orthros, in which both NAR and autoregressive (AR) decoders are jointly trained on the shared speech encoder. The latter is used for selecting better translation among various length candidates generated from the former, which dramatically improves the effectiveness of a large length beam with negligible overhead. We further investigate effective length prediction methods from speech inputs and the impact of vocabulary sizes. Experiments on four benchmarks show the effectiveness of the proposed method in improving inference speed while maintaining competitive translation quality compared to state-of-the-art AR E2E-ST systems. Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh, Tatsuya Kawahara, Shinji Watanabe 0001 |
ICASSP | 3 |
| 2021 | Proceedings of the 18th Biennial Machine Translation Summit (Volume 1: Research Track)
Kevin Duh, Francisco Guzmán |
MTSummit (1) | 1 |
| 2020 | CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information RetrievalabstractWe present CLIRMatrix, a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia.CLIR-Matrix comprises (1) BI-139, a bilingual dataset of queries in one language matched with relevant documents in another language for 139×138=19,182 language pairs, and (2) MULTI-8, a multilingual dataset of queries and documents jointly aligned in 8 different languages.In total, we mined 49 million unique queries and 34 billion (query, document, label) triplets, making it the largest and most comprehensive CLIR dataset to date.This collection is intended to support research in end-to-end neural information retrieval and is publicly available at https: //github.com/ssun32/CLIRMatrix.We provide baseline neural model results on BI-139, and evaluate MULTI-8 in both singlelanguage retrieval and mix-language retrieval settings. Kevin Duh |
EMNLP (1) | 2 |
| 2020 | Benchmarking Neural and Statistical Machine Translation on Low-Resource African LanguagesabstractResearch in machine translation (MT) is developing at a rapid pace. However, most work in the community has focused on languages where large amounts of digital resources are available. In this study, we benchmark state of the art statistical and neural machine translation systems on two African languages which do not have large amounts of resources: Somali and Swahili. These languages are of social importance and serve as test-beds for developing technologies that perform reasonably well despite the low-resource constraint. Our findings suggest that statistical machine translation (SMT) and neural machine translation (NMT) can perform similarly in low-resource scenarios, but neural systems require more careful tuning to match performance. We also investigate how to exploit additional data, such as bilingual text harvested from the web, or user dictionaries; we find that NMT can significantly improve in performance with the use of these additional data. Finally, we survey the landscape of machine translation resources for the languages of Africa and provide some suggestions for promising future research directions. Kevin Duh, Paul McNamee, Matt Post, Brian Thompson 0001 |
LREC | 1 |
| 2020 | Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data In Your Machine Translation System?abstractData privacy is an important issue for “machine learning as a service” providers. We focus on the problem of membership inference attacks: Given a data sample and black-box access to a model’s API, determine whether the sample existed in the model’s training data. Our contribution is an investigation of this problem in the context of sequence-to-sequence models, which are important in applications such as machine translation and video captioning. We define the membership inference problem for sequence generation, provide an open dataset based on state-of-the-art machine translation models, and report initial results on whether these models leak private information against several kinds of membership inference attacks. Sorami Hisamoto, Matt Post, Kevin Duh |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Reproducible and Efficient Benchmarks for Hyperparameter Optimization of Neural Machine Translation SystemsabstractHyperparameter selection is a crucial part of building neural machine translation (NMT) systems across both academia and industry. Fine-grained adjustments to a model’s architecture or training recipe can mean the difference between a positive and negative research result or between a state-of-the-art and underperforming system. While recent literature has proposed methods for automatic hyperparameter optimization (HPO), there has been limited work on applying these methods to neural machine translation (NMT), due in part to the high costs associated with experiments that train large numbers of model variants. To facilitate research in this space, we introduce a lookup-based approach that uses a library of pre-trained models for fast, low cost HPO experimentation. Our contributions include (1) the release of a large collection of trained NMT models covering a wide range of hyperparameters, (2) the proposal of targeted metrics for evaluating HPO methods on NMT, and (3) a reproducible benchmark of several HPO methods against our model library, including novel graph-based and multiobjective methods. Xuan Zhang 0008, Kevin Duh |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | AMR Parsing as Sequence-to-Graph TransductionabstractWe propose an attention-based model that treats AMR parsing as sequence-to-graph transduction.Unlike most AMR parsers that rely on pre-trained aligners, external semantic resources, or data augmentation, our proposed parser is aligner-free, and it can be effectively trained with limited amounts of labeled AMR data.Our experimental results outperform all previously reported SMATCH scores, on both AMR 2.0 (76.3% F1 on LDC2017T10) and AMR 1.0 (70.2% F1 on LDC2014T12). Sheng Zhang 0012, Xutai Ma, Kevin Duh, Benjamin Van Durme |
ACL (1) | 3 |
| 2019 | Multilingual End-to-End Speech TranslationabstractIn this paper, we propose a simple yet effective framework for multilingual end-to-end speech translation (ST), in which speech utterances in source languages are directly translated to the desired target languages with a universal sequence-to-sequence architecture. While multilingual models have shown to be useful for automatic speech recognition (ASR) and machine translation (MT), this is the first time they are applied to the end-to-end ST problem. We show the effectiveness of multilingual end-to-end ST in two scenarios: one-to-many and many-to-many translations with publicly available data. We experimentally confirm that multilingual end-to-end ST models significantly outperform bilingual ones in both scenarios. The generalization of multilingual training is also evaluated in a transfer learning scenario to a very low-resource language pair. All of our codes and the database are publicly available to encourage further research in this emergent multilingual ST topic11Available at https://github.com/espnet/espnet.. Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, Shinji Watanabe 0001 |
ASRU | 2 |
| 2019 | HABLex: Human Annotated Bilingual Lexicons for Experiments in Machine TranslationabstractBrian Thompson, Rebecca Knowles, Xuan Zhang, Huda Khayrallah, Kevin Duh, Philipp Koehn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Brian Thompson 0001, Rebecca Knowles, Xuan Zhang 0008, Huda Khayrallah, Kevin Duh, Philipp Koehn |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Broad-Coverage Semantic Parsing as TransductionabstractSheng Zhang, Xutai Ma, Kevin Duh, Benjamin Van Durme. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sheng Zhang 0012, Xutai Ma, Kevin Duh, Benjamin Van Durme |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Call for Prudent Choice of Subword Merge Operations in Neural Machine Translation
Shuoyang Ding, Adithya Renduchintala, Kevin Duh |
MTSummit (1) | 3 |
| 2019 | Identifying Fluently Inadequate Output in Neural and Statistical Machine Translation
Marianna J. Martindale, Marine Carpuat, Kevin Duh, Paul McNamee |
MTSummit (1) | 3 |
| 2019 | Character-Aware Decoder for Translation into Morphologically Rich Languages
Adithya Renduchintala, Pamela Shapiro, Kevin Duh, Philipp Koehn |
MTSummit (1) | 3 |
| 2019 | Robust Document Representations for Cross-Lingual Information Retrieval in Low-Resource Settings
Mahsa Yarmohammadi, Xutai Ma, Sorami Hisamoto, Muhammad Mahbubur Rahman 0001, Yiming Wang 0006, Hainan Xu, Daniel Povey, Philipp Koehn, Kevin Duh |
MTSummit (1) | 9 |
| 2019 | An Interactive Teaching Tool for Introducing Novices to Machine TranslationabstractThe first step in the research process is developing an understanding of the problem at hand. Novices may be interested in learning about machine translation (MT), but often lack experience and intuition about the task of translation (either by human or machine) and its challenges. The goal of this work is to allow students to interactively discover why MT is an open problem, and encourage them to ask questions, propose solutions, and test intuitions. We present a hands-on activity in which students build and evaluate their own MT systems using curated parallel texts. By having students hand-engineer MT system rules in a simple user interface, which they can then run on real data, they gain intuition about why early MT research took this approach, where it fails, and what features of language make MT a challenging problem even today. Developing translation rules typically strikes novices as an obvious approach that should succeed, but the idea quickly struggles in the face of natural language complexity. This interactive, intuition-building exercise can be augmented by a discussion of state-of-the-art MT techniques and challenges, focusing on areas or aspects of linguistic complexity that the students found difficult. We envision this lesson plan being used in the framework of a larger AI or natural language processing course (where only a small amount of time can be dedicated to MT) or as a standalone activity. We describe and release the tool that supports this lesson, as well as accompanying data. Huda Khayrallah, Rebecca Knowles, Kevin Duh, Matt Post |
SIGCSE | 3 |
| 2019 | Evolution-Strategy-Based Automation of System Development for High-Performance Speech RecognitionabstractThe state-of-the-art large vocabulary speech recognition systems consist of several components including hidden Markov model and deep neural network. To realize the highest recognition performance, numerous meta-parameters specifying the designs and training setups of these components must be optimized. A prominent obstacle in system development is the laborious effort required by human experts in tuning these meta-parameters. To automate the process, we propose to tune the meta-parameters of a whole large vocabulary speech recognition system using the evolution strategy with a multi-objective Pareto optimization. As the result of the evolution, the system is optimized for both low word error rate and compact model size. Since the approach requires repeated training and evaluation of the recognition systems that require large computation, we make use of parallel computation on cloud computers. Experimental results show the effectiveness of the proposed approach by discovering appropriate configuration for large vocabulary speech recognition systems automatically. Takafumi Moriya, Tomohiro Tanaka, Takahiro Shinozaki, Shinji Watanabe 0001, Kevin Duh |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Stochastic Answer Networks for Machine Reading ComprehensionabstractWe propose a simple yet robust stochastic answer network (SAN) that simulates multi-step reasoning in machine reading comprehension.Compared to previous work such as ReasoNet which used reinforcement learning to determine the number of steps, the unique feature is the use of a kind of stochastic prediction dropout on the answer module (final layer) of the neural network during the training.We show that this simple trick improves robustness and achieves results competitive to the state-of-the-art on the Stanford Question Answering Dataset (SQuAD), the Adversarial SQuAD, and the Microsoft MAchine Reading COmprehension Dataset (MS MARCO). Xiaodong Liu 0003, Yelong Shen, Kevin Duh, Jianfeng Gao 0001 |
ACL (1) | 3 |
| 2018 | Cross-lingual Decompositional Semantic ParsingabstractWe introduce the task of cross-lingual decompositional semantic parsing: mapping content provided in a source language into a decompositional semantic analysis based on a target language.We present: (1) a form of decompositional semantic analysis designed to allow systems to target varying levels of structural complexity (shallow to deep analysis), (2) an evaluation metric to measure the similarity between system output and reference semantic analysis, (3) an end-to-end model with a novel annotating mechanism that supports intra-sentential coreference, and (4) an evaluation dataset on which our model outperforms strong baselines by at least 1.75 F 1 score. Sheng Zhang 0012, Xutai Ma, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
EMNLP | 4 |
| 2018 | Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus ProgramabstractCurrently, datasets that support audio-visual recognition of people in videos are scarce and limited. In this paper, we introduce an expansion of video data from the IARPA Janus program to support this research area. We refer to the expanded set, which adds labels for voice to the already-existing face labels, as the Janus Multimedia dataset. We first describe the speaker labeling process, which involved a combination of automatic and manual criteria. We then discuss two evaluation settings for this data. In the core condition, the voice and face of the labeled individual are present in every video. In the full condition, no such guarantee is made. The power of audiovisual fusion is then shown using these publicly-available videos and labels, showing significant improvement over only recognizing voice or face alone. In addition to this work, several other possible paths for future research with this dataset are discussed. Gregory Sell, Kevin Duh, David Snyder, Dave Etter, Daniel Garcia-Romero |
ICASSP | 2 |
| 2018 | Bayesian Analysis in Natural Language Processing Shay Cohen (University of Edinburgh)Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 35), 2016, xxvii+246 pp; paperback, ISBN 9781627058735, $85.00; ebook, ISBN 9781627054218, $68.00; doi: 10.2200/S00719ED1V01Y201605HLT035abstractBayesian techniques are useful tools for modeling a wide range of data and phenomena. Natural language is no exception. The Bayesian approach works as follows:1. Model the data x probabilistically with p(x|θ), where θ are some unknown parameters. For example, this could be a generative story for a sentence x, based on some unknown context-free grammar parameters θ.2. Represent the uncertainty about θ by a prior distribution p(θ).3. Given data, apply Bayes theorem p(θ|x) ∝ p(x|θ)p(θ) to find the posterior distribution for the quantities of interest.This approach enables an elegant and unified way to incorporate prior knowledge and manage uncertainty over parameters. It can also be used to provide capacity control for complex models as an alternative to smoothing. There have been many successful applications of Bayesian techniques in natural language processing (NLP). Some examples include: word segmentation (Goldwater et al. 2009), syntax (Johnson et al. 2007), morphology (Snyder & Barzilay 2008), coreference resolution (Haghighi & Klein 2007), and machine translation (Blunsom et al. 2009).Cohen’s book provides an accessible yet in-depth introduction to Bayesian techniques. It is aimed at a researcher or student who is already familiar with statistical modeling in natural language (i.e., at the level of introductory books such as Manning & Schütze [1999], Jurafsky & Martin [2009]). The stated goal of the book is to “cover the methods and algorithms that are needed to fluently read Bayesian learning papers in NLP and to do research in the area.” I believe Cohen successfully achieves this goal, striking a nice balance between breadth and depth of material.Chapter 1 is a brief review of probability and statistics. It covers prerequisite concepts such as independence, conditional independence, and exchangeability of random variables. The differences between Bayesian and frequentist philosophies are discussed, albeit briefly. In general, the book maintains a pragmatic approach, focusing more on the mathematics and less on the philosophy.Chapter 2 motivates Bayesian techniques in NLP with two distinct examples: latent Dirichlet allocation (LDA) for unsupervised topic modeling and Bayesian linear regression for supervised text analytics. Although Bayesian techniques are most often used for unsupervised problems in NLP (e.g., for addressing problems where the Expectation-Maximization [EM] algorithm fails to find good solutions), the supervised example demonstrates their broader applicability. One aspect that I appreciate about Cohen’s exposition is that he strives to distinguish between what is technically feasible with Bayesian techniques vs. what is frequently used in research. This helps avoid potential misconceptions.Chapter 3 describes the priors p(θ) that are common in NLP: for example, the Dirichlet distribution, the logistic normal distribution, and non-informative priors such as the Jeffreys prior.Chapters 4–6 explain Bayesian inference in detail and form the core of the book. How does one combine the data likelihood and the prior to compute the posterior distribution p(θ|x) ∝ p(x|θ)p(θ)? This can be a computationally difficult problem. Chapter 4 discusses the maximum a posterior (MAP) estimation of p(θ|x). Chapter 5 covers sampling methods, particularly Gibbs sampling, Metropolis-Hastings sampling, and slice sampling. Chapter 6 describes variational inference, the mean-field approximation, and the variational EM algorithm. In each chapter, algorithm variants that are popular in NLP research are discussed in detail. An example is blocked Gibbs sampling using dynamic programming.Chapter 7 focuses on Bayesian nonparameterics. These models allow for an infinite-dimensional parameter space, but the actual number of parameters grows with sample size. This is an active field of research. Cohen does a laudable job of explaining the basic mathematics behind the Dirichlet process, which generalizes the Dirichlet distribution. He describes the stick-breaking and Chinese Restaurant Process viewpoints, and shows how the Dirichlet process can be used as a prior in a nonparametric mixture model and a hierarchical topic model.Chapter 8, the final chapter, demonstrates how Bayesian techniques can be applied to grammar models in NLP. It focuses on parametric and non-parameteric Bayesian models for probabilistic context-free grammars. I wished there was an additional chapter on other applications besides grammar models. However, I think a reader who completes this book would have gained the technical background to do such a survey by themselves.In summary, Cohen’s Bayesian Analysis in Natural Language Processing is a good starting point for a researcher or a student who wishes to learn more about Bayesian techniques. It covers the necessary and sufficient knowledge needed to understand papers in this area, and leaves the remaining details as references. It can be viewed as a concise introduction that complements the many excellent statistics textbooks on Bayesian techniques, which tend to be more detailed.I wish I had this book when I first started to learn Bayesian techniques back in 2009. I recall struggling through several advanced statistics textbooks, before being rescued by Kevin Knight’s classic workbook, “Bayesian Inference with Tears.”1 Cohen’s book might have saved me some Kleenex. Shay B. Cohen, Graeme Hirst, Kevin Duh |
Comput. Linguistics | 3 |
| 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural NetworkabstractLanguage processing mechanism by humans is generally more robust than computers. The Cmabrigde Uinervtisy (Cambridge University) effect from the psycholinguistics literature has demonstrated such a robust word processing mechanism, where jumbled words (e.g. Cmabrigde / Cambridge) are recognized with little cost. On the other hand, computational models for word recognition (e.g. spelling checkers) perform poorly on data with such noise. Inspired by the findings from the Cmabrigde Uinervtisy effect, we propose a word recognition model based on a semi-character level recurrent neural network (scRNN). In our experiments, we demonstrate that scRNN has significantly more robust performance in word spelling correction (i.e. word recognition) compared to existing spelling checkers and character-based convolutional neural network. Furthermore, we demonstrate that the model is cognitively plausible by replicating a psycholinguistics experiment about human reading difficulty using our model. Keisuke Sakaguchi, Kevin Duh, Matt Post, Benjamin Van Durme |
AAAI | 2 |
| 2017 | An Empirical Analysis of Multiple-Turn Reasoning Strategies in Reading Comprehension TasksabstractReading comprehension (RC) is a challenging task that requires synthesis of information across sentences and multiple turns of reasoning. Using a state-of-the-art RC model, we empirically investigate the performance of single-turn and multiple-turn reasoning on the SQuAD and MS MARCO datasets. The RC model is an end-to-end neural network with iterative attention, and uses reinforcement learning to dynamically control the number of turns. We find that multiple-turn reasoning outperforms single-turn reasoning for all question and answer types; further, we observe that enabling a flexible number of turns generally improves upon a fixed multiple-turn strategy. %across all question types, and is particularly beneficial to questions with lengthy, descriptive answers. We achieve results competitive to the state-of-the-art on these two datasets. Yelong Shen, Xiaodong Liu 0003, Kevin Duh, Jianfeng Gao 0001 |
IJCNLP(1) | 3 |
| 2017 | Inference is Everything: Recasting Semantic Resources into a Unified Evaluation FrameworkabstractWe propose to unify a variety of existing semantic classification tasks, such as semantic role labeling, anaphora resolution, and paraphrase detection, under the heading of Recognizing Textual Entailment (RTE). We present a general strategy to automatically generate one or more sentential hypotheses based on an input sentence and pre-existing manual semantic annotations. The resulting suite of datasets enables us to probe a statistical RTE model’s performance on different aspects of semantics. We demonstrate the value of this approach by investigating the behavior of a popular neural network RTE model. Aaron Steven White, Pushpendre Rastogi, Kevin Duh, Benjamin Van Durme |
IJCNLP(1) | 3 |
| 2017 | Selective Decoding for Cross-lingual Open Information ExtractionabstractCross-lingual open information extraction is the task of distilling facts from the source language into representations in the target language. We propose a novel encoder-decoder model for this problem. It employs a novel selective decoding mechanism, which explicitly models the sequence labeling process as well as the sequence generation process on the decoder side. Compared to a standard encoder-decoder model, selective decoding significantly increases the performance on a Chinese-English cross-lingual open IE dataset by 3.87-4.49 BLEU and 1.91-5.92 F1. We also extend our approach to low-resource scenarios, and gain promising improvement. Sheng Zhang 0012, Kevin Duh, Benjamin Van Durme |
IJCNLP(1) | 2 |
| 2017 | Ordinal Common-sense InferenceabstractHumans have the capacity to draw common-sense inferences from natural language: various things that are likely but not certain to hold based on established discourse, and are rarely stated explicitly. We propose an evaluation of automated common-sense inference based on an extension of recognizing textual entailment: predicting ordinal human responses on the subjective likelihood of an inference holding in a given context. We describe a framework for extracting common-sense knowledge from corpora, which is then used to construct a dataset for this ordinal entailment task. We train a neural sequence-to-sequence model on this dataset, which we use to score and generate possible inferences. Further, we annotate subsets of previously established datasets via our ordinal annotation protocol in order to then analyze the distinctions between these and what we have constructed. Sheng Zhang 0012, Rachel Rudinger, Kevin Duh, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | Non-Linear Similarity Learning for CompositionalityabstractMany NLP applications rely on the existence ofsimilarity measures over text data.Although word vector space modelsprovide good similarity measures between words,phrasal and sentential similarities derived from compositionof individual words remain as a difficult problem.In this paper, we propose a new method of ofnon-linear similarity learning for semantic compositionality.In this method, word representations are learnedthrough the similarity learning of sentencesin a high-dimensional space with kernel functions.On the task of predicting the semantic similarity oftwo sentences (SemEval 2014, Task 1),our method outperforms linear baselines,feature engineering approaches,recursive neural networks,and achieve competitive results with long short-term memory models. Masashi Tsubaki, Kevin Duh, Masashi Shimbo, Yuji Matsumoto 0001 |
AAAI | 2 |
| 2016 | Modelling the Usage of Discourse Connectives as Rational Speech ActsabstractDiscourse relations can either be implicit or explicitly expressed by markers, such as 'therefore' and 'but'.How a speaker makes this choice is a question that is not well understood.We propose a psycholinguistic model that predicts whether a speaker will produce an explicit marker given the discourse relation s/he wishes to express.Based on the framework of the Rational Speech Acts model, we quantify the utility of producing a marker based on the information-theoretic measure of surprisal, the cost of production, and a bias to maintain uniform information density throughout the utterance.Experiments based on the Penn Discourse Treebank show that our approach outperforms stateof-the-art approaches, while giving an explanatory account of the speaker's choice.1 'Speakers' and 'listeners' are interchangeably used with 'authors' and 'readers' in this article Frances Yung, Kevin Duh, Taku Komura, Yuji Matsumoto 0001 |
CoNLL | 2 |
| 2016 | A Generalized Framework for Hierarchical Word Sequence Language Model
Xiaoyi Wu, Kevin Duh, Yuji Matsumoto 0001 |
PACLIC | 2 |
| 2016 | Automated structure discovery and parameter tuning of neural network language model based on evolution strategyabstractLong short-term memory (LSTM) recurrent neural network based language models are known to improve speech recognition performance. However, significant effort is required to optimize network structures and training configurations. In this study, we automate the development process using evolutionary algorithms. In particular, we apply the covariance matrix adaptation-evolution strategy (CMA-ES), which has demonstrated robustness in other black box hyper-parameter optimization problems. By flexibly allowing optimization of various meta-parameters including layer wise unit types, our method automatically finds a configuration that gives improved recognition performance. Further, by using a Pareto based multi-objective CMA-ES, both WER and computational time were reduced jointly: after 10 generations, relative WER and computational time reductions for decoding were 4.1% and 22.7% respectively, compared to an initial baseline system whose WER was 8.7%. Tomohiro Tanaka, Takafumi Moriya, Takahiro Shinozaki, Shinji Watanabe 0001, Takaaki Hori, Kevin Duh |
SLT | 6 |
| 2016 | Transition-Based Dependency Parsing Exploiting SupertagsabstractLexical information, including surface word form and part-of-speech (POS) information, plays a crucial role when predicting ambiguous dependency relationships in dependency parsing. However, for resolving dependency ambiguities, surface word information may be too sparse, while POS information may be too coarse. Supertags, which are lexical templates that represent rich syntactic information, have been shown to provide effective features at an intermediate level on the coarse-to-fine scale. In this work, we present a supertag design framework that allows us to instantiate various supertag sets based on the dependency structures. Using this framework, we instantiate various supertag sets and utilize them as features in transition-based dependency parsing systems. Performing experiments on the Penn Treebank and Universal Dependencies data sets, we show that our supertags are effective for transition-based parsers in multilingual parsing as well as English parsing. The comparison of the results of the different supertag sets shows that it is crucial to incorporate the head directionality, head labels, and dependent possession information in supertags to improve the parser performance. Hiroki Ouchi, Kevin Duh, Hiroyuki Shindo, Yuji Matsumoto 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Joint Case Argument Identification for Japanese Predicate Argument Structure AnalysisabstractHiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Hiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto 0001 |
ACL (1) | 3 |
| 2015 | Automation of system building for state-of-the-art large vocabulary speech recognition using evolution strategyabstractWhen building a state-of-the-art speech recognition system, the laborious effort required by human experts in tuning numerous parameters remains a prominent obstacle. The goal of this paper is to automate the process. We propose to tune DNN-HMM based large vocabulary speech recognition systems using the covariance matrix adaptation evolution strategy (CMA-ES) with a multi-objective Pareto optimization. This optimizes systems to achieve both high-accuracy and compact model size. An additional advantage of our approach is that it is efficiently parallelizable and easily adapted to cloud computing services. We performed experiments on the Corpus of Spontaneous Japanese (CSJ) using the TSUBAME 2.5 supercomputer. Compared with a strong manually tuned configuration borrowed from a similar system, our approach automatically discovered systems with lower WER by 0.48%, and systems with 59% smaller model size while keeping WER constant. The optimized training script is released in the Kaldi speech recognition toolkit as the first publicly available recipe for Japanese large vocabulary speech recognition. Takafumi Moriya, Tomohiro Tanaka, Takahiro Shinozaki, Shinji Watanabe 0001, Kevin Duh |
ASRU | 5 |
| 2015 | Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information RetrievalabstractXiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, Ye-yi Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Xiaodong Liu 0003, Jianfeng Gao 0001, Xiaodong He 0001, Li Deng 0001, Kevin Duh, Ye-Yi Wang |
HLT-NAACL | 5 |
| 2015 | Multi-Target Machine Translation with Multi-Synchronous Context-free GrammarsabstractWe propose a method for simultaneously translating from a single source language to multiple target languages T1, T2, etc.The motivation behind this method is that if we only have a weak language model for T1 and translations in T1 and T2 are associated, we can use the information from a strong language model over T2 to disambiguate the translations in T1, providing better translation results.As a specific framework to realize multi-target translation, we expand the formalism of synchronous context-free grammars to handle multiple targets, and describe methods for rule extraction, scoring, pruning, and search with these models.Experiments find that multi-target translation with a strong language model in a similar second target language can provide gains of up to 0.8-1.5 BLEU points. 1 Graham Neubig, Philip Arthur, Kevin Duh |
HLT-NAACL | 3 |
| 2015 | A Hybrid Ranking Approach to Chinese Spelling CheckabstractWe propose a novel framework for Chinese Spelling Check (CSC), which is an automatic algorithm to detect and correct Chinese spelling errors. Our framework contains two key components: candidate generation and candidate ranking . Our framework differs from previous research, such as Statistical Machine Translation (SMT) based model or Language Model (LM) based model, in that we use both SMT and LM models as components of our framework for generating the correction candidates, in order to obtain maximum recall; to improve the precision, we further employ a Support Vector Machines (SVM) classifier to rank the candidates generated by the SMT and the LM. Experiments show that our framework outperforms other systems, which adopted the same or similar resources as ours in the SIGHAN 7 shared task; even comparing with the state-of-the-art systems, which used more resources, such as a considerable large dictionary, an idiom dictionary and other semantic information, our framework still obtains competitive results. Furthermore, to address the resource scarceness problem for training the SMT model, we generate around 2 million artificial training sentences using the Chinese character confusion sets, which include a set of Chinese characters with similar shapes and similar pronunciations, provided by the SIGHAN 7 shared task. Xiaodong Liu 0003, Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | Multilingual Topic Models for Bilingual Dictionary ExtractionabstractA machine-readable bilingual dictionary plays a crucial role in many natural language processing tasks, such as statistical machine translation and cross-language information retrieval. In this article, we propose a framework for extracting a bilingual dictionary from comparable corpora by exploiting a novel combination of topic modeling and word aligners such as the IBM models. Using a multilingual topic model, we first convert a comparable document -aligned corpus into a parallel topic -aligned corpus. This novel topic-aligned corpus is similar in structure to the sentence -aligned corpus frequently employed in statistical machine translation and allows us to extract a bilingual dictionary using a word alignment model. The main advantages of our framework is that (1) no seed dictionary is necessary for bootstrapping the process, and (2) multilingual comparable corpora in more than two languages can also be exploited. In our experiments on a large-scale Wikipedia dataset, we demonstrate that our approach can extract higher precision dictionaries compared to previous approaches and that our method improves further as we add more languages to the dataset. Xiaodong Liu 0003, Kevin Duh, Yuji Matsumoto 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2014 | Improving Dependency Parsers with SupertagsabstractTransition-based dependency parsing systems can utilize rich feature representations.However, in practice, features are generally limited to combinations of lexical tokens and part-of-speech tags.In this paper, we investigate richer features based on supertags, which represent lexical templates extracted from dependency structure annotated corpus.First, we develop two types of supertags that encode information about head position and dependency relations in different levels of granularity.Then, we propose a transition-based dependency parser that incorporates the predictions from a CRF-based supertagger as new features.On standard English Penn Treebank corpus, we show that our supertag features achieve parsing improvements of 1.3% in unlabeled attachment, 2.07% root attachment, and 3.94% in complete tree accuracy. Hiroki Ouchi, Kevin Duh, Yuji Matsumoto 0001 |
EACL | 2 |
| 2014 | Analysis and Prediction of Unalignable Words in Parallel TextabstractProfessional human translators usually do not employ the concept of word alignments, producing translations 'sense-forsense' instead of 'word-for-word'.This suggests that unalignable words may be prevalent in the parallel text used for machine translation (MT).We analyze this phenomenon in-depth for Chinese-English translation.We further propose a simple and effective method to improve automatic word alignment by pre-removing unalignable words, and show improvements on hierarchical MT systems in both translation directions. Frances Yung, Kevin Duh, Yuji Matsumoto 0001 |
EACL | 2 |
| 2014 | Parsing Chinese Synthetic Words with a Character-based Dependency Model
Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001 |
LREC | 2 |
| 2013 | Topic Models + Word Alignment = A Flexible Framework for Extracting Bilingual Dictionary from Comparable Corpus
Xiaodong Liu 0003, Kevin Duh, Yuji Matsumoto 0001 |
CoNLL | 2 |
| 2013 | Modeling and Learning Semantic Co-Compositionality through Prototype Projections and Neural NetworksabstractWe present a novel vector space model for semantic co-compositionality.Inspired by Generative Lexicon Theory (Pustejovsky, 1995), our goal is a compositional model where both predicate and argument are allowed to modify each others' meaning representations while generating the overall semantics.This readily addresses some major challenges with current vector space models, notably the polysemy issue and the use of one representation per word type.We implement cocompositionality using prototype projections on predicates/arguments and show that this is effective in adapting their word representations.We further cast the model as a neural network and propose an unsupervised algorithm to jointly train word representations with co-compositionality.The model achieves the best result to date (ρ = 0.47) on the semantic similarity task of transitive verbs (Grefenstette and Sadrzadeh, 2011). Masashi Tsubaki, Kevin Duh, Masashi Shimbo, Yuji Matsumoto 0001 |
EMNLP | 2 |
| 2013 | What Information is Helpful for Dependency Based Semantic Role Labeling
Yanyan Luo, Kevin Duh, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2013 | Multi-Metric Optimization Using Ensemble Tuning
Baskaran Sankaran, Anoop Sarkar, Kevin Duh |
HLT-NAACL | 3 |
| 2013 | Syntax-Based Post-Ordering for Efficient Japanese-to-English TranslationabstractThis article proposes a novel reordering method for efficient two-step Japanese-to-English statistical machine translation (SMT) that isolates reordering from SMT and solves it after lexical translation. This reordering problem, called post-ordering , is solved as an SMT problem from Head-Final English (HFE) to English. HFE is syntax-based reordered English that is very successfully used for reordering with English-to-Japanese SMT. The proposed method incorporates its advantage into the reverse direction, Japanese-to-English, and solves the post-ordering problem by accurate syntax-based SMT with target language syntax. Two-step SMT with the proposed post-ordering empirically reduces the decoding time of the accurate but slow syntax-based SMT by its good approximation using intermediate HFE. The proposed method improves the decoding speed of syntax-based SMT decoding by about six times with comparable translation accuracy in Japanese-to-English patent translation experiments. Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2012 | Learning to Translate with Multiple Objectives
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata |
ACL (1) | 1 |
| 2012 | Creating Stories: Social Curation of Twitter Messages
Kevin Duh, Tsutomu Hirao, Akisato Kimura, Katsuhiko Ishiguro, Tomoharu Iwata, Ching-man Au Yeung |
ICWSM | 1 |
| 2012 | Bidirectional Semi-supervised Learning with Graphs
Tomoharu Iwata, Kevin Duh |
ECML/PKDD (2) | 2 |
| 2012 | Flexible sample selection strategies for transfer learning in ranking
Kevin Duh, Akinori Fujino |
Inf. Process. Manag. | 1 |
| 2012 | HPSG-Based Preprocessing for English-to-Japanese TranslationabstractJapanese sentences have completely different word orders from corresponding English sentences. Typical phrase-based statistical machine translation (SMT) systems such as Moses search for the best word permutation within a given distance limit (distortion limit). For English-to-Japanese translation, we need a large distance limit to obtain acceptable translations, and the number of translation candidates is extremely large. Therefore, SMT systems often fail to find acceptable translations within a limited time. To solve this problem, some researchers use rule-based preprocessing approaches, which reorder English words just like Japanese by using dozens of rules. Our idea is based on the following two observations: (1) Japanese is a typical head-final language, and (2) we can detect heads of English sentences by a head-driven phrase structure grammar (HPSG) parser. The main contributions of this article are twofold: First, we demonstrate how off-the-shelf, state-of-the-art HPSG parser enables us to write the reordering rules in an abstract level and can easily improve the quality of English-to-Japanese translation. Second, we also show that syntactic heads achieve better results than semantic heads. The proposed method outperforms the best system of NTCIR-7 PATMT EJ task. Hideki Isozaki, Katsuhito Sudoh, Hajime Tsukada, Kevin Duh |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2011 | Providing Cross-Lingual Editing Assistance to Wikipedia Editors
Ching-man Au Yeung, Kevin Duh, Masaaki Nagata |
CICLing (2) | 2 |
| 2011 | Generalized Minimum Bayes Risk System Combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata |
IJCNLP | 1 |
| 2011 | Distributed Minimum Error Rate Training of SMT using Particle Swarm Optimization
Jun Suzuki 0001, Kevin Duh, Masaaki Nagata |
IJCNLP | 2 |
| 2011 | Extracting Pre-ordering Rules from Predicate-Argument Structures
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
IJCNLP | 3 |
| 2011 | Alignment Inference and Bayesian Adaptation for Machine Translation
Kevin Duh, Katsuhito Sudoh, Tomoharu Iwata, Hajime Tsukada |
MTSummit | 1 |
| 2011 | Post-ordering in Statistical Machine Translation
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
MTSummit | 3 |
| 2011 | Extracting Pre-ordering Rules from Chunk-based Dependency Trees for Japanese-to-English Translation
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
MTSummit | 3 |
| 2011 | Semi-supervised ranking for document retrieval
Kevin Duh, Katrin Kirchhoff |
Comput. Speech Lang. | 1 |
| 2010 | Hierarchical Phrase-based Machine Translation with Word-based Reordering Model
Katsuhiko Hayashi 0001, Hajime Tsukada, Katsuhito Sudoh, Kevin Duh, Seiichi Yamamoto |
COLING | 4 |
| 2010 | Automatic Evaluation of Translation Quality for Distant Language Pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, Hajime Tsukada |
EMNLP | 3 |
| 2008 | Learning to rank with partially-labeled dataabstractRanking algorithms, whose goal is to appropriately order a set of objects/documents, are an important component of information retrieval systems. Previous work on ranking algorithms has focused on cases where only labeled data is available for training (i.e. supervised learning). In this paper, we consider the question whether unlabeled (test) data can be exploited to improve ranking performance. We present a framework for transductive learning of ranking functions and show that the answer is affirmative. Our framework is based on generating better features from the test data (via KernelPCA) and incorporating such features via Boosting, thus learning different ranking functions adapted to the individual test queries. We evaluate this method on the LETOR (TREC, OHSUMED) dataset and demonstrate significant improvements. Kevin Duh, Katrin Kirchhoff |
SIGIR | 1 |
| 2006 | Lexicon Acquisition for Dialectal Arabic Using Transductive Learning
Kevin Duh, Katrin Kirchhoff |
EMNLP | 1 |
| 2006 | Multilingual Dependency Parsing using Bayes Point Machines
Simon Corston-Oliver, Anthony Aue, Kevin Duh, Eric K. Ringger |
HLT-NAACL | 3 |
| 2006 | Morphology-based language modeling for conversational Arabic speech recognition
Katrin Kirchhoff, Dimitra Vergyri, Jeff A. Bilmes, Kevin Duh, Andreas Stolcke |
Comput. Speech Lang. | 4 |
| 2005 | Jointly Labeling Multiple Sequences: A Factorial HMM Approach
Kevin Duh |
ACL | 1 |
| 2005 | Genetic triangulation of graphical models for speech and language processingabstractGraphical models are an increasingly popular approach for speech and language processing. As researchers design ever more complex models it becomes crucial to find triangulations that make inference problems tractable. This paper presents a genetic algorithm for triangulation search that is well-suited for speech and language graphical models. It is unique in two ways: First, it can find triangulations appropriate for graphs with a mix of stochastic and deterministic dependencies. Second, the search is guided by optimizing the inference speed (CPU runtime) on real data. We show results on 10 real-world speech and language graphs and demonstrate inference speed-ups over standard triangulation methods. 1. Chris D. Bartels, Kevin Duh, Jeff A. Bilmes, Katrin Kirchhoff, Simon King 0001 |
INTERSPEECH | 2 |
| 2004 | Automatic Learning of Language Model Structure
Kevin Duh, Katrin Kirchhoff |
COLING | 1 |
| 2004 | Morphology-based language modeling for arabic speech recognitionabstractLanguage modeling is a difficult problem for languages with rich morphology. In this paper we investigate the use of morphology-based language models at different stages in a speech recognition system for conversational Arabic. Class-based and single-stream factored language models using mor-phological word representations are applied within an N-best list rescoring framework. In addition, we explore the use of factored language models in first-pass recognition, which is fa-cilitated by two novel procedures: the data-driven optimization of a multi-stream language model structure, and the conversion of a factored language model to a standard word-based model. We evaluate these techniques on a large-vocabulary recognition task and demonstrate that they lead to perplexity and word error rate reductions. 1. Dimitra Vergyri, Katrin Kirchhoff, Kevin Duh, Andreas Stolcke |
INTERSPEECH | 3 |