EDBT 2026 Demo / reviewers in the wild / expert
Graham Neubig
dblp:03/8155
· DBLP profile ↗
336ranked-venue papers
15as first author
139since 2021 · last 2026
0000-0002-2072-3789ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 307 · 14 first-author · 132 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Massively Multilingual Joint Segmentation and GlossingabstractMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria R. Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer |
ACL (1) | 7 |
| 2026 | Gained in Translation: Privileged Pairwise Judges Enhance Multilingual ReasoningabstractWhen asked a question in a language less seen in its training data, current reasoning large language models (RLMs) often exhibit dramatically lower performance than when asked the same question in English.In response, we introduce SP3F (Self-Play with Privileged Pairwise Feedback), a two-stage framework for enhancing multilingual reasoning without any data in the target language(s).First, we supervise fine-tune (SFT) on translated versions of English question-answer pairs to raise base model correctness.Second, we perform RL with feedback from a pairwise judge in a self-play fashion (Swamy et al., 2024), with the judge receiving the English reference response as privileged information.Thus, even when none of the model's responses are completely correct, the privileged pairwise judge can still tell which response is better.End-to-end, SP3F greatly improves base model performance, even outperforming fully post-trained models on multiple math and non-math tasks with less than 1/8 of the training data across the single-language, multilingual, and generalization to unseen language settings.Our key insight is that we can use English reference responses during both SFT and RL by framing both learning problems in terms of translation.In particular, we use reference responses as data for translation during SFT and as privileged information for the pairwise judge during downstream RL.More explicitly, our contribution is three-fold: Lintang Sutawika, Gokul Swamy 0001, Steven Z. Wu, Graham Neubig |
ACL (1) | 4 |
| 2026 | Code with Me or for Me? How Increasing AI Automation Transforms Developer WorkflowsabstractDevelopers now have access to a growing array of increasingly autonomous AI tools for software development. While many studies examine copilots that provide chat assistance or code completions, evaluations of coding agents—which can automatically write files and run code—still rely on static benchmarks. We present the first controlled study of developer interactions with coding agents, characterizing how more autonomous AI tools affect productivity and experience. We evaluate two leading copilot and agentic coding assistants, recruiting participants who regularly use the former. Our results show agents can assist developers in ways that surpass copilots (e.g., completing tasks humans may not have accomplished) and reduce the effort required to finish tasks. Yet challenges remain for broader adoption, including ensuring users adequately understand agent behaviors. Our findings reveal how workflows shift with coding agents and how interactions differ from copilots, motivating recommendations for researchers and highlighting challenges in adopting agentic systems. Valerie Chen, Ameet Talwalkar, Robert Brennan, Graham Neubig |
CHI | 4 |
| 2026 | FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
Atsunori Moteki, Akiyoshi Uchida, Shoichi Masui, Fan Yang 0032, Kanji Uchino, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Koki Nakagawa, Shan Jiang 0006 |
ICPR (4) | 9 |
| 2026 | Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki |
LREC | 6 |
| 2026 | Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human ActivitiesabstractPeriodic human activities with implicit workflows are common in manufacturing, sports, and daily life. While short-term periodic activities—characterized by simple structures and high-contrast patterns—have been widely studied, long-term periodic workflows with low-contrast patterns remain largely underexplored. To bridge this gap, we introduce the first benchmark comprising 580 multimodal human activity sequences featuring long-term periodic workflows. The benchmark supports three evaluation tasks aligned with real-world applications: unsupervised periodic workflow detection, task completion tracking, and procedural anomaly detection. We also propose a lightweight, training-free baseline for modeling diverse periodic workflow patterns. Experiments show that: (i) our benchmark presents significant challenges to both unsupervised periodic detection methods and zero-shot approaches based on powerful large language models (LLMs); (ii) our baseline outperforms competing methods by a substantial margin in all evaluation tasks; and (iii) in real-world applications, our baseline demonstrates deployment advantages on par with traditional supervised workflow detection approaches, eliminating the need for annotation and retraining. Our project page is https://sites.google.com/view/periodicworkflow. Fan Yang 0032, Quanting Xie, Atsunori Moteki, Shoichi Masui, Shan Jiang 0006, Kanji Uchino, Yonatan Bisk, Graham Neubig |
WACV | 8 |
| 2026 | MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and LinkingabstractAbstract This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hindi, Japanese, Indonesian, Vietnamese, and Tamil, featuring over 7,000 named entity mentions linked to 2,500 unique Wikidata entities. We also include several benchmarks using multilingual and multimodal entity linking methods exploring different language models like LLaMa-2 and Aya-23. Our findings indicate that incorporating visual data improves the accuracy of entity linking, especially for entities where the textual context is ambiguous or insufficient, and particularly for models that do not have strong multilingual abilities. For the work, the dataset, methods are available online.1 Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 8 |
| 2025 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleabstractJiawei Guo, Tianyu Zheng, Yizhi Li, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyu Zheng, Yuelin Bai, Bo Li 0080, Yubo Wang 0019, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue |
ACL (1) | 8 |
| 2025 | Evaluating Language Models as Synthetic Data GeneratorsabstractSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan 0002, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig |
ACL (1) | 10 |
| 2025 | BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language ModelsabstractLanguage model evaluation is a daunting task: prompts are brittle, corpus-level perplexities are vague, and the choice of benchmarks are endless.Finding examples that show meaningful, generalizable differences between two LMs is crucial to understanding where one model succeeds and another fails.Can this process be done automatically?In this work, we propose methodology for automated comparison of language models that uses performance-aware contextual embeddings to find fine-grained features of text where one LM outperforms another.Our method, which we name BEHAVIORBOX, extracts coherent features that demonstrate differences with respect to the ease of generation between two LMs.Specifically, BEHAVIORBOX finds features that describe groups of words in fine-grained contexts, such as conditional 'were' in the phrase 'if you were' and exclamation marks after emotional statements, where one model outperforms another within a particular datatset.We apply BEHAVIORBOX to compare models that vary in size, model family, and post-training, and enumerate insights into specific contexts that illustrate meaningful differences in performance which cannot be found by measures such as corpus-level perplexity alone.1 Lindia Tjuatja, Graham Neubig |
ACL (1) | 2 |
| 2025 | Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse AttentionabstractMany-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks.However, this shifts the computational burden from training-time to inference-time, making deployment of many-shot ICL challenging to justify in-practice.This cost is further increased if a custom demonstration set is retrieved for each inference example.We present Dynamic Block-Sparse Attention, a training-free framework for retrieval-based many-shot in-context learning.By combining carefully designed blocksparse attention and retrieval of cached groups of demonstrations, we achieve comparable perexample latency to finetuning while maintaining on average >95% of the best method's accuracy across strong ICL and finetuning baselines.We hope that this will further enable the deployment of many-shot ICL at scale. 1 Emily Xiao, Chin-Jou Li, Graham Neubig, Amanda Bertsch |
ACL (1) | 4 |
| 2025 | MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkabstractXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig |
ACL (1) | 13 |
| 2025 | Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. UltimateabstractThis study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts. Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki |
CoG | 7 |
| 2025 | AutoPresent: Designing Structured Visuals from ScratchabstractDesigning structured visuals such as presentation slides is essential for communicative needs, necessitating both content creation and visual planning skills. In this work, we tackle the challenge of automated slide generation, where models produce slide presentations from natural language (NL) instructions. We first introduce the SlidesBench benchmark, the first benchmark for slide generation with 7k training and 585 testing examples derived from 310 slide decks across 10 domains. SlidesBench supports evaluations that are (i) reference-based to measure similarity to a target slide, and (ii) reference-free to measure the design quality of generated slides alone. We benchmark end-to-end image generation and program generation methods with a variety of models, and find that programmatic methods produce higher-quality slides in user-interactable formats. Built on the success of program generation, we create AutoPresent, an 8B LlaMa-based model trained on 7k pairs of instructions paired with code for slide generation, and achieve results comparable to the closed-source model GPT-4O. We further explore iterative design refinement where the model is tasked to self-refine its own output, and we found that this process improves the slide’s quality. We hope that our work will provide a basis for future work on generating structured visuals. Our code, data, demo, and video demonstrations are publicly available at https: //github.com/para-lost/AutoPresent Jiaxin Ge, Zhiruo Wang 0001, Yi-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, Trevor Darrell |
CVPR | 10 |
| 2025 | Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design DecisionsabstractEmmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Michael Chen, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig |
EMNLP | 12 |
| 2025 | Grounding Multilingual Multimodal LLMs With Cultural KnowledgeabstractMultimodal Large Language Models excel in high-resource settings, but often misinterpret long-tail cultural entities and underperform in low-resource languages.To address this gap, we propose a data-centric approach that directly grounds MLLMs in cultural knowledge.Leveraging a large scale knowledge graph from Wikidata, we collect images that represent culturally significant entities, and generate synthetic multilingual visual question answering data.The resulting dataset, CulturalGround, comprises 22 million high-quality, culturallyrich VQA pairs spanning 42 countries and 39 languages.We train an open-source MLLM CulturalPangea on CulturalGround, interleaving standard multilingual instruction-tuning data to preserve general abilities.Cultural-Pangea achieves state-of-the-art performance among open models on various culture-focused multilingual multimodal benchmarks, outperforming prior models by an average of +5.0% without degrading results on mainstream vision-language tasks.Our findings show that our targeted, culturally grounded approach could substantially narrow the cultural gap in MLLMs and offer a practical path towards globally inclusive multimodal systems. Jean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham Neubig |
EMNLP | 4 |
| 2025 | Harnessing Webpage UIs for Text-Rich Visual UnderstandingabstractText-rich visual understanding—the ability to interpret both textual content and visual elements within a scene—is crucial for multimodal large language models (MLLMs) to effectively interact with structured environments. We propose leveraging webpage UIs as a naturally structured and diverse data source to enhance MLLMs’ capabilities in this area. Existing approaches, such as rule-based extraction, multimodal model captioning, and rigid HTML parsing, are hindered by issues like noise, hallucinations, and limited generalization. To overcome these challenges, we introduce MultiUI, a dataset of 7.3 million samples spanning various UI types and tasks, structured using enhanced accessibility trees and task taxonomies. By scaling multimodal instructions from web UIs through LLMs, our dataset enhances generalization beyond web domains, significantly improving performance in document understanding, GUI comprehension, grounding, and advanced agent tasks. This demonstrates the potential of structured web data to elevate MLLMs’ proficiency in processing text-rich visual environments and generalizing across domains. Junpeng Liu 0001, Tianyue Ou, Yifan Song 0002, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, Xiang Yue |
ICLR | 8 |
| 2025 | Repetition Improves Language Model EmbeddingsabstractBidirectional models are considered essential for strong text embeddings. Recent approaches to adapt autoregressive language models (LMs) into strong text embedding models have largely had the requirement to modify the LM architecture to be bidirectional. We challenge this premise by introducing ``echo embeddings'' which converts autoregressive LMs into high quality text embedding models \emph{without} changing the architecture or requiring fine-tuning. By repeating the input and extracting embeddings from the repeated tokens—which have access to all original tokens—echo embeddings improve over classical LM embeddings by over 5\% in zero-shot settings. Our zero-shot embeddings nearly match those obtained by bidirectionally-converted LMs that undergo additional masked-language modeling training. Echo embeddings are also compatible with supervised fine-tuning, matching or outperforming bidirectionally-converted LMs in an apples-to-apples comparison, even with an identical compute budget during training and inference. Overall, repetition is a simple and effective strategy to circumvent the need for bidirectional attention in embedding models, paving the way towards a unified architecture for all NLP tasks. Jacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig, Aditi Raghunathan |
ICLR | 4 |
| 2025 | Better Instruction-Following Through Minimum Bayes RiskabstractGeneral-purpose LLM judges capable of human-level evaluation provide not only a scalable and accurate way of evaluating instruction-following LLMs but also new avenues for supervising and improving their performance. One promising way of leveraging LLM judges for supervision is through Minimum Bayes Risk (MBR) decoding, which uses a reference-based evaluator to select a high-quality output from amongst a set of candidate outputs. In the first part of this work, we explore using MBR decoding as a method for improving the test-time performance of instruction-following LLMs. We find that MBR decoding with reference-based LLM judges substantially improves over greedy decoding, best-of-N decoding with reference-free judges and MBR decoding with lexical and embedding-based metrics on AlpacaEval and MT-Bench. These gains are consistent across LLMs with up to 70B parameters, demonstrating that smaller LLM judges can be used to supervise much larger LLMs. Then, seeking to retain the improvements from MBR decoding while mitigating additional test-time costs, we explore iterative self-training on MBR-decoded outputs. We find that self-training using Direct Preference Optimisation leads to significant performance gains, such that the self-trained models with greedy decoding generally match and sometimes exceed the performance of their base models with MBR decoding. Ian Wu, Patrick Fernandes, Amanda Bertsch, Seungone Kim, Sina Khoshfetrat Pakazad, Graham Neubig |
ICLR | 6 |
| 2025 | Pangea: A Fully Open Multilingual Multimodal LLM for 39 LanguagesabstractDespite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented.
This paper introduces PANGEA, a multilingual multimodal LLM trained on PANGEAINS, a diverse 6M instruction dataset spanning 39 languages. PANGEAINS features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage.
To rigorously assess models' capabilities, we introduce PANGEABENCH, a holistic evaluation suite encompassing 14 datasets covering 47 languages.
Results show that PANGEA significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum. Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim, Jean de Dieu Nyandwi, Simran Khanuja, Anjali Kantharuban, Lintang Sutawika, Sathyanarayanan Ramamoorthy, Graham Neubig |
ICLR | 10 |
| 2025 | RAGGED: Towards Informed Design of Scalable and Stable RAG SystemsabstractRetrieval-augmented generation (RAG) enhances language models by integrating external knowledge, but its effectiveness is highly dependent on system configuration. Improper retrieval settings can degrade performance, making RAG less reliable than closed-book generation. In this work, we introduce RAGGED, a framework for systematically evaluating RAG systems across diverse retriever-reader configurations, retrieval depths, and datasets. Our analysis reveals that reader robustness to noise is the key determinant of RAG stability and scalability. Some readers benefit from increased retrieval depth, while others degrade due to their sensitivity to distracting content. Through large-scale experiments on open-domain, multi-hop, and specialized-domain datasets, we show that retrievers, rerankers, and prompts influence performance but do not fundamentally alter these reader-driven trends. By providing a principled framework and new metrics to assess RAG stability and scalability, RAGGED enables systematic evaluation of retrieval-augmented generation systems, guiding future research on optimizing retrieval depth and model robustness. Jennifer Hsia, Afreen Shaikh, Zhiruo Wang 0001, Graham Neubig |
ICML | 4 |
| 2025 | Training Software Engineering Agents and Verifiers with SWE-GymabstractWe present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories. Xingyao Wang 0002, Graham Neubig, Navdeep Jaitly, Heng Ji 0001, Alane Suhr, Yizhe Zhang 0002 |
ICML | 3 |
| 2025 | Overtrained Language Models Are Harder to Fine-TuneabstractLarge language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work, we challenge this assumption and show that extended pre-training can make models harder to fine-tune, leading to degraded final performance. We term this phenomenon \textbf{catastrophic overtraining}. For example, the instruction-tuned OLMo-1B model pre-trained on 3T tokens leads to over 2\% worse performance on multiple standard LLM benchmarks than its 2.3T token counterpart. Through controlled experiments and theoretical analysis, we show that catastrophic overtraining arises from a systematic increase in the broad sensitivity of pre-trained parameters to modifications, including but not limited to fine-tuning. Our findings call for a critical reassessment of pre-training design that considers the downstream adaptability of the model. Jacob Mitchell Springer, Sachin Goyal, Kaiyue Wen, Tanishq Kumar, Xiang Yue, Sadhika Malladi, Graham Neubig, Aditi Raghunathan |
ICML | 7 |
| 2025 | Agent Workflow MemoryabstractDespite the potential of language model-based agents to solve real-world tasks such as web navigation, current methods still struggle with long-horizon tasks with complex action trajectories. In contrast, humans can flexibly solve complex tasks by learning reusable task workflows from past experiences and using them to guide future actions. To build agents that can similarly benefit from this process, we introduce Agent Workflow Memory (AWM), a method for inducing commonly reused routines, i.e., workflows, and selectively providing workflows to the agent to guide subsequent generations. AWM flexibly applies to both offline and online scenarios, where agents induce workflows from training examples beforehand or from test queries on the fly. We experiment on two major web navigation benchmarks — Mind2Web and WebArena — that collectively cover 1000+ tasks from 200+ domains across travel, shopping, and social media, among others. AWM substantially improves the baseline results by 24.6% and 51.1% relative success rate on Mind2Web and WebArena while reducing the number of steps taken to solve WebArena tasks successfully. Furthermore, online AWM robustly generalizes in cross-task, website, and domain evaluations, surpassing baselines from 8.9 to 14.0 absolute points as train-test task distribution gaps widen. Zhiruo Wang 0001, Jiayuan Mao, Daniel Fried, Graham Neubig |
ICML | 4 |
| 2025 | Demystifying Long Chain-of-Thought ReasoningabstractScaling inference compute has become a key driver of advanced reasoning in large language models (LLMs). A proven approach for scaling inference compute is to generate long chains-of-thought (CoTs), enabling models to engage in structured reasoning strategies such as backtracking and error correction. Reinforcement learning (RL) has emerged as a crucial method for developing these capabilities, yet the conditions under which long CoTs emerge remain unclear, and RL training requires careful design choices. In this study, we systematically investigate the underlying mechanics of long CoT reasoning—examining the factors that enable models to generate extended reasoning trajectories. Through extensive supervised fine-tuning (SFT) and RL experiments, we identify three key findings: 1) while SFT is not strictly necessary, it significantly simplifies training and improves efficiency; 2) reasoning capabilities tend to emerge with increased training compute but are not guaranteed, making reward shaping essential for stabilizing CoT length growth; and 3) scaling verifiable reward signals is critical for RL, and we find that leveraging noisy, web-extracted solutions with filtering mechanisms shows promising potential, particularly in out-of-distribution (OOD) reasoning tasks such as STEM problem-solving. These insights provide practical guidance for optimizing training strategies to enhance long CoT reasoning in LLMs. Shiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig, Xiang Yue |
ICML | 4 |
| 2025 | In-Context Learning with Long-Context Models: An In-Depth ExplorationabstractAmanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon, Jonathan Berant, Matthew R. Gormley, Graham Neubig. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon 0002, Jonathan Berant, Matthew R. Gormley, Graham Neubig |
NAACL (Long Papers) | 7 |
| 2025 | Towards Automatic Evaluation for Image TranscreationabstractSimran Khanuja, Vivek Iyer, Xiaoyu He, Graham Neubig. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Simran Khanuja, Vivek Iyer, Graham Neubig |
NAACL (Long Papers) | 4 |
| 2025 | The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language ModelsabstractSeungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Choi 0001, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Y. Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
NAACL (Long Papers) | 29 |
| 2025 | JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware EvaluationabstractShota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa |
NAACL (Long Papers) | 7 |
| 2025 | What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and LengthabstractLindia Tjuatja, Graham Neubig, Tal Linzen, Sophie Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lindia Tjuatja, Graham Neubig, Tal Linzen, Sophie Hao |
NAACL (Long Papers) | 2 |
| 2025 | Benchmarking Failures in Tool-Augmented Language ModelsabstractEduardo Treviño, Hugo Contant, James Ngai, Graham Neubig, Zora Zhiruo Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Eduardo Treviño, Hugo Contant, James Ngai, Graham Neubig, Zhiruo Wang 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Checklists Are Better Than Reward Models For Aligning Language ModelsabstractLanguage models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this —typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item—using both AI judges and specialized verifier programs—then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods on top of a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks — RLCF is the only method to help on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. We show that RLCF can also be used off-policy to improve Llama 3.1 8B Instruct and OLMo 2 7B Instruct. These results establish rubrics as a key tool for improving language models' support of queries that express a multitude of needs. We release our our dataset of rubrics (WildChecklists), models, and code to the public. Vijay Viswanathan 0002, Yanchao Sun, Xiang Kong, Graham Neubig, Sherry Tongshuang Wu |
NeurIPS | 5 |
| 2025 | TheAgentCompany: Benchmarking LLM Agents on Consequential Real World TasksabstractWe interact with computers on an everyday basis, be it in everyday life or work, and many aspects of work can be done entirely with access to a computer and the Internet. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and affect change in their surrounding environments. But how performant are AI agents at helping to accelerate or even autonomously perform work-related tasks? The answer to this question has important implications for both industry looking to adopt AI into their workflows, and for economic policy to understand the effects that adoption of AI may have on the labor market. To measure the progress of these LLM agents' performance on performing real-world professional tasks, in this paper, we introduce TheAgentCompany, an extensible benchmark for evaluating AI agents that interact with the world in similar ways to those of a digital worker: by browsing the Web, writing code, running programs, and communicating with other coworkers. We build a self-contained environment with internal web sites and data that mimics a small software company environment, and create a variety of tasks that may be performed by workers in such a company. We test baseline agents powered by both closed API-based and open-weights language models (LMs), and find that with the most competitive agent, 30% of the tasks can be completed autonomously. This paints a nuanced picture on task automation with LM agents -- in a setting simulating a real workplace, a good portion of simpler tasks could be solved autonomously, but more difficult long-horizon tasks are still beyond the reach of current systems. For more information and demos, refer to https://the-agent-company.com. Frank F. Xu, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zhiruo Wang 0001, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, Graham Neubig |
NeurIPS | 21 |
| 2024 | Wav2Gloss: Generating Interlinear Glossed Text from SpeechabstractTaiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel Robinson, Jiatong Shi, Shinji Watanabe, Graham Neubig, David Mortensen, Lori Levin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Taiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel R. Robinson, Jiatong Shi, Shinji Watanabe 0001, Graham Neubig, David R. Mortensen, Lori S. Levin |
ACL (1) | 7 |
| 2024 | Instruction-tuned Language Models are Better Knowledge LearnersabstractZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001 |
ACL (1) | 6 |
| 2024 | VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksabstractJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, Daniel Fried. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, Daniel Fried |
ACL (1) | 7 |
| 2024 | SOTOPIA-π: Interactive Learning of Socially Intelligent Language AgentsabstractRuiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, Hao Zhu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ruiyi Wang, Haofei Yu, Wenxin Zhang 0004, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, Hao Zhu 0011 |
ACL (1) | 7 |
| 2024 | HILITE: Human-in-the-loop Interactive Tool for Image EditingabstractImage editing tools have a plethora of commercial and creative applications — content-creation, digital photography, advertisements, graphic design, and development of educational media. The shortcomings of image editing software include difficulty of use and, for AI-based software, reliance on single image editing models, which often poses the dilemma of a tradeoff between image editing quality and user-friendliness. While the performances of individual image editing models have improved with their evolution over time, these singular models are often specialized on specific image editing tasks. In this work, we introduce HILITE, an open-source interactive image editing platform with a human-in-the-loop design that combines six diffusion-based image editing models. For one, HILITE’s accessible and easily-understandable user interface provides a straightforward user workflow from image input and prompt entry to selection of desired output. Secondly, the combination of several models with diverse specializations in turn allows HILITE to generalize on a wide variety of image editing tasks, essentially creating a "one-stop shop" for image editing. Third, HILITE iteratively takes user feedback, which both enhances the user experience and enables collection of crowd-sourced data for image editing. HILITE outperforms two major image editing softwares, OpenAI’s DALL•E 3 and Google’s Imagen 3, across two widely-user quantitative metrics for image editing evaluation. Considering the growing demand for readily-available and high-performing image editing tools, HILITE provides a novel platform design with multifaceted use cases in both business and academia. The platform can be found at https://platform.opennlplabs.org/ or https://platform-deployment.vercel.app/. Arya Pasumarthi, Armaan Sharma, Jainish H. Patel, Ayush Bheemaiah, Subhadra Vadlamannati, Seth Chang, Sophia Li, Eshaan Barkataki, Yutong Zhang 0011, Diyi Yang, Graham Neubig, Simran Khanuja |
IEEE Big Data | 11 |
| 2024 | Evaluating Text-to-Visual Generation with Image-to-Text Generation
Zhiqiu Lin, Deepak Pathak, Xide Xia, Graham Neubig, Pengchuan Zhang, Deva Ramanan |
ECCV (9) | 6 |
| 2024 | VIMI: Grounding Video Generation through Multi-modal InstructionabstractYuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov |
EMNLP | 7 |
| 2024 | GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed TextabstractLanguage documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format.However, there are few existing resources providing large amounts of standardized, easily accessible IGT data, limiting their applicability to linguistic research, and making it difficult to use such data in NLP modeling.We compile the largest existing corpus of IGT data from a variety of sources, covering over 450k examples across 1.8k languages, to enable research on crosslingual transfer and IGT generation.We normalize much of our data to follow a standard set of labels across languages.Furthermore, we explore the task of automatically generating IGT in order to aid documentation projects.As many languages lack sufficient monolingual data, we pretrain a large multilingual model on our corpus.We demonstrate the utility of this model by finetuning it on monolingual corpora, outperforming SOTA models by up to 6.6%.Our pretrained model and dataset are available on Hugging Face. Michael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice, Graham Neubig, Alexis Palmer, Lori S. Levin |
EMNLP | 5 |
| 2024 | An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevanceabstractGiven the rise of multimedia content, human translators increasingly focus on culturally adapting not only words but also other modalities such as images to convey the same meaning. While several applications stand to benefit from this, machine translation systems remain confined to dealing with language in speech and text. In this work, we introduce a new task of translating images to make them culturally relevant. First, we build three pipelines comprising state-of-the-art generative models to do the task. Next, we build a two-part evaluation dataset – (i) concept: comprising 600 images that are cross-culturally coherent, focusing on a single concept per image; and (ii) application: comprising 100 images curated from real-world applications. We conduct a multi-faceted human evaluation of translated images to assess for cultural relevance and meaning preservation. We find that as of today, image-editing models fail at this task, but can be improved by leveraging LLMs and retrievers in the loop. Best pipelines can only translate 5% of images for some countries in the easier concept dataset and no translation is successful for some countries in the application dataset, highlighting the challenging nature of the task. Our project webpage is here: https://machine-transcreation.github.io/image-transcreation and our code, data and model outputs can be found here: https://github.com/simran-khanuja/image-transcreation. Simran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, Graham Neubig |
EMNLP | 4 |
| 2024 | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsabstractSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Y. Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
EMNLP | 7 |
| 2024 | Learning Performance-Improving Code EditsabstractWith the decline of Moore's law, optimizing program performance has become a major focus of software research. However, high-level optimizations such as API and algorithm changes remain elusive due to the difficulty of understanding the semantics of code. Simultaneously, pretrained large language models (LLMs) have demonstrated strong capabilities at solving a wide range of programming tasks. To that end, we introduce a framework for adapting LLMs to high-level program optimization. First, we curate a dataset of performance-improving edits made by human programmers of over 77,000 competitive C++ programming submission pairs, accompanied by extensive unit tests. A major challenge is the significant variability of measuring performance on commodity hardware, which can lead to spurious "improvements." To isolate and reliably evaluate the impact of program optimizations, we design an environment based on the gem5 full system simulator, the de facto simulator used in academia and industry. Next, we propose a broad range of adaptation strategies for code optimization; for prompting, these include retrieval-based few-shot prompting and chain-of-thought, and for finetuning, these include performance-conditioned generation and synthetic data augmentation based on self-play. A combination of these techniques achieves a mean speedup of 6.86$\times$ with eight generations, higher than average optimizations from individual programmers (3.66$\times$). Using our model's fastest generations, we set a new upper limit on the fastest speedup possible for our dataset at 9.64$\times$ compared to using the fastest human submissions available (9.56$\times$). Alexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon 0002, Jacob R. Gardner, Yiming Yang 0002, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, Amir Yazdanbakhsh |
ICLR | 8 |
| 2024 | SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agentsabstract*Humans are social beings*; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and *interact* under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents. Hao Zhu 0011, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap |
ICLR | 10 |
| 2024 | WebArena: A Realistic Web Environment for Building Autonomous AgentsabstractWith advances in generative AI, there is now potential for autonomous agents to manage daily tasks via natural language commands. However, current agents are primarily created and tested in simplified synthetic environments, leading to a disconnect with real-world scenarios. In this paper, we build an environment for language-guided agents that is highly realistic and reproducible. Specifically, we focus on agents that perform tasks on the web, and create an environment with fully functional websites from four common domains: e-commerce, social forum discussions, collaborative software development, and content management. Our environment is enriched with tools (e.g., a map) and external knowledge bases (e.g., user manuals) to encourage human-like task-solving. Building upon our environment, we release a set of benchmark tasks focusing on evaluating the functional correctness of task completions. The tasks in our benchmark are diverse, long-horizon, and designed to emulate tasks that humans routinely perform on the internet. We experiment with several baseline agents, integrating recent techniques such as reasoning before acting. The results demonstrate that solving complex tasks is challenging: our best GPT-4-based agent only achieves an end-to-end task success rate of 14.41%, significantly lower than the human performance of 78.24%. These results highlight the need for further development of robust agents, that current state-of-the-art large language models are far from perfect performance in these real-life tasks, and that \ours can be used to measure such progress.\footnote{Code, data, environment reproduction instructions, video demonstrations are available in the supplementary.} Shuyan Zhou, Frank F. Xu, Hao Zhu 0011, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon 0002, Graham Neubig |
ICLR | 12 |
| 2024 | TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic TasksabstractLanguage models (LMs) can solve tasks such as answering questions about tables or images by writing programs. However, using primitive functions often leads to verbose and error-prone programs, and higher-level functions require expert design. To enable better solutions without human labor, we ask code LMs to curate reusable high-level functions, and use them to write solutions. We present TROVE, a training-free method of inducing a verifiable and efficient toolbox of functions, by generating via using, growing, and periodically trimming the toolbox. On 11 datasets from math, table question answering, and image reasoning tasks, TROVE consistently yields simpler solutions with higher accuracy than baselines using CodeLLaMa and previous methods using GPT, while using 79-98% smaller toolboxes. TROVE further enables 31% faster and 13% more accurate human verification than baselines. With the same pipeline, it creates diverse functions for varied tasks and datasets, providing insights into their individual characteristics. Zhiruo Wang 0001, Graham Neubig, Daniel Fried |
ICML | 2 |
| 2024 | Program-Aided Reasoners (Better) Know What They KnowabstractAnubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Anubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig |
NAACL-HLT | 6 |
| 2024 | DeMuX: Data-efficient Multilingual LearningabstractSimran Khanuja, Srinivas Gowriraj, Lucio Dery, Graham Neubig. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Simran Khanuja, Srinivas Gowriraj, Lucio M. Dery, Graham Neubig |
NAACL-HLT | 4 |
| 2024 | Alignment for HonestyabstractRecent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness. In this paper, we argue for the importance of alignment for \emph{honesty}, ensuring that LLMs proactively refuse to answer questions when they lack knowledge, while still not being overly conservative. However, a pivotal aspect of alignment for honesty involves discerning an LLM's knowledge boundaries, which demands comprehensive solutions in terms of metric development, benchmark creation, and training methodologies. We address these challenges by first establishing a precise problem definition and defining ``honesty'' inspired by the Analects of Confucius. This serves as a cornerstone for developing metrics that effectively measure an LLM's honesty by quantifying its progress post-alignment. Furthermore, we introduce a flexible training framework which is further instantiated by several efficient fine-tuning techniques that emphasize honesty without sacrificing performance on other tasks. Our extensive experiments reveal that these aligned models show a marked increase in honesty, as indicated by our proposed metrics. We open-source all relevant resources to facilitate future research at \url{https://github.com/GAIR-NLP/alignment-for-honesty}. Yuqing Yang 0004, Ethan Chern, Xipeng Qiu, Graham Neubig, Pengfei Liu 0003 |
NeurIPS | 4 |
| 2024 | NaturalBench: Evaluating Vision-Language Models on Natural Adversarial SamplesabstractVision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs still struggle with natural images and questions that humans can easily answer, which we term $\textbf{natural adversarial samples}$. We also find it surprisingly easy to generate these VQA samples from natural image-text corpora using off-the-shelf models like CLIP and ChatGPT. We propose a semi-automated approach to collect a new benchmark, ${\bf NaturalBench}$, for reliably evaluating VLMs with 10,000 human-verified VQA samples. Crucially, we adopt a $\textbf{vision-centric}$ design by pairing each question with two images that yield different answers, preventing ``blind'' solutions from answering without using the images. This makes NaturalBench more challenging than previous benchmarks that can largely be solved with language priors like commonsense knowledge. We evaluate ${\bf 53}$ state-of-the-art VLMs on NaturalBench, showing that models like BLIP-3, LLaVA-OneVision, Cambrian-1, InternLM-XC2, Llama3.2-Vision, Molmo, Qwen2-VL, and even the (closed-source) GPT-4o lag 50%-70% behind human performance (which is above 90%). We analyze why NaturalBench is hard from two angles: (1) ${\bf Compositionality:}$ Solving NaturalBench requires diverse visio-linguistic skills, including understanding attribute bindings, object relationships, and advanced reasoning like logic and counting. To this end, unlike prior work that uses a single tag per sample, we tag each NaturalBench sample with 1 to 8 skill tags for fine-grained evaluation. (2) ${\bf Biases: }$ NaturalBench exposes severe biases in VLMs, as models often choose the same answer regardless of the image. We show that debiasing can be crucial for VLM performance. Lastly, we apply our benchmark curation method to diverse data sources, including long captions (over 100 words) and non-English languages like Chinese and Hindi, highlighting its potential for dynamic evaluations of VLMs. Zhiqiu Lin, Wenxuan Peng, Jean de Dieu Nyandwi, Zixian Ma, Simran Khanuja, Ranjay Krishna, Graham Neubig, Deva Ramanan |
NeurIPS | 9 |
| 2024 | MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesabstractEvaluating large language models (LLMs) is challenging. Traditional ground-truth- based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User- facing evaluation, such as Chatbot Arena, provides reliable signals but is costly and slow. In this work, we propose MixEval, a new paradigm for establishing efficient, gold-standard LLM evaluation by strategically mixing off-the-shelf bench- marks. It bridges (1) comprehensive and well-distributed real-world user queries and (2) efficient and fairly-graded ground-truth-based benchmarks, by matching queries mined from the web with similar queries from existing benchmarks. Based on MixEval, we further build MixEval-Hard, which offers more room for model improvement. Our benchmarks’ advantages lie in (1) a 0.96 model ranking correlation with Chatbot Arena arising from the highly impartial query distribution and grading mechanism, (2) fast, cheap, and reproducible execution (6% of the time and cost of MMLU), and (3) dynamic evaluation enabled by the rapid and stable data update pipeline. We provide extensive meta-evaluation and analysis for our and existing LLM benchmarks to deepen the community’s understanding of LLM evaluation and guide future research directions. Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, Yang You 0001 |
NeurIPS | 7 |
| 2024 | Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at ScaleabstractLLMs can now act as autonomous agents that interact with digital environments and complete specific objectives (e.g., arranging an online meeting). However, accuracy is still far from satisfactory, partly due to a lack of large-scale, direct demonstrations for digital tasks. Obtaining supervised data from humans is costly, and automatic data collection through exploration or reinforcement learning relies on complex environmental and content setup, resulting in datasets that lack comprehensive coverage of various scenarios. On the other hand, there is abundant knowledge that may indirectly assist task completion, such as online tutorials that were created for human consumption. In this work, we present Synatra, an approach that effectively transforms this indirect knowledge into direct supervision at scale. We define different types of indirect knowledge, and carefully study the available sources to obtain it, methods to encode the structure of direct demonstrations, and finally methods to transform indirect knowledge into direct demonstrations. We use 100k such synthetically-created demonstrations to finetune a 7B CodeLlama, and demonstrate that the resulting agent surpasses all comparably sized models on three web-based task benchmarks Mind2Web, MiniWoB++ and WebArena, as well as surpassing GPT-3.5 on WebArena and Mind2Web. In addition, while synthetic demonstrations prove to be only 3% the cost of human demonstrations (at $0.031 each), we show that the synthetic demonstrations can be more effective than an identical number of human demonstrations collected from limited domains. Tianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu 0004, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth 0001, Graham Neubig, Shuyan Zhou |
NeurIPS | 9 |
| 2024 | Divergences between Language Models and Human BrainsabstractDo machines and humans process language in similar ways? Recent research has hinted at the affirmative, showing that human neural activity can be effectively predicted using the internal representations of language models (LMs). Although such results are thought to reflect shared computational principles between LMs and human brains, there are also clear differences in how LMs and humans represent and use language. In this work, we systematically explore the divergences between human and machine language processing by examining the differences between LM representations and human brain responses to language as measured by Magnetoencephalography (MEG) across two datasets in which subjects read and listened to narrative stories. Using an LLM-based data-driven approach, we identify two domains that LMs do not capture well: social/emotional intelligence and physical commonsense. We validate these findings with human behavioral experiments and hypothesize that the gap is due to insufficient representations of social/emotional and physical knowledge in LMs. Our results show that fine-tuning LMs on these domains can improve their alignment with human brain responses. Yuchen Zhou 0004, Emmy Liu, Graham Neubig, Michael J. Tarr, Leila Wehbe |
NeurIPS | 3 |
| 2024 | Large Language Models Enable Few-Shot ClusteringabstractAbstract Unlike traditional unsupervised clustering, semi-supervised clustering allows users to provide meaningful structure to the data, which helps the clustering algorithm to match the user’s intent. Existing approaches to semi-supervised clustering require a significant amount of feedback from an expert to improve the clusters. In this paper, we ask whether a large language model (LLM) can amplify an expert’s guidance to enable query-efficient, few-shot semi-supervised text clustering. We show that LLMs are surprisingly effective at improving clustering. We explore three stages where LLMs can be incorporated into clustering: before clustering (improving input features), during clustering (by providing constraints to the clusterer), and after clustering (using LLMs post-correction). We find that incorporating LLMs in the first two stages routinely provides significant improvements in cluster quality, and that LLMs enable a user to make trade-offs between cost and accuracy to produce desired clusters. We release our code and LLM prompts for the public to use.1 Vijay Viswanathan 0002, Kiril Gashteovski, Carolin Lawrence, Sherry Tongshuang Wu, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey DesignabstractAbstract One widely cited barrier to the adoption of LLMs as proxies for humans in subjective tasks is their sensitivity to prompt wording—but interestingly, humans also display sensitivities to instruction changes in the form of response biases. We investigate the extent to which LLMs reflect human response biases, if at all. We look to survey design, where human response biases caused by changes in the wordings of “prompts” have been extensively explored in social psychology literature. Drawing from these works, we design a dataset and framework to evaluate whether LLMs exhibit human-like response biases in survey questionnaires. Our comprehensive evaluation of nine models shows that popular open and commercial LLMs generally fail to reflect human-like behavior, particularly in models that have undergone RLHF. Furthermore, even if a model shows a significant change in the same direction as humans, we find that they are sensitive to perturbations that do not elicit significant changes in humans. These results highlight the pitfalls of using LLMs as human proxies, and underscore the need for finer-grained characterizations of model behavior.1 Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkwar, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsabstractModern machine learning relies on datasets to develop and validate research ideas.Given the growth of publicly available data, finding the right dataset to use is increasingly difficult.Any research question imposes explicit and implicit constraints on how well a given dataset will enable researchers to answer this question, such as dataset size, modality, and domain.We operationalize the task of recommending datasets given a short natural language description of a research idea, to help people find relevant datasets for their needs.Dataset recommendation poses unique challenges as an information retrieval problem; datasets are hard to directly index for search and there are no corpora readily available for this task.To facilitate this task, we build the DataFinder Dataset which consists of a larger automatically-constructed training set (17.5K queries) and a smaller expertannotated evaluation set (392 queries).Using this data, we compare various information retrieval algorithms on our test set and present a superior bi-encoder retriever for text-based dataset recommendation.This system, trained on the DataFinder Dataset, finds more relevant search results than existing third-party dataset search engines.To encourage progress on dataset recommendation, we release our dataset and models to the public.1 Vijay Viswanathan 0002, Luyu Gao, Sherry Tongshuang Wu, Pengfei Liu 0003, Graham Neubig |
ACL (1) | 5 |
| 2023 | When Does Translation Require Context? A Data-driven, Multilingual ExplorationabstractAlthough proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics.Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way.In this paper, we develop the Multilingual Discourse-Aware (MUDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset.The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context.We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed.We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively.We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.1 Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig |
ACL (1) | 5 |
| 2023 | Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalabstractJohn Wieting, Jonathan Clark, William Cohen, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. John Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 4 |
| 2023 | EXCALIBUR: Encouraging and Evaluating Embodied ExplorationabstractExperience precedes understanding. Humans constantly explore and learn about their environment out of curiosity, gather information, and update their models of the world. On the other hand, machines are either trained to learn passively from static and fixed datasets, or taught to complete specific goal-conditioned tasks. To encourage the development of exploratory interactive agents, we present the EXCALIBUR benchmark. EXCALIBUR allows agents to explore their environment for long durations and then query their understanding of the physical world via inquiries like: “is the small heavy red bowl made from glass?” or “is there a silver spoon heavier than the egg?”. This design encourages agents to perform free-form home exploration without myopia induced by goal conditioning. Once the agents have answered a series of questions, they can renter the scene to refine their knowledge, update their beliefs, and improve their performance on the questions. Our experiments demonstrate the challenges posed by this dataset for the present-day state-of-the-art embodied systems and the headroom afforded to develop new innovative methods. Finally, we present a virtual reality interface that enables humans to seamlessly interact within the simulated world and use it to gather human performance measures. EXCALIBUR affords unique challenges in comparison to presentday benchmarks and represents the next frontier for embodied AI research. Hao Zhu 0011, Raghav Kapoor, So Yeon Min, Winson Han, Jiatai Li, Kaiwen Geng, Graham Neubig, Yonatan Bisk, Aniruddha Kembhavi, Luca Weihs |
CVPR | 7 |
| 2023 | CTC Alignments Improve Autoregressive TranslationabstractBrian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W Black, Shinji Watanabe. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W. Black, Shinji Watanabe 0001 |
EACL | 4 |
| 2023 | Active Retrieval Augmented GenerationabstractZhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu 0033, Jane Dwivedi-Yu, Yiming Yang 0002, Jamie Callan, Graham Neubig |
EMNLP | 9 |
| 2023 | Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingabstractIdioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts.Despite significant advances, machine translation systems still struggle to translate idiomatic expressions.We provide a simple characterization of idiomatic translation and related issues.This allows us to conduct a synthetic experiment revealing a tipping point at which transformer-based machine translation models correctly default to idiomatic translations.To expand multilingual resources, we compile a dataset of ∼ 4k natural sentences containing idiomatic expressions in French, Finnish, and Japanese.To improve translation of natural idioms, we introduce two straightforward yet effective techniques: the strategic upweighting of training loss on potentially idiomatic sentences, and using retrievalaugmented models.This not only improves the accuracy of a strong pretrained MT model on idiomatic sentences by up to 13% in absolute accuracy, but also holds potential benefits for non-idiomatic sentences.1 Emmy Liu, Aditi Chaudhary, Graham Neubig |
EMNLP | 3 |
| 2023 | GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingabstractYueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yueqi Song, Simran Khanuja, Pengfei Liu 0003, Fahim Faisal, Alissa Ostapenko, Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig |
EMNLP | 11 |
| 2023 | CodeBERTScore: Evaluating Code Generation with Pretrained Models of CodeabstractSince the rise of neural natural-language-tocode models (NL→Code) that can generate long expressions and statements rather than a single next-token, one of the major problems has been reliably evaluating their generated output.In this paper, we propose CodeBERTScore: an evaluation metric for code generation, which builds on BERTScore (Zhang et al., 2020).Instead of encoding only the generated tokens as in BERTScore, CodeBERTScore also encodes the natural language input preceding the generated code, thus modeling the consistency between the generated code and its given natural language context as well.We perform an extensive evaluation of CodeBERTScore across four programming languages.We find that Code-BERTScore achieves a higher correlation with human preference and with functional correctness than all existing metrics.That is, generated code that receives a higher score by Code-BERTScore is more likely to be preferred by humans, as well as to function correctly when executed.We release five language-specific pretrained models to use with our publicly available code.Our language-specific models have been downloaded more than 1,000,000 times from the Huggingface Hub. 1 Shuyan Zhou, Uri Alon 0002, Sumit Agarwal, Graham Neubig |
EMNLP | 4 |
| 2023 | AANG : Automating Auxiliary Learning
Lucio M. Dery, Paul Michel, Mikhail Khodak, Graham Neubig, Ameet Talwalkar |
ICLR | 4 |
| 2023 | Computational Language Acquisition with Theory of Mind
Andy Liu, Hao Zhu 0011, Emmy Liu, Yonatan Bisk, Graham Neubig |
ICLR | 5 |
| 2023 | Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, Luke Zettlemoyer |
ICLR | 6 |
| 2023 | DiffusER: Diffusion via Edit-based Reconstruction
Machel Reid, Vincent J. Hellendoorn, Graham Neubig |
ICLR | 3 |
| 2023 | DocPrompting: Generating Code by Retrieving the Docs
Shuyan Zhou, Uri Alon 0002, Frank F. Xu, Zhengbao Jiang, Graham Neubig |
ICLR | 5 |
| 2023 | PAL: Program-aided Language ModelsabstractLarge language models (LLMs) have demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time ("few-shot prompting"). Much of this success can be attributed to prompting methods such as "chain-of-thought", which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the problem. While LLMs seem to be adept at this sort of step-by-step decomposition, LLMs often make logical and arithmetic mistakes in the solution part, even when the problem is decomposed correctly. In this paper, we present Program-Aided Language models (PAL): a novel approach that uses the LLM to read natural language problems and generate programs as the intermediate reasoning steps, but offloads the solution step to a runtime such as a Python interpreter. With PAL, decomposing the natural language problem into runnable steps remains the only learning task for the LLM, while solving is delegated to the interpreter. We demonstrate this synergy between a neural LLM and a symbolic interpreter across 13 mathematical, symbolic, and algorithmic reasoning tasks from BIG-Bench Hard and others. In all these natural language reasoning tasks, generating code using an LLM and reasoning using a Python interpreter leads to more accurate results than much larger models. For example, PAL using Codex achieves state-of-the-art few-shot accuracy on GSM8K, surpassing PaLM which uses chain-of-thought by absolute 15% top-1. Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 0002, Pengfei Liu 0003, Yiming Yang 0002, Jamie Callan, Graham Neubig |
ICML | 8 |
| 2023 | Cross-Modal Fine-Tuning: Align then RefineabstractFine-tuning large-scale pretrained models has led to tremendous progress in well-studied modalities such as vision and NLP. However, similar gains have not been observed in many other modalities due to a lack of relevant pretrained models. In this work, we propose ORCA, a general cross-modal fine-tuning framework that extends the applicability of a single large-scale pretrained model to diverse modalities. ORCA adapts to a target task via an align-then-refine workflow: given the target input, ORCA first learns an embedding network that aligns the embedded feature distribution with the pretraining modality. The pretrained model is then fine-tuned on the embedded data to exploit the knowledge shared across modalities. Through extensive experiments, we show that ORCA obtains state-of-the-art results on 3 benchmarks containing over 60 datasets from 12 modalities, outperforming a wide range of hand-designed, AutoML, general-purpose, and task-specific cross-modal methods. We highlight the importance of data alignment via a series of ablation studies and exemplify ORCA's utility in data-limited regimes. Junhong Shen, Liam Li, Lucio M. Dery, Corey Staten, Mikhail Khodak, Graham Neubig, Ameet Talwalkar |
ICML | 6 |
| 2023 | Why do Nearest Neighbor Language Models Work?abstractLanguage models (LMs) compute the probability of a text by sequentially computing a representation of an already-seen context and using this representation to predict the next word. Currently, most LMs calculate these representations through a neural network consuming the immediate previous context. However recently, retrieval-augmented LMs have shown to improve over standard neural LMs, by accessing information retrieved from a large datastore, in addition to their standard, parametric, next-word prediction. In this paper, we set out to understand why retrieval-augmented language models, and specifically why k-nearest neighbor language models (kNN-LMs) perform better than standard parametric LMs, even when the k-nearest neighbor component retrieves examples from the same training set that the LM was originally trained on. To this end, we perform analysis of various dimensions over which kNN-LM diverges from standard LMs, and investigate these dimensions one by one. Empirically, we identify three main reasons why kNN-LM performs better than standard LMs: using a different input representation for predicting the next tokens, approximate kNN search, and the importance of softmax temperature for the kNN distribution. Further, we incorporate some insights into the standard parametric LM, improving performance without the need for an explicit retrieval component. The code is available at https://github.com/frankxu2004/knnlm-why. Frank F. Xu, Uri Alon 0002, Graham Neubig |
ICML | 3 |
| 2023 | Unlimiformer: Long-Range Transformers with Unlimited Length InputabstractSince the proposal of transformers, these models have been limited to bounded input lengths, because of their need to attend to every token in the input. In this work, we propose Unlimiformer: a general approach that wraps any existing pretrained encoder-decoder transformer, and offloads the cross-attention computation to a single $k$-nearest-neighbor ($k$NN) index, while the returned $k$NN distances are the attention dot-product scores. This $k$NN index can be kept on either the GPU or CPU memory and queried in sub-linear time; this way, we can index practically unlimited input sequences, while every attention head in every decoder layer retrieves its top-$k$ keys, instead of attending to every key. We evaluate Unlimiformer on several long-document and book-summarization benchmarks, showing that it can process even **500k** token-long inputs from the BookSum dataset, without any input truncation at test time. We demonstrate that Unlimiformer improves pretrained models such as BART and Longformer by extending them to unlimited inputs without additional learned weights and without modifying their code. Our code and models are publicly available at https://github.com/abertsch72/unlimiformer , and support LLaMA-2 as well. Amanda Bertsch, Uri Alon 0002, Graham Neubig, Matthew R. Gormley |
NeurIPS | 3 |
| 2023 | Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language GenerationabstractAbstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info. Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 10 |
| 2023 | DIRE and its Data: Neural Decompiled Variable Renamings with Respect to Software ClassabstractThe decompiler is one of the most common tools for examining executable binaries without the corresponding source code. It transforms binaries into high-level code, reversing the compilation process. Unfortunately, decompiler output is far from readable because the decompilation process is often incomplete. State-of-the-art techniques use machine learning to predict missing information like variable names. While these approaches are often able to suggest good variable names in context, no existing work examines how the selection of training data influences these machine learning models. We investigate how data provenance and the quality of training data affect performance, and how well, if at all, trained models generalize across software domains. We focus on the variable renaming problem using one such machine learning model, DIRE . We first describe DIRE in detail and the accompanying technique used to generate training data from raw code. We also evaluate DIRE ’s overall performance without respect to data quality. Next, we show how training on more popular, possibly higher quality code (measured using GitHub stars) leads to a more generalizable model because popular code tends to have more diverse variable names. Finally, we evaluate how well DIRE predicts domain-specific identifiers, propose a modification to incorporate domain information, and show that it can predict identifiers in domain-specific scenarios 23% more frequently than the original DIRE model. Luke Dramko, Jeremy Lacomis, Edward J. Schwartz, Miltiadis Allamanis, Graham Neubig, Bogdan Vasilescu, Claire Le Goues |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2022 | Explain, Edit, and Understand: Rethinking User Study Design for Evaluating Model ExplanationsabstractIn attempts to "explain" predictions of machine learning models, researchers have proposed hundreds of techniques for attributing predictions to features that are deemed important. While these attributions are often claimed to hold the potential to improve human "understanding" of the models, surprisingly little work explicitly evaluates progress towards this aspiration. In this paper, we conduct a crowdsourcing study, where participants interact with deception detection models that have been trained to distinguish between genuine and fake hotel reviews. They are challenged both to simulate the model on fresh reviews, and to edit reviews with the goal of lowering the probability of the originally predicted class. Successful manipulations would lead to an adversarial example. During the training (but not the test) phase, input spans are highlighted to communicate salience. Through our evaluation, we observe that for a linear bag-of-words model, participants with access to the feature coefficients during training are able to cause a larger reduction in model confidence in the testing phase when compared to the no-explanation control. For the BERT-based classifier, popular local explanations do not improve their ability to reduce the model confidence over the no-explanation case. Remarkably, when the explanation for the BERT model is given by the (global) attributions of a linear model trained to imitate the BERT model, people can effectively manipulate the model. Siddhant Arora, Danish Pruthi, Norman M. Sadeh, William W. Cohen, Zachary C. Lipton, Graham Neubig |
AAAI | 6 |
| 2022 | DEEP: DEnoising Entity Pre-training for Neural Machine TranslationabstractIt has been shown that machine translation models usually generate poor translations for named entities that are infrequent in the training corpus.Earlier named entity translation methods mainly focus on phonetic transliteration, which ignores the sentence context for translation and is limited in domain and language coverage.To address this limitation, we propose DEEP, a DEnoising Entity Pretraining method that leverages large amounts of monolingual data and a knowledge base to improve named entity translation accuracy within sentences.Besides, we investigate a multi-task learning strategy that finetunes a pre-trained neural machine translation model on both entity-augmented monolingual data and parallel data to further improve entity translation.Experimental results on three language pairs demonstrate that DEEP results in significant improvements over strong denoising autoencoding baselines, with a gain of up to 1.3 BLEU and up to 9.2 entity accuracy points for English-Russian translation. 1 Junjie Hu 0001, Hiroaki Hayashi, Kyunghyun Cho, Graham Neubig |
ACL (1) | 4 |
| 2022 | Systematic Inequalities in Language Technology Performance across the World's LanguagesabstractNatural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development.While the performance of NLP methods has grown enormously over the last decade, this progress has been restricted to a minuscule subset of the world's ≈6,500 languages.We introduce a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.Our analyses involve the field at large, but also more in-depth studies on both user-facing technologies (machine translation, language understanding, question answering, text-to-speech synthesis) as well as foundational NLP tasks (dependency parsing, morphological inflection).In the process, we (1) quantify disparities in the current state of NLP research, (2) explore some of its associated societal and academic factors, and (3) produce tailored recommendations for evidencebased policy making aimed at promoting more global and equitable language technologies.1 Damián E. Blasi, Antonios Anastasopoulos, Graham Neubig |
ACL (1) | 3 |
| 2022 | AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource LanguagesabstractAbteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann |
ACL (1) | 13 |
| 2022 | BRIO: Bringing Order to Abstractive SummarizationabstractAbstractive summarization models are commonly trained using maximum likelihood estimation, which assumes a deterministic (onepoint) target distribution in which an ideal model will assign all the probability mass to the reference summary.This assumption may lead to performance degradation during inference, where the model needs to compare several system-generated (candidate) summaries that have deviated from the reference summary.To address this problem, we propose a novel training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality.Our method achieves a new state-of-the-art result on the CNN/DailyMail (47.78 ROUGE-1) and XSum (49.07 ROUGE-1) datasets.Further analysis also shows that our model can estimate probabilities of candidate summaries that are more correlated with their level of quality. 1 Yixin Liu 0003, Pengfei Liu 0003, Dragomir R. Radev, Graham Neubig |
ACL (1) | 4 |
| 2022 | Expanding Pretrained Models to Thousands More Languages via Lexicon-based AdaptationabstractThe performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text present in a target language.Thus, the majority of the world's languages cannot benefit from recent progress in NLP as they have no or limited textual data.To expand possibilities of using NLP technology in these under-represented languages, we systematically study strategies that relax the reliance on conventional language resources through the use of bilingual lexicons, an alternative resource with much better language coverage.We analyze different strategies to synthesize textual or labeled data using lexicons, and how this data can be combined with monolingual or parallel text when available.For 19 under-represented languages across 3 tasks, our methods lead to consistent improvements of up to 5 and 15 points with and without extra monolingual text respectively.Overall, our study highlights how NLP methods can be adapted to thousands more languages that are under-served by current technology. 1 Xinyi Wang 0001, Sebastian Ruder, Graham Neubig |
ACL (1) | 3 |
| 2022 | On The Ingredients of an Effective Zero-shot Semantic ParserabstractSemantic parsers map natural language utterances into meaning representations (e.g.programs).Such models are typically bottlenecked by the paucity of training data due to the laborious annotation efforts.Recent studies have performed zero-shot learning by synthesizing training examples of canonical utterances and programs from a grammar, and further paraphrasing these utterances to improve linguistic diversity.However, such synthetic examples cannot fully capture patterns in real data.In this paper we analyze zero-shot parsers through the lenses of the language and logical gaps (Herzig and Berant, 2019), which quantify the discrepancy of language and programmatic patterns between the synthetic canonical examples and real-world user-issued ones.We propose bridging these gaps using improved grammars, stronger paraphrasers, and efficient learning methods using canonical examples that most likely reflect real user intents.Our model achieves strong results on the SCHOLAR and GEO benchmarks with zero labeled data. 1 John Wieting, Avirup Sil, Graham Neubig |
ACL (1) | 4 |
| 2022 | Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web DataabstractShuyan Zhou, Li Zhang, Yue Yang, Qing Lyu, Pengcheng Yin, Chris Callison-Burch, Graham Neubig. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shuyan Zhou, Li Zhang 0039, Yue Yang 0006, Qing Lyu 0001, Chris Callison-Burch, Graham Neubig |
ACL (1) | 7 |
| 2022 | Language Learning from Communicative Goals and Linguistic Input
Hao Zhu 0011, Yonatan Bisk, Graham Neubig |
CogSci | 3 |
| 2022 | Understanding and Improving Zero-shot Multi-hop Reasoning in Generative Question AnsweringabstractGenerative question answering (QA) models generate answers to questions either solely based on the parameters of the model (the closed-book setting) or additionally retrieving relevant evidence (the open-book setting). Generative QA models can answer some relatively complex questions, but the mechanism through which they do so is still poorly understood. We perform several studies aimed at better understanding the multi-hop reasoning capabilities of generative QA models. First, we decompose multi-hop questions into multiple corresponding single-hop questions, and find marked inconsistency in QA models’ answers on these pairs of ostensibly identical question chains. Second, we find that models lack zero-shot multi-hop reasoning ability: when trained only on single-hop questions, models generalize poorly to multi-hop questions. Finally, we demonstrate that it is possible to improve models’ zero-shot multi-hop reasoning capacity through two methods that approximate real multi-hop natural language (NL) questions by training on either concatenation of single-hop questions or logical forms (SPARQL). In sum, these results demonstrate that multi-hop reasoning does not emerge naturally in generative QA models, but can be encouraged by advances in training or modeling techniques. Code is available at https://github.com/jzbjyb/multihop. Zhengbao Jiang, Jun Araki, Haibo Ding, Graham Neubig |
COLING | 4 |
| 2022 | MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionabstractDavid Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, Dietrich Klakow. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing K. Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia V. Taylor, Fatoumata Ouoba Kabore, Chris C. Emezue, Aremu Anuoluwapo, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Tapo, Tebogo Macucwa, Vukosi Marivate, Elvis Mboning, Tajuddeen Rabiu Gwadabe, Tosin P. Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo L. Mokono, Ignatius Ezeani, Chiamaka Ijeoma Chukwuneke, Mofe Adeyemi, Gilles Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu Ngoli, Dietrich Klakow |
EMNLP | 2 |
| 2022 | Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single TransformerabstractSystems for knowledge-intensive tasks such as open-domain question answering (QA) usually consist of two stages: efficient retrieval of relevant documents from a large corpus and detailed reading of the selected documents to generate answers.Retrievers and readers are usually modeled separately, which necessitates a cumbersome implementation and is hard to train and adapt in an end-to-end fashion.In this paper, we revisit this design and eschew the separate architecture and training in favor of a single Transformer that performs Retrieval as Attention (ReAtt), and end-to-end training solely based on supervision from the end QA task.We demonstrate for the first time that a single model trained end-to-end can achieve both competitive retrieval and QA performance, matching or slightly outperforming state-of-the-art separately trained retrievers and readers.Moreover, end-to-end adaptation significantly boosts its performance on out-of-domain datasets in both supervised and unsupervised settings, making our model a simple and adaptable solution for knowledgeintensive tasks.Code and models are available at https://github.com/jzbjyb/ReAtt. Zhengbao Jiang, Luyu Gao, Zhiruo Wang 0001, Jun Araki, Haibo Ding, Jamie Callan, Graham Neubig |
EMNLP | 7 |
| 2022 | Are representations built from the ground up? An empirical examination of local composition in language modelsabstractCompositionality, the phenomenon where the meaning of a phrase can be derived from its constituent parts, is a hallmark of human language.At the same time, many phrases are non-compositional, carrying a meaning beyond that of each part in isolation.Representing both of these types of phrases is critical for language understanding, but it is an open question whether modern language models (LMs) learn to do so; in this work we examine this question.We first formulate a problem of predicting the LM-internal representations of longer phrases given those of their constituents.We find that the representation of a parent phrase can be predicted with some accuracy given an affine transformation of its children.While we would expect the predictive accuracy to correlate with human judgments of semantic compositionality, we find this is largely not the case, indicating that LMs may not accurately distinguish between compositional and non-compositional phrases.We perform a variety of analyses, shedding light on when different varieties of LMs do and do not generate compositional representations, and discuss implications for future modeling work. 1 Emmy Liu, Graham Neubig |
EMNLP | 2 |
| 2022 | Language Models of Code are Few-Shot Commonsense LearnersabstractWe address the general task of structured commonsense reasoning: given a natural language input, the goal is to generate a graph such as an event or a reasoning-graph.To employ large language models (LMs) for this task, existing approaches "serialize" the output graph as a flat list of nodes and edges.Although feasible, these serialized graphs strongly deviate from the natural language corpora that LMs were pre-trained on, hindering LMs from generating them correctly.In this paper, we show that when we instead frame structured commonsense reasoning tasks as code generation tasks, pre-trained LMs of code are better structured commonsense reasoners than LMs of natural language, even when the downstream task does not involve source code at all.We demonstrate our approach across three diverse structured commonsense reasoning tasks.In all these natural language tasks, we show that using our approach, a code generation LM (CODEX) outperforms natural-LMs that are fine-tuned on the target task (e.g., T5) and other strong LMs such as GPT-3 in the few-shot setting.Our code and data are available at https: //github.com/madaan/CoCoGen . Aman Madaan, Shuyan Zhou, Uri Alon 0002, Yiming Yang 0002, Graham Neubig |
EMNLP | 5 |
| 2022 | English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsabstractUniversal cross-lingual sentence embeddings map semantically similar cross-lingual sentences into a shared embedding space.Aligning cross-lingual sentence embeddings usually requires supervised cross-lingual parallel sentences.In this work, we propose mSimCSE, which extends SimCSE (Gao et al., 2021) to multilingual settings and reveal that contrastive learning on English data can surprisingly learn high-quality universal cross-lingual sentence embeddings without any parallel data.In unsupervised and weakly supervised settings, mSim-CSE significantly improves previous sentence embedding methods on cross-lingual retrieval and multilingual STS tasks.The performance of unsupervised mSimCSE is comparable to fully supervised methods in retrieving lowresource languages and multilingual STS.The performance can be further enhanced when cross-lingual NLI data is available.1 Yau-Shian Wang, Ashley Wu, Graham Neubig |
EMNLP | 3 |
| 2022 | Interpreting Language Models with Contrastive ExplanationsabstractModel interpretability methods are often used to explain NLP model decisions on tasks such as text classification, where the output space is relatively small.However, when applied to language generation, where the output space often consists of tens of thousands of tokens, these methods are unable to provide informative explanations.Language models must consider various features to predict a token, such as its part of speech, number, tense, or semantics.Existing explanation methods conflate evidence for all these features into a single explanation, which is less interpretable for human understanding.To disentangle the different decisions in language modeling, we focus on explaining language models contrastively: we look for salient input tokens that explain why the model predicted one token instead of another.We demonstrate that contrastive explanations are quantifiably better than non-contrastive explanations in verifying major grammatical phenomena, and that they significantly improve contrastive model simulatability for human observers.We also identify groups of contrastive decisions where the model uses similar evidence, and we are able to characterize what input tokens models use during various language generation decisions.1 Kayo Yin, Graham Neubig |
EMNLP | 2 |
| 2022 | Should We Be Pre-training? An Argument for End-task Aware Training as an Alternative
Lucio M. Dery, Paul Michel, Ameet Talwalkar, Graham Neubig |
ICLR | 4 |
| 2022 | Towards a Unified View of Parameter-Efficient Transfer Learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig |
ICLR | 5 |
| 2022 | Distributionally Robust Models with Parametric Likelihood Ratios
Paul Michel, Tatsunori B. Hashimoto, Graham Neubig |
ICLR | 3 |
| 2022 | Capturing Structural Locality in Non-parametric Language Models
Frank F. Xu, Junxian He, Graham Neubig, Vincent J. Hellendoorn |
ICLR | 3 |
| 2022 | Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalabstractRetrieval-based language models (R-LM) model the probability of natural language text by combining a standard language model (LM) with examples retrieved from an external datastore at test time. While effective, a major bottleneck of using these models in practice is the computationally costly datastore search, which can be performed as frequently as every time step. In this paper, we present RetoMaton - retrieval automaton - which approximates the datastore search, based on (1) saving pointers between consecutive datastore entries, and (2) clustering of entries into "states". This effectively results in a weighted finite automaton built on top of the datastore, instead of representing the datastore as a flat list. The creation of the automaton is unsupervised, and a RetoMaton can be constructed from any text collection: either the original training corpus or from another domain. Traversing this automaton at inference time, in parallel to the LM inference, reduces its perplexity by up to 1.85, or alternatively saves up to 83% of the nearest neighbor searches over $k$NN-LM (Khandelwal et al., 2020) without hurting perplexity. Our code and trained models are available at https://github.com/neulab/retomaton . Uri Alon 0002, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth 0001, Graham Neubig |
ICML | 6 |
| 2022 | Symmetric Machine Theory of MindabstractTheory of mind, the ability to model others’ thoughts and desires, is a cornerstone of human social intelligence. This makes it an important challenge for the machine learning community, but previous works mainly attempt to design agents that model the "mental state" of others as passive observers or in specific predefined roles, such as in speaker-listener scenarios. In contrast, we propose to model machine theory of mind in a more general symmetric scenario. We introduce a multi-agent environment SymmToM where, like in real life, all agents can speak, listen, see other agents, and move freely through the world. Effective strategies to maximize an agent’s reward require it to develop a theory of mind. We show that reinforcement learning agents that model the mental states of others achieve significant performance improvements over agents with no such theory of mind model. Importantly, our best agents still fail to achieve performance comparable to agents with access to the gold-standard mental state of other agents, demonstrating that the modeling of theory of mind in multi-agent scenarios is very much an open challenge. Melanie Sclar, Graham Neubig, Yonatan Bisk |
ICML | 2 |
| 2022 | VarCLR: Variable Semantic Representation Pre-training via Contrastive LearningabstractVariable names are critical for conveying intended program behavior. Machine learning-based program analysis methods use variable name representations for a wide range of tasks, such as suggesting new variable names and bug detection. Ideally, such methods could capture semantic relationships between names beyond syntactic similarity, e.g., the fact that the names average and mean are similar. Unfortunately, previous work has found that even the best of previous representation approaches primarily capture "relatedness" (whether two variables are linked at all), rather than "similarity" (whether they actually have the same meaning). Jeremy Lacomis, Edward J. Schwartz, Graham Neubig, Bogdan Vasilescu, Claire Le Goues |
ICSE | 4 |
| 2022 | Building African Voices
Perez Ogayo, Graham Neubig, Alan W. Black |
INTERSPEECH | 2 |
| 2022 | Quality-Aware Decoding for Neural Machine TranslationabstractPatrick Fernandes, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, Andre Martins. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Patrick Fernandes, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, Perez Ogayo, Graham Neubig, André F. T. Martins |
NAACL-HLT | 6 |
| 2022 | OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question AnsweringabstractZhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhengbao Jiang, Graham Neubig, Weizhu Chen |
NAACL-HLT | 4 |
| 2022 | Testing the Ability of Language Models to Interpret Figurative LanguageabstractEmmy Liu, Chenxuan Cui, Kenneth Zheng, Graham Neubig. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Emmy Liu, Chenxuan Cui, Kenneth Zheng, Graham Neubig |
NAACL-HLT | 4 |
| 2022 | Learning to Scaffold: Optimizing Model Explanations for TeachingabstractModern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior. However, what is the precise goal of providing such explanations, and how can we demonstrate that explanations achieve this goal? Some research argues that explanations should help teach a student (either human or machine) to simulate the model being explained, and that the quality of explanations can be measured by the simulation accuracy of students on unexplained examples. In this work, leveraging meta-learning techniques, we extend this idea to improve the quality of the explanations themselves, specifically by optimizing explanations such that student models more effectively learn to simulate the original model. We train models on three natural language processing and computer vision tasks, and find that students trained with explanations extracted with our framework are able to simulate the teacher significantly more effectively than ones produced with previous methods. Through human annotations and a user study, we further find that these learned explanations more closely align with how humans would explain the required decisions in these tasks. Our code is available at https://github.com/coderpat/learning-scaffold. Patrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins, Graham Neubig |
NeurIPS | 5 |
| 2022 | Augmenting Decompiler Output with Learned Variable Names and Types
Jeremy Lacomis, Edward J. Schwartz, Claire Le Goues, Graham Neubig, Bogdan Vasilescu |
USENIX Security Symposium | 5 |
| 2022 | Can We Automate Scientific Reviewing?abstractThe rapid development of science and technology has been accompanied by an exponential growth in peer-reviewed scientific publications. At the same time, the review of each paper is a laborious process that must be carried out by subject matter experts. Thus, providing high-quality reviews of this growing number of papers is a significant challenge. In this work, we ask the question “can we automate scientific reviewing? ”, discussing the possibility of using natural language processing (NLP) models to generate peer reviews for scientific papers. Because it is non-trivial to define what a “good” review is in the first place, we first discuss possible evaluation metrics that could be used to judge success in this task. We then focus on the machine learning domain and collect a dataset of papers in the domain, annotate them with different aspects of content covered in each review, and train targeted summarization models that take in papers as input and generate reviews as output. Comprehensive experimental results on the test set show that while system-generated reviews are comprehensive, touching upon more aspects of the paper than human-written reviews, the generated texts are less constructive and less factual than human-written reviews for all aspects except the explanation of the core ideas of the papers, which are largely factually correct. Given these results, we pose eight challenges in the pursuit of a good review generation system together with potential solutions, which, hopefully, will inspire more future research in this direction. We make relevant resource publicly available for use by future research: https://github. com/neulab/ReviewAdvisor. In addition, while our conclusion is that the technology is not yet ready for use in high-stakes review settings we provide a system demo, ReviewAdvisor (http://review.nlpedia.ai/), showing the current capabilities and failings of state-of-the-art NLP models at this task (see demo screenshot in A.2). A review of this paper written by the system proposed in this paper can be found in A.1. Weizhe Yuan, Pengfei Liu 0003, Graham Neubig |
J. Artif. Intell. Res. | 3 |
| 2022 | Evaluating Explanations: How Much Do Explanations from the Teacher Aid Students?abstractAbstract While many methods purport to explain predictions by highlighting salient features, what aims these explanations serve and how they ought to be evaluated often go unstated. In this work, we introduce a framework to quantify the value of explanations via the accuracy gains that they confer on a student model trained to simulate a teacher model. Crucially, the explanations are available to the student during training, but are not available at test time. Compared with prior proposals, our approach is less easily gamed, enabling principled, automatic, model-agnostic evaluation of attributions. Using our framework, we compare numerous attribution methods for text classification and question answering, and observe quantitative differences that are consistent (to a moderate to high degree) across different student model architectures and learning strategies.1 Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio B. Soares, Michael Collins 0001, Zachary C. Lipton, Graham Neubig, William W. Cohen |
Trans. Assoc. Comput. Linguistics | 7 |
| 2022 | In-IDE Code Generation from Natural Language: Promise and ChallengesabstractA great part of software development involves conceptualizing or communicating the underlying procedures and logic that needs to be expressed in programs. One major difficulty of programming is turning concept into code , especially when dealing with the APIs of unfamiliar libraries. Recently, there has been a proliferation of machine learning methods for code generation and retrieval from natural language queries , but these have primarily been evaluated purely based on retrieval accuracy or overlap of generated code with developer-written code, and the actual effect of these methods on the developer workflow is surprisingly unattested. In this article, we perform the first comprehensive investigation of the promise and challenges of using such technology inside the PyCharm IDE, asking, “At the current state of technology does it improve developer productivity or accuracy, how does it affect the developer experience, and what are the remaining gaps and challenges?” To facilitate the study, we first develop a plugin for the PyCharm IDE that implements a hybrid of code generation and code retrieval functionality, and we orchestrate virtual environments to enable collection of many user events (e.g., web browsing, keystrokes, fine-grained code edits). We ask developers with various backgrounds to complete 7 varieties of 14 Python programming tasks ranging from basic file manipulation to machine learning or data visualization, with or without the help of the plugin. While qualitative surveys of developer experience are largely positive, quantitative results with regards to increased productivity, code quality, or program correctness are inconclusive. Further analysis identifies several pain points that could improve the effectiveness of future machine learning-based code generation/retrieval developer assistants and demonstrates when developers prefer code generation over code retrieval and vice versa. We release all data and software to pave the road for future empirical studies on this topic, as well as development of better code generation models. Frank F. Xu, Bogdan Vasilescu, Graham Neubig |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2021 | Measuring and Increasing Context Usage in Context-Aware Machine TranslationabstractPatrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Patrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins |
ACL/IJCNLP (1) | 3 |
| 2021 | CitationIE: Leveraging the Citation Graph for Scientific Information ExtractionabstractVijay Viswanathan, Graham Neubig, Pengfei Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vijay Viswanathan 0002, Graham Neubig, Pengfei Liu 0003 |
ACL/IJCNLP (1) | 2 |
| 2021 | Do Context-Aware Translation Models Pay the Right Attention?abstractKayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig |
ACL/IJCNLP (1) | 6 |
| 2021 | Dependency Induction Through the Lens of Visual PerceptionabstractMost previous work on grammar induction focuses on learning phrasal or dependency structure purely from text.However, because the signal provided by text alone is limited, recently introduced visually grounded syntax models make use of multimodal information leading to improved performance in constituency grammar induction.However, as compared to dependency grammars, constituency grammars do not provide a straightforward way to incorporate visual information without enforcing language-specific heuristics.In this paper, we propose an unsupervised grammar induction model that leverages word concreteness and a structural vision-based heuristic to jointly learn constituency-structure and dependency-structure grammars.Our experiments find that concreteness is a strong indicator for learning dependency grammars, improving the direct attachment score (DAS) by over 50% as compared to state-of-the-art models trained on pure text.Next, we propose an extension of our model that leverages both word concreteness and visual semantic role labels in constituency and dependency parsing.Our experiments show that the proposed extension outperforms the current state-of-the-art visually grounded models in constituency parsing even with a smaller grammar size. 1 Ruisi Su, Shruti Rijhwani, Hao Zhu 0011, Junxian He, Yonatan Bisk, Graham Neubig |
CoNLL | 7 |
| 2021 | Word Alignment by Fine-tuning Embeddings on Parallel CorporaabstractWord alignment over parallel corpora has a wide variety of applications, including learning translation lexicons, cross-lingual transfer of language processing tools, and automatic evaluation or analysis of translation outputs.The great majority of past work on word alignment has worked by performing unsupervised learning on parallel text.Recently, however, other work has demonstrated that pre-trained contextualized word embeddings derived from multilingually trained language models (LMs) prove an attractive alternative, achieving competitive results on the word alignment task even in the absence of explicit training on parallel data.In this paper, we examine methods to marry the two approaches: leveraging pre-trained LMs but finetuning them on parallel text with objectives designed to improve alignment quality, and proposing methods to effectively extract alignments from these fine-tuned models.We perform experiments on five language pairs and demonstrate that our model can consistently outperform previous state-of-the-art models of all varieties.In addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs.Our aligner, AWE-SOME (Aligning Word Embedding Spaces Of Multilingual Encoders), with pre-trained models is available at https://github. com/neulab/awesome-align. Zi-Yi Dou, Graham Neubig |
EACL | 2 |
| 2021 | Towards More Fine-grained and Reliable NLP Performance PredictionabstractPerformance prediction, the task of estimating a system's performance without performing experiments, allows us to reduce the experimental burden caused by the combinatorial explosion of different datasets, languages, tasks, and models.In this paper, we make two contributions to improving performance prediction for NLP tasks.First, we examine performance predictors not only for holistic measures of accuracy like F1 or BLEU, but also fine-grained performance measures such as accuracy over individual classes of examples.Second, we propose methods to understand the reliability of a performance prediction model from two angles: confidence intervals and calibration.We perform an analysis of four types of NLP tasks, and both demonstrate the feasibility of fine-grained performance prediction and the necessity to perform reliability analysis for performance prediction methods in the future.We make our code publicly available Zihuiwen Ye, Pengfei Liu 0003, Jinlan Fu, Graham Neubig |
EACL | 4 |
| 2021 | When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical SelectionabstractLearning fine-grained distinctions between vocabulary items is a key challenge in learning a new language.For example, the noun "wall" has different lexical manifestations in Spanish -"pared" refers to an indoor wall while "muro" refers to an outside wall.However, this variety of lexical distinction may not be obvious to non-native learners unless the distinction is explained in such a way.In this work, we present a method for automatically identifying fine-grained lexical distinctions, and extracting concise descriptions explaining these distinctions in a human-and machine-readable format.We confirm the quality of these extracted descriptions in a language learning setup for two languages, Spanish and Greek, where we use them to teach non-native speakers when to translate a given ambiguous word into its different possible translations.Code and data are publicly released here.1 Aditi Chaudhary, Kayo Yin, Antonios Anastasopoulos, Graham Neubig |
EMNLP (1) | 4 |
| 2021 | Efficient Nearest Neighbor Language ModelsabstractNon-parametric neural language models (NLMs) learn predictive distributions of text utilizing an external datastore, which allows them to learn through explicitly memorizing the training datapoints.While effective, these models often require retrieval from a large datastore at test time, significantly increasing the inference overhead and thus limiting the deployment of non-parametric NLMs in practical applications.In this paper, we take the recently proposed k-nearest neighbors language model (Khandelwal et al., 2019) as an example, exploring methods to improve its efficiency along various dimensions.Experiments on the standard WikiText-103 benchmark and domain-adaptation datasets show that our methods are able to achieve up to a 6x speed-up in inference speed while retaining comparable performance.The empirical analysis we present may provide guidelines for future research seeking to develop or deploy more efficient non-parametric NLMs. 1 Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP (1) | 2 |
| 2021 | Evaluating the Morphosyntactic Well-formedness of Generated TextsabstractAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov |
EMNLP (1) | 6 |
| 2021 | AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African LanguagesabstractReproducible benchmarks are crucial in driving progress of machine translation research.However, existing machine translation benchmarks have been mostly limited to highresource or well-represented languages.Despite an increasing interest in low-resource machine translation, there are no standardized reproducible benchmarks for many African languages, many of which are used by millions of speakers but have less digitized textual data.To tackle these challenges, we propose AFROMT, a standardized, clean, and reproducible machine translation benchmark for eight widely spoken African languages.We also develop a suite of analysis tools for system diagnosis taking into account unique properties of these languages.Furthermore, we explore the newly considered case of low-resource focused pretraining and develop two novel data augmentation-based strategies, leveraging word-level alignment information and pseudo-monolingual data for pretraining multilingual sequence-to-sequence models.We demonstrate significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines.We also show gains of up to 12 BLEU points over cross-lingual transfer baselines in data-constrained scenarios.All code and pretrained models will be released as further steps towards larger reproducible benchmarks for African languages.1 Machel Reid, Junjie Hu 0001, Graham Neubig, Yutaka Matsuo |
EMNLP (1) | 3 |
| 2021 | XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationabstractSebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson |
EMNLP (1) | 10 |
| 2021 | Distributionally Robust Multilingual Machine TranslationabstractMultilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memoryefficiency of deployed models.However, the heavy data imbalance between languages hinders the model from performing uniformly across language pairs.In this paper, we propose a new learning objective for MNMT based on distributionally robust optimization, which minimizes the worst-case expected loss over the set of language pairs.We further show how to practically optimize this objective for large translation corpora using an iterated best response scheme, which is both effective and incurs negligible additional computational cost compared to standard empirical risk minimization.We perform extensive experiments on three sets of languages from two datasets and show that our method consistently outperforms strong baseline methods in terms of average and per-language performance under both many-to-one and one-to-many translation settings.1 Chunting Zhou, Daniel Levy 0002, Xian Li 0003, Marjan Ghazvininejad, Graham Neubig |
EMNLP (1) | 5 |
| 2021 | Modeling the Second Player in Distributionally Robust Optimization
Paul Michel, Tatsunori B. Hashimoto, Graham Neubig |
ICLR | 3 |
| 2021 | Meta Back-Translation
Xinyi Wang 0001, Yiming Yang 0002, Graham Neubig |
ICLR | 4 |
| 2021 | Learning Structural Edits via Incremental Tree Transformations
Ziyu Yao 0002, Frank F. Xu, Huan Sun 0001, Graham Neubig |
ICLR | 5 |
| 2021 | Examining and Combating Spurious Features under Distribution ShiftabstractA central goal of machine learning is to learn robust representations that capture the fundamental relationship between inputs and output labels. However, minimizing training errors over finite or biased datasets results in models latching on to spurious correlations between the training input/output pairs that are not fundamental to the problem at hand. In this paper, we define and analyze robust and spurious representations using the information-theoretic concept of minimal sufficient statistics. We prove that even when there is only bias of the input distribution (i.e. covariate shift), models can still pick up spurious features from their training data. Group distributionally robust optimization (DRO) provides an effective tool to alleviate covariate shift by minimizing the worst-case training losses over a set of pre-defined groups. Inspired by our analysis, we demonstrate that group DRO can fail when groups do not directly account for various spurious correlations that occur in the data. To address this, we further propose to minimize the worst-case losses over a more flexible set of distributions that are defined on the joint distribution of groups and instances, instead of treating each group as a whole at optimization time. Through extensive experiments on one image and two language tasks, we show that our model is significantly more robust than comparable baselines under various partitions. Chunting Zhou, Xuezhe Ma, Paul Michel, Graham Neubig |
ICML | 4 |
| 2021 | Few-shot Language Coordination by Modeling Theory of MindabstractNo man is an island. Humans develop the ability to communicate with a large community by coordinating with different interlocutors within short conversations. This ability is largely understudied by the research on building neural language communicative agents. We study the task of few-shot language coordination: agents quickly adapting to their conversational partners’ language abilities. Different from current communicative agents trained with self-play, we in- investigate this more general paradigm by requiring the lead agent to coordinate with a population of agents each of whom has different linguistic abilities. This leads to a general agent able to quickly adapt to communicating with unseen agents in the population. Unlike prior work, success here requires the ability to model the partner’s beliefs, a vital component of human communication. Drawing inspiration from the study of theory-of-mind (ToM; Premack & Woodruff (1978)), we study the effect of the speaker explicitly modeling the listener’s mental state. Learning by communicating with a population, the speakers, as shown in our experiments, acquire the ability to learn to predict the reactions of their partner upon various messages on-the-fly. The speaker’s predictions for the future actions help it generate the best instructions in order to maximize communicative goal with message costs. To examine our hypothesis that the instructions generated with ToM modeling yield better communication per- performance, we employ our agents in both a referential game and a language navigation task. Positive results from our experiments also hint at the importance of explicitly modeling language acquisition as a socio-pragmatic progress. Hao Zhu 0011, Graham Neubig, Yonatan Bisk |
ICML | 2 |
| 2021 | Phoneme Recognition Through Fine Tuning of Phonetic Representations: A Case Study on Luhya Language VarietiesabstractModels pre-trained on multiple languages have shown significant promise for improving speech recognition, particularly for low-resource languages. In this work, we focus on phoneme recognition using Allosaurus, a method for multilingual recognition based on phonetic annotation, which incorporates phonological knowledge through a language-dependent allophone layer that associates a universal narrow phone-set with the phonemes that appear in each language. To evaluate in a challenging real-world scenario, we curate phone recognition datasets for Bukusu and Saamia, two varieties of the Luhya language cluster of western Kenya and eastern Uganda. To our knowledge, these datasets are the first of their kind. We carry out similar experiments on the dataset of an endangered Tangkhulic language, East Tusom, a Tibeto-Burman language variety spoken mostly in India. We explore both zero-shot and few-shot recognition by fine-tuning using datasets of varying sizes (10 to 1000 utterances). We find that fine-tuning of Allosaurus, even with just 100 utterances, leads to significant improvements in phone error rates. Kathleen Siminyu, Antonios Anastasopoulos, David R. Mortensen, Michael R. Marlo, Graham Neubig |
Interspeech | 6 |
| 2021 | GSum: A General Framework for Guided Neural Abstractive SummarizationabstractZi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zi-Yi Dou, Pengfei Liu 0003, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig |
NAACL-HLT | 5 |
| 2021 | Explicit Alignment Objectives for Multilingual Bidirectional EncodersabstractJunjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Junjie Hu 0001, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig |
NAACL-HLT | 5 |
| 2021 | Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language ModelsabstractPo-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alexander Hauptmann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Po-Yao Huang 0001, Mandela Patrick, Junjie Hu 0001, Graham Neubig, Florian Metze, Alex Hauptmann 0001 |
NAACL-HLT | 4 |
| 2021 | On Learning Text Style Transfer with Direct RewardsabstractIn most cases, the lack of parallel corpora makes it impossible to directly train supervised models for the text style transfer task.In this paper, we explore training algorithms that instead optimize reward functions that explicitly consider different aspects of the styletransferred outputs.In particular, we leverage semantic similarity metrics originally used for fine-tuning neural machine translation models to explicitly assess the preservation of content between system outputs and input texts.We also investigate the potential weaknesses of the existing automatic metrics and propose efficient strategies of using these metrics for training.The experimental results show that our model provides significant gains in both automatic and human evaluation over strong baselines, indicating the effectiveness of our proposed methods and training strategies. 1 Yixin Liu 0003, Graham Neubig, John Wieting |
NAACL-HLT | 2 |
| 2021 | Multi-view Subword RegularizationabstractMultilingual pretrained representations generally rely on subword segmentation algorithms to create a shared multilingual vocabulary.However, standard heuristic algorithms often lead to sub-optimal segmentation, especially for languages with limited amounts of data.In this paper, we take two major steps towards alleviating this problem.First, we demonstrate empirically that applying existing subword regularization methods (Kudo, 2018;Provilkov et al., 2020) during fine-tuning of pre-trained multilingual representations improves the effectiveness of cross-lingual transfer.Second, to take full advantage of different possible input segmentations, we propose Multi-view Subword Regularization (MVR), a method that enforces the consistency between predictions of using inputs tokenized by the standard and probabilistic segmentations.Results on the XTREME multilingual benchmark (Hu et al., 2020) show that MVR brings consistent improvements of up to 2.5 points over using standard segmentation algorithms. 1 Xinyi Wang 0001, Sebastian Ruder, Graham Neubig |
NAACL-HLT | 3 |
| 2021 | MetaXL: Meta Representation Transformation for Low-resource Cross-lingual LearningabstractMengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Mengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, Ahmed Awadallah 0001 |
NAACL-HLT | 5 |
| 2021 | Compositional Generalization for Neural Semantic Parsing via Span-level Supervised AttentionabstractPengcheng Yin, Hao Fang, Graham Neubig, Adam Pauls, Emmanouil Antonios Platanios, Yu Su, Sam Thomson, Jacob Andreas. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hao Fang 0002, Graham Neubig, Adam Pauls, Emmanouil A. Platanios, Yu Su 0001, Sam Thomson, Jacob Andreas |
NAACL-HLT | 3 |
| 2021 | BARTScore: Evaluating Generated Text as Text GenerationabstractA wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate, or effective. In this work, we conceptualize the evaluation of generated text as a text generation problem, modeled using pre-trained sequence-to-sequence models. The general idea is that models trained to convert the generated text to/from a reference output or the source text will achieve higher scores when the generated text is better. We operationalize this idea using BART, an encoder-decoder based pre-trained model, and propose a metric BARTScore with a number of variants that can be flexibly applied in an unsupervised fashion to evaluation of text from different perspectives (e.g. informativeness, fluency, or factuality). BARTScore is conceptually simple and empirically effective. It can outperform existing top-scoring metrics in 16 of 22 test settings, covering evaluation of 16 datasets (e.g., machine translation, text summarization) and 7 different perspectives (e.g., informativeness, factuality). Code to calculate BARTScore is available at https://github.com/neulab/BARTScore, and we have released an interactive leaderboard for meta-evaluation at http://explainaboard.nlpedia.ai/leaderboard/task-meval/ on the ExplainaBoard platform, which allows us to interactively understand the strengths, weaknesses, and complementarity of each metric. Weizhe Yuan, Graham Neubig, Pengfei Liu 0003 |
NeurIPS | 2 |
| 2021 | MasakhaNER: Named Entity Recognition for African LanguagesabstractAbstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1 David Ifeoluwa Adelani, Jade Z. Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew 0002, Israel Abebe Azime, Shamsuddeen Hassan Muhammad, Chris C. Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba O. Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin P. Adewumi, Paul Rayson, Mofe Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Ijeoma Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane Mboup, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing K. Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima Diop, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, Salomey Osei |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Reducing Confusion in Active Learning for Part-Of-Speech Tagging
Aditi Chaudhary, Zaid Sheikh, Antonios Anastasopoulos, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | WikiAsp: A Dataset for Multi-domain Aspect-based SummarizationabstractAbstract Aspect-based summarization is the task of generating focused summaries based on specific points of interest. Such summaries aid efficient analysis of text, such as quickly understanding reviews or opinions from different angles. However, due to large differences in the type of aspects for different domains (e.g., sentiment, product features), the development of previous models has tended to be domain-specific. In this paper, we propose WikiAsp,1 a large-scale dataset for multi-domain aspect- based summarization that attempts to spur research in the direction of open-domain aspect-based summarization. Specifically, we build the dataset using Wikipedia articles from 20 different domains, using the section titles and boundaries of each article as a proxy for aspect annotation. We propose several straightforward baseline models for this task and conduct experiments on the dataset. Results highlight key challenges that existing summarization models face in this setting, such as proper pronoun handling of quoted sources and consistent explanation of time-sensitive events. Hiroaki Hayashi, Prashant Budania, Chris Ackerson, Raj Neervannan, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 6 |
| 2021 | How Can We Know When Language Models Know? On the Calibration of Language Models for Question AnsweringabstractAbstract Recent works have shown that language models (LM) capture different types of knowledge regarding facts or common sense. However, because no model is perfect, they still fail to provide appropriate answers in many cases. In this paper, we ask the question, “How can we know when language models know, with confidence, the answer to a particular query?” We examine this question from the point of view of calibration, the property of a probabilistic model’s predicted probabilities actually being well correlated with the probabilities of correctness. We examine three strong generative models—T5, BART, and GPT-2—and study whether their probabilities on QA tasks are well calibrated, finding the answer is a relatively emphatic no. We then examine methods to calibrate such models to make their confidence scores correlate better with the likelihood of correctness through fine-tuning, post-hoc probability modification, or adjustment of the predicted outputs or inputs. Experiments on a diverse range of datasets demonstrate the effectiveness of our methods. We also perform analysis to study the strengths and limitations of these methods, shedding light on further improvements that may be made in methods for calibrating LMs. We have released the code at https://github.com/jzbjyb/lm-calibration. Zhengbao Jiang, Jun Araki, Haibo Ding, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Lexically Aware Semi-Supervised Learning for OCR Post-CorrectionabstractAbstract Much of the existing linguistic data in many languages of the world is locked away in non- digitized books and documents. Optical character recognition (OCR) can be used to produce digitized text, and previous work has demonstrated the utility of neural post-correction methods that improve the results of general- purpose OCR systems on recognition of less- well-resourced languages. However, these methods rely on manually curated post- correction data, which are relatively scarce compared to the non-annotated raw images that need to be digitized. In this paper, we present a semi-supervised learning method that makes it possible to utilize these raw images to improve performance, specifically through the use of self-training, a technique where a model is iteratively trained on its own outputs. In addition, to enforce consistency in the recognized vocabulary, we introduce a lexically aware decoding method that augments the neural post-correction model with a count-based language model constructed from the recognized texts, implemented using weighted finite-state automata (WFSA) for efficient and effective decoding. Results on four endangered languages demonstrate the utility of the proposed method, with relative error reductions of 15%–29%, where we find the combination of self-training and lexically aware decoding essential for achieving consistent improvements.1 Shruti Rijhwani, Daisy Rosenblum, Antonios Anastasopoulos, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | Latent Relation Language ModelsabstractIn this paper, we propose Latent Relation Language Models (LRLMs), a class of language models that parameterizes the joint distribution over the words in a document and the entities that occur therein via knowledge graph relations. This model has a number of attractive properties: it not only improves language modeling performance, but is also able to annotate the posterior probability of entity spans for a given text through relations. Experiments demonstrate empirical improvements over both word-based language models and a previous approach that incorporates knowledge graph information. Qualitative analysis further demonstrates the proposed model's ability to learn to predict appropriate relations in context. 1 Hiroaki Hayashi, Zecong Hu, Chenyan Xiong, Graham Neubig |
AAAI | 4 |
| 2020 | What Makes A Good Story? Designing Composite Rewards for Visual StorytellingabstractPrevious storytelling approaches mostly focused on optimizing traditional metrics such as BLEU, ROUGE and CIDEr. In this paper, we re-examine this problem from a different angle, by looking deep into what defines a natural and topically-coherent story. To this end, we propose three assessment criteria: relevance, coherence and expressiveness, which we observe through empirical analysis could constitute a “high-quality” story to the human eye. We further propose a reinforcement learning framework, ReCo-RL, with reward functions designed to capture the essence of these quality criteria. Experiments on the Visual Storytelling Dataset (VIST) with both automatic and human evaluation demonstrate that our ReCo-RL model achieves better performance than state-of-the-art baselines on both traditional metrics and the proposed new criteria. Junjie Hu 0001, Yu Cheng 0001, Zhe Gan, Jingjing Liu 0001, Jianfeng Gao 0001, Graham Neubig |
AAAI | 6 |
| 2020 | Merging Weak and Active Supervision for Semantic ParsingabstractA semantic parser maps natural language commands (NLs) from the users to executable meaning representations (MRs), which are later executed in certain environment to obtain user-desired results. The fully-supervised training of such parser requires NL/MR pairs, annotated by domain experts, which makes them expensive to collect. However, weakly-supervised semantic parsers are learnt only from pairs of NL and expected execution results, leaving the MRs latent. While weak supervision is cheaper to acquire, learning from this input poses difficulties. It demands that parsers search a large space with a very weak learning signal and it is hard to avoid spurious MRs that achieve the correct answer in the wrong way. These factors lead to a performance gap between parsers trained in weakly- and fully-supervised setting. To bridge this gap, we examine the intersection between weak supervision and active learning, which allows the learner to actively select examples and query for manual annotations as extra supervision to improve the model trained under weak supervision. We study different active learning heuristics for selecting examples to query, and various forms of extra supervision for such queries. We evaluate the effectiveness of our method on two different datasets. Experiments on the WikiSQL show that by annotating only 1.8% of examples, we improve over a state-of-the-art weakly-supervised baseline by 6.4%, achieving an accuracy of 79.0%, which is only 1.3% away from the model trained with full supervision. Experiments on WikiTableQuestions with human annotators show that our method can improve the performance with only 100 active queries, especially for weakly-supervised parsers learnt from a cold start. 1 Ansong Ni, Graham Neubig |
AAAI | 3 |
| 2020 | Should All Cross-Lingual Embeddings Speak English?abstractMost of recent work in cross-lingual word embeddings is severely Anglocentric.The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is selected by default as the hub when learning in a multilingual setting.With this work, however, we challenge these practices.First, we show that the choice of hub language can significantly impact downstream lexicon induction and zero-shot POS tagging performance.Second, we both expand a standard Englishcentered evaluation dictionary collection to include all language pairs using triangulation, and create new dictionaries for under-represented languages.1 Evaluating established methods over all these language pairs sheds light into their suitability for aligning embeddings from distant languages and presents new challenges for the field.Finally, in our analysis we identify general guidelines for strong cross-lingual embedding baselines, that extend to language pairs that do not include English. Antonios Anastasopoulos, Graham Neubig |
ACL | 2 |
| 2020 | Generalizing Natural Language Analysis through Span-relation RepresentationsabstractNatural language processing covers a wide variety of tasks predicting syntax, semantics, and information content, and usually each type of output is generated with specially designed architectures.In this paper, we provide the simple insight that a great variety of tasks can be represented in a single unified format consisting of labeling spans and relations between spans, thus a single task-independent model can be used across different tasks.We perform extensive experiments to test this insight on 10 disparate tasks spanning dependency parsing (syntax), semantic role labeling (semantics), relation extraction (information content), aspect based sentiment analysis (sentiment), and many others, achieving performance comparable to state-of-the-art specialized models.We further demonstrate benefits of multi-task learning, and also show that the proposed method makes it easy to analyze differences and similarities in how the model handles different tasks.Finally, we convert these datasets into a unified format to build a benchmark, which provides a holistic testbed for evaluating future models for generalized natural language analysis. Zhengbao Jiang, Wei Xu 0004, Jun Araki, Graham Neubig |
ACL | 4 |
| 2020 | Weight Poisoning Attacks on Pretrained ModelsabstractRecently, NLP has seen a surge in the usage of large pre-trained models.Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of their choice.This raises the question of whether downloading untrusted pre-trained weights can pose a security threat.In this paper, we show that it is possible to construct "weight poisoning" attacks where pre-trained weights are injected with vulnerabilities that expose "backdoors" after fine-tuning, enabling the attacker to manipulate the model prediction simply by injecting an arbitrary keyword.We show that by applying a regularization method, which we call RIPPLe, and an initialization procedure, which we call Embedding Surgery, such attacks are possible even with limited knowledge of the dataset and finetuning procedure.Our experiments on sentiment classification, toxicity detection, and spam detection show that this attack is widely applicable and poses a serious threat.Finally, we outline practical defenses against such attacks.Code to reproduce our experiments is available at https://github.com/ neulab/RIPPLe. Keita Kurita, Paul Michel, Graham Neubig |
ACL | 3 |
| 2020 | Politeness Transfer: A Tag and Generate ApproachabstractAman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, Shrimai Prabhumoye. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos, Graham Neubig, Yiming Yang 0002, Ruslan Salakhutdinov, Alan W. Black, Shrimai Prabhumoye |
ACL | 5 |
| 2020 | Learning to Deceive with Attention-Based ExplanationsabstractAttention mechanisms are ubiquitous components in neural architectures applied to natural language processing.In addition to yielding gains in predictive accuracy, attention weights are often claimed to confer interpretability, purportedly useful both for providing insights to practitioners and for explaining why a model makes its decisions to stakeholders.We call the latter use of attention mechanisms into question by demonstrating a simple method for training models to produce deceptive attention masks.Our method diminishes the total weight assigned to designated impermissible tokens, even when the models can be shown to nevertheless rely on these features to drive predictions.Across multiple models and tasks, our approach manipulates attention weights while paying surprisingly little cost in accuracy.Through a human study, we show that our manipulated attention-based explanations deceive people into thinking that predictions from a model biased against gender minorities do not rely on the gender.Consequently, our results cast doubt on attention's reliability as a tool for auditing algorithms in the context of fairness and accountability.1 Danish Pruthi, Bhuwan Dhingra, Graham Neubig, Zachary C. Lipton |
ACL | 4 |
| 2020 | Soft Gazetteers for Low-Resource Named Entity RecognitionabstractTraditional named entity recognition models use gazetteers (lists of entities) as features to improve performance.Although modern neural network models do not require such handcrafted features for strong performance, recent work (Wu et al., 2018) has demonstrated their utility for named entity recognition on English data.However, designing such features for low-resource languages is challenging, because exhaustive entity gazetteers do not exist in these languages.To address this problem, we propose a method of "soft gazetteers" that incorporates ubiquitously available information from English knowledge bases, such as Wikipedia, into neural named entity recognition models through cross-lingual entity linking.Our experiments on four low-resource languages show an average improvement of 4 points in F1 score. 1 Shruti Rijhwani, Shuyan Zhou, Graham Neubig, Jaime G. Carbonell |
ACL | 3 |
| 2020 | Balancing Training for Multilingual Neural Machine TranslationabstractWhen training multilingual machine translation (MT) models that can translate to/from multiple languages, we are faced with imbalanced training sets: some languages have much more training data than others.Standard practice is to up-sample less resourced languages to increase representation, and the degree of up-sampling has a large effect on the overall performance.In this paper, we propose a method that instead automatically learns how to weight training data through a data scorer that is optimized to maximize performance on all test languages.Experiments on two sets of languages under both one-to-many and manyto-one MT settings show our method not only consistently outperforms heuristic baselines in terms of average performance, but also offers flexible control over the performance of which languages are optimized.1 Xinyi Wang 0001, Yulia Tsvetkov, Graham Neubig |
ACL | 3 |
| 2020 | Predicting Performance for Natural Language Processing TasksabstractGiven the complexity of combinations of tasks, languages, and domains in natural language processing (NLP) research, it is computationally prohibitive to exhaustively test newly proposed models on each possible experimental setting.In this work, we attempt to explore the possibility of gaining plausible judgments of how well an NLP model can perform under an experimental setting, without actually training or testing the model.To do so, we build regression models to predict the evaluation score of an NLP experiment given the experimental settings as input.Experimenting on 9 different NLP tasks, we find that our predictors can produce meaningful predictions over unseen languages and different modeling architectures, outperforming reasonable baselines as well as human experts.Going further, we outline how our predictor can be used to find a small subset of representative experiments that should be run in order to obtain plausible predictions for all other experimental settings.1 Mengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang 0002, Graham Neubig |
ACL | 5 |
| 2020 | Incorporating External Knowledge through Pre-training for Natural Language to Code GenerationabstractOpen-domain code generation aims to generate code in a general-purpose programming language (such as Python) from natural language (NL) intents.Motivated by the intuition that developers usually retrieve resources on the web when writing code, we explore the effectiveness of incorporating two varieties of external knowledge into NL-to-code generation: automatically mined NL-code pairs from the online programming QA forum StackOverflow and programming language API documentation.Our evaluations show that combining the two sources with data augmentation and retrieval-based data re-sampling improves the current state-of-the-art by up to 2.2% absolute BLEU score on the code generation testbed CoNaLa.The code and resources are available at https://github.com/ neulab/external-knowledge-codegen. Frank F. Xu, Zhengbao Jiang, Bogdan Vasilescu, Graham Neubig |
ACL | 5 |
| 2020 | TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataabstractRecent years have witnessed the burgeoning of pretrained language models (LMs) for textbased natural language (NL) understanding tasks.Such models are typically trained on free-form NL text, hence may not be suitable for tasks like semantic parsing over structured data, which require reasoning over both free-form NL questions and structured tabular data (e.g., database tables).In this paper we present TABERT, a pretrained LM that jointly learns representations for NL sentences and (semi-)structured tables.TABERT is trained on a large corpus of 26 million tables and their English contexts.In experiments, neural semantic parsers using TABERT as feature representation layers achieve new best results on the challenging weakly-supervised semantic parsing benchmark WIKITABLEQUESTIONS, while performing competitively on the text-to-SQL dataset SPIDER. 1 Graham Neubig, Scott Yih, Sebastian Riedel 0001 |
ACL | 2 |
| 2020 | Automatic Interlinear Glossing for Under-Resourced Languages Leveraging TranslationsabstractInterlinear Glossed Text (IGT) is a widely used format for encoding linguistic information in language documentation projects and scholarly papers.Manual production of IGT takes time and requires linguistic expertise.We attempt to address this issue by creating automatic glossing models, using modern multi-source neural models that additionally leverage easy-to-collect translations.We further explore cross-lingual transfer and a simple output length control mechanism, further refining our models.Evaluated on three challenging low-resource scenarios, our approach significantly outperforms a recent, state-of-the-art baseline, particularly improving on overall accuracy as well as lemma and tag recall. Xingyuan Zhao, Satoru Ozaki, Antonios Anastasopoulos, Graham Neubig, Lori S. Levin |
COLING | 4 |
| 2020 | Project MAIA: Multilingual AI Agent AssistantabstractThis paper presents the Multilingual Artificial Intelligence Agent Assistant (MAIA), a project led by Unbabel with the collaboration of CMU, INESC-ID and IT Lisbon. MAIA will employ cutting-edge machine learning and natural language processing technologies to build multilingual AI agent assistants, eliminating language barriers. MAIA’s translation layer will empower human agents to provide customer support in real-time, in any language, with human quality. André F. T. Martins, João Graça, Paulo Dimas, Helena Moniz, Graham Neubig |
EAMT | 5 |
| 2020 | Re-evaluating Evaluation in Text SummarizationabstractAutomated evaluation metrics as a stand-in for manual evaluation are an essential part of the development of text-generation tasks such as text summarization.However, while the field has progressed, our standard metrics have not -for nearly 20 years ROUGE has been the standard evaluation in most summarization papers.In this paper, we make an attempt to re-evaluate the evaluation method for text summarization: assessing the reliability of automatic metrics using top-scoring system outputs, both abstractive and extractive, on recently popular datasets for both systemlevel and summary-level evaluation settings.We find that conclusions about evaluation metrics on older datasets do not necessarily hold on modern datasets and systems.We release a dataset of human judgments that are collected from 25 top-scoring neural summarization systems (14 abstractive and 11 extractive): Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 0003, Graham Neubig |
EMNLP (1) | 5 |
| 2020 | Automatic Extraction of Rules Governing Morphological AgreementabstractAditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig |
EMNLP (1) | 7 |
| 2020 | Dynamic Data Selection and Weighting for Iterative Back-TranslationabstractBack-translation has proven to be an effective method to utilize monolingual data in neural machine translation (NMT), and iteratively conducting back-translation can further improve the model performance.Selecting which monolingual data to back-translate is crucial, as we require that the resulting synthetic data are of high quality and reflect the target domain.To achieve these two goals, data selection and weighting strategies have been proposed, with a common practice being to select samples close to the target domain but also dissimilar to the average general-domain text.In this paper, we provide insights into this commonly used approach and generalize it to a dynamic curriculum learning strategy, which is applied to iterative back-translation models.In addition, we propose weighting strategies based on both the current quality of the sentence and its improvement over the previous iteration.We evaluate our models on domain adaptation, low-resource, and high-resource MT settings and on two language pairs.Experimental results demonstrate that our methods achieve improvements of up to 1.8 BLEU points over competitive baselines.1 Zi-Yi Dou, Antonios Anastasopoulos, Graham Neubig |
EMNLP (1) | 3 |
| 2020 | Interpretable Multi-dataset Evaluation for Named Entity RecognitionabstractWith the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits.Simply looking at differences between holistic metrics such as accuracy, BLEU, or F1 does not tell us why or how particular methods perform differently and how diverse datasets influence the model design choices.In this paper, we present a general methodology for interpretable evaluation for the named entity recognition (NER) task.The proposed evaluation method enables us to interpret the differences in models and datasets, as well as the interplay between them, identifying the strengths and weaknesses of current systems.By making our analysis tool available, we make it easy for future researchers to run similar analyses and drive progress in this area: https: //github.com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Graham Neubig |
EMNLP (1) | 3 |
| 2020 | X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsabstractLanguage models (LMs) have proven surprisingly successful at capturing factual knowledge by completing cloze-style fill-in-theblank questions such as "Punta Cana is located in _."However, while knowledge is both written and queried in many languages, studies on LMs' factual representation ability have almost invariably been performed on English.To assess factual knowledge retrieval in LMs in different languages, we create a multilingual benchmark of cloze-style probes for 23 typologically diverse languages.To properly handle language variations, we expand probing methods from single-to multi-word entities, and develop several decoding algorithms to generate multi-token predictions.Extensive experimental results provide insights about how well (or poorly) current state-of-theart LMs perform at this task in languages with more or fewer available resources.We further propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge, and verify its effectiveness on several benchmark languages.Benchmark data and code have be released at https: //x-factr.github.io. Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, Graham Neubig |
EMNLP (1) | 5 |
| 2020 | OCR Post Correction for Endangered Language TextsabstractThere is little to no data available to build natural language processing models for most endangered languages.However, textual data in these languages often exists in formats that are not machine-readable, such as paper books and scanned images.In this work, we address the task of extracting text from these resources.We create a benchmark dataset of transcriptions for scanned books in three critically endangered languages and present a systematic analysis of how general-purpose OCR tools are not robust to the data-scarce setting of endangered languages.We develop an OCR postcorrection method tailored to ease training in this data-scarce setting, reducing the recognition error rate by 34% on average across the three languages.1 Shruti Rijhwani, Antonios Anastasopoulos, Graham Neubig |
EMNLP (1) | 3 |
| 2020 | A Bilingual Generative Transformer for Semantic Sentence EmbeddingabstractSemantic sentence embedding models encode natural language sentences into vectors, such that closeness in embedding space indicates closeness in the semantics between the sentences.Bilingual data offers a useful signal for learning such embeddings: properties shared by both sentences in a translation pair are likely semantic, while divergent properties are likely stylistic or language-specific.We propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors.Our proposed approach differs from past work on semantic sentence encoding in two ways.First, by using a variational probabilistic framework, we introduce priors that encourage source separation, and can use our model's posterior to predict sentence embeddings for monolingual data at test time.Second, we use high-capacity transformers as both data generating distributions and inference networkscontrasting with most past work on sentence embeddings.In experiments, our approach substantially outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations.Further, we demonstrate that our approach yields the largest gains on more difficult subsets of these evaluations where simple word overlap is not a good indicator of similarity. 1 John Wieting, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP (1) | 2 |
| 2020 | Universal Phone Recognition with a Multilingual Allophone SystemabstractMultilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1 Siddharth Dalmia, Juncheng Li 0001, Matthew Lee 0012, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze |
ICASSP | 9 |
| 2020 | Differentiable Reasoning over a Virtual Knowledge Base
Bhuwan Dhingra, Manzil Zaheer, Vidhisha Balachandran, Graham Neubig, Ruslan Salakhutdinov, William W. Cohen |
ICLR | 4 |
| 2020 | A Probabilistic Formulation of Unsupervised Text Style Transfer
Junxian He, Xinyi Wang 0001, Graham Neubig, Taylor Berg-Kirkpatrick |
ICLR | 3 |
| 2020 | Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework
Jiateng Xie, Ruochen Xu, Yiming Yang 0002, Graham Neubig, Jaime G. Carbonell |
ICLR | 5 |
| 2020 | Understanding Knowledge Distillation in Non-autoregressive Machine Translation
Chunting Zhou, Jiatao Gu, Graham Neubig |
ICLR | 3 |
| 2020 | XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationabstractMuch recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We will release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks. Junjie Hu 0001, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, Melvin Johnson |
ICML | 4 |
| 2020 | Optimizing Data Usage via Differentiable RewardsabstractTo acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content or practice problems. Similarly, a machine learning model could potentially be trained better with a scorer that “adapts” to its current learning state and estimates the importance of each training data instance. Training such an adaptive scorer efficiently is a challenging problem; in order to precisely quantify the effect of a data instance at a given time during the training, it is typically necessary to first complete the entire training process. To efficiently optimize data usage, we propose a reinforcement learning approach called Differentiable Data Selection (DDS). In DDS, we formulate a scorer network as a learnable function of the training data, which can be efficiently updated along with the main model being trained. Specifically, DDS updates the scorer with an intuitive reward signal: it should up-weigh the data that has a similar gradient with a dev set upon which we would finally like to perform well. Without significant computing overhead, DDS delivers strong and consistent improvements over several strong baselines on two very different tasks of machine translation and image classification. Xinyi Wang 0001, Paul Michel, Antonios Anastasopoulos, Jaime G. Carbonell, Graham Neubig |
ICML | 6 |
| 2020 | AlloVera: A Multilingual Allophone DatabaseabstractWe introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which are predictable from phonological context. While phonemic representations are language specific, phonetic representations (stated in terms of (allo)phones) are much closer to a universal (language-independent) transcription. AlloVera allows the training of speech recognition models that output phonetic transcriptions in the International Phonetic Alphabet (IPA), regardless of the input language. We show that a “universal” allophone model, Allosaurus, built with AlloVera, outperforms “universal” phonemic models and language-specific models on a speech-transcription task. We explore the implications of this technology (and related technologies) for the documentation of endangered and minority languages. We further explore other applications for which AlloVera will be suitable as it grows, including phonological typology. David R. Mortensen, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W. Black, Florian Metze, Graham Neubig |
LREC | 9 |
| 2020 | Learning Sparse Prototypes for Text GenerationabstractPrototype-driven text generation uses non-parametric models that first choose from a library of sentence "prototypes" and then modify the prototype to generate the output text. While effective, these methods are inefficient at test time as a result of needing to store and index the entire training corpus. Further, existing methods often require heuristics to identify which prototypes to reference at training time. In this paper, we propose a novel generative model that automatically learns a sparse prototype support set that, nonetheless, achieves strong language modeling performance. This is achieved by (1) imposing a sparsity-inducing prior on the prototype selection distribution, and (2) utilizing amortized variational inference to learn a prototype retrieval function. In experiments, our model outperforms previous prototype-driven language models while achieving up to a 1000x memory reduction, as well as a 1000x speed-up at test time. More interestingly, we show that the learned prototypes are able to capture semantics and syntax at different granularity as we vary the sparsity of prototype selection, and that certain sentence attributes can be controlled by specifying the prototype for generation. Junxian He, Taylor Berg-Kirkpatrick, Graham Neubig |
NeurIPS | 3 |
| 2020 | A Set of Recommendations for Assessing Human-Machine Parity in Language TranslationabstractThe quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et al.'s 2018 investigation into Chinese to English news translation, showing that the finding of human–machine parity was owed to weaknesses in the evaluation design—which is currently considered best practice in the field. We show that the professional human translations contained significantly fewer errors, and that perceived quality in human evaluation depends on the choice of raters, the availability of linguistic context, and the creation of reference translations. Our results call for revisiting current best practices to assess strong machine translation systems in general and human–machine parity in particular, for which we offer a set of recommendations based on our empirical findings. Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, Antonio Toral |
J. Artif. Intell. Res. | 3 |
| 2020 | Optimizing segmentation granularity for neural machine translation
Elizabeth Salesky, Andrew Runge, Alex Coda, Jan Niehues, Graham Neubig |
Mach. Transl. | 5 |
| 2020 | Improving neural machine translation through phrase-based soft forced decoding
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 4 |
| 2020 | How Can We Know What Language Models KnowabstractRecent work has presented intriguing results examining the knowledge contained in language models (LMs) by having the LM fill in the blanks of prompts such as “ Obama is a __ by profession”. These prompts are usually manually created, and quite possibly sub-optimal; another prompt such as “ Obama worked as a __ ” may result in more accurately predicting the correct profession. Because of this, given an inappropriate prompt, we might fail to retrieve facts that the LM does know, and thus any given prompt only provides a lower bound estimate of the knowledge contained in an LM. In this paper, we attempt to more accurately estimate the knowledge contained in LMs by automatically discovering better prompts to use in this querying process. Specifically, we propose mining-based and paraphrasing-based methods to automatically generate high-quality and diverse prompts, as well as ensemble methods to combine answers from different prompts. Extensive experiments on the LAMA benchmark for extracting relational knowledge from LMs demonstrate that our methods can improve accuracy from 31.1% to 39.6%, providing a tighter lower bound on what LMs know. We have released the code and the resulting LM Prompt And Query Archive (LPAQA) at https://github.com/jzbjyb/LPAQA . Zhengbao Jiang, Frank F. Xu, Jun Araki, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | Improving Candidate Generation for Low-resource Cross-lingual Entity LinkingabstractCross-lingual entity linking (XEL) is the task of finding referents in a target-language knowledge base (KB) for mentions extracted from source-language texts. The first step of (X)EL is candidate generation, which retrieves a list of plausible candidate entities from the target-language KB for each mention. Approaches based on resources from Wikipedia have proven successful in the realm of relatively high-resource languages, but these do not extend well to low-resource languages with few, if any, Wikipedia pages. Recently, transfer learning methods have been shown to reduce the demand for resources in the low-resource languages by utilizing resources in closely related languages, but the performance still lags far behind their high-resource counterparts. In this paper, we first assess the problems faced by current entity candidate generation methods for low-resource XEL, then propose three improvements that (1) reduce the disconnect between entity mentions and KB entries, and (2) improve the robustness of the model to low-resource scenarios. The methods are simple, but effective: We experiment with our approach on seven XEL datasets and find that they yield an average gain of 16.9% in Top-30 gold candidate recall, compared with state-of-the-art baselines. Our improved model also yields an average gain of 7.9% in in-KB accuracy of end-to-end XEL. 1 Shuyan Zhou, Shruti Rijhwani, John Wieting, Jaime G. Carbonell, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 5 |
| 2020 | The Return of Lexical Dependencies: Neural Lexicalized PCFGs
Hao Zhu 0011, Yonatan Bisk, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Multi-Source Neural Machine Translation With Missing DataabstractMachine translation is rife with ambiguities in word ordering and word choice, and even with the advent of machine-learning methods that learn to resolve this ambiguity based on statistics from large corpora, mistakes are frequent. Multi-source translation is an approach that attempts to resolve these ambiguities by exploiting multiple inputs (e.g. sentences in three different languages) to increase translation accuracy. These methods are trained on multilingual corpora, which include the multiple source languages and the target language, and then at test time uses information from both source languages while generating the target. While there are many of these multilingual corpora, such as multilingual translations of TED talks or European parliament proceedings, in practice, many multilingual corpora are not complete due to the difficulty to provide translations in all of the relevant languages. Existing studies on multi-source translation did not explicitly handle such situations, and thus are only applicable to complete corpora that have all of the languages of interest, severely limiting their practical applicability. In this article, we examine approaches for multi-source neural machine translation (NMT) that can learn from and translate such incomplete corpora. Specifically, we propose methods to deal with incomplete corpora at both training time and test time. For training time, we examine two methods: (1) a simple method that simply replaces missing source translations with a special NULL symbol, and (2) a data augmentation approach that fills in incomplete parts with source translations created from multi-source NMT. For test-time, we examine methods that use multi-source translation even when only a single source is provided by first translating into an additional auxiliary language using standard NMT, then using multi-source translation on the original source and this generated auxiliary language sentence. Extensive experiments demonstrate that the proposed training-time and test-time methods both significantly improve translation performance. Yuta Nishimura, Katsuhito Sudoh, Graham Neubig, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 16 |
| 2019 | Zero-Shot Neural Transfer for Cross-Lingual Entity LinkingabstractCross-lingual entity linking maps an entity mention in a source language to its corresponding entry in a structured knowledge base that is in a different (target) language. While previous work relies heavily on bilingual lexical resources to bridge the gap between the source and the target languages, these resources are scarce or unavailable for many low-resource languages. To address this problem, we investigate zero-shot cross-lingual entity linking, in which we assume no bilingual lexical resources are available in the source low-resource language. Specifically, we propose pivot-basedentity linking, which leverages information from a highresource “pivot” language to train character-level neural entity linking models that are transferred to the source lowresource language in a zero-shot manner. With experiments on 9 low-resource languages and transfer through a total of54 languages, we show that our proposed pivot-based framework improves entity linking accuracy 17% (absolute) on average over the baseline systems, for the zero-shot scenario.1 Further, we also investigate the use of language-universal phonological representations which improves average accuracy (absolute) by 36% when transferring between languages that use different scripts. Shruti Rijhwani, Jiateng Xie, Graham Neubig, Jaime G. Carbonell |
AAAI | 3 |
| 2019 | Cross-Lingual Syntactic Transfer through Unsupervised Adaptation of Invertible ProjectionsabstractCross-lingual transfer is an effective way to build syntactic analysis tools in low-resource languages.However, transfer is difficult when transferring to typologically distant languages, especially when neither annotated target data nor parallel corpora are available.In this paper, we focus on methods for cross-lingual transfer to distant languages and propose to learn a generative model with a structured prior that utilizes labeled source data and unlabeled target data jointly.The parameters of source model and target model are softly shared through a regularized log likelihood objective.An invertible projection is employed to learn a new interlingual latent embedding space that compensates for imperfect crosslingual word embedding input.We evaluate our method on two syntactic tasks: part-ofspeech (POS) tagging and dependency parsing.On the Universal Dependency Treebanks, we use English as the only source corpus and transfer to a wide range of target languages.On the 10 languages in this dataset that are distant from English, our method yields an average of 5.2% absolute improvement on POS tagging and 8.3% absolute improvement on dependency parsing over a direct transfer method using state-of-the-art discriminative models. 1 3 Following Ahmad et al. (2019), we use the offline pre-trained alignment matrix present in https://github.com/Babylonpartners/ fastText_multilingual, which contains alignment matrices for 78 languages, which also allows comparison with their numbers in Section 4.3. Junxian He, Zhisong Zhang, Taylor Berg-Kirkpatrick, Graham Neubig |
ACL (1) | 4 |
| 2019 | Domain Adaptation of Neural Machine Translation by Lexicon InductionabstractIt has been previously noted that neural machine translation (NMT) is very sensitive to domain shift.In this paper, we argue that this is a dual effect of the highly lexicalized nature of NMT, resulting in failure for sentences with large numbers of unknown words, and lack of supervision for domain-specific words.To remedy this problem, we propose an unsupervised adaptation method which finetunes a pre-trained out-of-domain NMT model using a pseudo-in-domain corpus.Specifically, we perform lexicon induction to extract an in-domain lexicon, and construct a pseudo-parallel in-domain corpus by performing word-for-word back-translation of monolingual in-domain target sentences.In five domains over twenty pairwise adaptation settings and two model architectures, our method achieves consistent improvements without using any in-domain parallel sentences, improving up to 14 BLEU over unadapted models, and up to 2 BLEU over strong back-translation baselines. Junjie Hu 0001, Mengzhou Xia, Graham Neubig, Jaime G. Carbonell |
ACL (1) | 3 |
| 2019 | Improving Open Information Extraction via Iterative Rank-Aware LearningabstractOpen information extraction (IE) is the task of extracting open-domain assertions from natural language sentences.A key step in open IE is confidence modeling, ranking the extractions based on their estimated quality to adjust precision and recall of extracted assertions.We found that the extraction likelihood, a confidence measure used by current supervised open IE systems, is not well calibrated when comparing the quality of assertions extracted from different sentences.We propose an additional binary classification loss to calibrate the likelihood to make it more globally comparable, and an iterative learning process, where extractions generated by the open IE model are incrementally included as training samples to help the model learn from trial and error.Experiments on OIE2016 demonstrate the effectiveness of our method.1 Zhengbao Jiang, Graham Neubig |
ACL (1) | 3 |
| 2019 | Choosing Transfer Languages for Cross-Lingual LearningabstractYu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig |
ACL (1) | 13 |
| 2019 | Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding SpacesabstractRecent work on bilingual lexicon induction (BLI) has frequently depended either on aligned bilingual lexicons or on distribution matching, often with an assumption about the isometry of the two spaces.We propose a technique to quantitatively estimate this assumption of the isometry between two embedding spaces and empirically show that this assumption weakens as the languages in question become increasingly etymologically distant.We then propose Bilingual Lexicon Induction with Semi-Supervision (BLISS) -a semi-supervised approach that relaxes the isometric assumption while leveraging both limited aligned bilingual lexicons and a larger set of unaligned word embeddings, as well as a novel hubness filtering technique.Our proposed method obtains state of the art results on 15 of 18 language pairs on the MUSE dataset, and does particularly well when the embedding spaces don't appear to be isometric.In addition, we also show that adding supervision stabilizes the learning procedure, and is effective even with minimal supervision.⇤ Barun Patra, Joel Ruben Antony Moniz, Sarthak Garg, Matthew R. Gormley, Graham Neubig |
ACL (1) | 5 |
| 2019 | Self-Attentional Models for Lattice InputsabstractLattices are an efficient and effective method to encode ambiguity of upstream systems in natural language processing tasks, for example to compactly capture multiple speech recognition hypotheses, or to represent multiple linguistic analyses.Previous work has extended recurrent neural networks to model lattice inputs and achieved improvements in various tasks, but these models suffer from very slow computation speeds.This paper extends the recently proposed paradigm of self-attention to handle lattice inputs.Self-attention is a sequence modeling technique that relates inputs to one another by computing pairwise similarities and has gained popularity for both its strong results and its computational efficiency.To extend such models to handle lattices, we introduce probabilistic reachability masks that incorporate lattice structure into the model and support lattice scores if available.We also propose a method for adapting positional embeddings to lattice structures.We apply the proposed model to a speech translation task and find that it outperforms all examined baselines while being much faster to compute than previous neural lattice models during both training and inference. Matthias Sperber, Graham Neubig, Ngoc-Quan Pham, Alex Waibel |
ACL (1) | 2 |
| 2019 | Target Conditioned Sampling: Optimizing Data Selection for Multilingual Neural Machine TranslationabstractTo improve low-resource Neural Machine Translation (NMT) with multilingual corpora, training on the most related high-resource language only is often more effective than using all data available (Neubig and Hu, 2018).However, it is possible that an intelligent data selection strategy can further improve lowresource NMT with data from other auxiliary languages.In this paper, we seek to construct a sampling distribution over all multilingual data, so that it minimizes the training loss of the low-resource language.Based on this formulation, we propose an efficient algorithm, Target Conditioned Sampling (TCS), which first samples a target sentence, and then conditionally samples its source sentence.Experiments show that TCS brings significant gains of up to 2 BLEU on three of four languages we test, with minimal training overhead 1 . Xinyi Wang 0001, Graham Neubig |
ACL (1) | 2 |
| 2019 | Beyond BLEU: Training Neural Machine Translation with Semantic SimilarityabstractWhile most neural machine translation (NMT) systems are still trained using maximum likelihood estimation, recent work has demonstrated that optimizing systems to directly improve evaluation metrics such as BLEU can substantially improve final translation accuracy.However, training with BLEU has some limitations: it doesn't assign partial credit, it has a limited range of output values, and it can penalize semantically correct hypotheses if they differ lexically from the reference.In this paper, we introduce an alternative reward function for optimizing NMT systems that is based on recent work in semantic similarity.We evaluate on four disparate languages translated to English, and find that training with our proposed metric results in better translations as evaluated by BLEU, semantic similarity, and human evaluation, and also that the optimization procedure converges faster.Analysis suggests that this is because the proposed metric is more conducive to optimization, assigning partial credit and providing more diversity in scores than BLEU. 1 John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, Graham Neubig |
ACL (1) | 4 |
| 2019 | Simple and Effective Paraphrastic Similarity from Parallel TranslationsabstractWe present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the timeconsuming intermediate step of creating paraphrase corpora.Further, we show that the resulting model can be applied to cross-lingual tasks where it both outperforms and is orders of magnitude faster than more complex stateof-the-art baselines.1 John Wieting, Kevin Gimpel, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 3 |
| 2019 | Generalized Data Augmentation for Low-Resource TranslationabstractTranslation to or from low-resource languages (LRLs) poses challenges for machine translation in terms of both adequacy and fluency.Data augmentation utilizing large amounts of monolingual data is regarded as an effective way to alleviate these problems.In this paper, we propose a general framework for data augmentation in low-resource machine translation that not only uses target-side monolingual data, but also pivots through a related highresource language (HRL).Specifically, we experiment with a two-step pivoting method to convert high-resource data to the LRL, making use of available resources to better approximate the true data distribution of the LRL.First, we inject LRL words into HRL sentences through an induced bilingual dictionary.Second, we further edit these modified sentences using a modified unsupervised machine translation framework.Extensive experiments on four low-resource datasets show that under extreme low-resource settings, our data augmentation techniques improve translation quality by up to 1.5 to 8 BLEU points compared to supervised back-translation baselines.1 Mengzhou Xia, Xiang Kong, Antonios Anastasopoulos, Graham Neubig |
ACL (1) | 4 |
| 2019 | Reranking for Neural Semantic ParsingabstractSemantic parsing considers the task of transducing natural language (NL) utterances into machine executable meaning representations (MRs).While neural network-based semantic parsers have achieved impressive improvements over previous methods, results are still far from perfect, and cursory manual inspection can easily identify obvious problems such as lack of adequacy or coherence of the generated MRs.This paper presents a simple approach to quickly iterate and improve the performance of an existing neural semantic parser by reranking an n-best list of predicted MRs, using features that are designed to fix observed problems with baseline models.We implement our reranker in a competitive neural semantic parser and test on four semantic parsing (GEO, ATIS) and Python code generation (DJANGO, CONALA) tasks, improving the strong baseline parser by up to 5.7% absolute in BLEU (CONALA) and 2.9% in accuracy (DJANGO), outperforming the best published neural parser results on all four datasets.1 Graham Neubig |
ACL (1) | 2 |
| 2019 | Comparing Top-Down and Bottom-Up Neural Generative Dependency ModelsabstractRecurrent neural network grammars (RNNGs) generate sentences using phrase-structure syntax and perform very well in terms of both language modeling and parsing performance.However, since dependency annotations are much more readily available than phrase structure annotations, we propose two new generative models of projective dependency syntax, so as to explore whether generative dependency models are similarly effective.Both models use RNNs to represent the derivation history with making any explicit independence assumptions, but they differ in how they construct the trees: one builds the tree bottom up and the other top down, which profoundly changes the estimation problem faced by the learner.We evaluate the two models on three typologically different languages: English, Arabic, and Japanese.We find that both generative models improve parsing performance over a discriminative baseline, but, in contrast to RNNGs, they are significantly less effective than non-syntactic LSTM language models.Little difference between the tree construction orders is observed for either parsing or language modeling. Austin Matthews, Graham Neubig, Chris Dyer |
CoNLL | 2 |
| 2019 | Pushing the Limits of Low-Resource Morphological InflectionabstractAntonios Anastasopoulos, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Antonios Anastasopoulos, Graham Neubig |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity RecognizersabstractAditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime Carbonell. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime G. Carbonell |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Unsupervised Domain Adaptation for Neural Machine Translation with Domain-Aware Feature EmbeddingsabstractZi-Yi Dou, Junjie Hu, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zi-Yi Dou, Junjie Hu 0001, Antonios Anastasopoulos, Graham Neubig |
EMNLP/IJCNLP (1) | 4 |
| 2019 | A Surprisingly Effective Fix for Deep Latent Variable Modeling of TextabstractBohan Li, Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang 0002 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative FlowabstractXuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xuezhe Ma, Chunting Zhou, Xian Li 0003, Graham Neubig, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Handling Syntactic Divergence in Low-resource Machine TranslationabstractChunting Zhou, Xuezhe Ma, Junjie Hu, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunting Zhou, Xuezhe Ma, Junjie Hu 0001, Graham Neubig |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
Junxian He, Daniel Spokoyny, Graham Neubig, Taylor Berg-Kirkpatrick |
ICLR (Poster) | 3 |
| 2019 | Multilingual Neural Machine Translation With Soft Decoupled Encoding
Xinyi Wang 0001, Philip Arthur, Graham Neubig |
ICLR (Poster) | 4 |
| 2019 | Learning to Represent Edits
Graham Neubig, Miltiadis Allamanis, Marc Brockschmidt, Alexander L. Gaunt |
ICLR (Poster) | 2 |
| 2019 | Mitigating Noisy Inputs for Question AnsweringabstractNatural language processing systems are often downstream of unreliable inputs: machine translation, optical character recognition, or speech recognition. For instance, virtual assistants can only answer your questions after understanding your speech. We investigate and mitigate the effects of noise from Automatic Speech Recognition systems on two factoid Question Answering (QA) tasks. Integrating confidences into the model and forced decoding of unknown words are empirically shown to improve the accuracy of downstream neural QA systems. We create and train models on a synthetic corpus of over 500,000 noisy sentences and evaluate on two human corpora from Quizbowl and Jeopardy! competitions. Denis Peskov, Joe Barrow, Pedro Rodríguez 0001, Graham Neubig, Jordan L. Boyd-Graber |
INTERSPEECH | 4 |
| 2019 | DIRE: A Neural Approach to Decompiled Identifier NamingabstractThe decompiler is one of the most common tools for examining binaries without corresponding source code. It transforms binaries into high-level code, reversing the compilation process. Decompilers can reconstruct much of the information that is lost during the compilation process (e.g., structure and type information). Unfortunately, they do not reconstruct semantically meaningful variable names, which are known to increase code understandability. We propose the Decompiled Identifier Renaming Engine (DIRE), a novel probabilistic technique for variable name recovery that uses both lexical and structural information recovered by the decompiler. We also present a technique for generating corpora suitable for training and evaluating models of decompiled code renaming, which we use to create a corpus of 164,632 unique x86-64 binaries generated from C projects mined from GitHub. Our results show that on this corpus DIRE can predict variable names identical to the names in the original source code up to 74.3% of the time. Jeremy Lacomis, Edward J. Schwartz, Miltiadis Allamanis, Claire Le Goues, Graham Neubig, Bogdan Vasilescu |
ASE | 6 |
| 2019 | Are Sixteen Heads Really Better than One?abstractMulti-headed attention is a driving force behind recent state-of-the-art NLP models. By applying multiple attention mechanisms in parallel, it can express sophisticated functions beyond the simple weighted average. However we observe that, in practice, a large proportion of attention heads can be removed at test time without significantly impacting performance, and that some layers can even be reduced to a single head. Further analysis on machine translation models reveals that the self-attention layers can be significantly pruned, while the encoder-decoder layers are more dependent on multi-headedness. Paul Michel, Omer Levy, Graham Neubig |
NeurIPS | 3 |
| 2019 | Contextualized Representations for Low-resource Utterance TaggingabstractUtterance-level analysis of the speaker's intentions and emotions is a core task in conversational understanding.Depending on the end objective of the conversational understanding task, different categorical dialog-act or affect labels are expertly designed to cover specific aspects of the speakers' intentions or emotions respectively.Accurately annotating with these labels requires a high level of human expertise, and thus applying this process to a large conversation corpus or new domains is prohibitively expensive.The resulting paucity of data limits the use of sophisticated neural models.In this paper, we tackle these limitations by performing unsupervised training of utterance representations from a large corpus of spontaneous dialogue data.Models initialized with these representations achieve competitive performance on utterance-level dialogueact recognition and emotion classification, especially in low-resource settings encountered when analyzing conversations in new domains. Bhargavi Paranjape, Graham Neubig |
SIGdial | 2 |
| 2019 | Attention-Passing Models for Robust and Data-Efficient End-to-End Speech TranslationabstractSpeech translation has traditionally been approached through cascaded models consisting of a speech recognizer trained on a corpus of transcribed speech, and a machine translation system trained on parallel texts. Several recent works have shown the feasibility of collapsing the cascade into a single, direct model that can be trained in an end-to-end fashion on a corpus of translated speech. However, experiments are inconclusive on whether the cascade or the direct model is stronger, and have only been conducted under the unrealistic assumption that both are trained on equal amounts of data, ignoring other available speech recognition and machine translation corpora. In this paper, we demonstrate that direct speech translation models require more data to perform well than cascaded models, and although they allow including auxiliary data through multi-task training, they are poor at exploiting such data, putting them at a severe disadvantage. As a remedy, we propose the use of end- to-end trainable models with two attention mechanisms, the first establishing source speech to source text alignments, the second modeling source to target text alignment. We show that such models naturally decompose into multi-task–trainable recognition and translation tasks and propose an attention-passing technique that alleviates error propagation issues in a previous formulation of a model with two attention stages. Our proposed model outperforms all examined baselines and is able to exploit auxiliary training data much more effectively than direct attentional models. Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence ModelsabstractBeam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam search decoding procedure.In experiments, we show that optimizing this new training objective yields substantially better results on two sequence tasks (Named Entity Recognition and CCG Supertagging) when compared with both cross entropy trained greedy decoding and cross entropy trained beam decoding baselines. Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick |
AAAI | 2 |
| 2018 | Stack-Pointer Networks for Dependency ParsingabstractWe introduce a novel architecture for dependency parsing: stack-pointer networks (STACKPTR).Combining pointer networks (Vinyals et al., 2015) with an internal stack, the proposed model first reads and encodes the whole sentence, then builds the dependency tree top-down (from root-to-leaf) in a depth-first fashion.The stack tracks the status of the depthfirst search and the pointer networks select one child for the word at the top of the stack at each step.The STACKPTR parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the leftto-right restriction in classical transitionbased parsers.Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with O(n 2 ) time complexity.We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-theart performance on 21 of them. Xuezhe Ma, Zecong Hu, Jingzhou Liu, Nanyun Peng 0001, Graham Neubig, Eduard H. Hovy |
ACL (1) | 5 |
| 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum DataabstractHarsh Jhamtani, Varun Gangal, Eduard Hovy, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 4 |
| 2018 | Neural Factor Graph Models for Cross-lingual Morphological TaggingabstractMorphological analysis involves predicting the syntactic traits of a word (e.g.{POS: Noun, Case: Acc, Gender: Fem}).Previous work in morphological tagging improves performance for low-resource languages (LRLs) through cross-lingual training with a high-resource language (HRL) from the same family, but is limited by the strict-often false-assumption that tag sets exactly overlap between the HRL and LRL.In this paper we propose a method for cross-lingual morphological tagging that aims to improve information sharing between languages by relaxing this assumption.The proposed model uses factorial conditional random fields with neural network potentials, making it possible to (1) utilize the expressive power of neural network representations to smooth over superficial differences in the surface forms, (2) model pairwise and transitive relationships between tags, and (3) accurately generate tag sets that are unseen or rare in the training data.Experiments on four languages from the Universal Dependencies Treebank (Nivre et al., 2017) demonstrate superior tagging accuracies over existing cross-lingual approaches. 1 Chaitanya Malaviya, Matthew R. Gormley, Graham Neubig |
ACL (1) | 3 |
| 2018 | StructVAE: Tree-structured Latent Variable Models for Semi-supervised Semantic ParsingabstractSemantic parsing is the task of transducing natural language (NL) utterances into formal meaning representations (MRs), commonly represented as tree structures.Annotating NL utterances with their corresponding MRs is expensive and timeconsuming, and thus the limited availability of labeled data often becomes the bottleneck of data-driven, supervised models.We introduce STRUCTVAE, a variational auto-encoding model for semisupervised semantic parsing, which learns both from limited amounts of parallel data, and readily-available unlabeled NL utterances.STRUCTVAE models latent MRs not observed in the unlabeled data as treestructured latent variables.Experiments on semantic parsing on the ATIS domain and Python code generation show that with extra unlabeled data, STRUCTVAE outperforms strong supervised models. 1 Chunting Zhou, Junxian He, Graham Neubig |
ACL (1) | 4 |
| 2018 | Stress Test Evaluation for Natural Language InferenceabstractNatural language inference (NLI) is the task of determining if a natural language hypothesis can be inferred from a given premise in a justifiable manner. NLI was proposed as a benchmark task for natural language understanding. Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text. However, the extent to which these models understand the semantic content of sentences is unclear. In this work, we propose an evaluation methodology consisting of automatically constructed “stress tests” that allow us to examine whether systems have the ability to make real inferential decisions. Our evaluation of six sentence-encoder models on these stress tests reveals strengths and weaknesses of these models with respect to challenging linguistic phenomena, and suggests important directions for future work in this area. Aakanksha Naik, Abhilasha Ravichander, Norman M. Sadeh, Carolyn P. Rosé, Graham Neubig |
COLING | 5 |
| 2018 | Adapting Word Embeddings to New Languages with Morphological and Phonological Subword RepresentationsabstractMuch work in Natural Language Processing (NLP) has been for resource-rich languages, making generalization to new, less-resourced languages challenging.We present two approaches for improving generalization to lowresourced languages by adapting continuous word representations using linguistically motivated subword units: phonemes, morphemes and graphemes.Our method requires neither parallel corpora nor bilingual dictionaries and provides a significant gain in performance over previous methods relying on these resources.We demonstrate the effectiveness of our approaches on Named Entity Recognition for four languages, namely Uyghur, Turkish, Bengali and Hindi, of which Uyghur and Bengali are low resource languages, and also perform experiments on Machine Translation.Exploiting subwords with transfer learning gives us a boost of +15.2 NER F1 for Uyghur and +9.7 F1 for Bengali.We also show improvements in the monolingual setting where we achieve (avg.)+3 F1 and (avg.)+1.35 BLEU. Aditi Chaudhary, Chunting Zhou, Lori S. Levin, Graham Neubig, David R. Mortensen, Jaime G. Carbonell |
EMNLP | 4 |
| 2018 | Retrieval-Based Neural Code GenerationabstractIn models to generate program source code from natural language, representing this code in a tree structure has been a common approach.However, existing methods often fail to generate complex code correctly due to a lack of ability to memorize large and complex structures.We introduce RECODE, a method based on subtree retrieval that makes it possible to explicitly reference existing code examples within a neural code generation model.First, we retrieve sentences that are similar to input sentences using a dynamicprogramming-based sentence similarity scoring method.Next, we extract n-grams of action sequences that build the associated abstract syntax tree.Finally, we increase the probability of actions that cause the retrieved n-gram action subtree to be in the predicted code.We show that our approach improves the performance on two code generation tasks by up to +2.6 BLEU. 1 Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Anthony Tomasic, Graham Neubig |
EMNLP | 6 |
| 2018 | Unsupervised Learning of Syntactic Structure with Invertible Neural ProjectionsabstractUnsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters.In most cases, these models have not leveraged continuous word representations.In this work, we propose a novel generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a structured generative prior.We show that the invertibility condition allows for efficient exact inference and marginal likelihood computation in our model so long as the prior is well-behaved.In experiments we instantiate our approach with both Markov and tree-structured priors, evaluating on two tasks: part-of-speech (POS) induction, and unsupervised dependency parsing without gold POS annotation.On the Penn Treebank, our Markov-structured model surpasses state-of-the-art results on POS induction.Similarly, we find that our tree-structured model achieves state-of-the-art performance on unsupervised dependency parsing for the difficult training condition where neither gold POS annotation nor punctuation-based constraints are available. Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP | 2 |
| 2018 | MTNT: A Testbed for Machine Translation of Noisy TextabstractNoisy or non-standard input text can cause disastrous mistranslations in most modern Machine Translation (MT) systems, and there has been growing research interest in creating noise-robust MT systems.However, as of yet there are no publicly available parallel corpora of with naturally occurring noisy inputs and translations, and thus previous work has resorted to evaluating on synthetically created datasets.In this paper, we propose a benchmark dataset for Machine Translation of Noisy Text (MTNT), consisting of noisy comments on Reddit 1 and professionally sourced translations.We commissioned translations of English comments into French and Japanese, as well as French and Japanese comments into English, on the order of 7k-37k sentences per language pair.We qualitatively and quantitatively examine the types of noise included in this dataset, then demonstrate that existing MT models fail badly on a number of noise-related phenomena, even after performing adaptation on a small training set of in-domain data.This indicates that this dataset can provide an attractive testbed for methods tailored to handling noisy text in MT. 2 Paul Michel, Graham Neubig |
EMNLP | 2 |
| 2018 | Rapid Adaptation of Neural Machine Translation to New LanguagesabstractThis paper examines the problem of adapting neural machine translation systems to new, low-resourced languages (LRLs) as effectively and rapidly as possible.We propose methods based on starting with massively multilingual "seed models", which can be trained ahead-of-time, and then continuing training on data related to the LRL.We contrast a number of strategies, leading to a novel, simple, yet effective method of "similar-language regularization", where we jointly train on both a LRL of interest and a similar high-resourced language to prevent over-fitting to small LRL data.Experiments demonstrate that massively multilingual models, even without any explicit adaptation, are surprisingly effective, achieving BLEU scores of up to 15.5 with no data from the LRL, and that the proposed similarlanguage regularization method improves over other adaptation methods by 1.7 BLEU points average over 4 LRL settings.1 Graham Neubig, Junjie Hu 0001 |
EMNLP | 1 |
| 2018 | Contextual Parameter Generation for Universal Neural Machine TranslationabstractWe propose a simple modification to existing neural machine translation (NMT) models that enables using a single universal model to translate between multiple languages while allowing for language specific parameterization, and that can also be used for domain adaptation.Our approach requires no changes to the model architecture of a standard NMT system, but instead introduces a new component, the contextual parameter generator (CPG), that generates the parameters of the system (e.g., weights in a neural network).This parameter generator accepts source and target language embeddings as input, and generates the parameters for the encoder and the decoder, respectively.The rest of the model remains unchanged and is shared across all languages.We show how this simple modification enables the system to use monolingual data for training and also perform zero-shot translation.We further show it is able to surpass state-of-theart performance for both the IWSLT-15 and IWSLT-17 datasets and that the learned language embeddings are able to uncover interesting relationships between languages. Emmanouil A. Platanios, Mrinmaya Sachan, Graham Neubig, Tom M. Mitchell |
EMNLP | 3 |
| 2018 | SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine TranslationabstractIn this work, we examine methods for data augmentation for text-based tasks such as neural machine translation (NMT).We formulate the design of a data augmentation policy with desirable properties as an optimization problem, and derive a generic analytic solution.This solution not only subsumes some existing augmentation schemes, but also leads to an extremely simple data augmentation strategy for NMT: randomly replacing words in both the source sentence and the target sentence with other random words from their corresponding vocabularies.We name this method SwitchOut.Experiments on three translation datasets of different scales show that SwitchOut yields consistent improvements of about 0.5 BLEU, achieving better or comparable performances to strong alternatives such as word dropout (Sennrich et al., 2016a).Code to implement this method is included in the appendix. Xinyi Wang 0001, Zihang Dai, Graham Neubig |
EMNLP | 4 |
| 2018 | A Tree-based Decoder for Neural Machine TranslationabstractRecent advances in Neural Machine Translation (NMT) show that adding syntactic information to NMT systems can improve the quality of their translations.Most existing work utilizes some specific types of linguisticallyinspired tree structures, like constituency and dependency parse trees.This is often done via a standard RNN decoder that operates on a linearized target tree structure.However, it is an open question of what specific linguistic formalism, if any, is the best structural representation for NMT.In this paper, we (1) propose an NMT model that can naturally generate the topology of an arbitrary tree structure on the target side, and (2) experiment with various target tree structures.Our experiments show the surprising result that our model delivers the best improvements with balanced binary trees constructed without any linguistic knowledge; this model outperforms standard seq2seq models by up to 2.1 BLEU points, and other methods for incorporating target-side syntax by up to 0.7 BLEU. 1 Xinyi Wang 0001, Graham Neubig |
EMNLP | 4 |
| 2018 | Neural Cross-lingual Named Entity Recognition with Minimal ResourcesabstractFor languages with no annotated resources, unsupervised transfer of natural language processing models such as named-entity recognition (NER) from resource-rich languages would be an appealing capability.However, differences in words and word order across languages make it a challenging problem.To improve mapping of lexical items across languages, we propose a method that finds translations based on bilingual word embeddings.To improve robustness to word order differences, we propose to use self-attention, which allows for a degree of flexibility with respect to word order.We demonstrate that these methods achieve state-of-the-art or competitive NER performance on commonly tested languages under a cross-lingual setting, with much lower resource requirements than past approaches.We also evaluate the challenges of applying these methods to Uyghur, a lowresource language.1 Jiateng Xie, Zhilin Yang 0001, Graham Neubig, Noah A. Smith, Jaime G. Carbonell |
EMNLP | 3 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 6 |
| 2018 | Self-Attentional Acoustic ModelsabstractSelf-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial to apply to acoustic modeling due to computational and modeling issues. In this paper, we apply self-attention to acoustic modeling, proposing several improvements to mitigate these issues: First, self-attention memory grows quadratically in the sequence length, which we address through a downsampling technique. Second, we find that previous approaches to incorporate position information into the model are unsuitable and explore other representations and hybrid models to this end. Third, to stress the importance of local context in the acoustic signal, we propose a Gaussian biasing approach that allows explicit control over the context range. Experiments find that our model approaches a strong baseline based on LSTMs with network-in-network connections while being much faster to compute. Besides speed, we find that interpretability is a strength of self-attentional acoustic models, and demonstrate that self-attention heads learn a linguistically plausible division of labor. Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 3 |
| 2018 | Evaluation Phonemic Transcription of Low-Resource Tonal Languages for Language Documentation
Oliver Adams, Trevor Cohn, Graham Neubig, Hilaria Cruz, Steven Bird, Alexis Michaud |
LREC | 3 |
| 2018 | Learning to mine aligned code and natural language pairs from stack overflowabstractFor tasks like code synthesis from natural language, code retrieval, and code summarization, data-driven models have shown great promise. However, creating these models require parallel data between natural language (NL) and code with fine-grained alignments. Stack Overflow (SO) is a promising source to create such a data set: the questions are diverse and most of them have corresponding answers with high quality code snippets. However, existing heuristic methods (e.g., pairing the title of a post with the code in the accepted answer) are limited both in their coverage and the correctness of the NL-code pairs obtained. In this paper, we propose a novel method to mine high-quality aligned data from SO using two sets of features: hand-crafted features considering the structure of the extracted snippets, and correspondence features obtained by training a probabilistic model to capture the correlation between NL and code using neural networks. These features are fed into a classifier that determines the quality of mined NL-code pairs. Experiments using Python and Java as test beds show that the proposed method greatly expands coverage and accuracy over existing mining methods, even when using only a small number of labeled examples. Further, we find that reasonable results are achieved even when training the classifier on one language and testing on another, showing promise for scaling NL-code mining to a wide variety of programming languages beyond those for which we are able to annotate data. Edgar Chen, Bogdan Vasilescu, Graham Neubig |
MSR | 5 |
| 2018 | Attentive Interaction Model: Modeling Changes in View in ArgumentationabstractYohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn Rosé, Graham Neubig. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn P. Rosé, Graham Neubig |
NAACL-HLT | 6 |
| 2018 | Handling Homographs in Neural Machine TranslationabstractFrederick Liu, Han Lu, Graham Neubig. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Frederick Liu, Graham Neubig |
NAACL-HLT | 3 |
| 2018 | Using Morphological Knowledge in Open-Vocabulary Neural Language ModelsabstractAustin Matthews, Graham Neubig, Chris Dyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Austin Matthews, Graham Neubig, Chris Dyer |
NAACL-HLT | 2 |
| 2018 | Guiding Neural Machine Translation with Retrieved Translation PiecesabstractJingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, Satoshi Nakamura. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
NAACL-HLT | 4 |
| 2018 | Cavs: An Efficient Runtime System for Dynamic Neural Networks
Shizhen Xu, Hao Zhang 0025, Graham Neubig, Wei Dai 0003, Jin Kyu Kim, Zhijie Deng, Qirong Ho, Eric P. Xing |
USENIX ATC | 3 |
| 2018 | An end-to-end model for cross-lingual transformation of paralinguistic information
Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
Mach. Transl. | 4 |
| 2018 | Neural Lattice Language ModelsabstractIn this work, we propose a new language modeling paradigm that has the ability to perform both prediction and moderation of information flow at multiple granularities: neural lattice language models. These models construct a lattice of possible paths through a sentence and marginalize across this lattice to calculate sequence probabilities or optimize parameters. This approach allows us to seamlessly incorporate linguistic intuitions — including polysemy and the existence of multiword lexical items — into our language model. Experiments on multiple language modeling tasks show that English neural lattice language models that utilize polysemous embeddings are able to improve perplexity by 9.95% relative to a word-level baseline, and that a Chinese model that handles multi-character tokens is able to improve perplexity by 20.94% relative to a character-level baseline. Jacob Buckman, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Learning Character-level Compositionality with Visual FeaturesabstractPrevious work has modeled the compositionality of words by creating characterlevel models of meaning, reducing problems of sparsity for rare words.However, in many writing systems compositionality has an effect even on the character-level: the meaning of a character is derived by the sum of its parts.In this paper, we model this effect by creating embeddings for characters based on their visual characteristics, creating an image for the character and running it through a convolutional neural network to produce a visual character embedding.Experiments on a text classification task demonstrate that such model allows for better processing of instances with rare characters in languages such as Chinese, Japanese, and Korean.Additionally, qualitative analyses demonstrate that our proposed model learns to focus on the parts of characters that carry semantic content, resulting in embeddings that are coherent in visual space. Frederick Liu, Chieh Lo, Graham Neubig |
ACL (1) | 4 |
| 2017 | Neural Machine Translation via Binary Code PredictionabstractIn this paper, we propose a new method for calculating the output layer in neural machine translation systems.The method is based on predicting a binary code for each word and can reduce computation time/memory requirements of the output layer to be logarithmic in vocabulary size in the best case.In addition, we also introduce two advanced approaches to improve the robustness of the proposed model: using error-correcting codes and combining softmax and binary codes.Experiments on two English ↔ Japanese bidirectional translation tasks show proposed models achieve BLEU scores that approach the softmax, while reducing memory usage to the order of less than 1/10 and improving decoding speed on CPUs by x5 to x10. Yusuke Oda, Philip Arthur, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001 |
ACL (1) | 3 |
| 2017 | A Syntactic Neural Model for General-Purpose Code GenerationabstractWe consider the problem of parsing natural language descriptions into source code written in a general-purpose programming language like Python.Existing datadriven methods treat this problem as a language generation task without considering the underlying syntax of the target programming language.Informed by previous work in semantic parsing, in this paper we propose a novel neural architecture powered by a grammar model to explicitly capture the target syntax as prior knowledge.Experiments find this an effective way to scale up to generation of complex programs from natural language descriptions, achieving state-of-the-art results that well outperform previous code generation and semantic parsing approaches. Graham Neubig |
ACL (1) | 2 |
| 2017 | Multi-space Variational Encoder-Decoders for Semi-supervised Labeled Sequence TransductionabstractLabeled sequence transduction is a task of transforming one sequence into another sequence that satisfies desiderata specified by a set of labels.In this paper we propose multi-space variational encoderdecoders, a new model for labeled sequence transduction with semi-supervised learning.The generative model can use neural networks to handle both discrete and continuous latent variables to exploit various features of data.Experiments show that our model provides not only a powerful supervised framework but also can effectively take advantage of the unlabeled data.On the SIGMORPHON morphological inflection benchmark, our model outperforms single-model state-ofart results by a large margin for the majority of languages.1 Chunting Zhou, Graham Neubig |
ACL (1) | 2 |
| 2017 | Cross-Lingual Word Embeddings for Low-Resource Language ModelingabstractOliver Adams, Adam Makarucha, Graham Neubig, Steven Bird, Trevor Cohn. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Oliver Adams, Adam J. Makarucha, Graham Neubig, Steven Bird, Trevor Cohn |
EACL (1) | 3 |
| 2017 | Learning to Translate in Real-time with Neural Machine TranslationabstractTranslating in real-time, a.k.a.simultaneous translation, outputs translation words before the input sentence ends, which is a challenging problem for conventional machine translation methods.We propose a neural machine translation (NMT) framework for simultaneous translation in which an agent learns to make decisions on when to translate from the interaction with a pre-trained NMT environment.To trade off quality and delay, we extensively explore various targets for delay and design a method for beam-search applicable in the simultaneous MT setting.Experiments against state-of-the-art baselines on two language pairs demonstrate the efficacy of the proposed framework both quantitatively and qualitatively. 1 Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O. K. Li |
EACL (1) | 2 |
| 2017 | What Do Recurrent Neural Network Grammars Learn About Syntax?abstractAdhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith |
EACL (1) | 5 |
| 2017 | Charmanteau: Character Embedding Models For Portmanteau CreationabstractPortmanteaus are a word formation phenomenon where two words are combined to form a new word.We propose character-level neural sequence-tosequence (S2S) methods for the task of portmanteau generation that are end-toend-trainable, language independent, and do not explicitly use additional phonetic information.We propose a noisy-channelstyle model, which allows for the incorporation of unsupervised word lists, improving performance over a standard sourceto-target model.This model is made possible by an exhaustive candidate generation strategy specifically enabled by the features of the portmanteau task.Experiments find our approach superior to a state-of-the-art FST-based baseline with respect to ground truth accuracy and human evaluation. Varun Gangal, Harsh Jhamtani, Graham Neubig, Eduard H. Hovy, Eric Nyberg |
EMNLP | 3 |
| 2017 | Learning Language Representations for Typology PredictionabstractOne central mystery of neural NLP is what neural models "know" about their subject matter.When a neural machine translation system learns to translate from one language to another, does it learn the syntax or semantics of the languages?Can this knowledge be extracted from the system to fill holes in human scientific knowledge?Existing typological databases contain relatively full feature specifications for only a few hundred languages.Exploiting the existence of parallel texts in more than a thousand languages, we build a massive many-to-one neural machine translation (NMT) system from 1017 languages into English, and use this to predict information missing from typological databases.Experiments show that the proposed method is able to infer not only syntactic, but also phonological and phonetic inventory features, and improves over a baseline that has access to information about the languages' geographic and phylogenetic neighbors.1 Chaitanya Malaviya, Graham Neubig, Patrick Littell |
EMNLP | 2 |
| 2017 | Neural Lattice-to-Sequence Models for Uncertain InputsabstractThe input to a neural sequence-tosequence model is often determined by an up-stream system, e.g. a word segmenter, part of speech tagger, or speech recognizer.These up-stream models are potentially error-prone.Representing inputs through word lattices allows making this uncertainty explicit by capturing alternative sequences and their posterior probabilities in a compact form.In this work, we extend the TreeLSTM (Tai et al., 2015) into a LatticeLSTM that is able to consume word lattices, and can be used as encoder in an attentional encoderdecoder model.We integrate lattice posterior scores into this architecture by extending the TreeLSTM's child-sum and forget gates and introducing a bias term into the attention mechanism.We experiment with speech translation lattices and report consistent improvements over baselines that translate either the 1-best hypothesis or the lattice without posterior scores. Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel |
EMNLP | 2 |
| 2017 | Improving Neural Machine Translation through Phrase-based Forced DecodingabstractCompared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an existing phrase-based SMT model to compute the phrase-based decoding cost for an NMT output and then using the phrase-based decoding cost to rerank the n-best NMT outputs. The main challenge in implementing this approach is that NMT outputs may not be in the search space of the standard phrase-based decoding algorithm, because the search space of phrase-based SMT is limited by the phrase-based translation rule table. We propose a soft forced decoding algorithm, which can always successfully find a decoding path for any NMT output. We show that using the forced decoding cost to rerank the NMT outputs can successfully improve translation quality on four different language pairs. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
IJCNLP(1) | 4 |
| 2017 | Semi-Supervised Learning of a Pronunciation Dictionary from Disjoint Phonemic Transcripts and Text
Takahiro Shinozaki, Shinji Watanabe 0001, Daichi Mochihashi, Graham Neubig |
INTERSPEECH | 4 |
| 2017 | Adaptive Spelling Error Correction Models for Learner EnglishabstractSpelling errors are a characteristic of learner English and degrade the performances of natural language processing systems targeting English learners. This paper describes a method specially designed for automatically correcting spelling errors in learner English that reduces the effects from noise (e.g., grammatical and spelling errors) by adaptively creating spelling error correction models from raw learner corpora. An evaluation shows that the proposed method outperforms previous edit-distance-based and language-model-based methods. We also report results of an investigation into what types of spelling errors English learners tend to make, using the spelling error models created by the proposed method as a tool for our analysis. Ryo Nagata, Hiroya Takamura, Graham Neubig |
KES | 3 |
| 2017 | On-the-fly Operation Batching in Dynamic Computation GraphsabstractDynamic neural networks toolkits such as PyTorch, DyNet, and Chainer offer more flexibility for implementing models that cope with data of varying dimensions and structure, relative to toolkits that operate on statically declared computations (e.g., TensorFlow, CNTK, and Theano). However, existing toolkits - both static and dynamic - require that the developer organize the computations into the batches necessary for exploiting high-performance data-parallel algorithms and hardware. This batching task is generally difficult, but it becomes a major hurdle as architectures become complex. In this paper, we present an algorithm, and its implementation in the DyNet toolkit, for automatically batching operations. Developers simply write minibatch computations as aggregations of single instance computations, and the batching algorithm seamlessly executes them, on the fly, in computationally efficient batches. On a variety of tasks, we obtain throughput similar to manual batches, as well as comparable speedups over single-instance learning on architectures that are impractical to batch manually. Graham Neubig, Yoav Goldberg, Chris Dyer |
NIPS | 1 |
| 2017 | Controllable Invariance through Adversarial Feature LearningabstractLearning meaningful representations that maintain the content necessary for a particular task while filtering away detrimental variations is a problem of great interest in machine learning. In this paper, we tackle the problem of learning representations invariant to a specific factor or trait of data. The representation learning process is formulated as an adversarial minimax game. We analyze the optimal equilibrium of such a game and find that it amounts to maximizing the uncertainty of inferring the detrimental factor given the representation while maximizing the certainty of making task-specific predictions. On three benchmark tasks, namely fair and bias-free classification, language-independent generation, and lighting-independent image classification, we show that the proposed framework induces an invariant representation, and leads to better generalization evidenced by the improved performance. Qizhe Xie, Zihang Dai, Yulun Du, Eduard H. Hovy, Graham Neubig |
NIPS | 5 |
| 2017 | How Would You Say It? Eliciting Lexically Diverse Dialogue for Supervised Semantic ParsingabstractBuilding dialogue interfaces for realworld scenarios often entails training semantic parsers starting from zero examples.How can we build datasets that better capture the variety of ways users might phrase their queries, and what queries are actually realistic?Wang et al. (2015) proposed a method to build semantic parsing datasets by generating canonical utterances using a grammar and having crowdworkers paraphrase them into natural wording.A limitation of this approach is that it induces bias towards using similar language as the canonical utterances.In this work, we present a methodology that elicits meaningful and lexically diverse queries from users for semantic parsing tasks.Starting from a seed lexicon and a generative grammar, we pair logical forms with mixed text-image representations and ask crowdworkers to paraphrase and confirm the plausibility of the queries that they generated.We use this method to build a semantic parsing dataset from scratch for a dialog agent in a smart-home simulation.We find evidence that this dataset, which we have named SMARTHOME, is demonstrably more lexically diverse and difficult to parse than existing domain-specific semantic parsing datasets. Abhilasha Ravichander, Thomas Manzini, Matthias Grabmair, Graham Neubig, Jonathan Francis, Eric Nyberg |
SIGDIAL Conference | 4 |
| 2017 | Transcribing against time
Matthias Sperber, Graham Neubig, Jan Niehues, Satoshi Nakamura 0001, Alex Waibel |
Speech Commun. | 2 |
| 2017 | Preserving Word-Level Emphasis in Speech-to-Speech TranslationabstractSpeech-to-speech translation (S2ST) is a technology that translates speech across languages, which can remove barriers in cross-lingual communication. In the conventional S2ST systems, the linguistic meaning of speech was translated, but paralinguistic information conveying other features of the speech such as emotion or emphasis were ignored. In this paper, we propose a method to translate paralinguistic information, specifically focusing on emphasis. The method consists of a series of components that can accurately translate emphasis using all acoustic features of speech. First, linear-regression hidden semi-Markov models (LRHSMMs) are used to estimate a real-numbered emphasis value for every word in an utterance, resulting in a sequence of values for the utterance. After that the emphasis translation module translates the estimated emphasis sequence into a target language emphasis sequence using a conditional random field model considering the features of emphasis levels, words, and part-of-speech tags. Finally, the speech synthesis module synthesizes emphasized speech with LR-HSMMs, taking into account the translated emphasis sequence and transcription. The results indicate that our translation model can translate emphasis information, correctly emphasizing words in the target language with 91.6% F-measure by objective evaluation. A listening test with human subjects further showed that they could identify the emphasized words with 87.8% F-measure, and that the naturalness of the audio was preserved. Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A Continuous Space Rule Selection Model for Syntax-based Statistical Machine Translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
ACL (1) | 4 |
| 2016 | Lightly Supervised Quality EstimationabstractEvaluating the quality of output from language processing systems such as machine translation or speech recognition is an essential step in ensuring that they are sufficient for practical use. However, depending on the practical requirements, evaluation approaches can differ strongly. Often, reference-based evaluation measures (such as BLEU or WER) are appealing because they are cheap and allow rapid quantitative comparison. On the other hand, practitioners often focus on manual evaluation because they must deal with frequently changing domains and quality standards requested by customers, for which reference-based evaluation is insufficient or not possible due to missing in-domain reference data (Harris et al., 2016). In this paper, we attempt to bridge this gap by proposing a framework for lightly supervised quality estimation. We collect manually annotated scores for a small number of segments in a test corpus or document, and combine them with automatically predicted quality scores for the remaining segments to predict an overall quality estimate. An evaluation shows that our framework estimates quality more reliably than using fully automatic quality estimation approaches, while keeping annotation effort low by not requiring full references to be available for the particular domain. Matthias Sperber, Graham Neubig, Jan Niehues, Sebastian Stüker, Alex Waibel |
COLING | 2 |
| 2016 | Learning a Lexicon and Translation Model from Phoneme LatticesabstractLanguage documentation begins by gathering speech.Manual or automatic transcription at the word level is typically not possible because of the absence of an orthography or prior lexicon, and though manual phonemic transcription is possible, it is prohibitively slow.On the other hand, translations of the minority language into a major language are more easily acquired.We propose a method to harness such translations to improve automatic phoneme recognition.The method assumes no prior lexicon or translation model, instead learning them from phoneme lattices and translations of the speech being transcribed.Experiments demonstrate phoneme error rate improvements against two baselines and the model's ability to learn useful bilingual lexical entries. Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird, Quoc Truong Do, Satoshi Nakamura 0001 |
EMNLP | 2 |
| 2016 | Incorporating Discrete Translation Lexicons into Neural Machine TranslationabstractNeural machine translation (NMT) often makes mistakes in translating low-frequency content words that are essential to understanding the meaning of the sentence.We propose a method to alleviate this problem by augmenting NMT systems with discrete translation lexicons that efficiently encode translations of these low-frequency words.We describe a method to calculate the lexicon probability of the next word in the translation candidate by using the attention vector of the NMT model to select which source word lexical probabilities the model should focus on.We test two methods to combine this probability with the standard NMT probability: (1) using it as a bias, and (2) linear interpolation.Experiments on two corpora show an improvement of 2.0-2.3BLEU and 0.13-0.44NIST score, and faster convergence time. 1 Philip Arthur, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 2 |
| 2016 | Controlling Output Length in Neural Encoder-DecodersabstractNeural encoder-decoder models have shown great success in many sequence generation tasks.However, previous work has not investigated situations in which we would like to control the length of encoder-decoder outputs.This capability is crucial for applications such as text summarization, in which we have to generate concise summaries with a desired length.In this paper, we propose methods for controlling the output sequence length for neural encoder-decoder models: two decoding-based methods and two learning-based methods. 1 Results show that our learning-based methods have the capability to control length without degrading summary quality in a summarization task. Yuta Kikuchi, Graham Neubig, Ryohei Sasano, Hiroya Takamura, Manabu Okumura |
EMNLP | 2 |
| 2016 | Generalizing and Hybridizing Count-based and Neural Language ModelsabstractLanguage models (LMs) are statistical models that calculate probabilities over sequences of words or other discrete symbols.Currently two major paradigms for language modeling exist: count-based n-gram models, which have advantages of scalability and test-time speed, and neural LMs, which often achieve superior modeling performance.We demonstrate how both varieties of models can be unified in a single modeling framework that defines a set of probability distributions over the vocabulary of words, and then dynamically calculates mixture weights over these distributions.This formulation allows us to create novel hybrid models that combine the desirable features of count-based and neural LMs, and experiments demonstrate the advantages of these approaches. 1 Graham Neubig, Chris Dyer |
EMNLP | 1 |
| 2016 | Personalized unknown word detection in non-native language reading using eye gazeabstractThis paper proposes a method to detect unknown words during natural reading of non-native language text by using eye-tracking features. A previous approach utilizes gaze duration and word rarity features to perform this detection. However, while this system can be used by trained users, its performance is not sufficient during natural reading by untrained users. In this paper, we 1) apply support vector machines (SVM) with novel eye movement features that were not considered in the previous work and 2) examine the effect of personalization. The experimental results demonstrate that learning using SVMs and proposed eye movement features improves detection performance as measured by F-measure and that personalization further improves results. Rui Hiraoka, Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
ICMI | 4 |
| 2016 | Learning a Translation Model from Word LatticesabstractTranslation models have been used to improve automatic speech recognition when speech input is paired with a written translation, primarily for the task of computer-aided translation. Existing approaches require large amounts of parallel text for training the translation models, but for many language pairs this data is not available. We propose a model for learning lexical translation parameters directly from the word lattices for which a transcription is sought. The model is expressed through composition of each lattice with a weighted finite-state transducer representing the translation model, where inference is performed by sampling paths through the composed finitestate transducer. We show consistent word error rate reductions in two datasets, using between just 20 minutes and 4 hours of speech input, additionally outperforming a translation model trained on the 1-best path. Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird |
INTERSPEECH | 2 |
| 2016 | Transferring Emphasis in Speech Translation Using Hard-Attentional Neural Network Models
Quoc Truong Do, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2016 | A Hybrid System for Continuous Word-Level Emphasis Modeling Based on HMM State Clustering and Adaptive Training
Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2016 | Unsupervised Joint Estimation of Grapheme-to-Phoneme Conversion Systems and Acoustic Model Adaptation for Non-Native Speech Recognition
Satoshi Tsujioka, Sakriani Sakti, Koichiro Yoshino, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2016 | Unsupervised Phoneme Segmentation of Previously Unseen Languages
Marco Vetter, Markus Müller 0001, Fatima Hamlaoui, Graham Neubig, Satoshi Nakamura 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 4 |
| 2016 | Optimizing Computer-Assisted Transcription Quality with Iterative User Interfaces
Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
LREC | 2 |
| 2016 | Morphological Inflection Generation Using Character Sequence to Sequence LearningabstractManaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Manaal Faruqui, Yulia Tsvetkov, Graham Neubig, Chris Dyer |
HLT-NAACL | 3 |
| 2016 | Selecting Syntactic, Non-redundant Segments in Active Learning for Machine TranslationabstractAkiva Miura, Graham Neubig, Michael Paul, Satoshi Nakamura. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Akiva Miura, Graham Neubig, Michael Paul, Satoshi Nakamura 0001 |
HLT-NAACL | 2 |
| 2016 | Analyzing the Effect of Entrainment on Dialogue ActsabstractEntrainment is a factor in dialogue that affects not only human-human but also human-machine interaction. While entrainment on the lexical level is well documented, less is known about how entrainment affects dialogue on a more abstract, structural level. In this paper, we investigate the effect of entrainment on dialogue acts and on lexical choice given dialogue acts, as well as how entrainment changes during a dialogue. We also define a novel measure of entrainment to measure these various types of entrainment. These results may serve as guidelines for dialogue systems that would like to entrain with users in a similar manner. Masahiro Mizukami, Koichiro Yoshino, Graham Neubig, David R. Traum, Satoshi Nakamura 0001 |
SIGDIAL Conference | 3 |
| 2016 | Deep bottleneck features and sound-dependent i-vectors for simultaneous recognition of speech and environmental soundsabstractIn speech interfaces, it is often necessary to understand the overall auditory environment, not only recognizing what is being said, but also being aware of the location or actions surrounding the utterance. However, automatic speech recognition (ASR) becomes difficult when recognizing speech with environmental sounds. Standard solutions treat environmental sounds as noise, and remove them to improve ASR performance. On the other hand, most studies on environmental sounds construct classifiers for environmental sounds only, without interference of spoken utterances. But, in reality, such separate situations almost never exist. This study attempts to address the problem of simultaneous recognition of speech and environmental sounds. Particularly, we examine the possibility of using deep neural network (DNN) techniques to recognize speech and environmental sounds simultaneously, and improve the accuracy of both tasks under respective noisy conditions. First, we investigate DNN architectures including two parallel single-task DNNs, and a single multi-task DNN. However, we found direct multi-task learning of simultaneous speech and environmental recognition to be difficult. Therefore, we further propose a method that combines bottleneck features and sound-dependent i-vectors within this framework. Experimental evaluation results reveal that the utilizing bottleneck features and i-vectors as the input of DNNs can help to improve accuracy of each recognition task. Sakriani Sakti, Seiji Kawanishi, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001 |
SLT | 3 |
| 2016 | Optimization for Statistical Machine Translation: A SurveyabstractIn statistical machine translation (SMT), the optimization of the system parameters to maximize translation accuracy is now a fundamental part of virtually all modern systems. In this article, we survey 12 years of research on optimization for SMT, from the seminal work on discriminative models (Och and Ney 2002) and minimum error rate training (Och 2003), to the most recent advances. Starting with a brief introduction to the fundamentals of SMT systems, we follow by covering a wide variety of optimization algorithms for use in both batch and online optimization. Specifically, we discuss losses based on direct error minimization, maximum likelihood, maximum margin, risk minimization, ranking, and more, along with the appropriate methods for minimizing these losses. We also cover recent topics, including large-scale optimization, nonlinear models, domain-dependent optimization, and the effect of MT evaluation measures or search on optimization. Finally, we discuss the current state of affairs in MT optimization, and point out some unresolved problems that will likely be the target of further research in optimization for MT. Graham Neubig, Taro Watanabe |
Comput. Linguistics | 1 |
| 2016 | Learning local word reorderings for hierarchical phrase-based statistical machine translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 5 |
| 2016 | Learning cooperative persuasive dialogue policies using framing
Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
Speech Commun. | 2 |
| 2016 | Postfilters to Modify the Modulation Spectrum for Statistical Parametric Speech SynthesisabstractThis paper presents novel approaches based on modulation spectrum (MS) for high-quality statistical parametric speech synthesis, including text-to-speech (TTS) and voice conversion (VC). Although statistical parametric speech synthesis offers various advantages over concatenative speech synthesis, the synthetic speech quality is still not as good as that of concatenative speech synthesis or the quality of natural speech. One of the biggest issues causing the quality degradation is the over-smoothing effect often observed in the generated speech parameter trajectories. Global variance (GV) is known as a feature well correlated with the over-smoothing effect, and the effectiveness of keeping the GV of the generated speech parameter trajectories similar to those of natural speech has been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we propose using the MS of the generated speech parameter trajectories as a new feature to effectively quantify the over-smoothing effect. Moreover, we propose postfilters to modify the MS utterance by utterance or segment by segment to make the MS of synthetic speech close to that of natural speech. The proposed postfilters are applicable to various synthesizers based on statistical parametric speech synthesis. We first perform an evaluation of the proposed method in the framework of hidden Markov model (HMM)-based TTS, examining its properties from different perspectives. Furthermore, effectiveness of the proposed postfilters are also evaluated in Gaussian mixture model (GMM)-based VC and classification and regression trees (CART)-based TTS (a.k.a., CLUSTERGEN). The experimental results demonstrate that 1) the proposed utterance-level postfilter achieves quality comparable to the conventional generation algorithm considering the GV, and yields significant improvements by applying to the GV-based generation algorithm in HMM-based TTS, 2) the proposed segment-level postfilter capable of achieving low-delay synthesis also yields significant improvements in synthetic speech quality, and 3) the proposed postfilters are also effective in not only HMM-based TTS but also GMM-based VC and CLUSTERGEN. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Teaching Social Communication Skills Through Human-Agent InteractionabstractThere are a large number of computer-based systems that aim to train and improve social skills. However, most of these do not resemble the training regimens used by human instructors. In this article, we propose a computer-based training system that follows the procedure of social skills training (SST), a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. We attempt to automate the process of SST by developing a dialogue system named the automated social skills trainer , which teaches social communication skills through human-agent interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users. Its design is based on conventional SST performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. We performed a series of three experiments investigating (1) the advantages of using computer-based training systems compared to human-human interaction (HHI) by subjectively evaluating nervousness, ease of talking, and ability to talk well; (2) the relationship between speech language features and human social skills; and (3) the effect of computer-based training using our proposed system. Results of our first experiment show that interaction with an avatar decreases nervousness and increases the user's subjective impression of his or her ability to talk well compared to interaction with an unfamiliar person. The experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with social skills. Finally, experiments measuring the effect of performing SST with the proposed application show that participants significantly improve their skill, as assessed by separate evaluators, by using the system for 50 minutes. A user survey also shows that the users thought our system is useful and easy to use, and that interaction with the avatar felt similar to HHI. Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001 |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Syntax-based Simultaneous Translation through Prediction of Unseen Syntactic ConstituentsabstractYusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ACL (1) | 2 |
| 2015 | The NAIST ASR system for the 2015 Multi-Genre Broadcast challenge: On combination of deep learning systems using a rank-score functionabstractThe Multi-Genre Broadcast challenge is an official challenge of the IEEE Automatic Speech Recognition and Understanding Workshop. This paper presents NAISTs contribution to the premiere of this challenge. The presented speech-to-text system for English makes use of various front-ends (e.g., MFCC, i-vector and FBANK), DNN acoustic models and several language models for decoding and rescoring (N-gram, RNNLM). Subsets of the training data with varying sizes were evaluated with respect to the overall training quality. Two speech segmentation systems were developed for the challenge, based on DNNs and GMM-HMMs. Recognition was performed in three stages: Decoding, lattice rescoring and system combination. This paper focuses on the system combination experiments and presents a rank-score based system weighting approach, which gave better performance compared to a normal system combination strategy. The DNN based ASR system trained on MFCC + i-vector features with the sMBR training criterion gives the best performance of 27.8% WER, and thus significantly outperforms the baseline DNN-HMM sMBR yielding 33.7% WER. Quoc Truong Do, Michael Heck, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 4 |
| 2015 | A study of social-affective communication: Automatic prediction of emotion triggers and responses in television talk showsabstractAdvancements in spoken language technologies have allowed users to interact with computers in an increasingly natural manner. However, most conversational agents or dialogue systems are yet to consider emotional awareness in interaction. To consider emotion in these situations, social-affective knowledge in conversational agents is essential. In this paper, we present a study of the social-affective process in natural conversation from television talk shows. We analyze occurrences of emotion (emotional responses), and the events that elicit them (emotional triggers). We then utilize our analysis for prediction to model the ability of a dialogue system to decide an action and response in an affective interaction. This knowledge has great potential to incorporate emotion into human-computer interaction. Experiments in two languages, English and Indonesian, show that automatic prediction performance surpasses random guessing accuracy. Nurul Lubis, Sakriani Sakti, Graham Neubig, Koichiro Yoshino, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2015 | Adaptive selection from multiple response candidates in example-based dialogueabstractIn spoken dialogue systems, dialogue modeling is one of the most important factors for contributing to user satisfaction improvement. Especially in Example-Based Dialogue Modeling (EBDM), effective methods to build dialogue example databases and to select response utterances from examples are the keys for improving dialogue quality. In dialogue corpora, it often have plural appropriate responses for one utterance. However, the system merges these plural appropriate responses into the one system response, thus, it does not try to use plural responses properly by user preference. In fact, responses that each user thinks to be preferable are different. In this paper, we propose a framework that select an appropriate response from plural appropriate response candidates satisfies users. It has a multi-response example database, and selects an appropriate response based on collaborative filtering. Experimental results showed that the proposed framework were successfully choosing appropriate responses, considering multi-response candidates improves user satisfaction to 4.1 from 3.7 of single response, and the adaptive response selection method increased user satisfaction from 3.7 to 4.3. Masahiro Mizukami, Hideaki Kizuki, Toshio Nomura, Graham Neubig, Koichiro Yoshino, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 4 |
| 2015 | Incremental sentence compression using LSTM recurrent networksabstractMany of the current sentence compression techniques attempt to produce a shortened form of a sentence by relying on syntactic structure such as dependency tree representations. While the performance of sentence compression has been improving, these approaches require a full parse of the sentence before performing sentence compression, making it difficult to perform compression in real time. In this paper, we examine the possibilities of performing incremental sentence compression using long short-term memory (LSTM) recurrent neural networks (RNN). The decision of whether to remove a word is done at each time step, without waiting for the end of the sentence. Various RNN parameters are investigated, including the number of layers and network connections. Furthermore, we also propose using a pretraining method in which the network is pretrained as an autoencoder. Experimental results reveal that our method obtains compression rates similar to human references and a better accuracy than the state-of-the-art tree transduction models. Sakriani Sakti, Faiz Ilham, Graham Neubig, Tomoki Toda, Ayu Purwarianti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2015 | An Enhanced Electrolarynx with Automatic Fundamental Frequency Control based on Statistical PredictionabstractAn electrolarynx is a type of speaking aid device which is able to mechanically generate excitation sounds to help laryngectomees produce electrolaryngeal (EL) speech. Although EL speech is quite intelligible, its naturalness suffers from monotonous fundamental frequency patterns of the mechanical excitation sounds. To make it possible to generate more natural excitation sounds, we have proposed a method to automatically control the fundamental frequency of the sounds generated by the electrolarynx based on a statistical prediction model, which predicts the fundamental frequency patterns from the produced EL speech in real-time. In this paper, we develop a prototype system by implementing the proposed control method in an actual, physical electrolarynx and evaluate its performance. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ASSETS | 3 |
| 2015 | A Binarized Neural Network Joint Model for Machine TranslationabstractThe neural network joint model (NNJM), which augments the neural network language model (NNLM) with an m-word source context window, has achieved large gains in machine translation accuracy, but also has problems with high normalization cost when using large vocabularies.Training the NNJM with noise-contrastive estimation (NCE), instead of standard maximum likelihood estimation (MLE), can reduce computation cost.In this paper, we propose an alternative to NCE, the binarized NNJM (BNNJM), which learns a binary classifier that takes both the context and target words as input, and can be efficiently trained using MLE.We compare the BNNJM and NNJM trained by NCE on various translation tasks. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 4 |
| 2015 | EEG signal enhancement using multi-channel wiener filter with a spatial correlation priorabstractEvent-related potentials (ERPs) of electroencephalogram (EEG) are often used as features for brain machine interfaces or for analysis of brain activities. However, as EEG signals easily suffer from various artifacts, ERPs are often collapsed and hard to observe. There are several attempts at using multi-channel EEG signals to enhance EEG signals of interest and make ERPs more clearly observed. For example, a previous work has proposed a blind EEG signal separation method using a multi-channel Wiener filter designed with a probabilistic generative model of observed EEG signals. This method copes with the under-determination of EEG signal separation by assuming sparseness of each EEG component in the time-frequency domain. Although this method blindly separates EEG signals into individual EEG components using time-varying scaled spatial correlation matrices, target EEG components, such as P300 of ERP, are often known in advance in some applications. In this paper, inspired by this previous work, we propose a probabilistic EEG signal enhancement method using a multi-channel Wiener filter, newly incorporating prior information of the spatial correlation matrices related to the target EEG component in the probabilistic generative model to improve performance of EEG signal enhancement. An experimental evaluation for P300 enhancement shows that the proposed method significantly reduces artifacts. Hayato Maki, Tomoki Toda, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2015 | Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASRabstractThis paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks or spectrogram features. In particular, we combine within low and high levels of CNN framework which we refer to as low-level and high-level feature combination. As comparison, we also construct the similar configuration with deep neural network (DNN). Performance was evaluated in the framework of hybrid neural network - hidden Markov model (NN-HMM) system on TIMIT phoneme sequence recognition task. The results reveal that cochleogram-spectrogram feature combination provides significant advantages. The best accuracy was obtained by high-level combination of two dimensional cochleogram-spectrogram features using CNN, achieved up to 8.2% relative phoneme error rate (PER) reduction from CNN single features or 19.7% relative PER reduction from DNN single features. Andros Tjandra, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2015 | Preserving word-level emphasis in speech-to-speech translation using linear regression HSMMsabstractIn speech, emphasis is an important type of paralinguistic information that helps convey the focus of an utterance, new information, and emotion. If emphasis can be incorporated into a speech-to-speech (S2S) translation system, it will be possible to convey this information across the language barrier. However, previous related work focuses only on the translation of particular prosodic features, such as F0, or works with emphasis but focuses on extremely small vocabularies, such as the 10 digits. In this paper, we describe a new S2S method that is able to translate the emphasis across languages and consider multiple features of emphasis such as power, F0, and duration over larger vocabularies. We do so by introducing two new components: word-level emphasis estimation using linear regression hidden semi-Markov models, and emphasis translation that translates the word-level emphasis to the target language with conditional random fields. The text-to-speech synthesis system is also modified to be able to synthesize emphasized speech. The result shows that our system can translate the emphasis correctly with 91.6% F -measure for objective test, and 87.8% for subjective test. Quoc Truong Do, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2015 | Statistical singing voice conversion based on direct waveform modification with global varianceabstractThis paper presents techniques to improve the quality of voices generated through statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method makes it possible to convert singing voice characteristics of a source singer into those of a target singer without using vocoder-based waveform generation. However, quality of the converted singing voice still degrades compared to that of a natural singing voice due to various factors, such as the over-smoothing of the converted spectral parameter trajectory. To alleviate this over-smoothing, we propose a technique to restore the global variance of the converted spectral parameter trajectory within the framework of the DIFFSVC method. We also propose another technique to specifically avoid over-smoothing at unvoiced frames. Results of subjective and objective evaluations demonstrate that the proposed techniques significantly improve speech quality of the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional DIFFSVC. Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2015 | Speed or accuracy? a study in evaluation of simultaneous speech translationabstractSimultaneous speech translation is a technology that attempts to reduce the delay inherent in speech translation by beginning translation before the end of explicit sentence boundaries. Despite best efforts, there is still often a trade-off between speed and accuracy in these systems, with systems with less delay also achieving lower accuracy. However, somewhat surprisingly, there is no previous work examining the relative importance of speed and accuracy, and thus given two systems with various speeds and accuracies, it is difficult to say with certainty which is better. In this paper, we make the first steps towards evaluation of simultaneous speech translation systems in consideration of both speed and accuracy. We collect user evaluations of speech translation results with different levels of accuracy and delay, and using this data to learn the parameters of an evaluation measure that can judge the trade-off between these two factors. Based on these results, we find that considering both accuracy and delay in the evaluation of speech translation results helps improve correlations with human judgements, and that users placed higher relative importance on reducing delay when results were presented through text, rather than speech. Takashi Mieno, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2015 | A latent variable model for joint pause prediction and dependency parsingabstractThe prosody of speech is closely related to syntactic structure of the spoken sentence, and thus analysis models that jointly consider these two types of information are promising. However, manual annotation of syntactic information and prosodic information such as pauses is laborious, and thus it can be difficult to obtain sufficient data to train such joint models. In this paper, we tackle this problem by introducing a joint pause prediction and dependency parsing model that treats pauses between consecutive words as latent variables. Using this model, it is possible to learn from not only data labeled with both syntax and pause information, but also data labeled with only syntactic information, which can be obtained in larger quantities. Experiments find that a joint pause prediction and dependency parsing model obtains better pause prediction F-measure than a decision-tree-based baseline trained on the same data, and that the addition of more data using the proposed latent variable model leads for further gains of up to 11.6 points in F-measure. The Tung Nguyen, Graham Neubig, Hiroyuki Shindo, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2015 | Non-native speech synthesis preserving speaker individuality based on partial correction of prosodic and phonetic characteristicsabstractThis paper presents a novel non-native speech synthesis technique that preserves the individuality of a non-native speaker. Cross-lingual speech synthesis based on voice conversion or HMM-based speech synthesis, which synthesizes foreign language speech of a specific non-native speaker reflecting the speaker-dependent acoustic characteristics extracted from the speaker’s natural speech in his/her mother tongue, tends to cause a degradation of speaker individuality in synthetic speech compared to intra-lingual speech synthesis. This paper proposes a new approach to cross-lingual speech synthesis that preserves speaker individuality by explicitly using non-native speech spoken by the target speaker. Although the use of nonnative speech makes it possible to preserve the speaker individuality in the synthesized target speech, naturalness is significantly degraded as the speech is directly affected by unnatural prosody and pronunciation often caused by differences in the linguistic systems of the source and target languages. To improve naturalness while preserving speaker individuality, we propose (1) a prosodic correction method based on model adaptation, and (2) a phonetic correction method based on spectrum replacement for unvoiced consonants. The experimental results demonstrate that these proposed methods are capable of significantly improving naturalness while preserving the speaker individuality in synthetic speech. Yuji Oshima, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2015 | Non-audible murmur enhancement based on statistical conversion using air- and body-conductive microphones in noisy environmentsabstractNon-Audible Murmur (NAM) is an extremely soft whispered voice detected by a special body-conductive microphone called a NAM microphone. Although NAM is a promising medium for silent speech communication, its quality is significantly degraded by its faint volume and spectral changes caused by body-conductive recording. To improve the quality of NAM, several enhancement methods based on statistical voice conversion (VC) techniques have been proposed, and their effectiveness has been confirmed in quiet environments. However, it can be expected that NAM will be used not only in quiet, but also in noisy environments, and it is thus necessary to develop enhancement methods that will also work in these cases. In this paper, we propose a framework for NAM enhancement using not only the NAM microphone but also an air-conductive microphone. Airand body-conducted NAM signals are used as the input of VC to estimate a more naturally sounding speech signal. To clarify adverse effects of external noises on the performance of the proposed framework and investigate a possibility to alleviate them by revising VC models, we also implement noise-dependent VC models within the proposed framework. Experimental results demonstrate that the proposed framework yields significant improvements in the spectral conversion accuracy and listenability of enhanced speech under both quiet and noisy environments. Yusuke Tajiri, Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2015 | Articulatory controllable speech modification based on Gaussian mixture models with direct waveform modification using spectrum differentialabstractIn our previous work, we have developed a speech modification system capable of manipulating unobserved articulatory movements by sequentially performing speech-to-articulatory inversion mapping and articulatory-to-speech production mapping based on a Gaussian mixture model (GMM)-based statistical feature mapping technique. One of the biggest issues to be addressed in this system is quality degradation of the synthetic speech caused by modeling and conversion errors in a vocoderbased waveform generation framework. To address this issue, we propose several implementation methods of direct waveform modification. The proposed methods directly filter an input speech waveform with a time sequence of spectral differential parameters calculated between unmodified and modified spectral envelop parameters in order to avoid using vocoderbased excitation signal generation. The experimental results show that the proposed direct waveform modification methods yield significantly larger quality improvements in the synthetic speech while also keeping a capability of intuitively modifying phoneme sounds by manipulating the unobserved articulatory movements. Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2015 | Automated Social Skills TrainerabstractSocial skills training is a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. In this paper, we attempt to automate the process of social skills training by developing a dialogue system named "automated social skills trainer," which provides social skills training through human-computer interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users to improve their social skills. Its design is based on conventional social skills training performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. An experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with autistic traits. Additional experiments measuring the effect of performing social skills training with the proposed application show that most participants improve their skill by using the system for 50 minutes. Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001 |
IUI | 3 |
| 2015 | Pseudogen: A Tool to Automatically Generate Pseudo-Code from Source CodeabstractUnderstanding the behavior of source code written in an unfamiliar programming language is difficult. One way to aid understanding of difficult code is to add corresponding pseudo-code, which describes in detail the workings of the code in a natural language such as English. In spite of its usefulness, most source code does not have corresponding pseudo-code because it is tedious to create. This paper demonstrates a tool Pseudogen that makes it possible to automatically generate pseudo-code from source code using statistical machine translation (SMT). Pseudogen currently supports generation of English or Japanese pseudo-code from Python source code, and the SMT framework makes it easy for users to create new generators for their preferred source code/pseudo-code pairs. Hiroyuki Fudaba, Yusuke Oda, Koichi Akabe, Graham Neubig, Hideaki Hata, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASE | 4 |
| 2015 | Learning to Generate Pseudo-Code from Source Code Using Statistical Machine Translation (T)abstractPseudo-code written in natural language can aid the comprehension of source code in unfamiliar programming languages. However, the great majority of source code has no corresponding pseudo-code, because pseudo-code is redundant and laborious to create. If pseudo-code could be generated automatically and instantly from given source code, we could allow for on-demand production of pseudo-code without human effort. In this paper, we propose a method to automatically generate pseudo-code from source code, specifically adopting the statistical machine translation (SMT) framework. SMT, which was originally designed to translate between two natural languages, allows us to automatically learn the relationship between source code/pseudo-code pairs, making it possible to create a pseudo-code generator with less human effort. In experiments, we generated English or Japanese pseudo-code from Python statements using SMT, and find that the generated pseudo-code is largely accurate, and aids code understanding. Yusuke Oda, Hiroyuki Fudaba, Graham Neubig, Hideaki Hata, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASE | 3 |
| 2015 | Multi-Target Machine Translation with Multi-Synchronous Context-free GrammarsabstractWe propose a method for simultaneously translating from a single source language to multiple target languages T1, T2, etc.The motivation behind this method is that if we only have a weak language model for T1 and translations in T1 and T2 are associated, we can use the information from a strong language model over T2 to disambiguate the translations in T1, providing better translation results.As a specific framework to realize multi-target translation, we expand the formalism of synchronous context-free grammars to handle multiple targets, and describe methods for rule extraction, scoring, pruning, and search with these models.Experiments find that multi-target translation with a strong language model in a similar second target language can provide gains of up to 0.8-1.5 BLEU points. 1 Graham Neubig, Philip Arthur, Kevin Duh |
HLT-NAACL | 1 |
| 2015 | Ckylark: A More Robust PCFG-LA ParserabstractYusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
HLT-NAACL | 2 |
| 2015 | Semantic Parsing of Ambiguous Input through Paraphrasing and VerificationabstractWe propose a new method for semantic parsing of ambiguous and ungrammatical input, such as search queries. We do so by building on an existing semantic parsing framework that uses synchronous context free grammars (SCFG) to jointly model the input sentence and output meaning representation. We generalize this SCFG framework to allow not one, but multiple outputs. Using this formalism, we construct a grammar that takes an ambiguous input string and jointly maps it into both a meaning representation and a natural language paraphrase that is less ambiguous than the original input. This paraphrase can be used to disambiguate the meaning representation via verification using a language model that calculates the probability of each paraphrase. Philip Arthur, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Discriminative Language Models as a Tool for Machine Translation Error Analysis
Koichi Akabe, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
COLING | 2 |
| 2014 | Reinforcement Learning of Cooperative Persuasive Dialogue Policies using Framing
Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
COLING | 2 |
| 2014 | Acquiring a Dictionary of Emotion-Provoking EventsabstractHoa Trong Vu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Hoa Trong Vu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
EACL | 2 |
| 2014 | Regression approaches to perceptual age control in singing voice conversionabstractThe perceptual age of a singing voice is the age of the singer as perceived by the listener, and is one of the notable characteristics that determines perceptions of a song. In this paper, we describe a novel voice timbre control technique based on the perceptual age for singing voice conversion (SVC). Singers can sing expressively by controlling prosody and voice timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome the limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice timbre of an arbitrary source singer into that of an arbitrary target singer. However, it is still difficult to intuitively control singing voice characteristics by manipulating parameters corresponding to specific physical traits, such as gender and age. In this paper, we develop a technique for controlling the voice timbre based on perceptual age that maintains the singer's individuality. The experimental results show that the proposed voice timbre control method makes it possible to change the singer's perceptual age while not having an adverse effect on the perceived individuality. Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2014 | Narrow Adaptive Regularization of weights for grapheme-to-phoneme conversionabstractAs the speech recognition field proceeds to open domain and multilingual tasks, the need for robust g2p conversion has been increasing. Towards this objective, we propose a new g2p conversion training method based on the Narrow Adaptive Regularization of Weights (NAROW) online learning algorithm. NAROW improves over its predecessor AROW by automatically adjusting hyperparameters to reduce mistake bounds, and ensuring that the learning rate is not updated when features for the input data have already been updated enough. The contribution of this paper is first to extend NAROW to structured learning, and show the inequality to bound the maximum number of errors in structured NAROW. In experiments, our proposed approach significantly improved over MIRA with consistent phoneme error rate reductions of 1.3-3.8% on a variety of dictionaries. Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2014 | A postfilter to modify the modulation spectrum in HMM-based speech synthesisabstractIn this paper, we propose a postfilter to compensate modulation spectrum in HMM-based speech synthesis. In order to alleviate over-smoothing effects which is a main cause of quality degradation in HMM-based speech synthesis, it is necessary to consider features that can capture over-smoothing. Global Variance (GV) is one well-known example of such a feature, and the effectiveness of parameter generation algorithm considering GV have been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we introduce the Modulation Spectrum (MS) of speech parameter trajectory as a new feature to effectively capture the over-smoothing effect, and we propose a postfilter based on the MS. The MS is represented as a power spectrum of the parameter trajectory. The generated speech parameter sequence is filtered to ensure that its MS has a pattern similar to natural speech. Experimental results show quality improvements when the proposed methods are applied to spectral and F0components, compared with conventional methods considering GV. Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2014 | An evaluation of excitation feature prediction in a hybrid approach to electrolaryngeal speech enhancementabstractWe implement removing micro-prosody with low-pass filtering and avoiding Unvoiced/Voiced (U/V) prediction as part of a hybrid approach to improve statistical excitation prediction in electrolaryngeal (EL) speech enhancement. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, in our previous work, we proposed a hybrid method using a noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. In this paper, we evaluate the effect of removing micro-prosody with low-pass filtering and avoiding U/V prediction in the hybrid enhancement process. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2014 | A hearing impairment simulation method using audiogram-based approximation of auditory charatecteristicsabstractHearing impairment simulation is an effective technique to educate normal-hearing people about auditory perception of the hearing-impaired. Because auditory characteristics of the hearing impaired vary greatly from person-to-person, personalization of the hearing impairment simulation systems is essential to accurately simulate these individual differences. However, measurement of auditory characteristics of individuals is time-consuming work. In this paper, we propose a hearing impairment simulation method that is easily applied to individual hearing-impaired persons. Auditory filter characteristics and gain characteristics are estimated from easily measurable audiograms of each individual. We also implement a method for manually adjusting the hearing impairment level to improve accuracy of the proposed hearing impairment simulation. An experimental evaluation is conducted to compare intelligibility between hearing-impaired and normal-hearing persons with the proposed hearing impairment simulation. The experimental results show that the proposed method effectively makes the word correct rate and phoneme confusion tendency of the normal hearing persons similar to those of the hearing impaired persons. Index Terms: hearing-impairment simulation, personalization, auditory filter characteristics, gain characteristics, audiogram, Nozomi Jinbo, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2014 | Statistical singing voice conversion with direct waveform modification based on the spectrum differentialabstractThis paper presents a novel statistical singing voice conversion (SVC) technique with direct waveform modification based on the spectrum differential that can convert voice timbre of a source singer into that of a target singer without using a vocoder to generate converted singing voice waveforms. SVC makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, speech quality of the converted singing voice is significantly degraded compared to that of a natural singing voice due to various factors, such as analysis and modeling errors in the vocoderbased framework. To alleviate this degradation, we propose a statistical conversion process that directly modifies the signal in the waveform domain by estimating the difference in the spectra of the source and target singers’ singing voices. The differential spectral feature is directly estimated using a differential Gaussian mixture model (GMM) that is analytically derived from the traditional GMM used as a conversion model in the conventional SVC. The experimental results demonstrate that the proposed method makes it possible to significantly improve speech quality in the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional SVC. Index Terms: singing voice, statistical voice conversion, vocoder, Gaussian mixture model, differential spectral compensation Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2014 | Structured soft margin confidence weighted learning for grapheme-to-phoneme conversionabstractIn recent years, structured online discriminative learning methods using second order statistics have been shown to outperform conventional generative and discriminative models in the grapheme-to-phoneme (g2p) conversion task. However, these methods update the parameters by sequentially using N-best hypotheses predicted with the current parameters. Thus, the parameters appearing in early hypotheses are overfitted compared with those in later hypotheses. In this paper, we propose a novel method called structured soft margin confidence weighted learning, which extends multi-class confidence weighted learning to structured learning. The proposed method extends multiclass CW in two ways, allowing for improved robustness to overfitting: (1) regularization inspired by soft margin support vector machines, allowing for margin error, and (2) update using N-best hypotheses simultaneously and interdependently. In an evaluation experiment on the g2p conversion task, the proposed method improved over all other approaches in terms of phoneme error rate with a significant difference. Index Terms: g2p conversion, out-of-vocabulary word, online discriminative training, structured learning, confidence weighted algorithm Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2014 | Data-driven generation of text balloons based on linguistic and acoustic features of a comics-anime corpusabstractMost automatic speech recognition systems existing today are still limited to recognizing what is being said, without being concerned with how it is being said. On the other hand, research on emotion recognition from speech has recently gained considerable interest, but how those emotions could be expressed in text-based communication has not been widely investigated. Our long-term goal is to construct expressive speech-to-text systems that conveys all information from acoustic speech, including verbal message, emotional state, speaker condition, and background noise, into unified text-based communication. In this preliminary study, we start with developing a system that can convey emotional speech into text-based communication by way of text balloons. As there exist many possible ways to generate the text balloons, we propose to utilize linguistic and acoustic features based on comic books and anime films. Experimental results reveal that expressive text is more preferable than static text, and the system is able to estimate the shape of text balloons with 87.01% accuracy. Index Terms: data-driven approaches, expressive text generation, linguistic and acoustic features Sho Matsumiya, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2014 | Direct F0 control of an electrolarynx based on statistical excitation feature prediction and its evaluation through simulationabstractAn electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce electrolaryngeal (EL) speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. To address this issue, we have proposed several EL speech enhancement methods using statistical voice conversion and showed that statistical prediction of excitation parameters, such as F0 patterns, was essential to significantly improve naturalness of EL speech. In these methods, the original EL speech is recorded with a microphone and the enhanced EL speech is presented from a loudspeaker in real time. This framework is effective for telecommunication but it is not suitable to face-to-face conversation because both the original EL speech and the enhanced EL speech are presented to listeners. In this paper, we propose direct F0 control of the electrolarynx based on statistical excitation prediction to develop an EL speech enhancement technique also effective for face-to-face conversation. F0 patterns of excitation signals produced by the electrolarynx are predicted in real time from the EL speech produced by the laryngectomee’s articulation of the excitation signals with previously predicted F0 values. A simulation experiment is conducted to evaluate the effectiveness of the proposed method. The experimental results demonstrate that the proposed method yields significant improvements in naturalness of EL speech while keeping its intelligibility high enough. Index Terms: laryngectomee, electrolarynx, electrolaryngeal speech, statistical excitation prediction, simulation evaluation Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2014 | Articulatory controllable speech modification based on statistical feature mapping with Gaussian mixture modelsabstractThis paper presents a novel speech modification method capable of controlling unobservable articulatory parameters based on a statistical feature mapping technique with Gaussian Mixture Models (GMMs). In previous work [1], the GMM-based statistical feature mapping was successfully applied to acousticto-articulatory inversion mapping and articulatory-to-acoustic production mapping separately. In this paper, these two mapping frameworks are integrated to a unified framework to develop a novel speech modification system. The proposed system sequentially performs the inversion and the production mapping, making it possible to modify phonemic sounds of an input speech signal by intuitively manipulating articulatory parameters estimated from the input speech signal. We also propose a manipulation method to automatically compensate for unmodified articulatory movements considering inter-dimensional correlation of the articulatory parameters. The proposed system is implemented for a single English speaker and its effectiveness is evaluated experimentally. The experimental results demonstrate that the proposed system is capable of modifying phonemic sounds by manipulating the estimated articulatory movements and higher speech quality is achieved by considering the inter-dimensional correlation in the manipulation. Index Terms: speech modification, acoustic-to-articulatory inversion mapping, articulatory-to-acoustic production mapping, Gaussian mixture model, inter-dimensional correlation Patrick Lumban Tobing, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001, Ayu Purwarianti |
INTERSPEECH | 3 |
| 2014 | Language Resource Addition: Dictionary or Corpus?
Shinsuke Mori, Graham Neubig |
LREC | 2 |
| 2014 | Towards Multilingual Conversations in the Medical Domain: Development of Multilingual Medical Data and A Network-based ASR System
Sakriani Sakti, Keigo Kubo, Sho Matsumiya, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001, Fumihiro Adachi, Ryosuke Isotani |
LREC | 4 |
| 2014 | Collection of a Simultaneous Translation Corpus for Comparative Analysis
Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
LREC | 2 |
| 2014 | Improving the robustness of example-based dialog retrieval using recursive neural network paraphrase identificationabstractPrevious work on example-based chat-oriented dialog systems utilizing real human-to-human conversation has shown promising results. However, most previous methods use relatively simple retrieval techniques, resulting in weakness to out of vocabulary (OOV) database queries and inadequate handling of interactions between words in the sentence. To overcome this problem, in this paper we propose a method to utilize recursive neural network paraphrase identification to improve the accuracy and robustness of example-based dialog response retrieval. We model our dialog-pair database and user input query with distributed word representations, and employ recursive autoencoders and dynamic pooling to determine whether two sentences with arbitrary length have the same meaning. The distributed representations have the potential to improve handling of OOV cases, and the recursive structure can reduce confusion in example matching. We evaluate the system performance based on objective and subjective metrics. Lasguido Nio, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
SLT | 3 |
| 2014 | On-the-fly user modeling for cost-sensitive correction of speech transcriptsabstractWe propose an on-the-fly updating framework for cost-sensitive manual correction of automatically recognized speech transcripts. This framework trains cost-models during the transcription process, and does not require the transcriber enrollment necessary in previous work. We use a baseline method that optimizes a segmentation into segments to supervise or not to supervise in a cost-sensitive fashion that minimizes human effort, and introduce a much faster algorithm for computing such a segmentation that can be used for on-the-fly updates. Besides removing the need to carry out enrollments, experiments show that our updating framework results in 28% higher human supervision efficiency than previous cost-sensitive approaches. Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
SLT | 2 |
| 2014 | Segmentation for Efficient Supervised Language Annotation with an Explicit Cost-Utility TradeoffabstractIn this paper, we study the problem of manually correcting automatic annotations of natural language in as efficient a manner as possible. We introduce a method for automatically segmenting a corpus into chunks such that many uncertain labels are grouped into the same chunk, while human supervision can be omitted altogether for other segments. A tradeoff must be found for segment sizes. Choosing short segments allows us to reduce the number of highly confident labels that are supervised by the annotator, which is useful because these labels are often already correct and supervising correct labels is a waste of effort. In contrast, long segments reduce the cognitive effort due to context switches. Our method helps find the segmentation that optimizes supervision efficiency by defining user models to predict the cost and utility of supervising each segment and solving a constrained optimization problem balancing these contradictory objectives. A user study demonstrates noticeable gains over pre-segmented, confidence-ordered baselines on two natural language processing tasks: speech transcription and word segmentation. Matthias Sperber, Mirjam Simantzik, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
Trans. Assoc. Comput. Linguistics | 3 |
| 2013 | Dialogue management for leading the conversation in persuasive dialogue systemsabstractIn this research, we propose a probabilistic dialogue modeling method for persuasive dialogue systems that interact with the user based on a specific goal, and lead the user to take actions that the system intends from candidate actions satisfying the user's needs. As a baseline system, we develop a dialogue model assuming the user makes decisions based on preference. Then we improve the model by introducing methods to guide the user from topic to topic. We evaluate the system knowledge and dialogue manager in a task that tests the system's persuasive power, and find that the proposed method is effective in this respect. Takuya Hiraoka, Yuki Yamauchi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2013 | Simple, lexicalized choice of translation timing for simultaneous speech translation
Tomoki Fujita, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2013 | Generalizing continuous-space translation of paralinguistic informationabstractIn previous work, we proposed a model for speech-to-speech translation that is sensitive to paralinguistic information such as duration and power of spoken words [1]. This model uses linear regression to map source acoustic features to target acoustic features directly and in continuous space. However, while the model is effective, it faces scalability issues as a single model must be trained for every word, which makes it difficult to generalize to words for which we do not have parallel speech. In this work we first demonstrate that simply training a linear regression model on all words is not sufficient to express paralinguistic translation. We next describe a neural network model that has sufficient expressive power to perform paralinguistic translation with a single model. We evaluate the proposed method on a digit translation task and show that we achieve similar results with a single neural network-based model as previous work did using word-dependent models. Index Terms: speech translation, paralinguistic information, linear regression, neural network Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2013 | An investigation of acoustic features for singing voice conversion based on perceptual ageabstractIn this paper, we investigate the acoustic features that can be modified to control the perceptual age of a singing voice. Singers can sing expressively by controlling prosody and vocal timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome this limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, it is still difficult to intu-itively control singing voice characteristics by manipulating pa-rameters corresponding to specific physical traits, such as gen-der and age. In this paper, we focus on controlling the perceived age of the singer and, as a first step, perform an investigation of the factors that play a part in the listener’s perception of the singer’s age. The experimental results demonstrate that 1) the perceptual age of singing voices corresponds relatively well to the actual age of the singer, 2) speech analysis/synthesis pro-cessing and statistical voice conversion processing don’t cause adverse effects on the perceptual age of singing voices, and 3) prosodic features have a larger effect on the perceptual age than spectral features. Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2013 | Grapheme-to-phoneme conversion based on adaptive regularization of weight vectorsabstractThe current state-of-the-art approach in grapheme-to-phoneme (g2p) conversion is structured learning based on the Margin Infused Relaxed Algorithm (MIRA), which is an online discriminative training method for multiclass classification. However, it is known that the aggressive weight update method of MIRA is prone to overfitting, even if the current example is an outlier or noisy. Adaptive Regularization of Weight Vectors (AROW) has been proposed to resolve this problem for binary classification. In addition, AROW’s update rule is simpler and more efficient than that of MIRA, allowing for more efficient training. Although AROW has these advantages, it has not been applied to g2p conversion yet. In this paper, we first apply AROW to g2p conversion which is structured learning problem. In an evaluation that employed a dataset including noisy data our proposed approach achieves a 5.3% error reduction rate compared to MIRA implemented in DirecTL+ in terms of phoneme error rate while requiring only 78% the training time. Index Terms:g2p conversion, out-of-vocabulary word, online discriminative training, structured learning, AROW Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2013 | A digital signal processor implementation of silent/electrolaryngeal speech enhancement based on real-time statistical voice conversionabstractIn this paper, we present a digital signal processor (DSP) implementation of real-time statistical voice conversion (VC) for silent speech enhancement and electrolaryngeal speech enhancement. As a silent speech interface, we focus on nonaudible murmur (NAM), which can be used in situations where audible speech is not acceptable. Electrolaryngeal speech is one of the typical types of alaryngeal speech produced by an alternative speaking method for laryngectomees. However, the sound quality of NAM and electrolaryngeal speech suffers from lack of naturalness. VC has proven to be one of the promising approaches to address this problem, and it has been successfully implemented on devices with sufficient computational resources. An implementation on devices that are highly portable but have limited computational resources would greatly contribute to its practical use. In this paper we further implement real-time VC on a DSP. To implement the two speech enhancement systems based on real-time VC, one from NAM to a whispered voice and the other from electrolaryngeal speech to a natural voice, we propose several methods for reducing computational cost while preserving conversion accuracy. We conduct experimental evaluations and show that real-time VC is capable of running on a DSP with little degradation. Index Terms: statistical voice conversion, real-time processing, reduction of computational cost, DSP, non-audible murmur, electrolaryngeal speech Takuto Moriguchi, Tomoki Toda, Motoaki Sano, Hiroshi Sato 0002, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | An empirical comparison of joint optimization techniques for speech translationabstractSpeech translation (ST) systems consist of three major components: automatic speech recognition (ASR), machine translation (MT), and speech synthesis (SS). In general the ASR system is tuned independently to minimize word error rate (WER), but previous research has shown that ASR and MT can be jointly optimized to improve translation quality [1]. Independently, many techniques have recently been proposed for the optimization of MT, such as empirical comparison of joint optimization using minimum error rate training (MERT) [2], pairwise ranking optimization (PRO) [3] and the batch margin infused relaxed algorithm (MIRA) [4]. The first contribution of this paper is an empirical comparison of these techniques in the context of joint optimization. As the last two methods are able to use sparse features, we also introduce lexicalized features using the frequencies of recognized words. In addition, motivated by initial results, we propose a hybrid optimization method that changes the translation evaluation measure depending on the features to be optimized. Experimental results for the best combination of algorithm and features show a gain of 1.3 BLEU points at 27% of the computational cost of previous joint optimization methods. Index Terms: speech translation, machine translation, automatic speech recognition, joint optimization Masaya Ohgushi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2013 | Efficient speech transcription through respeakingabstractWe propose a method for efficient off-line speech transcription through respeaking. Speech is segmented into smaller utterances using an initial automatic transcript. Respeaking is performed segment by segment, while confidence filtering helps save supervision effort. We conduct detailed experiments comparing speaking vs. typing, sequential vs. confidence-ordered supervision, and examine the effect of the respeaking word error rate on correction efficiency. Our results demonstrate that the proposed method can not only outperform typing in terms of correction efficiency, but is also much less demanding for the respeakers than traditional respeaking methods, consequently helping to keep costs down. Matthias Sperber, Graham Neubig, Christian Fügen, Satoshi Nakamura 0001, Alex Waibel |
INTERSPEECH | 2 |
| 2013 | Improvements to HMM-based speech synthesis based on parameter generation with rich context modelsabstractIn this paper, we improve parameter generation with rich context models by modifying an initialization method and further apply it to both spectral and F0 components in HMM-based speech synthesis. To alleviate over-smoothing effects caused by the traditional parameter generation methods, we have previously proposed an iterative parameter generation method with rich context models. It has been reported that this method yields quality improvements in synthetic speech but there are still limitations. This is because 1) this generation method still suffers from the over-smoothing effect, as it uses the parameters generated by the traditional method as an initial parameters, which strongly affect on the finally generated parameters and 2) it is applied to only the spectral component. To address these issues, we propose 1) an initialization method to generate less smoothed but more discontinuous initial parameters that tend to yield better generated parameters, and 2) a parameter generation method with rich context models for the F0 component. Experimental results show that the proposed methods yield significant improvements in quality of synthetic speech. Index Terms: HMM-based speech synthesis, rich context models, GMM, context clustering, over-smoothing, MSD-HMM Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | A hybrid approach to electrolaryngeal speech enhancement based on spectral subtraction and statistical voice conversionabstractWe present a hybrid approach to improving naturalness of electrolaryngeal (EL) speech while minimizing degradation in intelligibility. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, previous work has proposed methods for EL speech enhancement through either noise reduction or voice conversion. The former usually causes no degradation in intelligibility but yields only small improvements in naturalness as the mechanical excitation sounds remain essentially unchanged. On the other hand, the latter method significantly improves naturalness of EL speech using spectral and excitation parameters of natural voices converted from acoustic parameters of EL speech, but it usually causes degradation in intelligibility owing to errors in conversion. We propose a hybrid method using the noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. The experimental results demonstrate the proposed method yields significant improvements in naturalness compared with EL speech while keeping intelligibility high enough. Index Terms: speaking-aid, electrolaryngeal speech, spectral subtraction, voice conversion, hybrid approach Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2013 | Substring-based machine translation
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara |
Mach. Transl. | 1 |
| 2012 | Machine Translation without Words through Substring Alignment
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara |
ACL (1) | 1 |
| 2012 | Inducing a Discriminative Parser to Optimize Machine Translation Reordering
Graham Neubig, Taro Watanabe, Shinsuke Mori |
EMNLP-CoNLL | 1 |
| 2012 | A monotonic statistical machine translation approach to speaking style transformation
Graham Neubig, Yuya Akita, Shinsuke Mori, Tatsuya Kawahara |
Comput. Speech Lang. | 1 |
| 2011 | An Unsupervised Model for Joint Phrase Alignment and Extraction
Graham Neubig, Taro Watanabe, Eiichiro Sumita, Shinsuke Mori, Tatsuya Kawahara |
ACL | 1 |
| 2011 | Training Dependency Parsers from Partially Annotated Corpora
Daniel Flannery, Yusuke Miyao, Graham Neubig, Shinsuke Mori |
IJCNLP | 3 |
| 2011 | Safety Information Mining - What can NLP do in a disaster -
Graham Neubig, Yuichiroh Matsubayashi, Masato Hagiwara, Koji Murakami |
IJCNLP | 1 |
| 2011 | A Pointwise Approach to Pronunciation Estimation for a TTS Front-EndabstractIn this paper, we propose a pointwise approach to the Japanese TTS front-end. In this approach, phoneme sequence estimation of sentences is decomposed into two tasks: word segmentation of the input sentence and phoneme estimation of each word. Then these two tasks are solved by pointwise classifiers without referring to the neighboring classification results. In contrast to existing sequence-based methods, an n-gram model based on sequences of word-phoneme pairs for example, this framework enables us to use various language resources such as sentences in which only a few words are annotated, or an unsegmented list of compound words, among others. In the experiments, we compared a joint tri-gram model with the combination of a pointwise word segmenter and a pointwise phoneme sequence estimator. The results showed that our framework successfully enables a TTS front-end to refer to a partially annotated corpus and/or a word sequence list annotated with phoneme sequences to realize a far larger improvement in accuracy. Shinsuke Mori, Graham Neubig |
INTERSPEECH | 2 |
| 2011 | Searching Translation Memories for Paraphrases
Masao Utiyama, Graham Neubig, Takashi Onishi, Eiichiro Sumita |
MTSummit | 2 |
| 2010 | Improved statistical models for SMT-based speaking style transformationabstractAutomatic speech recognition (ASR) results contain not only ASR errors, but also disfluencies and colloquial expressions that must be corrected to create readable transcripts. We take the approach of statistical machine translation (SMT) to “translate” from ASR results into transcript-style text. We introduce two novel modeling techniques in this framework: a context-dependent translation model, which allows for usage of context to accurately model translation probabilities, and log-linear interpolation of conditional and joint probabilities, which allows for frequently observed translation patterns to be given higher priority. The system is implemented using weighted finite state transducers (WFST). On an evaluation using ASR results and manual transcripts of meetings of the Japanese Diet (national congress), the proposed methods showed a significant increase in accuracy over traditional modeling techniques. Graham Neubig, Yuya Akita, Shinsuke Mori, Tatsuya Kawahara |
ICASSP | 1 |
| 2010 | Semi-automated update of automatic transcription system for the Japanese national congressabstractUpdate of acoustic and language models is vital to maintain performance of automatic speech recognition (ASR) systems. To alleviate efforts for updating models, we propose a “semi-automated ” framework for the ASR system of the Japanese National Congress. The framework consists of our speaking-style transformation (SST) and lightly-supervised training (LSV) approaches, which can automatically generate spoken-style training texts and labels from documents like meeting minutes. An experimental evaluation demonstrated that this update framework improved the ASR performance for the latest meeting data. We also address an estimation method of the ASR accuracy based on SST, which uses minutes as reference texts and does not require verbatim transcripts. Index Terms: Spontaneous speech recognition, congressional speech, lightly-supervised training, speaking-style transformation 1. Yuya Akita, Masato Mimura, Graham Neubig, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2010 | Learning a language model from continuous speechabstractThis paper presents a new approach to language model construction, learning a language model not from text, but directly from continuous speech. A phoneme lattice is created using acoustic model scores, and Bayesian techniques are used to robustly learn a language model from this noisy input. A novel sampling technique is devised that allows for the integrated learning of word boundaries and an n-gram language model with no prior linguistic knowledge. The proposed techniques were used to learn a language model directly from continuous, potentially large-vocabulary speech. This language model was able to significantly reduce the ASR phoneme error rate over a separate set of test data, and the proposed lattice processing and lexical acquisition techniques were found to be important factors in this improvement. Index Terms: language acquisition, word segmentation, Pitman-Yor language model, Bayesian learning Graham Neubig, Masato Mimura, Shinsuke Mori, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2010 | Word-based Partial Annotation for Efficient Corpus Construction
Graham Neubig, Shinsuke Mori |
LREC | 1 |
| 2009 | A WFST-based log-linear framework for speaking-style transformationabstractWhen attempting to make transcripts from automatic speech recognition results, disfluency deletion, transformation of colloquial expressions, and insertion of dropped words must be performed to ensure that the final product is clean transcriptstyle text. This paper introduces a system for the automatic transformation of the spoken word to transcript-style language that enables not only deletion of disfluencies, but also substitutions of colloquial expressions and insertion of dropped words. A number of potentially useful features are combined in a loglinear probabilistic framework, and the utility of each is examined. The system is implemented using weighted finite state transducers (WFSTs) to allow for easy combination of features and integration with other WFST-based systems. On evaluation, the best system achieved a 5.37 % word error rate, a 5.49% absolute gain over a rule-based baseline and a 1.54 % absolute gain over a simple noisy-channel model. Index Terms: speaking style transformation, disfluency detection, weighted finite state transducers, log-linear model Graham Neubig, Shinsuke Mori, Tatsuya Kawahara |
INTERSPEECH | 1 |