Taro Watanabe

dblp:50/4741 · DBLP profile ↗
← Back
92ranked-venue papers
10as first author
49since 2021 · last 2026
0000-0001-8349-3522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 90 · 10 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms
abstract
Yuto Nishida, Naoki Shikoda, Yosuke Kishinami, Ryo Fujii, Makoto Morishita, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuto Nishida, Naoki Shikoda, Yosuke Kishinami, Ryo Fujii, Makoto Morishita, Hidetaka Kamigaito, Taro Watanabe
ACL (1)7
2026 HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
abstract
Recently, we have often observed hallucinated citations or references that do not correspond to any existing work in papers under review, preprints, or published papers.Such hallucinated citations pose a serious concern to scientific reliability.When they appear in accepted papers, they may also negatively affect the credibility of conferences.In this study, we refer to hallucinated citations as "HalluCitation" and systematically investigate their prevalence and impact.We analyze all papers published at ACL, NAACL, and EMNLP in 2024 and 2025, including main conference, Findings, and workshop papers.Our analysis reveals that over 300 papers contain at least one HalluCitation, most of which were published in 2025.Notably, half of these papers were identified at EMNLP 2025, the most recent conference, indicating that this issue is rapidly increasing.Moreover, more than 100 such papers were accepted as main conference and Findings papers at EMNLP 2025, affecting the credibility.
Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
ACL (1)3
2026 An AI-Assisted Co-planning System for Early English Reading Practice
Justin Vasselli, Adam Nohejl, Taro Watanabe
AIED3
2026 A Benchmark Corpus for the Diagnostic Assessment of Content in L2 English Speech
Kosuke Doi, Justin Vasselli, Taro Watanabe
LREC3
2026 A Large-Scale Dataset for Linking-Based Geocoding
Hibiki Nakatani, Yuichiro Yasui, Ryosuke Wakamoto, Masayuki Ishii, Tetsuhisa Suizu, Hiroki Ouchi, Taro Watanabe
LREC7
2026 Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
abstract
Abstract Automatic evaluation in grammatical error correction (GEC) is crucial for selecting the best-performing systems. Currently, reference-based metrics are a popular choice, which basically measure the similarity between hypothesis and reference sentences. However, similarity measures based on embeddings, such as BERTScore, are often ineffective, since many words in the source sentences remain unchanged in both the hypothesis and the reference. This study focuses on edits specifically designed for GEC, i.e., ERRANT, and computes similarity measured over the edits from the source sentence. To this end, we propose edit vector, a representation for an edit, and introduce a new metric, UOT-ERRANT, which transports these edit vectors from hypothesis to reference using unbalanced optimal transport. Experiments with SEEDA meta-evaluation show that UOT-ERRANT improves evaluation performance, particularly in the +Fluency domain where many edits occur. Moreover, our method is highly interpretable because the transport plan can be interpreted as a soft edit alignment, making UOT-ERRANT a useful metric for both system ranking and analyzing GEC systems. Our code is available from https://github.com/gotutiyan/uot-errant.
Takumi Goto, Yusuke Sakai 0010, Taro Watanabe
Trans. Assoc. Comput. Linguistics3
2025 Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine Translation
abstract
Zhi Qu, Yiran Wang, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhi Qu 0001, Yiran Wang 0006, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe
ACL (1)7
2025 Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
abstract
In generative commonsense reasoning tasks such as CommonGen, generative large language models (LLMs) compose sentences that include all given concepts.However, when focusing on instruction-following capabilities, if a prompt specifies a concept order, LLMs must generate sentences that adhere to the specified order.To address this, we propose Ordered CommonGen, a benchmark designed to evaluate the compositional generalization and instruction-following abilities of LLMs.This benchmark measures ordered coverage to assess whether concepts are generated in the specified order, enabling a simultaneous evaluation of both abilities.We conducted a comprehensive analysis using 36 LLMs and found that, while LLMs generally understand the intent of instructions, biases toward specific concept order patterns often lead to low-diversity outputs or identical results even when the concept order is altered.Moreover, even the most instructioncompliant LLM achieved only about 75% ordered coverage, highlighting the need for improvements in both instruction-following and compositional generalization capabilities. Concepts Coverage (↑) Similarlity (↓)Diversity (↑) Perplexity (↓) w/o order w/ order Ordered Rate pBLEU pBLEURT Distinct Diverse Rate
Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
ACL (1)3
2025 CoAM: Corpus of All-Type Multiword Expressions
abstract
Yusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman, Justin Vasselli, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman, Justin Vasselli, Hidetaka Kamigaito, Taro Watanabe
ACL (1)7
2025 Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
abstract
Hidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hidetaka Kamigaito, Hiroyuki Deguchi 0002, Yusuke Sakai 0010, Katsuhiko Hayashi 0001, Taro Watanabe
ACL (1)5
2025 Graph-Structured Trajectory Extraction from Travelogues
abstract
Aitaro Yamamoto, Hiroyuki Otomo, Hiroki Ouchi, Shohei Higashiyama, Hiroki Teranishi, Hiroyuki Shindo, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Aitaro Yamamoto, Hiroyuki Otomo, Hiroki Ouchi, Shohei Higashiyama, Hiroki Teranishi, Hiroyuki Shindo, Taro Watanabe
ACL (1)7
2025 HLU: Human Vs LLM Generated Text Detection Dataset for Urdu at Multiple Granularities
abstract
The rise of large language models (LLMs) generating human-like text has raised concerns about misuse, especially in low-resource languages like Urdu. To address this gap, we introduce the HLU dataset, which consists of three datasets: Document, Paragraph, and Sentence level. The document-level dataset contains 1,014 instances of human-written and LLM-generated articles across 13 domains, while the paragraph and sentence-level datasets each contain 667 instances. We conducted both human and automatic evaluations. In the human evaluation, the average accuracy at the document level was 35%, while at the paragraph and sentence levels, accuracies were 75.68% and 88.45%, respectively. For automatic evaluation, we finetuned the XLMRoBERTa model for both monolingual and multilingual settings achieving consistent results in both. Additionally, we assessed the performance of GPT4 and Claude3Opus using zero-shot prompting. Our experiments and evaluations indicate that distinguishing between human and machine-generated text is challenging for both humans and LLMs, marking a significant step in addressing this issue in Urdu.
Iqra Ali, Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe
COLING4
2025 IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
abstract
Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the same image, depending on the context, has not been sufficiently explored. To address this, we propose IRR: Image Review Rank, a novel evaluation framework designed to assess critic review texts from multiple perspectives. IRR evaluates LVLMs by measuring how closely their judgments align with human interpretations. We validate it using a dataset of images from 15 categories, each with five critic review texts and annotated rankings in both English and Japanese, totaling over 2,000 data instances. Our results indicate that, although LVLMs exhibited consistent performance across languages, their correlation with human annotations was insufficient, highlighting the need for further advancements. These findings highlight the limitations of current evaluation methods and the need for approaches that better capture human reasoning in Vision & Language tasks.
Kazuki Hayashi, Kazuma Onishi, Toma Suzuki, Yusuke Ide, Seiji Gobara, Shigeki Saito, Yusuke Sakai 0010, Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Taro Watanabe
COLING10
2025 A Text Embedding Model with Contrastive Example Mining for Point-of-Interest Geocoding
abstract
Geocoding is a fundamental technique that links location mentions to their geographic positions, which is important for understanding texts in terms of where the described events occurred. Unlike most geocoding studies that targeted coarse-grained locations, we focus on geocoding at a fine-grained point-of-interest (POI) level. To address the challenge of finding appropriate geo-database entries from among many candidates with similar POI names, we develop a text embedding-based geocoding model and investigate (1) entry encoding representations and (2) hard negative mining approaches suitable for enhancing the model’s disambiguation ability. Our experiments show that the second factor significantly impact the geocoding accuracy of the model.
Hibiki Nakatani, Hiroki Teranishi, Shohei Higashiyama, Yuya Sawada, Hiroki Ouchi, Taro Watanabe
COLING6
2025 Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?
abstract
Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words even in the era of large language models (LLMs). Frequency in film subtitles has proved to be a particularly good approximation of everyday language exposure. For many languages, however, film subtitles are not easily available, or are overwhelmingly translated from English. We demonstrate that frequencies extracted from carefully processed YouTube subtitles provide an approximation comparable to, and often better than, the best currently available resources. Moreover, they are available for languages for which a high-quality subtitle or speech corpus does not exist. We use YouTube subtitles to construct frequency norms for five diverse languages, Chinese, English, Indonesian, Japanese, and Spanish, and evaluate their correlation with lexical decision time, word familiarity, and lexical complexity. In addition to being strongly correlated with two psycholinguistic variables, a simple linear regression on the new frequencies achieves a new high score on a lexical complexity prediction task in English and Japanese, surpassing both models trained on film subtitle frequencies and the LLM GPT-4. We publicly release our code, the frequency lists, fastText word embeddings, and statistical language models.
Adam Nohejl, Frederikus Hudi, Eunike Andriani Kardinata, Shintaro Ozaki, Maria Angelica Riera Machin, Justin Vasselli, Taro Watanabe
COLING8
2025 Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
abstract
Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluating the robustness of reference-free dialogue metrics against four categories of adversarial attacks: speaker tag prefixes, static responses, ungrammatical responses, and repeated conversational context. We analyze metrics such as DialogRPT, UniEval, and PromptEval—a prompt-based method leveraging LLMs—across grounded and ungrounded datasets. By examining both their correlation with human judgment and susceptibility to adversarial attacks, we find that these two axes are not always aligned; metrics that appear to be equivalent when judged by traditional benchmarks may, in fact, vary in their scores of adversarial responses. These findings motivate the development of nuanced evaluation frameworks to address real-world dialogue challenges.
Justin Vasselli, Adam Nohejl, Taro Watanabe
COLING3
2025 SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
abstract
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai 0010, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe
EMNLP9
2025 LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructions
abstract
Hongyu Sun, Yusuke Sakai, Haruki Sakajo, Shintaro Ozaki, Kazuki Hayashi, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yusuke Sakai 0010, Haruki Sakajo, Shintaro Ozaki, Kazuki Hayashi, Hidetaka Kamigaito, Taro Watanabe
EMNLP7
2025 Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
abstract
Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that reduce their naturalness and cultural appropriateness.This work proposes Dialogue Act Script (DAS), a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.Rather than translating dialogue utterances directly, DAS enables the generation of new dialogues in the target language that are culturally and contextually appropriate.By using structured dialogue act representations, DAS supports flexible localization across languages, mitigating translationese and enabling more fluent, naturalistic conversations.Human evaluations across Italian, German, and Chinese show that DASgenerated dialogues consistently outperform those produced by both machine and human translators on measures of cultural relevance, coherence, and situational appropriateness. 1
Justin Vasselli, Eunike Andriani Kardinata, Yusuke Sakai 0010, Taro Watanabe
EMNLP4
2025 J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
abstract
We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support three critical perception tasks, object identification, reference resolution, and next-action prediction, by leveraging a comprehensive template of attributes (e.g., category, color, shape, size, material, and spatial relations). Extensive evaluations with both proprietary and open-source Vision Language Models (VLMs) reveal that incorporating detailed object attributes substantially improves multimodal perception performance compared to without object attributes. Despite the improvement, we find that there still exists a gap between proprietary and open-source VLMs. In addition, our analysis of object affordances demonstrates varying abilities in understanding object functionality and contextual relationships across different VLMs. These findings underscore the importance of rich, context-sensitive attribute annotations in advancing robot perception in dynamic environments. Code and data available at https://github.com/jatuhurrra/J-ORA.
Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino
IROS3
2025 Languages Transferred Within the Encoder: On Representation Transfer in Zero-Shot Multilingual Translation
abstract
Understanding representation transfer in multilingual neural machine translation (MNMT) can reveal the reason for the zero-shot translation deficiency. In this work, we systematically analyze the representational issue of MNMT models. We first introduce the identity pair, translating a sentence to itself, to address the lack of the base measure in multilingual investigations, as the identity pair can reflect the representation of a language within the model. Then, we demonstrate that the encoder transfers the source language to the representational subspace of the target language instead of the language-agnostic state. Thus, the zero-shot translation deficiency arises because the representation of a translation is entangled with other languages and not transferred to the target language effectively. Based on our findings, we propose two methods: 1) low-rank language-specific embedding at the encoder, and 2) language-specific contrastive learning of the representation at the decoder. The experimental results on Europarl-15, TED-19, and OPUS-100 datasets show that our methods substantially enhance the performance of zero-shot translations without sacrifices in supervised directions by improving language transfer capacity, thereby providing practical evidence to support our conclusions. Codes are available at https://github.com/zhiqu22/ZeroTrans.
Zhi Qu 0001, Chenchen Ding, Taro Watanabe
MTSummit (1)3
2025 How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
abstract
Yusuke Ide, Yuto Nishida, Justin Vasselli, Miyu Oba, Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yusuke Ide, Yuto Nishida, Justin Vasselli, Miyu Oba, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
NAACL (Long Papers)7
2025 Tonguescape: Exploring Language Models Understanding of Vowel Articulation
abstract
Haruki Sakajo, Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Haruki Sakajo, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
NAACL (Long Papers)4
2025 WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
abstract
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Stephanie Yulia Salim, Yi Zhou 0019, Yinxuan Gui, David Ifeoluwa Adelani, Annie En-Shiun Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Wijaya, Alice Oh, Chong-Wah Ngo
NAACL (Long Papers)48
2025 AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
abstract
Peinan Zhang, Yusuke Sakai, Masato Mita, Hiroki Ouchi, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Peinan Zhang, Yusuke Sakai 0010, Masato Mita, Hiroki Ouchi, Taro Watanabe
NAACL (Long Papers)5
2024 Monolingual Paraphrase Detection Corpus for Low Resource Pashto Language at Sentence Level
abstract
Paraphrase detection is a task to identify if two sentences are semantically similar or not. It plays an important role in maintaining the integrity of written work such as plagiarism detection and text reuse detection. Formerly, researchers focused on developing large corpora for English. However, no research has been conducted on sentence-level paraphrase detection in low-resource Pashto language. To bridge this gap, we introduce the first fully manually annotated Pashto sentential paraphrase detection corpus collected from authentic cases in journalism covering 10 different domains, including Sports, Health, Environment, and more. Our proposed corpus contains 6,727 sentences, encompassing 3,687 paraphrased and 3,040 non-paraphrased. Experimental findings reveal that our proposed corpus is sufficient to train XLM-RoBERTa to accurately detect paraphrased sentence pairs in Pashto with an F1 score of 84%. To compare our corpus with those in other languages, we also applied our fine-tuned model to the Indonesian and English paraphrase datasets in a zero-shot manner, achieving F1 scores of 82% and 78%, respectively. This result indicates that the quality of our corpus is not less than commonly used datasets. It‘s a pioneering contribution to the field. We will publicize a subset of 1,800 instances from our corpus, free from any licensing issues.
Iqra Ali, Hidetaka Kamigaito, Taro Watanabe
LREC/COLING3
2024 Disentangling Pretrained Representation to Leverage Low-Resource Languages in Multilingual Machine Translation
abstract
Multilingual neural machine translation aims to encapsulate multiple languages into a single model. However, it requires an enormous dataset, leaving the low-resource language (LRL) underdeveloped. As LRLs may benefit from shared knowledge of multilingual representation, we aspire to find effective ways to integrate unseen languages in a pre-trained model. Nevertheless, the intricacy of shared representation among languages hinders its full utilisation. To resolve this problem, we employed target language prediction and a central language-aware layer to improve representation in integrating LRLs. Focusing on improving LRLs in the linguistically diverse country of Indonesia, we evaluated five languages using a parallel corpus of 1,000 instances each, with experimental results measured by BLEU showing zero-shot improvement of 7.4 from the baseline score of 7.1 to a score of 15.5 at best. Further analysis showed that the gains in performance are attributed more to the disentanglement of multilingual representation in the encoder with the shift of the target language-specific representation in the decoder.
Frederikus Hudi, Zhi Qu 0001, Hidetaka Kamigaito, Taro Watanabe
LREC/COLING4
2024 Constructing Indonesian-English Travelogue Dataset
abstract
Research in low-resource language is often hampered due to the under-representation of how the language is being used in reality. This is particularly true for Indonesian language because there is a limited variety of textual datasets, and majority were acquired from official sources with formal writing style. All the more for the task of geoparsing, which could be implemented for navigation and travel planning applications, such datasets are rare, even in the high-resource languages, such as English. Being aware of the need for a new resource in both languages for this specific task, we constructed a new dataset comprising both Indonesian and English from personal travelogue articles. Our dataset consists of 88 articles, exactly half of them written in each language. We covered both named and nominal expressions of four entity types related to travel: location, facility, transportation, and line. We also conducted experiments by training classifiers to recognise named entities and their nominal expressions. The results of our experiments showed a promising future use of our dataset as we obtained F1-score above 0.9 for both languages.
Eunike Andriani Kardinata, Hiroki Ouchi, Taro Watanabe
LREC/COLING3
2024 JDocQA: Japanese Document Question Answering Dataset for Generative Language Models
abstract
Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common in our society. This is known as a quite challenging task because it requires not only text understanding but also understanding of figures and tables, and hence visual question answering (VQA) methods are often examined in addition to textual approaches. We introduce Japanese Document Question Answering (JDocQA), a large-scale document-based QA dataset, essentially requiring both visual and textual information to answer questions, which comprises 5,504 documents in PDF format and annotated 11,600 question-and-answer instances in Japanese. Each QA instance includes references to the document pages and bounding boxes for the answer clues. We incorporate multiple categories of questions and unanswerable questions from the document for realistic question-answering applications. We empirically evaluate the effectiveness of our dataset with text-based large language models (LLMs) and multimodal models. Incorporating unanswerable questions in finetuning may contribute to harnessing the so-called hallucination generation.
Eri Onami, Shuhei Kurita, Taiki Miyanishi, Taro Watanabe
LREC/COLING4
2024 Detector-Corrector: Edit-Based Automatic Post Editing for Human Post Editing
abstract
Post-editing is crucial in the real world because neural machine translation (NMT) sometimes makes errors.Automatic post-editing (APE) attempts to correct the outputs of an MT model for better translation quality.However, many APE models are based on sequence generation, and thus their decisions are harder to interpret for actual users.In this paper, we propose “detector–corrector”, an edit-based post-editing model, which breaks the editing process into two steps, error detection and error correction.The detector model tags each MT output token whether it should be corrected and/or reordered while the corrector model generates corrected words for the spans identified as errors by the detector.Experiments on the WMT’20 English–German and English–Chinese APE tasks showed that our detector–corrector improved the translation edit rate (TER) compared to the previous edit-based model and a black-box sequence-to-sequence APE model, in addition, our model is more explainable because it is based on edit operations.
Hiroyuki Deguchi 0002, Masaaki Nagata, Taro Watanabe
EAMT (1)3
2024 Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
abstract
In Simultaneous Machine Translation (SiMT), training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems.However, constructing such a corpus is challenging due to high costs, and limitations in annotator capabilities, and as a result, existing SI corpora are limited.Therefore, we propose a method to convert existing speech translation (ST) corpora into interpretation-style corpora, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus).We demonstrated that fine-tuning SiMT models using the LLM-SI-Corpus reduces latencies while achieving better quality compared to models fine-tuned with other corpora in both speechto-text and text-to-text settings.
Yusuke Sakai 0010, Mana Makinae, Hidetaka Kamigaito, Taro Watanabe
EMNLP4
2024 Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation models support fine-tuning hundreds of languages simultaneously.However, fine-tuning on full parameters solely is inefficient potentially leading to negative interactions among languages.In this work, we demonstrate that the fine-tuning for a language occurs in its intrinsic languagespecific subspace with a tiny fraction of entire parameters.Thus, we propose languagespecific LoRA to isolate intrinsic languagespecific subspaces.Furthermore, we propose architecture learning techniques and introduce a gradual pruning schedule during fine-tuning to exhaustively explore the optimal setting and the minimal intrinsic subspaces for each language, resulting in a lightweight yet effective fine-tuning procedure.The experimental results on a 12-language subset and a 30language subset of FLORES-101 show that our methods not only outperform full-parameter fine-tuning up to 2.25 spBLEU scores but also reduce trainable parameters to 0.4% for high and medium-resource languages and 1.6% for low-resource ones.Codes are available at https://github.com/Spike0924/LSLo.
Zhe Cao 0002, Zhi Qu 0001, Hidetaka Kamigaito, Taro Watanabe
EMNLP4
2024 Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
abstract
Scaling the context size of large language models (LLMs) enables them to perform various new tasks, e.g., book summarization.However, the memory cost of the Key and Value (KV) cache in attention significantly limits the practical applications of LLMs.Recent works have explored token pruning for KV cache reduction in LLMs, relying solely on attention scores as a token importance indicator.However, our investigation into value vector norms revealed a notably non-uniform pattern questioning their reliance only on attention scores.Inspired by this, we propose a new method: Value-Aware Token Pruning (VATP) which uses both attention scores and the ℓ 1 norm of value vectors to evaluate token importance.Extensive experiments on LLaMA2-7B-chat and Vicuna-v1.5-7Bacross 16 LongBench tasks demonstrate that VATP outperforms attention-score-only baselines in over 12 tasks, confirming the effectiveness of incorporating value vector norms into token importance evaluation of LLMs. 1
Zhiyu Guo, Hidetaka Kamigaito, Taro Watanabe
EMNLP3
2024 Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
abstract
This work investigates the impact of data augmentation on confidence calibration and uncertainty estimation in Named Entity Recognition (NER) tasks.For the future advance of NER in safety-critical fields like healthcare and finance, it is essential to achieve accurate predictions with calibrated confidence when applying Deep Neural Networks (DNNs), including Pretrained Language Models (PLMs), as a realworld application.However, DNNs are prone to miscalibration, which limits their applicability.Moreover, existing methods for calibration and uncertainty estimation are computational expensive.Our investigation in NER found that data augmentation improves calibration and uncertainty in cross-genre and cross-lingual setting, especially in-domain setting.Furthermore, we showed that the calibration for NER tends to be more effective when the perplexity of the sentences generated by data augmentation is lower, and that increasing the size of the augmentation further improves calibration and uncertainty.
Wataru Hashimoto 0002, Hidetaka Kamigaito, Taro Watanabe
EMNLP3
2024 Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language Model
abstract
Simultaneous Speech Translation (SiST) begins translating before the entire source input is received, making it crucial to balance quality and latency.In real interpreting situations, interpreters manage this simultaneity by breaking sentences into smaller segments and translating them while maintaining the source order as much as possible.SiST could benefit from this approach to balance quality and latency.However, current corpora used for simultaneous tasks often involve significant word reordering in translation, which is not ideal given that interpreters faithfully follow source syntax as much as possible.Inspired by conference interpreting by humans utilizing the salami technique, we introduce the Simul-MuST-C 1 , a dataset created by leveraging the Large Language Model (LLM), specifically GPT-4o, which aligns the target text as closely as possible to the source text by using minimal chunks that contain enough information to be interpreted.Experiments on three language pairs show that the effectiveness of segmentedbase monotonicity in training data varies with the grammatical distance between the source and the target, with grammatically distant language pairs benefiting the most in achieving quality while minimizing latency.
Mana Makinae, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
EMNLP4
2024 Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
abstract
What kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability?Recent language models still have much room for improvement in their data efficiency compared to humans.This paper investigates whether language models efficiently use indirect data (indirect evidence), from which they infer sentence acceptability.In contrast, humans use indirect evidence efficiently, which is considered one of the inductive biases contributing to efficient language acquisition.To explore this question, we introduce the Wug In-Direct Evidence Test (WIDET), a dataset consisting of training instances inserted into the pre-training data and evaluation instances.We inject synthetic instances with newly coined wug words into pretraining data and explore the model's behavior on evaluation data that assesses grammatical acceptability regarding those words.We prepare the injected instances by varying their levels of indirectness and quantity.Our experiments surprisingly show that language models do not induce grammatical knowledge even after repeated exposure to instances with the same structure but differing only in lexical items from evaluation instances in certain language phenomena.Our findings suggest a potential direction for future research: developing models that use latent indirect evidence to induce grammatical knowledge.
Miyu Oba, Yohei Oseki, Akiyo Fukatsu, Akari Haga, Hiroki Ouchi, Taro Watanabe, Saku Sugawara
EMNLP6
2024 Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
abstract
Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yusuke Sakai 0010, Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Taro Watanabe
NAACL-HLT4
2024 Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?
Seiji Gobara, Hidetaka Kamigaito, Taro Watanabe
PACLIC3
2024 Context-Aware Machine Translation with Source Coreference Explanation
abstract
Abstract Despite significant improvements in enhancing the quality of translation, context-aware machine translation (MT) models underperform in many cases. One of the main reasons is that they fail to utilize the correct features from context when the context is too long or their models are overly complex. This can lead to the explain-away effect, wherein the models only consider features easier to explain predictions, resulting in inaccurate translations. To address this issue, we propose a model that explains the decisions made for translation by predicting coreference features in the input. We construct a model for input coreference by exploiting contextual features from both the input and translation output representations on top of an existing MT model. We evaluate and analyze our method in the WMT document-level translation task of English-German dataset, the English-Russian dataset, and the multilingual TED talk dataset, demonstrating an improvement of over 1.0 BLEU score when compared with other context-aware models.
Huy-Hien Vu, Hidetaka Kamigaito, Taro Watanabe
Trans. Assoc. Comput. Linguistics3
2023 Subset Retrieval Nearest Neighbor Machine Translation
abstract
Hiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita
ACL (1)2
2023 Model-based Subsampling for Knowledge Graph Completion
abstract
Xincan Feng, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xincan Feng, Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Taro Watanabe
IJCNLP (1)4
2023 24-bit Languages
abstract
Yiran Wang, Taro Watanabe, Masao Utiyama, Yuji Matsumoto. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yiran Wang 0006, Taro Watanabe, Masao Utiyama, Yuji Matsumoto 0001
IJCNLP (1)2
2023 Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective
abstract
There are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning.
Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su
NeurIPS7
2023 Switching to Discriminative Image Captioning by Relieving a Bottleneck of Reinforcement Learning
abstract
Discriminativeness is a desirable feature of image captions: captions should describe the characteristic details of input images. However, recent high-performing captioning models, which are trained with reinforcement learning (RL), tend to generate overly generic captions despite their high performance in various other criteria. First, we investigate the cause of the unexpectedly low discriminativeness and show that RL has a deeply rooted side effect of limiting the output words to high-frequency words. The limited vocabulary is a severe bottleneck for discriminativeness as it is difficult for a model to describe the details beyond its vocabulary. Then, based on this identification of the bottleneck, we drastically recast discriminative image captioning as a much simpler task of encouraging low-frequency word generation. Hinted by long-tail classification and debiasing methods, we propose methods that easily switch off-the-shelf RL models to discriminativeness-aware models with only a single-epoch fine-tuning on the part of the parameters. Extensive experiments demonstrate that our methods significantly enhance the discriminative-ness of off-the-shelf RL models and even outperform previous discriminativeness-aware methods with much smaller computational costs. Detailed analysis and human evaluation also verify that our methods boost the discriminativeness without sacrificing the overall quality of captions.1
Ukyo Honda, Taro Watanabe, Yuji Matsumoto 0001
WACV2
2022 Adapting to Non-Centered Languages for Zero-shot Multilingual Translation
abstract
Multilingual neural machine translation can translate unseen language pairs during training, i.e. zero-shot translation. However, the zero-shot translation is always unstable. Although prior works attributed the instability to the domination of central language, e.g. English, we supplement this viewpoint with the strict dependence of non-centered languages. In this work, we propose a simple, lightweight yet effective language-specific modeling method by adapting to non-centered languages and combining the shared information and the language-specific information to counteract the instability of zero-shot translation. Experiments with Transformer on IWSLT17, Europarl, TED talks, and OPUS-100 datasets show that our method not only performs better than strong baselines in centered data conditions but also can easily fit non-centered data conditions. By further investigating the layer attribution, we show that our proposed method can disentangle the coupled representation in the correct direction.
Zhi Qu 0001, Taro Watanabe
COLING2
2022 Law Retrieval with Supervised Contrastive Learning Using the Hierarchical Structure of Law
Jungmin Choi, Ukyo Honda, Taro Watanabe, Hiroki Ouchi, Kentaro Inui
PACLIC3
2021 Nested Named Entity Recognition via Explicitly Excluding the Influence of the Best Path
abstract
Yiran Wang, Hiroyuki Shindo, Yuji Matsumoto, Taro Watanabe. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yiran Wang 0006, Hiroyuki Shindo, Yuji Matsumoto 0001, Taro Watanabe
ACL/IJCNLP (1)4
2021 Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image Captioning
abstract
Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto, Taro Watanabe, Yuji Matsumoto. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Taro Watanabe, Yuji Matsumoto 0001
EACL4
2021 User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization
abstract
Shohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita
NAACL-HLT3
2020 Coordination Boundary Identification without Labeled Data for Compound Terms Disambiguation
abstract
Yuya Sawada, Takashi Wada, Takayoshi Shibahara, Hiroki Teranishi, Shuhei Kondo, Hiroyuki Shindo, Taro Watanabe, Yuji Matsumoto. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Yuya Sawada, Takashi Wada 0001, Takayoshi Shibahara, Hiroki Teranishi, Shuhei Kondo, Hiroyuki Shindo, Taro Watanabe, Yuji Matsumoto 0001
COLING7
2020 Overhead Work Assist with Passive Gravity Compensation Mechanism and Horizontal Link Mechanism for Agriculture
abstract
Busy agricultural seasons involve long-term continuous work and heavy labor. Particularly during overhead work, such as harvesting, gibberellin treatment, and bagging, workers need to consistently raise upper limb weights of approximately 2 to 4 kg with their own muscular strength, resulting in a high work burden. For long-duration work in the field, a passive and robust assist system is advantageous. Therefore, we propose an assistance device named TasKi that uses self-weight compensation mechanisms and horizontal link mechanisms to reduce the burden on a worker's upper limbs during overhead work. TasKi can compensate for upper limb weight by using the force of a spring in various postures of the upper limbs without battery support. In this report, we describe the design of the TasKi mechanisms that achieve the upward work assist in actual agriculture with a simple structure. The mechanism of self-weight compensation and the degree of freedom and parameters of the link mechanism are studied.
Yasuyuki Yamada, Hirokazu Arakawa, Taro Watanabe, Shunya Fukuyama, Rie Nishihama, Isao Kikutani, Taro Nakamura 0001
RO-MAN3
2016 Phrase-based Machine Translation using Multiple Preordering Candidates
abstract
In this paper, we propose a new decoding method for phrase-based statistical machine translation which directly uses multiple preordering candidates as a graph structure. Compared with previous phrase-based decoding methods, our method is based on a simple left-to-right dynamic programming in which no decoding-time reordering is performed. As a result, its runtime is very fast and implementing the algorithm becomes easy. Our system does not depend on specific preordering methods as long as they output multiple preordering candidates, and it is trivial to employ existing preordering methods into our system. In our experiments for translating diverse 11 languages into English, the proposed method outperforms conventional phrase-based decoder in terms of translation qualities under comparable or faster decoding time.
Yusuke Oda, Taku Kudo, Tetsuji Nakagawa, Taro Watanabe
COLING4
2016 Optimization for Statistical Machine Translation: A Survey
abstract
In statistical machine translation (SMT), the optimization of the system parameters to maximize translation accuracy is now a fundamental part of virtually all modern systems. In this article, we survey 12 years of research on optimization for SMT, from the seminal work on discriminative models (Och and Ney 2002) and minimum error rate training (Och 2003), to the most recent advances. Starting with a brief introduction to the fundamentals of SMT systems, we follow by covering a wide variety of optimization algorithms for use in both batch and online optimization. Specifically, we discuss losses based on direct error minimization, maximum likelihood, maximum margin, risk minimization, ranking, and more, along with the appropriate methods for minimizing these losses. We also cover recent topics, including large-scale optimization, nonlinear models, domain-dependent optimization, and the effect of MT evaluation measures or search on optimization. Finally, we discuss the current state of affairs in MT optimization, and point out some unresolved problems that will likely be the target of further research in optimization for MT.
Graham Neubig, Taro Watanabe
Comput. Linguistics2
2015 Transition-based Neural Constituent Parsing
abstract
Taro Watanabe, Eiichiro Sumita. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Taro Watanabe, Eiichiro Sumita
ACL (1)1
2015 Hierarchical Back-off Modeling of Hiero Grammar based on Non-parametric Bayesian Model
abstract
In hierarchical phrase-based machine translation, a rule table is automatically learned by heuristically extracting syn-chronous rules from a parallel corpus. As a result, spuriously many rules are extracted which may be composed of various incorrect rules. The larger rule table incurs more run time for decoding and may result in lower translation quality. To resolve the problems, we propose a hierarchical back-off model for Hiero grammar, an instance of a synchronous context free grammar (SCFG), on the basis of the hierarchical Pitman-Yor process. The model can extract a compact rule and phrase table without resorting to any heuristics by hierarchically backing off to smaller phrases under SCFG. Inference is efficiently carried out using two-step synchronous parsing of Xiao et al., (2012) combined with slice sampling. In our experiments, the proposed model achieved higher or at least comparable translation quality against a previous Bayesian model on various language pairs; German/French/Spanish/Japanese-English. When compared against heuristic models, our model achieved comparable translation quality on a full size German-English language pair in Europarl v7 corpus with significantly smaller grammar size; less than 10 % of that for heuristic model. 1
Hidetaka Kamigaito, Taro Watanabe, Hiroya Takamura, Manabu Okumura, Eiichiro Sumita
EMNLP2
2015 Leave-one-out Word Alignment without Garbage Collector Effects
abstract
Expectation-maximization algorithms, such as those implemented in GIZA++ pervade the field of unsupervised word alignment.However, these algorithms have a problem of over-fitting, leading to "garbage collector effects," where rare words tend to be erroneously aligned to untranslated words.This paper proposes a leave-one-out expectationmaximization algorithm for unsupervised word alignment to address this problem.The proposed method excludes information derived from the alignment of a sentence pair from the alignment models used to align it.This prevents erroneous alignments within a sentence pair from supporting themselves.Experimental results on Chinese-English and Japanese-English corpora show that the F 1 , precision and recall of alignment were consistently increased by 5.0% -17.2%, and BLEU scores of end-to-end translation were raised by 0.03 -1.30.The proposed method also outperformed l 0 -normalized GIZA++ and Kneser-Ney smoothed GIZA++.
Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Taro Watanabe, Eiichiro Sumita
EMNLP4
2014 Recurrent Neural Networks for Word Alignment Model
abstract
This study proposes a word alignment model based on a recurrent neural network (RNN), in which an unlimited alignment history is represented by recurrently connected hidden layers.We perform unsupervised learning using noise-contrastive estimation (Gutmann and Hyvärinen, 2010;Mnih and Teh, 2012), which utilizes artificially generated negative samples.Our alignment model is directional, similar to the generative IBM models (Brown et al., 1993).To overcome this limitation, we encourage agreement between the two directional models by introducing a penalty function that ensures word embedding consistency across two directional models during training.The RNN-based model outperforms the feed-forward neural network-based model (Yang et al., 2013) as well as the IBM Model 4 under Japanese-English and French-English word alignment tasks, and achieves comparable translation performance to those baselines for Japanese-English and Chinese-English translation tasks.
Akihiro Tamura, Taro Watanabe, Eiichiro Sumita
ACL (1)2
2014 Recurrent Neural Network-based Tuple Sequence Model for Machine Translation
Youzheng Wu, Taro Watanabe, Chiori Hori
COLING2
2014 Unsupervised Word Alignment Using Frequency Constraint in Posterior Regularized EM
abstract
Generative word alignment models, such as IBM Models, are restricted to oneto-many alignment, and cannot explicitly represent many-to-many relationships in a bilingual text.The problem is partially solved either by introducing heuristics or by agreement constraints such that two directional word alignments agree with each other.In this paper, we focus on the posterior regularization framework (Ganchev et al., 2010) that can force two directional word alignment models to agree with each other during training, and propose new constraints that can take into account the difference between function words and content words.Experimental results on French-to-English and Japanese-to-English alignment tasks show statistically significant gains over the previous posterior regularization baseline.We also observed gains in Japanese-to-English translation tasks, which prove the effectiveness of our methods under grammatically different language pairs.
Hidetaka Kamigaito, Taro Watanabe, Hiroya Takamura, Manabu Okumura
EMNLP2
2014 Syntax-Augmented Machine Translation using Syntax-Label Clustering
abstract
Recently, syntactic information has helped significantly to improve statistical ma-chine translation. However, the use of syn-tactic information may have a negative im-pact on the speed of translation because of the large number of rules, especially when syntax labels are projected from a parser in syntax-augmented machine translation. In this paper, we propose a syntax-label clus-tering method that uses an exchange algo-rithm in which syntax labels are clustered together to reduce the number of rules. The proposed method achieves clustering by directly maximizing the likelihood of synchronous rules, whereas previous work considered only the similarity of proba-bilistic distributions of labels. We tested the proposed method on Japanese-English and Chinese-English translation tasks and found order-of-magnitude higher cluster-ing speeds for reducing labels and gains in translation quality compared with pre-vious clustering method. 1
Hideya Mino, Taro Watanabe, Eiichiro Sumita
EMNLP2
2014 Discriminative Training for Log-Linear Based SMT: Global or Local Methods
abstract
In statistical machine translation, the standard methods such as MERT tune a single weight with regard to a given development data. However, these methods suffer from two problems due to the diversity and uneven distribution of source sentences. First, their performance is highly dependent on the choice of a development set, which may lead to an unstable performance for testing. Second, the sentence level translation quality is not assured since tuning is performed on the document level rather than on sentence level. In contrast with the standard global training in which a single weight is learned, we propose novel local training methods to address these two problems. We perform training and testing in one step by locally learning the sentence-wise weight for each input sentence. Since the time of each tuning step is unnegligible and learning sentence-wise weights for the entire test set means many passes of tuning, it is a great challenge for the efficiency of local training. We propose an efficient two-phase method to put the local training into practice by employing the ultraconservative update. On NIST Chinese-to-English translation tasks with both medium and large scales of training data, our local training methods significantly outperform standard methods with the maximal improvements up to 2.0 BLEU points, meanwhile their efficiency is comparable to that of the standard methods.
Lemao Liu, Tiejun Zhao, Taro Watanabe, Hailong Cao, Conghui Zhu
ACM Trans. Asian Lang. Inf. Process.3
2013 Additive Neural Networks for Statistical Machine Translation
Lemao Liu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)2
2013 Part-of-Speech Induction in Dependency Trees for Statistical Machine Translation
Akihiro Tamura, Taro Watanabe, Eiichiro Sumita, Hiroya Takamura, Manabu Okumura
ACL (1)2
2013 Hierarchical Phrase Table Combination for Machine Translation
Conghui Zhu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)2
2013 Tuning SMT with a Large Number of Features via Online Feature Grouping
Lemao Liu, Tiejun Zhao, Taro Watanabe, Eiichiro Sumita
IJCNLP3
2013 Substring-based machine translation
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara
Mach. Transl.2
2012 Head-driven Transition-based Parsing with Top-down Prediction
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
ACL (1)2
2012 Machine Translation without Words through Substring Alignment
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara
ACL (1)2
2012 Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu
EMNLP-CoNLL3
2012 Inducing a Discriminative Parser to Optimize Machine Translation Reordering
Graham Neubig, Taro Watanabe, Shinsuke Mori
EMNLP-CoNLL2
2012 Bilingual Lexicon Extraction from Comparable Corpora Using Label Propagation
Akihiro Tamura, Taro Watanabe, Eiichiro Sumita
EMNLP-CoNLL2
2012 Optimized Online Rank Learning for Machine Translation
Taro Watanabe
HLT-NAACL1
2011 An Unsupervised Model for Joint Phrase Alignment and Extraction
Graham Neubig, Taro Watanabe, Eiichiro Sumita, Shinsuke Mori, Tatsuya Kawahara
ACL2
2011 Machine Translation System Combination by Confusion Forest
Taro Watanabe, Eiichiro Sumita
ACL1
2011 Third-order Variational Reranking on Packed-Shared Dependency Forests
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
EMNLP2
2007 Online Large-Margin Training for Statistical Machine Translation
Taro Watanabe, Jun Suzuki 0001, Hajime Tsukada, Hideki Isozaki
EMNLP-CoNLL1
2006 Left-to-Right Target Generation for Hierarchical Phrase-Based Translation
abstract
We present a hierarchical phrase-based statistical machine translation in which a target sentence is efficiently generated in left-to-right order. The model is a class of synchronous-CFG with a Greibach Normal Form-like structure for the projected production rule: The paired target-side of a production rule takes a phrase prefixed form. The decoder for the target-normalized form is based on an Early-style top down parser on the source side. The target-normalized form coupled with our top down parser implies a left-to-right generation of translations which enables us a straightforward integration with ngram language models. Our model was experimented on a Japanese-to-English newswire translation task, and showed statistically significant performance improvements against a phrase-based translation system.
Taro Watanabe, Hajime Tsukada, Hideki Isozaki
ACL1
2005 Empirical Study of Utilizing Morph-Syntactic Information in SMT
Young-Sook Hwang, Taro Watanabe, Yutaka Sasaki
IJCNLP2
2004 Example-based Machine Translation Based on Syntactic Transfer with Statistical Models
Kenji Imamura, Hideo Okuma, Taro Watanabe, Eiichiro Sumita
COLING3
2004 Reordering Constraints for Phrase-Based Statistical Machine Translation
Richard Zens, Hermann Ney, Taro Watanabe, Eiichiro Sumita
COLING3
2004 A Unified Approach in Speech-to-Speech Translation: Integrating Features of Speech recognition and Machine Translation
Ruiqiang Zhang, Gen-ichiro Kikui, Hirofumi Yamamoto, Frank K. Soong, Taro Watanabe, Wai Kit Lo
COLING5
2004 Improved spoken language translation using n-best speech recognition hypotheses
abstract
We intended to demonstrate the effect of using N-best speech recognition hypotheses for improving speech translation performance. A log-linear model, which integrated features from speech recognition and statistical machine translation, was used to rescore the translation candidates. Model parameters were estimated by optimizing an objectively measurable but subjectively relevant translation quality metric. Experimental results have shown that the proposed N-best approach improved translation quality over the conventional single-best approach. The improvements were confirmed consistently by several automatic translation evaluation metrics. 1.
Ruiqiang Zhang, Gen-ichiro Kikui, Hirofumi Yamamoto, Frank K. Soong, Taro Watanabe, Eiichiro Sumita, Wai Kit Lo
INTERSPEECH5
2003 Chunk-Based Statistical Translation
abstract
This paper describes an alternative translation model based on a text chunk under the framework of statistical machine translation. The translation model suggested here first performs chunking. Then, each word in a chunk is translated. Finally, translated chunks are reordered. Under this scenario of translation modeling, we have experimented on a broad-coverage Japanese-English traveling corpus and achieved improved performance.
Taro Watanabe, Eiichiro Sumita, Hiroshi G. Okuno
ACL1
2003 A corpus-centered approach to spoken language translation
Eiichiro Sumita, Yasuhiro Akiba, Takao Doi, Andrew M. Finch, Kenji Imamura, Michael Paul, Mitsuo Shimohata, Taro Watanabe
EACL8
2003 Example-based decoding for statistical machine translation
abstract
This paper presents a decoder for statistical machine translation that can take advantage of the example-based machine translation framework. The decoder presented here is based on the greedy approach to the decoding problem, but the search is initiated from a similar translation extracted from a bilingual corpus. The experiments on multilingual translations showed that the proposed method was far superior to a word-by-word generation beam search algorithm.
Taro Watanabe, Eiichiro Sumita
MTSummit1
2002 Using Language and Translation Models to Select the Best among Outputs from Multiple MT Systems
Yasuhiro Akiba, Taro Watanabe, Eiichiro Sumita
COLING2
2002 Language Model Adaptation with Additional Text Generated by Machine Translation
Hideharu Nakajima, Hirofumi Yamamoto, Taro Watanabe
COLING3
2002 Bidirectional Decoding for Statistical Machine Translation
Taro Watanabe, Eiichiro Sumita
COLING1
2002 Statistical machine translation decoder based on phrase
abstract
This paper describes a decoding algorithm for statistical machine translation based on phrases. In the past, the solution to the decoding problem were inspired from that of speech recognizers, translating each input word into one or more output words generating in left-to-right direction. The algorithm presented here iteratively constructs phrases or chunks of cepts until all the input words are consumed. This behavior resulted in computational complexity higher than those with left-to-right constraints, though the translation accuracy is better from the Japanese-to-English translation experiments. 1.
Taro Watanabe, Eiichiro Sumita
INTERSPEECH1
2002 Statistical Machine Translation on Paraphrased Corpora
Taro Watanabe, Mitsuo Shimohata, Eiichiro Sumita
LREC1
2000 Lessons Learned from a Task-based Evaluation of Speech-to-Speech Machine Translation
Lori S. Levin, Boris Bartlog, Ariadna Font Llitjós, Donna Gates, Alon Lavie, Dorcas Wallace, Taro Watanabe, Monika Woszczyna
LREC7
1994 A cooperative man-machine dialogue model for problem solving
Masahiro Araki, Taro Watanabe, Felix C. M. Quimbo, Shuji Doshita
ICSLP2