VLDB 2026 Research / reviewers in the wild / expert
Eiichiro Sumita
dblp:95/5465
· DBLP profile ↗
188ranked-venue papers
11as first author
19since 2021 · last 2023
0000-0002-1028-4399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 186 · 11 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Language Model Pre-training on True NegativesabstractDiscriminative pre-trained language models (PrLMs) learn to predict original texts from intentionally corrupted ones. Taking the former text as positive and the latter as negative samples, the PrLM can be trained effectively for contextualized representation. However, the training of such a type of PrLMs highly relies on the quality of the automatically constructed samples. Existing PrLMs simply treat all corrupted texts as equal negative without any examination, which actually lets the resulting model inevitably suffer from the false negative issue where training is carried out on pseudo-negative data and leads to less efficiency and less robustness in the resulting PrLMs. In this work, on the basis of defining the false negative issue in discriminative PrLMs that has been ignored for a long time, we design enhanced pre-training methods to counteract false negative predictions and encourage pre-training language models on true negatives by correcting the harmful gradient updates subject to false negative predictions. Experimental results on GLUE and SQuAD benchmarks show that our counter-false-negative pre-training methods indeed bring about better performance together with stronger robustness. Zhuosheng Zhang 0001, Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita |
AAAI | 4 |
| 2023 | Subset Retrieval Nearest Neighbor Machine TranslationabstractHiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita |
ACL (1) | 6 |
| 2023 | Pivot Translation for Zero-resource Language Pairs Based on a Multilingual Pretrained ModelabstractA multilingual translation model enables a single model to handle multiple languages. However, the translation qualities of unlearned language pairs (i.e., zero-shot translation qualities) are still poor. By contrast, pivot translation translates source texts into target ones via a pivot language such as English, thus enabling machine translation without parallel texts between the source and target languages. In this paper, we perform pivot translation using a multilingual model and compare it with direct translation. We improve the translation quality without using parallel texts of direct translation by fine-tuning the model with machine-translated pseudo-translations. We also discuss what type of parallel texts are suitable for effectively improving the translation quality in multilingual pivot translation. Kenji Imamura, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 3 |
| 2023 | Universal Multimodal Representation for Language UnderstandingabstractRepresentation learning is the foundation of natural language processing (NLP). This work presents new methods to employ visual information as assistant signals to general NLP tasks. For each sentence, we first retrieve a flexible number of images either from a light topic-image lookup table extracted over the existing sentence-image pairs or a shared cross-modal embedding space that is pre-trained on out-of-shelf text-image pairs. Then, the text and images are encoded by a Transformer encoder and convolutional neural network, respectively. The two sequences of representations are further fused by an attention layer for the interaction of the two modalities. In this study, the retrieval process is controllable and flexible. The universal visual representation overcomes the lack of large-scale bilingual sentence-image pairs. Our method can be easily applied to text-only tasks without manually annotated multimodal parallel corpora. We apply the proposed method to a wide range of natural language generation and understanding tasks, including neural machine translation, natural language inference, and semantic similarity. Experimental results show that our method is generally effective for different tasks and languages. Analysis indicates that the visual signals enrich textual representations of content words, provide fine-grained grounding information about the relationship between concepts and events, and potentially conduce to disambiguation. Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Low-resource Multilingual Neural Translation Using Linguistic Feature-based Relevance MechanismsabstractThis article investigates approaches to effectively harness source-side linguistic features for low-resource multilingual neural machine translation (MNMT). Previous works focus on using various features of a word such as lemma, part-of-speech tag, dependency label, and so on, to improve translation quality in a low-resource scenario. However, these studies deal with bilingual translation and do not focus on using features in multilingual training setups. Our work focuses on this particular point and experiments with low-resource multilingual models incorporating source-side linguistic features. Although techniques for integrating features into an NMT model such as concatenation and feature relevance perform quite well in bilingual settings, they do not work well in multilingual settings. To remedy this, we propose the use of dummy features and language indicator features in MNMT models. Experiments are conducted on English to Asian language translation on a multilingual, multi-parallel corpus spanning English and eight Asian languages where for each language pair, the training data size does not exceed 20,000 parallel sentences. After establishing strong bilingual baselines using feature relevance mechanisms and multilingual baselines without any features, we show that our proposed dummy features and language indicator features, in combination with feature relevance mechanisms, yield significant improvements in BLEU points for all language pairs. We then analyze our models from the perspectives of model sizes, the impact of individual linguistic features, validation perplexity computed during training, visualization of the activations of the relevance mechanisms, and exhaustive tuning of hyperparameters. We also report preliminary results for multilingual multi-way models using linguistic features. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | SelfSeg: A Self-supervised Sub-word Segmentation Method for Neural Machine TranslationabstractSub-word segmentation is an essential pre-processing step for Neural Machine Translation (NMT). Existing work has shown that neural sub-word segmenters are better than Byte-Pair Encoding (BPE), however, they are inefficient, as they require parallel corpora, days to train, and hours to decode. This article introduces SelfSeg, a self-supervised neural sub-word segmentation method that is much faster to train/decode and requires only monolingual dictionaries instead of parallel corpora. SelfSeg takes as input a word in the form of a partially masked character sequence, optimizes the word generation probability, and generates the segmentation with the maximum posterior probability, which is calculated using a dynamic programming algorithm. The training time of SelfSeg depends on word frequencies, and we explore several word frequency normalization strategies to accelerate the training phase. Additionally, we propose a regularization mechanism that allows the segmenter to generate various segmentations for one word. To show the effectiveness of our approach, we conduct MT experiments in low-, middle-, and high-resource scenarios, where we compare the performance of using different segmentation methods. The experimental results demonstrate that, on the low-resource ALT dataset, our method achieves more than 1.2 BLEU score improvement compared with BPE and SentencePiece, and a 1.1 score improvement over Dynamic Programming Encoding (DPE) and Vocabulary Learning via Optimal Transport (VOLT), on average. The regularization method achieves approximately a 4.3 BLEU score improvement over BPE and a 1.2 BLEU score improvement over BPE-dropout, the regularized version of BPE. We also observed significant improvements on IWSLT15 Vi→En, WMT16 Ro→En, and WMT15 Fi→En datasets and competitive results on the WMT14 De→En and WMT14 Fr→En datasets. Furthermore, our method is 17.8× faster during training and up to 36.8× faster during decoding in a high-resource scenario compared to DPE. We provide extensive analysis, including why monolingual word-level data is enough to train SelfSeg. Haiyue Song, Raj Dabre, Chenhui Chu, Sadao Kurohashi, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | FeatureBART: Feature Based Sequence-to-Sequence Pre-Training for Low-Resource NMTabstractIn this paper we present FeatureBART, a linguistically motivated sequence-to-sequence monolingual pre-training strategy in which syntactic features such as lemma, part-of-speech and dependency labels are incorporated into the span prediction based pre-training framework (BART). These automatically extracted features are incorporated via approaches such as concatenation and relevance mechanisms, among which the latter is known to be better than the former. When used for low-resource NMT as a downstream task, we show that these feature based models give large improvements in bilingual settings and modest ones in multilingual settings over their counterparts that do not use features. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Eiichiro Sumita |
COLING | 6 |
| 2022 | Effective Graph Context Representation for Document-level Machine TranslationabstractDocument-level neural machine translation (DocNMT) universally encodes several local sentences or the entire document. Thus, DocNMT does not consider the relevance of document-level contextual information, for example, some context (i.e., content words, logical order, and co-occurrence relation) is more effective than another auxiliary context (i.e., functional and auxiliary words). To address this issue, we first utilize the word frequency information to recognize content words in the input document, and then use heuristical relations to summarize content words and sentences as a graph structure without relying on external syntactic knowledge. Furthermore, we apply graph attention networks to this graph structure to learn its feature representation, which allows DocNMT to more effectively capture the document-level context. Experimental results on several widely-used document-level benchmarks demonstrated the effectiveness of the proposed approach. Kehai Chen, Muyun Yang, Masao Utiyama, Eiichiro Sumita, Rui Wang 0015, Min Zhang 0005 |
IJCAI | 4 |
| 2022 | Explicit Alignment Learning for Neural Machine TranslationabstractEven though neural machine translation (NMT) has become the state-of-the-art solution for end-to-end translation, it still suffers from a lack of translation interpretability, which may be conveniently enhanced by explicit alignment learning (EAL), as performed in traditional statistical machine translation (SMT). To provide the benefits of both NMT and SMT, this paper presents a novel model design that enhances NMT with an additional training process for EAL, in addition to the end-to-end translation training. Thus, we propose two approaches an explicit alignment learning approach, in which we further remove the need for the additional alignment model, and perform embedding mixup with the alignment based on encoder--decoder attention weights in the NMT model. We conducted experiments on both small-scale (IWSLT14 De->En and IWSLT13 Fr->En) and large-scale (WMT14 En->De, En->Fr, WMT17 Zh->En) benchmarks. Evaluation results show that our EAL methods significantly outperformed strong baseline methods, which shows the effectiveness of EAL. Further explorations show that the translation improvements are due to a better spatial alignment of the source and target language embeddings. Our method improves translation performance without the need to increase model parameters and training data, which verifies that the idea of incorporating techniques of SMT into NMT is worthwhile. Zuchao Li, Hai Zhao 0001, Fengshun Xiao, Masao Utiyama, Eiichiro Sumita |
IJCAI | 5 |
| 2022 | Text Compression-Aided Transformer EncodingabstractText encoding is one of the most important steps in Natural Language Processing (NLP). It has been done well by the self-attention mechanism in the current state-of-the-art Transformer encoder, which has brought about significant improvements in the performance of many NLP tasks. Though the Transformer encoder may effectively capture general information in its resulting representations, the backbone information, meaning the gist of the input text, is not specifically focused on. In this paper, we propose explicit and implicit text compression approaches to enhance the Transformer encoding and evaluate models using this approach on several typical downstream tasks that rely on the encoding heavily. Our explicit text compression approaches use dedicated models to compress text, while our implicit text compression approach simply adds an additional module to the main model to handle text compression. We propose three ways of integration, namely backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the backbone information into Transformer-based models for various downstream tasks. Our evaluation on benchmark datasets shows that the proposed explicit and implicit text compression approaches improve results in comparison to strong baselines. We therefore conclude, when comparing the encodings to the baseline models, text compression helps the encoders to learn better language representations. Zuchao Li, Zhuosheng Zhang 0001, Hai Zhao 0001, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Integrating Prior Translation Knowledge Into Neural Machine TranslationabstractNeural machine translation (NMT), which is an encoder-decoder joint neural language model with an attention mechanism, has achieved impressive results on various machine translation tasks in the past several years. However, the language model attribute of NMT tends to produce fluent yet sometimes unfaithful translations, which hinders the improvement of translation capacity. In response to this problem, we propose a simple and efficient method to integrate prior translation knowledge into NMT in a universal manner that is compatible with neural networks. Meanwhile, it enables NMT to consider the crossing language translation knowledge from the source-side of the training pipeline of NMT, thereby making full use of the prior translation knowledge to enhance the performance of NMT. The experimental results on two large-scale benchmark translation tasks demonstrated that our approach achieved a significant improvement over a strong baseline. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Smoothing Dialogue States for Open Conversational Machine ReadingabstractConversational machine reading (CMR) requires machines to communicate with humans through multi-turn interactions between two salient dialogue states of decision making and question generation processes.In open CMR settings, as the more realistic scenario, the retrieved background knowledge would be noisy, which results in severe challenges in the information transmission.Existing studies commonly train independent or pipeline systems for the two subtasks.However, those methods are trivial by using hard-label decisions to activate question generation, which eventually hinders the model performance.In this work, we propose an effective gating strategy by smoothing the two dialogue states in only one decoder and bridge decision making and question generation to provide a richer dialogue state reference.Experiments on the OR-ShARC dataset show the effectiveness of our method, which achieves new state-of-the-art results. Zhuosheng Zhang 0001, Siru Ouyang, Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita |
EMNLP (1) | 5 |
| 2021 | Unsupervised Neural Machine Translation with Universal GrammarabstractMachine translation usually relies on parallel corpora to provide parallel signals for training.The advent of unsupervised machine translation has brought machine translation away from this reliance, though performance still lags behind traditional supervised machine translation.In unsupervised machine translation, the model seeks symmetric language similarities as a source of weak parallel signal to achieve translation.Chomsky's Universal Grammar theory postulates that grammar is an innate form of knowledge to humans and is governed by universal principles and constraints.Therefore, in this paper, we seek to leverage such shared grammar clues to provide more explicit language parallel signals to enhance the training of unsupervised machine translation models.Through experiments on multiple typical language pairs, we demonstrate the effectiveness of our proposed approaches. Zuchao Li, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP (1) | 3 |
| 2021 | User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical NormalizationabstractShohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita |
NAACL-HLT | 4 |
| 2021 | Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data ScenariosabstractHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
NAACL-HLT | 5 |
| 2021 | Context-aware positional representation for self-attention networks
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
Neurocomputing | 4 |
| 2021 | Towards Tokenization and Part-of-Speech Tagging for Khmer: Data and DiscussionabstractAs a highly analytic language, Khmer has considerable ambiguities in tokenization and part-of-speech (POS) tagging processing. This topic is investigated in this study. Specifically, a 20,000-sentence Khmer corpus with manual tokenization and POS-tagging annotation is released after a series of work over the last 4 years. This is the largest morphologically annotated Khmer dataset as of 2020, when this article was prepared. Based on the annotated data, experiments were conducted to establish a comprehensive benchmark on the automatic processing of tokenization and POS-tagging for Khmer. Specifically, a support vector machine, a conditional random field (CRF) , a long short-term memory (LSTM) -based recurrent neural network, and an integrated LSTM-CRF model have been investigated and discussed. As a primary conclusion, processing at morpheme-level is satisfactory for the provided data. However, it is intrinsically difficult to identify further grammatical constituents of compounds or phrases because of the complex analytic features of the language. Syntactic annotation and automatic parsing for Khmer will be scheduled in the near future. Hour Kaing, Chenchen Ding, Masao Utiyama, Eiichiro Sumita, Sam Sethserey, Sopheap Seng, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical StudyabstractUnsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines. Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2021 | Modeling Future Cost for Neural Machine TranslationabstractExisting neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline. Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Explicit Sentence Compression for Neural Machine TranslationabstractState-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT English-to-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines. Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
AAAI | 5 |
| 2020 | Content Word Aware Neural Machine TranslationabstractNeural machine translation (NMT) encodes the source sentence in a universal way to generate the target sentence word-byword.However, NMT does not consider the importance of word in the sentence meaning, for example, some words (i.e., content words) express more important meaning than others (i.e., function words).To address this limitation, we first utilize word frequency information to distinguish between content and function words in a sentence, and then design a content word-aware NMT to improve translation performance.Empirical results on the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chineseto-English translation tasks show that the proposed methods can significantly improve the performance of Transformer-based NMT. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL | 4 |
| 2020 | A Three-Parameter Rank-Frequency Relation in Natural LanguagesabstractWe present that, the rank-frequency relation in textual data follows f ∝ r -α (r + γ) -β , where f is the token frequency and r is the rank by frequency, with (α, β, γ) as parameters.The formulation is derived based on the empirical observation that d 2 (x+y)/dx 2 is a typical impulse function, where (x, y) = (log r, log f ).The formulation is the power law when β = 0 and the Zipf-Mandelbrot law when α = 0. We illustrate that α is related to the analytic features of syntax and β + γ to those of morphology in natural languages from an investigation of multilingual corpora. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACL | 3 |
| 2020 | Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL | 5 |
| 2020 | Improving Low-Resource NMT through Relevance Based Linguistic Features IncorporationabstractIn this study, linguistic knowledge at different levels are incorporated into the neural machine translation (NMT) framework to improve translation quality for language pairs with extremely limited data.Integrating manually designed or automatically extracted features into the NMT framework is known to be beneficial.However, this study emphasizes that the relevance of the features is crucial to the performance.Specifically, we propose two methods, 1) self relevance and 2) word-based relevance, to improve the representation of features for NMT.Experiments are conducted on translation tasks from English to eight Asian languages, with no more than twenty thousand sentences for training.The proposed methods improve translation quality for all tasks by up to 3.09 BLEU points.Discussions with visualization provide the explainability of the proposed methods where we show that the relevance methods provide weights to features thereby enhancing their impact on low-resource machine translation. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
COLING | 5 |
| 2020 | Bilingual Subword Segmentation for Neural Machine TranslationabstractThis paper proposed a new subword segmentation method for neural machine translation, "Bilingual Subword Segmentation," which tokenizes sentences to minimize the difference between the number of subword units in a sentence and that of its translation.While existing subword segmentation methods tokenize a sentence without considering its translation, the proposed method tokenizes a sentence by using subword units induced from bilingual sentences; this method could be more favorable to machine translation.Evaluations on WAT Asian Scientific Paper Excerpt Corpus (ASPEC) English-to-Japanese and Japanese-to-English translation tasks and WMT14 English-to-German and German-to-English translation tasks show that our bilingual subword segmentation improves the performance of Transformer neural machine translation (up to +0.81 BLEU). Hiroyuki Deguchi 0002, Masao Utiyama, Akihiro Tamura, Takashi Ninomiya, Eiichiro Sumita |
COLING | 5 |
| 2020 | Intermediate Self-supervised Learning for Machine Translation Quality EstimationabstractPre-training sentence encoders is effective in many natural language processing tasks including machine translation (MT) quality estimation (QE), due partly to the scarcity of annotated QE data required for supervised learning.In this paper, we investigate the use of an intermediate self-supervised learning task for sentence encoder aiming at improving QE performances at the sentence and word levels.Our approach is motivated by a problem inherent to QE: mistakes in translation caused by wrongly inserted and deleted tokens.We modify the translation language model (TLM) training objective of the cross-lingual language model (XLM) to orientate the pretrained model towards the target task.The proposed method does not rely on annotated data and is complementary to QE methods involving pre-trained sentence encoders and domain adaptation.Experiments on English-to-German and English-to-Russian translation directions show that intermediate learning improves over domain adaptated models.Additionally, our method reaches results in par with state-of-the-art QE models without requiring the combination of several approaches and outperforms similar methods based on pre-trained sentence encoders. Raphaël Rubino, Eiichiro Sumita |
COLING | 2 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 6 |
| 2020 | Neural Machine Translation with Universal Visual Representation
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
ICLR | 5 |
| 2020 | Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ICLR | 5 |
| 2020 | A Myanmar (Burmese)-English Named Entity Transliteration DictionaryabstractTransliteration is generally a phonetically based transcription across different writing systems. It is a crucial task for various downstream natural language processing applications. For the Myanmar (Burmese) language, robust automatic transliteration for borrowed English words is a challenging task because of the complex Myanmar writing system and the lack of data. In this study, we constructed a Myanmar-English named entity dictionary containing more than eighty thousand transliteration instances. The data have been released under a CC BY-NC-SA license. We evaluated the automatic transliteration performance using statistical and neural network-based approaches based on the prepared data. The neural network model outperformed the statistical model significantly in terms of the BLEU score on the character level. Different units used in the Myanmar script for processing were also compared and discussed. Aye Myat Mon, Chenchen Ding, Hour Kaing, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
LREC | 6 |
| 2020 | Agreement on Target-Bidirectional Recurrent Neural Networks for Sequence-to-Sequence LearningabstractRecurrent neural networks are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a shortcoming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus performance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional RNNs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of either sequence level or non-sequence level metrics. Extensive experiments were performed on three standard sequence-to-sequence transduction tasks: machine transliteration, grapheme-to-phoneme transformation and machine translation. The results show that the proposed approach achieves consistent and substantial improvements, compared to many state-of-the-art systems. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
J. Artif. Intell. Res. | 4 |
| 2020 | Extremely low-resource neural machine translation for Asian languagesabstractAbstract This paper presents a set of effective approaches to handle extremely low-resource language pairs for self-attention based neural machine translation (NMT) focusing on English and four Asian languages. Starting from an initial set of parallel sentences used to train bilingual baseline models, we introduce additional monolingual corpora and data processing techniques to improve translation quality. We describe a series of best practices and empirically validate the methods through an evaluation conducted on eight translation directions, based on state-of-the-art NMT approaches such as hyper-parameter search, data augmentation with forward and backward translation in combination with tags and noise, as well as joint multilingual training. Experiments show that the commonly used default architecture of self-attention NMT models does not reach the best results, validating previous work on the importance of hyper-parameter tuning. Additionally, empirical results indicate the amount of synthetic data required to efficiently increase the parameters of the models leading to the best translation quality measured by automatic metrics. We show that the best NMT models trained on large amount of tagged back-translations outperform three other synthetic data generation approaches. Finally, comparison with statistical machine translation (SMT) indicates that extremely low-resource NMT requires a large amount of synthetic parallel data obtained with back-translation in order to close the performance gap with the preceding SMT approach. Raphaël Rubino, Benjamin Marie, Raj Dabre, Atsushi Fujita, Masao Utiyama, Eiichiro Sumita |
Mach. Transl. | 6 |
| 2020 | Improving neural machine translation through phrase-based soft forced decoding
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 3 |
| 2020 | Towards Burmese (Myanmar) Morphological Analysis: Syllable-based Tokenization and Part-of-speech TaggingabstractThis article presents a comprehensive study on two primary tasks in Burmese (Myanmar) morphological analysis: tokenization and part-of-speech (POS) tagging. Twenty thousand Burmese sentences of newswire are annotated with two-layer tokenization and POS-tagging information, as one component of the Asian Language Treebank Project. The annotated corpus has been released under a CC BY-NC-SA license, and it is the largest open-access database of annotated Burmese when this manuscript was prepared in 2017. Detailed descriptions of the preparation, refinement, and features of the annotated corpus are provided in the first half of the article. Facilitated by the annotated corpus, experiment-based investigations are presented in the second half of the article, wherein the standard sequence-labeling approach of conditional random fields and a long short-term memory (LSTM)-based recurrent neural network (RNN) are applied and discussed. We obtained several general conclusions, covering the effect of joint tokenization and POS-tagging and importance of ensemble from the viewpoint of stabilizing the performance of LSTM-based RNN. This study provides a solid basis for further studies on Burmese processing. Chenchen Ding, Hnin Thu Zar Aye, Win Pa Pa, Khin Thandar Nwet, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2020 | A Burmese (Myanmar) Treebank: Guideline and AnalysisabstractA 20,000-sentence Burmese (Myanmar) treebank on news articles has been released under a CC BY-NC-SA license. Complete phrase structure annotation was developed for each sentence from the morphologically annotated data prepared in previous work of Ding et al. [1]. As the final result of the Burmese component in the Asian Language Treebank Project , this is the first large-scale, open-access treebank for the Burmese language. The annotation details and features of this treebank are presented. Chenchen Ding, Sann Su Su Yee, Win Pa Pa, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2020 | Towards More Diverse Input Representation for Neural Machine TranslationabstractSource input information plays a very important role in the Transformer-based translation system. In practice, word embedding and positional embedding of each word are added as the input representation. Then self-attention networks are used to encode the global dependencies in the input representation to generate a source representation. However, this processing on the source representation only adopts a single source feature and excludes richer and more diverse features such as recurrence features, local features, and syntactic features, which results in tedious representation and thereby hinders the further translation performance improvement. In this paper, we introduce a simple and efficient method to encode more diverse source features into the input representation simultaneously, and thereby learning an effective source representation by self-attention networks. In particular, the proposed grouped strategy is only applied to the input representation layer, to keep the diversity of translation information and the efficiency of the self-attention networks at the same time. Experimental results show that our approach improves the translation performance over the state-of-the-art baselines of Transformer in regard to WMT14 English-to-German and NIST Chinese-to-English machine translation tasks. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao, Muyun Yang, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Unsupervised Neural Machine Translation With Cross-Lingual Language Representation AgreementabstractUnsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Neural Machine Translation with Reordering EmbeddingsabstractThe reordering model plays an important role in phrase-based statistical machine translation.However, there are few works that exploit the reordering information in neural machine translation.In this paper, we propose a reordering mechanism to learn the reordering embedding of a word based on its contextual information.These reordering embeddings are stacked together with self-attention networks to learn sentence representation for machine translation.The reordering mechanism can be easily integrated into both the encoder and the decoder in the Transformer translation system.Experimental results on WMT'14 English-to-German, NIST Chinese-to-English, and WAT ASPEC Japanese-to-English translation tasks demonstrate that the proposed methods can significantly improve the performance of the Transformer translation system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL (1) | 4 |
| 2019 | Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine TranslationabstractUnsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 5 |
| 2019 | Sentence-Level Agreement for Neural Machine TranslationabstractThe training objective of neural machine translation (NMT) is to minimize the loss between the words in the translated sentences and those in the references. In NMT, there is a natural correspondence between the source sentence and the target sentence. However, this relationship has only been represented using the entire neural network and the training objective is computed in word-level. In this paper, we propose a sentence-level agreement module to directly minimize the difference between the representation of source and target sentence. The proposed agreement module can be integrated into NMT as an additional training objective function and can also be used to enhance the representation of the source sentences. Empirical results on the NIST Chinese-to-English and WMT English-to-German tasks show the proposed agreement module can significantly improve the NMT performance. Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Min Zhang 0005, Tiejun Zhao |
ACL (1) | 5 |
| 2019 | Recurrent Positional Embedding for Neural Machine TranslationabstractKehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Online Sentence Segmentation for Simultaneous Interpretation using Multi-Shifted Recurrent Neural Network
Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 3 |
| 2019 | NOVA: A Feasible and Flexible Annotation System for Joint Tokenization and Part-of-Speech TaggingabstractA feasible and flexible annotation system is designed for joint tokenization and part-of-speech (POS) tagging to annotate those languages without natural definitions of words . This design was motivated by the fact that word separators are not used in many highly analytic East and Southeast Asian languages. Although several of the languages are well-studied, e.g., Chinese and Japanese, many are understudied with low resources, e.g., Burmese (Myanmar) and Khmer. In the first part of the article, the proposed annotation system, named nova, is introduced. nova contains only four basic tags (n, v, a, and o); these tags can be further modified and combined to adapt complex linguistic phenomena in tokenization and POS tagging. In the second part of the article, the feasibility and flexibility of nova is illustrated from the annotation practice on Burmese and Khmer. The relation between nova and two universal POS tagsets is discussed in the final part of the article. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Neural Machine Translation With Sentence-Level Topic ContextabstractTraditional neural machine translation (NMT) methods use the word-level context to predict target language translation while neglecting the sentence-level context, which has been shown to be beneficial for translation prediction in statistical machine translation. This paper represents the sentence-level context as latent topic representations by using a convolution neural network, and designs a topic attention to integrate source sentence-level topic context information into both attention-based and Transformer-based NMT. In particular, our method can improve the performance of NMT by modeling source topics and translations jointly. Experiments on the large-scale LDC Chinese-to-English translation tasks and WMT'14 English-to-German translation tasks show that the proposed approach can achieve significant improvements compared with baseline systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Syntax-Directed Attention for Neural Machine TranslationabstractAttention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
AAAI | 4 |
| 2018 | Forest-Based Neural Machine TranslationabstractTree-based neural machine translation (NMT) approaches, although achieved impressive performance, suffer from a major drawback: they only use the 1best parse tree to direct the translation, which potentially introduces translation mistakes due to parsing errors.For statistical machine translation (SMT), forestbased methods have been proven to be effective for solving this problem, while for NMT this kind of approach has not been attempted.This paper proposes a forest-based NMT method that translates a linearized packed forest under a simple sequence-to-sequence framework (i.e., a forest-to-string NMT model).The BLEU score of the proposed method is higher than that of the string-to-string NMT, treebased NMT, and forest-based SMT systems. Chunpeng Ma, Akihiro Tamura, Masao Utiyama, Tiejun Zhao, Eiichiro Sumita |
ACL (1) | 5 |
| 2018 | Exploring Recombination for Efficient Decoding of Neural Machine TranslationabstractIn Neural Machine Translation (NMT), the decoder can capture the features of the entire prediction history with neural connections and representations.This means that partial hypotheses with different prefixes will be regarded differently no matter how similar they are.However, this might be inefficient since some partial hypotheses can contain only local differences that will not influence future predictions.In this work, we introduce recombination in NMT decoding based on the concept of the "equivalence" of partial hypotheses.Heuristically, we use a simple n-gram suffix based equivalence function and adapt it into beam search decoding.Through experiments on large-scale Chinese-to-English and English-to-Germen translation tasks, we show that the proposed method can obtain similar translation quality with a smaller beam size, making NMT decoding more efficient. Zhisong Zhang, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 4 |
| 2018 | Multilingual Parallel Corpus for Global Communication Plan
Kenji Imamura, Eiichiro Sumita |
LREC | 2 |
| 2018 | Guiding Neural Machine Translation with Retrieved Translation PiecesabstractJingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, Satoshi Nakamura. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
NAACL-HLT | 3 |
| 2018 | Graph-Based Bilingual Word Embedding for Statistical Machine TranslationabstractBilingual word embedding has been shown to be helpful for Statistical Machine Translation (SMT). However, most existing methods suffer from two obvious drawbacks. First, they only focus on simple contexts such as an entire document or a fixed-sized sliding window to build word embedding and ignore latent useful information from the selected context. Second, the word sense but not the word should be the minimal semantic unit; however, most existing methods still use word representation. To overcome these drawbacks, this article presents a novel Graph-Based Bilingual Word Embedding (GBWE) method that projects bilingual word senses into a multidimensional semantic space. First, a bilingual word co-occurrence graph is constructed using the co-occurrence and pointwise mutual information between the words. Then, maximum complete subgraphs (cliques), which play the role of a minimal unit for bilingual sense representation, are dynamically extracted according to the contextual information. Consequently, correspondence analysis, principal component analyses, and neural networks are used to summarize the clique-word matrix into lower dimensions to build the embedding model. Without contextual information, the proposed GBWE can be applied to lexical translation. In addition, given contextual information, GBWE is able to give a dynamic solution for bilingual word representations, which can be applied to phrase translation and generation. Empirical results show that GBWE can enhance the performance of lexical translation, as well as Chinese/French-to-English and Chinese-to-Japanese phrase-based SMT tasks (IWSLT, NTCIR, NIST, and WAT). Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2018 | A Neural Approach to Source Dependence Based Context Model for Statistical Machine TranslationabstractIn statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods. Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain AdaptationabstractNeural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines. Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2017 | Deterministic Attention for Sequence-to-Sequence Constituent ParsingabstractThe sequence-to-sequence model is proven to be extremely successful in constituent parsing. It relies on one key technique, the probabilistic attention mechanism, to automatically select the context for prediction. Despite its successes, the probabilistic attention model does not always select the most important context. For example, the headword and boundary words of a subtree have been shown to be critical when predicting the constituent label of the subtree, but this contextual information becomes increasingly difficult to learn as the length of the sequence increases. In this study, we proposed a deterministic attention mechanism that deterministically selects the important context and is not affected by the sequence length. We implemented two different instances of this framework. When combined with a novel bottom-up linearization method, our parser demonstrated better performance than that achieved by the sequence-to-sequence parser with probabilistic attention mechanism. Chunpeng Ma, Lemao Liu, Akihiro Tamura, Tiejun Zhao, Eiichiro Sumita |
AAAI | 5 |
| 2017 | Neural Machine Translation with Source Dependency RepresentationabstractSource dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao |
EMNLP | 6 |
| 2017 | Instance Weighting for Neural Machine Translation Domain AdaptationabstractInstance weighting has been widely applied to phrase-based machine translation domain adaptation.However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.In this paper, two instance weighting technologies, i.e., sentence weighting and domain weighting with a dynamic weight learning strategy, are proposed for NMT domain adaptation.Empirical results on the IWSLT English-German/French tasks show that the proposed methods can substantially improve NMT performance by up to 2.7-6.7 BLEU points, outperforming the existing baselines by up to 1.6-3.6BLEU points. Rui Wang 0015, Masao Utiyama, Lemao Liu, Kehai Chen, Eiichiro Sumita |
EMNLP | 5 |
| 2017 | Context-Aware Smoothing for Neural Machine TranslationabstractIn Neural Machine Translation (NMT), each word is represented as a low-dimension, real-value vector for encoding its syntax and semantic information. This means that even if the word is in a different sentence context, it is represented as the fixed vector to learn source representation. Moreover, a large number of Out-Of-Vocabulary (OOV) words, which have different syntax and semantic information, are represented as the same vector representation of “unk”. To alleviate this problem, we propose a novel context-aware smoothing method to dynamically learn a sentence-specific vector for each word (including OOV words) depending on its local context words in a sentence. The learned context-aware representation is integrated into the NMT to improve the translation performance. Empirical results on NIST Chinese-to-English translation task show that the proposed approach achieves 1.78 BLEU improvements on average over a strong attentional NMT, and outperforms some existing systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IJCNLP(1) | 4 |
| 2017 | Improving Neural Machine Translation through Phrase-based Forced DecodingabstractCompared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an existing phrase-based SMT model to compute the phrase-based decoding cost for an NMT output and then using the phrase-based decoding cost to rerank the n-best NMT outputs. The main challenge in implementing this approach is that NMT outputs may not be in the search space of the standard phrase-based decoding algorithm, because the search space of phrase-based SMT is limited by the phrase-based translation rule table. We propose a soft forced decoding algorithm, which can always successfully find a decoding path for any NMT output. We show that using the forced decoding cost to rerank the NMT outputs can successfully improve translation quality on four different language pairs. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
IJCNLP(1) | 3 |
| 2017 | Empirical Study of Dropout Scheme for Neural Machine Translation
Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 3 |
| 2017 | A Target Attention Model for Neural Machine Translation
Hideya Mino, Andrew M. Finch, Eiichiro Sumita |
MTSummit (1) | 3 |
| 2017 | Inducing a Bilingual Lexicon from Short Parallel Multiword SequencesabstractThis article proposes a technique for mining bilingual lexicons from pairs of parallel short word sequences. The technique builds a generative model from a corpus of training data consisting of such pairs. The model is a hierarchical nonparametric Bayesian model that directly induces a bilingual lexicon while training. The model learns in an unsupervised manner and is designed to exploit characteristics of the language pairs being mined. The proposed model is capable of utilizing commonly used word-pair frequency information and additionally can employ the internal character alignments within the words themselves. It is thereby capable of mining transliterations and can use reliably aligned transliteration pairs to support the mining of other words in their context. The model is also capable of performing word reordering and word deletion during the alignment process, and it is furthermore capable of operating in the absence of full segmentation information. In this work, we study two mining tasks based on English-Japanese and English-Chinese language pairs, and compare the proposed approach to baselines based on a simpler models that use only word-pair frequency information. Our results show that the proposed method is able to mine bilingual word pairs at higher levels of precision and recall than the baselines. Andrew M. Finch, Taisuke Harada, Kumiko Tanaka-Ishii, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2017 | Translation Quality Estimation Using Only Bilingual CorporaabstractIn computer-aided translation scenarios, quality estimation of machine translation hypotheses plays a critical role. Existing methods for word-level translation quality estimation (TQE) rely on the availability of manually annotated TQE training data obtained via direct annotation or postediting. However, due to the cost of human labor, such data are either limited in size or is only available for few tasks in practice. To avoid the reliance on such annotated TQE data, this paper proposes an approach to train word-level TQE models using bilingual corpora, which are typically used in machine translation training and is relatively easier to access. We formalize the training of our proposed method under the framework of maximum marginal likelihood estimation. To avoid degenerated solutions, we propose a novel regularized training objective whose optimization is achieved by an efficient approximation. Extensive experiments on both written and spoken language datasets empirically show that our approach yields comparable performance to the standard training on annotated data. Lemao Liu, Atsushi Fujita, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence LearningabstractRecurrent neural networks, particularly the long short- term memory networks, are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a fundamental short- coming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus perfor- mance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional LSTMs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of sequence-level losses. Extensive experiments were performed on two standard sequence-to-sequence trans- duction tasks: machine transliteration and grapheme-to- phoneme transformation. The results show that the proposed approach achieves consistent and substantial im- provements, compared to six state-of-the-art systems. In particular, our approach outperforms the best reported error rates by a margin (up to 9% relative gains) on the grapheme-to-phoneme task. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
AAAI | 4 |
| 2016 | Bilingual Segmented Topic ModelabstractThis study proposes the bilingual segmented topic model (BiSTM), which hierarchically models documents by treating each document as a set of segments, e.g., sections.While previous bilingual topic models, such as bilingual latent Dirichlet allocation (BiLDA) (Mimno et al., 2009;Ni et al., 2009), consider only cross-lingual alignments between entire documents, the proposed model considers cross-lingual alignments between segments in addition to document-level alignments and assigns the same topic distribution to aligned segments.This study also presents a method for simultaneously inferring latent topics and segmentation boundaries, incorporating unsupervised topic segmentation (Du et al., 2013) into BiSTM.Experimental results show that the proposed model significantly outperforms BiLDA in terms of perplexity and demonstrates improved performance in translation pair extraction (up to +0.083 extraction accuracy). Akihiro Tamura, Eiichiro Sumita |
ACL (1) | 2 |
| 2016 | A Continuous Space Rule Selection Model for Syntax-based Statistical Machine Translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
ACL (1) | 3 |
| 2016 | Neural Machine Translation with Supervised AttentionabstractThe attention mechanism is appealing for neural machine translation, since it is able to dynamically encode a source sentence by generating a alignment between a target word and source words. Unfortunately, it has been proved to be worse than conventional alignment models in alignment accuracy. In this paper, we analyze and explain this issue from the point view of reordering, and propose a supervised attention which is learned with guidance from conventional alignment models. Experiments on two Chinese-to-English translation tasks show that the supervised attention mechanism yields better alignments leading to substantial gains over the standard attention based NMT. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
COLING | 4 |
| 2016 | Connecting Phrase based Statistical Machine Translation AdaptationabstractAlthough more additional corpora are now available for Statistical Machine Translation (SMT), only the ones which belong to the same or similar domains of the original corpus can indeed enhance SMT performance directly. A series of SMT adaptation methods have been proposed to select these similar-domain data, and most of them focus on sentence selection. In comparison, phrase is a smaller and more fine grained unit for data selection, therefore we propose a straightforward and efficient connecting phrase based adaptation method, which is applied to both bilingual phrase pair and monolingual n-gram adaptation. The proposed method is evaluated on IWSLT/NIST data sets, and the results show that phrase based SMT performances are significantly improved (up to +1.6 in comparison with phrase based SMT baseline system and +0.9 in comparison with existing methods). Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
COLING | 5 |
| 2016 | Unsupervised Word Alignment by Agreement Under ITG Constraint
Hidetaka Kamigaito, Akihiro Tamura, Hiroya Takamura, Manabu Okumura, Eiichiro Sumita |
EMNLP | 5 |
| 2016 | Assessing Translation Ability through Vocabulary Ability Assessment
Yo Ehara, Yukino Baba, Masao Utiyama, Eiichiro Sumita |
IJCAI | 4 |
| 2016 | ASPEC: Asian Scientific Paper Excerpt Corpus
Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, Hitoshi Isahara |
LREC | 5 |
| 2016 | Introducing the Asian Language Treebank (ALT)
Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
LREC | 5 |
| 2016 | Agreement on Target-bidirectional Neural Machine TranslationabstractLemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
HLT-NAACL | 4 |
| 2016 | Interlocking Phrases in Phrase-based Statistical Machine TranslationabstractThis paper presents an study of the use of interlocking phrases in phrase-based statistical machine translation. We examine the effect on translation quality when the translation units used in the translation hypotheses are allowed to overlap on the source side, on the target side and on both sides. A large-scale evaluation on 380 language pairs was conducted. Our results show that overall the use of overlapping phrases improved translation quality by 0.3 BLEU points on average. Further analysis revealed that language pairs requiring a larger amount of re-ordering benefited the most from our approach. When the evaluation was restricted to such pairs, the average improvement increased to up to 0.75 BLEU points with over 97% of the pairs improving. Our approach requires only a simple modification to the decoding algorithm and we believe it should be generally applicable to improve the performance of phrase-based decoders. Ye Kyaw Thu, Andrew M. Finch, Eiichiro Sumita |
HLT-NAACL | 3 |
| 2016 | Learning local word reorderings for hierarchical phrase-based statistical machine translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 3 |
| 2016 | Word Segmentation for Burmese (Myanmar)abstractExperiments on various word segmentation approaches for the Burmese language are conducted and discussed in this note. Specifically, dictionary-based, statistical, and machine learning approaches are tested. Experimental results demonstrate that statistical and machine learning approaches perform significantly better than dictionary-based approaches. We believe that this note, based on an annotated corpus of relatively considerable size (containing approximately a half million words), is the first systematic comparison of word segmentation approaches for Burmese. This work aims to discover the properties and proper approaches to Burmese textual processing and to promote further researches on this understudied language. Chenchen Ding, Ye Kyaw Thu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2016 | Converting Continuous-Space Language Models into N-gram Language Models with Efficient Bilingual Pruning for Statistical Machine TranslationabstractThe Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N -grams based Converting (NNGC) method for transforming a Continuous-Space Language Model (CSLM) to a Back-off N -gram Language Model (BNLM). Furthermore, a Bilingual LM Pruning (BLMP) approach is developed for enhancing LMs in SMT decoding and speeding up CSLM converting. The proposed pruning and converting methods can convert a large LM efficiently by working jointly. That is, a LM can be effectively pruned before it is converted from CSLM without sacrificing performance, and further improved if an additional corpus contains out-of-domain information. For different SMT tasks, our experimental results indicate that the proposed NNGC and BLMP methods outperform the existing counterpart approaches significantly in BLEU and computational cost. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2015 | Transition-based Neural Constituent ParsingabstractTaro Watanabe, Eiichiro Sumita. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Taro Watanabe, Eiichiro Sumita |
ACL (1) | 2 |
| 2015 | Improving fast_align by Reorderingabstractfast align is a simple, fast, and efficient approach for word alignment based on the IBM model 2. fast align performs well for language pairs with relatively similar word orders; however, it does not perform well for language pairs with drastically different word orders.We propose a segmenting-reversing reordering process to solve this problem by alternately applying fast align and reordering source sentences during training.Experimental results with Japanese-English translation demonstrate that the proposed approach improves the performance of fast align significantly without the loss of efficiency.Experiments using other languages are also reported. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
EMNLP | 3 |
| 2015 | Hierarchical Phrase-based Stream DecodingabstractThis paper proposes a method for hierarchical phrase-based stream decoding.A stream decoder is able to take a continuous stream of tokens as input, and segments this stream into word sequences that are translated and output as a stream of target word sequences.Phrase-based stream decoding techniques have been shown to be effective as a means of simultaneous interpretation.In this paper we transfer the essence of this idea into the framework of hierarchical machine translation.The hierarchical decoding framework organizes the decoding process into a chart; this structure is naturally suited to the process of stream decoding, leading to an efficient stream decoding algorithm that searches a restricted subspace containing only relevant hypotheses.Furthermore, the decoder allows more explicit access to the word re-ordering process that is of critical importance in decoding while interpreting.The decoder was evaluated on TED talk data for English-Spanish and English-Chinese.Our results show that like the phrase-based stream decoder, the hierarchical is capable of approaching the performance of the underlying hierarchical phrase-based machine translation decoder, at useful levels of latency.In addition the hierarchical approach appeared to be robust to the difficulties presented by the more challenging English-Chinese task. Andrew M. Finch, Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
EMNLP | 4 |
| 2015 | Hierarchical Back-off Modeling of Hiero Grammar based on Non-parametric Bayesian ModelabstractIn hierarchical phrase-based machine translation, a rule table is automatically learned by heuristically extracting syn-chronous rules from a parallel corpus. As a result, spuriously many rules are extracted which may be composed of various incorrect rules. The larger rule table incurs more run time for decoding and may result in lower translation quality. To resolve the problems, we propose a hierarchical back-off model for Hiero grammar, an instance of a synchronous context free grammar (SCFG), on the basis of the hierarchical Pitman-Yor process. The model can extract a compact rule and phrase table without resorting to any heuristics by hierarchically backing off to smaller phrases under SCFG. Inference is efficiently carried out using two-step synchronous parsing of Xiao et al., (2012) combined with slice sampling. In our experiments, the proposed model achieved higher or at least comparable translation quality against a previous Bayesian model on various language pairs; German/French/Spanish/Japanese-English. When compared against heuristic models, our model achieved comparable translation quality on a full size German-English language pair in Europarl v7 corpus with significantly smaller grammar size; less than 10 % of that for heuristic model. 1 Hidetaka Kamigaito, Taro Watanabe, Hiroya Takamura, Manabu Okumura, Eiichiro Sumita |
EMNLP | 5 |
| 2015 | Leave-one-out Word Alignment without Garbage Collector EffectsabstractExpectation-maximization algorithms, such as those implemented in GIZA++ pervade the field of unsupervised word alignment.However, these algorithms have a problem of over-fitting, leading to "garbage collector effects," where rare words tend to be erroneously aligned to untranslated words.This paper proposes a leave-one-out expectationmaximization algorithm for unsupervised word alignment to address this problem.The proposed method excludes information derived from the alignment of a sentence pair from the alignment models used to align it.This prevents erroneous alignments within a sentence pair from supporting themselves.Experimental results on Chinese-English and Japanese-English corpora show that the F 1 , precision and recall of alignment were consistently increased by 5.0% -17.2%, and BLEU scores of end-to-end translation were raised by 0.03 -1.30.The proposed method also outperformed l 0 -normalized GIZA++ and Kneser-Ney smoothed GIZA++. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Taro Watanabe, Eiichiro Sumita |
EMNLP | 5 |
| 2015 | A Binarized Neural Network Joint Model for Machine TranslationabstractThe neural network joint model (NNJM), which augments the neural network language model (NNLM) with an m-word source context window, has achieved large gains in machine translation accuracy, but also has problems with high normalization cost when using large vocabularies.Training the NNJM with noise-contrastive estimation (NCE), instead of standard maximum likelihood estimation (MLE), can reduce computation cost.In this paper, we propose an alternative to NCE, the binarized NNJM (BNNJM), which learns a binary classifier that takes both the context and target words as input, and can be efficiently trained using MLE.We compare the BNNJM and NNJM trained by NCE on various translation tasks. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 3 |
| 2015 | HMM based myanmar text to speech system
Ye Kyaw Thu, Win Pa Pa, Jinfu Ni, Yoshinori Shiga, Andrew M. Finch, Chiori Hori, Hisashi Kawai, Eiichiro Sumita |
INTERSPEECH | 8 |
| 2015 | Patent claim translation based on sublanguage-specific sentence structure
Masaru Fuji, Atsushi Fujita, Masao Utiyama, Eiichiro Sumita, Yuji Matsumoto 0001 |
MTSummit | 4 |
| 2015 | Learning bilingual phrase representations with recurrent neural networks
Hideya Mino, Andrew M. Finch, Eiichiro Sumita |
MTSummit | 3 |
| 2015 | A Large-scale Study of Statistical Machine Translation Methods for Khmer Language
Ye Kyaw Thu, Vichet Chea, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
PACLIC | 5 |
| 2015 | Preordering using a Target-Language Parser via Cross-Language Syntactic Projection for Statistical Machine TranslationabstractWhen translating between languages with widely different word orders, word reordering can present a major challenge. Although some word reordering methods do not employ source-language syntactic structures, such structures are inherently useful for word reordering. However, high-quality syntactic parsers are not available for many languages. We propose a preordering method using a target-language syntactic parser to process source-language syntactic structures without a source-language syntactic parser. To train our preordering model based on ITG, we produced syntactic constituent structures for source-language training sentences by (1) parsing target-language training sentences, (2) projecting constituent structures of the target-language sentences to the corresponding source-language sentences, (3) selecting parallel sentences with highly synchronized parallel structures, (4) producing probabilistic models for parsing using the projected partial structures and the Pitman-Yor process, and (5) parsing to produce full binary syntactic structures maximally synchronized with the corresponding target-language syntactic structures, using the constraints of the projected partial structures and the probabilistic models. Our ITG-based preordering model is trained using the produced binary syntactic structures and word alignments. The proposed method facilitates the learning of ITG by producing highly synchronized parallel syntactic structures based on cross-language syntactic projection and sentence selection. The preordering model jointly parses input sentences and identifies their reordered structures. Experiments with Japanese--English and Chinese--English patent translation indicate that our method outperforms existing methods, including string-to-tree syntax-based SMT, a preordering method that does not require a parser, and a preordering method that uses a source-language dependency parser. Isao Goto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | Bilingual Continuous-Space Language Model Growing for Statistical Machine TranslationabstractLarger n-gram language models (LMs) perform better in statistical machine translation (SMT). However, the existing approaches have two main drawbacks for constructing larger LMs: 1) it is not convenient to obtain larger corpora in the same domain as the bilingual parallel corpora in SMT; 2) most of the previous studies focus on monolingual information from the target corpora only, and redundant n-grams have not been fully utilized in SMT. Nowadays, continuous-space language model (CSLM), especially neural network language model (NNLM), has been shown great improvement in the estimation accuracies of the probabilities for predicting the target words. However, most of these CSLM and NNLM approaches still consider monolingual information only or require additional corpus. In this paper, we propose a novel neural network based bilingual LM growing method. Compared to the existing approaches, the proposed method enables us to use bilingual parallel corpus for LM growing in SMT. The results show that our new method outperforms the existing approaches on both SMT performance and computational efficiency significantly. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2014 | Recurrent Neural Networks for Word Alignment ModelabstractThis study proposes a word alignment model based on a recurrent neural network (RNN), in which an unlimited alignment history is represented by recurrently connected hidden layers.We perform unsupervised learning using noise-contrastive estimation (Gutmann and Hyvärinen, 2010;Mnih and Teh, 2012), which utilizes artificially generated negative samples.Our alignment model is directional, similar to the generative IBM models (Brown et al., 1993).To overcome this limitation, we encourage agreement between the two directional models by introducing a penalty function that ensures word embedding consistency across two directional models during training.The RNN-based model outperforms the feed-forward neural network-based model (Yang et al., 2013) as well as the IBM Model 4 under Japanese-English and French-English word alignment tasks, and achieves comparable translation performance to those baselines for Japanese-English and Chinese-English translation tasks. Akihiro Tamura, Taro Watanabe, Eiichiro Sumita |
ACL (1) | 3 |
| 2014 | Syntax-Augmented Machine Translation using Syntax-Label ClusteringabstractRecently, syntactic information has helped significantly to improve statistical ma-chine translation. However, the use of syn-tactic information may have a negative im-pact on the speed of translation because of the large number of rules, especially when syntax labels are projected from a parser in syntax-augmented machine translation. In this paper, we propose a syntax-label clus-tering method that uses an exchange algo-rithm in which syntax labels are clustered together to reduce the number of rules. The proposed method achieves clustering by directly maximizing the likelihood of synchronous rules, whereas previous work considered only the similarity of proba-bilistic distributions of labels. We tested the proposed method on Japanese-English and Chinese-English translation tasks and found order-of-magnitude higher cluster-ing speeds for reducing labels and gains in translation quality compared with pre-vious clustering method. 1 Hideya Mino, Taro Watanabe, Eiichiro Sumita |
EMNLP | 3 |
| 2014 | Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine TranslationabstractLanguages that have no explicit word delimiters often have to be segmented for statistical machine translation (SMT).This is commonly performed by automated segmenters trained on manually annotated corpora.However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT.An analysis was performed to test this hypothesis using a manually annotated word alignment (WA) corpus for Chinese-English SMT.An analysis revealed that 74.60% of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA annotations.We formulated an approach based on word splitting with reference to the annotated WA to alleviate these conflicts.Experimental results show that the refined WS reduced word alignment error rate by 6.82% and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine translation (OpenMT) corpora compared to related work. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
EMNLP | 4 |
| 2014 | Neural Network Based Bilingual Language Model Growing for Statistical Machine TranslationabstractSince larger n-gram Language Model (LM) usually performs better in Statistical Machine Translation (SMT), how to construct efficient large LM is an important topic in SMT.However, most of the existing LM growing methods need an extra monolingual corpus, where additional LM adaption technology is necessary.In this paper, we propose a novel neural network based bilingual LM growing method, only using the bilingual parallel corpus in SMT.The results show that our method can improve both the perplexity score for LM evaluation and BLEU score for SMT, and significantly outperforms the existing LM growing methods without extra corpus. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
EMNLP | 5 |
| 2014 | Learning Hierarchical Translation SpansabstractWe propose a simple and effective approach to learn translation spans for the hierarchical phrase-based translation model.Our model evaluates if a source span should be covered by translation rules during decoding, which is integrated into the translation system as soft constraints.Compared to syntactic constraints, our model is directly acquired from an aligned parallel corpus and does not require parsers.Rich source side contextual features and advanced machine learning methods were utilized for this learning task.The proposed approach was evaluated on NTCIR-9 Chinese-English and Japanese-English translation tasks and showed significant improvement over the baseline system. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 3 |
| 2014 | Distortion Model Based on Word Sequence Labeling for Statistical Machine TranslationabstractThis article proposes a new distortion model for phrase-based statistical machine translation. In decoding, a distortion model estimates the source word position to be translated next (subsequent position; SP) given the last translated source word position (current position; CP). We propose a distortion model that can simultaneously consider the word at the CP, the word at an SP candidate, the context of the CP and an SP candidate, relative word order among the SP candidates, and the words between the CP and an SP candidate. These considered elements are called rich context . Our model considers rich context by discriminating label sequences that specify spans from the CP to each SP candidate. It enables our model to learn the effect of relative word order among SP candidates as well as to learn the effect of distances from the training data. In contrast to the learning strategy of existing methods, our learning strategy is that the model learns preference relations among SP candidates in each sentence of the training data. This leaning strategy enables consideration of all of the rich context simultaneously. In our experiments, our model had higher BLUE and RIBES scores for Japanese-English, Chinese-English, and German-English translation compared to the lexical reordering models. Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2013 | Distortion Model Considering Rich Context for Statistical Machine Translation
Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACL (1) | 3 |
| 2013 | Additive Neural Networks for Statistical Machine Translation
Lemao Liu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 3 |
| 2013 | Part-of-Speech Induction in Dependency Trees for Statistical Machine Translation
Akihiro Tamura, Taro Watanabe, Eiichiro Sumita, Hiroya Takamura, Manabu Okumura |
ACL (1) | 3 |
| 2013 | Hierarchical Phrase Table Combination for Machine Translation
Conghui Zhu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 3 |
| 2013 | Building a Bilingual Dictionary from a Japanese-Chinese Patent Corpus
Keiji Yasuda, Eiichiro Sumita |
CICLing (2) | 2 |
| 2013 | An Empirical Study on Word Segmentation for Chinese Machine Translation
Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita, Bao-Liang Lu |
CICLing (2) | 3 |
| 2013 | Converting Continuous-Space Language Models into N-Gram Language Models for Statistical Machine TranslationabstractNeural network language models, or continuous-space language models (CSLMs), have been shown to improve the performance of statistical machine translation (SMT) when they are used for reranking n-best translations.However, CSLMs have not been used in the first pass decoding of SMT, because using CSLMs in decoding takes a lot of time.In contrast, we propose a method for converting CSLMs into back-off n-gram language models (BNLMs) so that we can use converted CSLMs in decoding.We show that they outperform the original BNLMs and are comparable with the traditional use of CSLMs in reranking. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
EMNLP | 4 |
| 2013 | Tuning SMT with a Large Number of Features via Online Feature Grouping
Lemao Liu, Tiejun Zhao, Taro Watanabe, Eiichiro Sumita |
IJCNLP | 4 |
| 2013 | Multilingual Speech-to-Speech Translation System: VoiceTraabstractThis study presents an overview of VoiceTra, which was developed by NICT and released as the world's first network-based multilingual speech-to-speech translation system for smartphones, and describes in detail its multilingual speech recognition, its multilingual translation, and its multilingual speech synthesis in regards to field experiments. We show the effects of system updates using the data collected from field experiments to improve our acoustic and language models. Shigeki Matsuda, Xinhui Hu, Yoshinori Shiga, Hideki Kashioka, Chiori Hori, Keiji Yasuda, Hideo Okuma, Masao Uchiyama, Eiichiro Sumita, Hisashi Kawai, Satoshi Nakamura 0001 |
MDM (2) | 9 |
| 2013 | Inducing Romanization Systems
Keiko Taguchi, Andrew M. Finch, Seiichi Yamamoto, Eiichiro Sumita |
MTSummit | 4 |
| 2013 | A-STAR: Toward translating Asian spoken languages
Sakriani Sakti, Michael Paul, Andrew M. Finch, Shinsuke Sakai, Thang Tat Vu, Noriyuki Kimura, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
Comput. Speech Lang. | 8 |
| 2013 | A Bayesian Alignment Approach to Transliteration MiningabstractIn this article we present a technique for mining transliteration pairs using a set of simple features derived from a many-to-many bilingual forced-alignment at the grapheme level to classify candidate transliteration word pairs as correct transliterations or not. We use a nonparametric Bayesian method for the alignment process, as this process rewards the reuse of parameters, resulting in compact models that align in a consistent manner and tend not to over-fit. Our approach uses the generative model resulting from aligning the training data to force-align the test data. We rely on the simple assumption that correct transliteration pairs would be well modeled and generated easily, whereas incorrect pairs---being more random in character---would be more costly to model and generate. Our generative model generates by concatenating bilingual grapheme sequence pairs. The many-to-many generation process is essential for handling many languages with non-Roman scripts, and it is hard to train well using a maximum likelihood techniques, as these tend to over-fit the data. Our approach works on the principle that generation using only grapheme sequence pairs that are in the model results in a high probability derivation, whereas if the model is forced to introduce a new parameter in order to explain part of the candidate pair, the derivation probability is substantially reduced and severely reduced if the new parameter corresponds to a sequence pair composed of a large number of graphemes. The features we extract from the alignment of the test data are not only based on the scores from the generative model, but also on the relative proportions of each sequence that are hard to generate. The features are used in conjunction with a support vector machine classifier trained on known positive examples together with synthetic negative examples to determine whether a candidate word pair is a correct transliteration pair. In our experiments, we used all data tracks from the 2010 Named-Entity Workshop (NEWS’10) and use the performance of the best system for each language pair as a reference point. Our results show that the new features we propose are powerfully predictive, enabling our approach to achieve levels of performance on this task that are comparable to the state of the art. Takaaki Fukunishi, Andrew M. Finch, Seiichi Yamamoto, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2013 | Post-Ordering by Parsing with ITG for Japanese-English Statistical Machine TranslationabstractWord reordering is a difficult task for translation between languages with widely different word orders, such as Japanese and English. A previously proposed post-ordering method for Japanese-to-English translation first translates a Japanese sentence into a sequence of English words in a word order similar to that of Japanese, then reorders the sequence into an English word order. We employed this post-ordering framework and improved upon its reordering method. The existing post-ordering method reorders the sequence of English words via SMT, whereas our method reorders the sequence by (1) parsing the sequence using ITG to obtain syntactic structures which are similar to Japanese syntactic structures, and (2) transferring the obtained syntactic structures into English syntactic structures according to the ITG. The experiments using Japanese-to-English patent translation demonstrated the effectiveness of our method and showed that both the RIBES and BLEU scores were improved over compared methods. Isao Goto, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2013 | How to Choose the Best Pivot Language for Automatic Translation of Low-Resource LanguagesabstractRecent research on multilingual statistical machine translation focuses on the usage of pivot languages in order to overcome language resource limitations for certain language pairs. Due to the richness of available language resources, English is, in general, the pivot language of choice. However, factors like language relatedness can also effect the choice of the pivot language for a given language pair, especially for Asian languages, where language resources are currently quite limited. In this article, we provide new insights into what factors make a pivot language effective and investigate the impact of these factors on the overall pivot translation performance for translation between 22 Indo-European and Asian languages. Experimental results using state-of-the-art statistical machine translation techniques revealed that the translation quality of 54.8% of the language pairs improved when a non-English pivot language was chosen. Moreover, 81.0% of system performance variations can be explained by a combination of factors such as language family, vocabulary, sentence length, language perplexity, translation model entropy, reordering, monotonicity, and engine performance. Michael Paul, Andrew M. Finch, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2013 | picoTrans: An intelligent icon-driven interface for cross-lingual communicationabstractpicoTrans is a prototype system that introduces a novel icon-based paradigm for cross-lingual communication on mobile devices. Our approach marries a machine translation system with the popular picture book. Users interact with picoTrans by pointing at pictures as if it were a picture book; the system generates natural language from these icons and the user is able to interact with the icon sequence to refine the meaning of the words that are generated. When users are satisfied that the sentence generated represents what they wish to express, they tap a translate button and picoTrans displays the translation. Structuring the process of communication in this way has many advantages. First, tapping icons is a very natural method of user input on mobile devices; typing is cumbersome and speech input errorful. Second, the sequence of icons which is annotated both with pictures and bilingually with words is meaningful to both users, and it opens up a second channel of communication between them that conveys the gist of what is being expressed. We performed a number of evaluations of picoTrans to determine: its coverage of a corpus of in-domain sentences; the input efficiency in terms of the number of key presses required relative to text entry; and users' overall impressions of using the system compared to using a picture book. Our results show that we are able to cover 74% of the expressions in our test corpus using a 2000-icon set; we believe that this icon set size is realistic for a mobile device. We also found that picoTrans requires fewer key presses than typing the input and that the system is able to predict the correct, intended natural language sentence from the icon sequence most of the time, making user interaction with the icon sequence often unnecessary. In the user evaluation, we found that in general users prefer using picoTrans and are able to communicate more rapidly and expressively. Furthermore, users had more confidence that they were able to communicate effectively using picoTrans. Andrew M. Finch, Kumiko Tanaka-Ishii, Keiji Yasuda, Eiichiro Sumita |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2012 | Phrasal Syntactic Category Sequence Model for Phrase-Based MT
Hailong Cao, Eiichiro Sumita, Tiejun Zhao, Sheng Li 0003 |
CICLing (2) | 2 |
| 2012 | Method to Build a Bilingual Lexicon for Speech-to-Speech Translation Systems
Keiji Yasuda, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 3 |
| 2012 | Crowd-based MT Evaluation for non-English Target Languages
Michael Paul, Eiichiro Sumita, Luisa Bentivogli, Marcello Federico |
EAMT | 2 |
| 2012 | Bilingual Lexicon Extraction from Comparable Corpora Using Label Propagation
Akihiro Tamura, Taro Watanabe, Eiichiro Sumita |
EMNLP-CoNLL | 3 |
| 2011 | An Unsupervised Model for Joint Phrase Alignment and Extraction
Graham Neubig, Taro Watanabe, Eiichiro Sumita, Shinsuke Mori, Tatsuya Kawahara |
ACL | 3 |
| 2011 | Machine Translation System Combination by Confusion Forest
Taro Watanabe, Eiichiro Sumita |
ACL | 2 |
| 2011 | Word Segmentation for Dialect Translation
Michael Paul, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 3 |
| 2011 | A Method to Measure the Reading Difficulty of Japanese Words
Keiji Yasuda, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 3 |
| 2011 | Rule-based Reordering Constraints for Phrase-based SMT
Chooi-Ling Goh, Takashi Onishi, Eiichiro Sumita |
EAMT | 3 |
| 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile DevicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved: the picture-book, in which the user simply points to multiple picture icons representing what they want to say, and the statistical machine translation (SMT) system that can translate arbitrary word sequences. Our prototype system tightly couples both processes within a translation framework that inherits many of the the positive features of both approaches, while at the same time mitigating their main weaknesses. Our system differs from traditional approaches in that its mode of input is a sequence of pictures, rather than text or speech. Text in the source language is generated automatically, and is used as a detailed representation of the intended meaning. The picture sequence which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the translation, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IJCAI | 4 |
| 2011 | Translation Quality Indicators for Pivot-based Statistical MT
Michael Paul, Eiichiro Sumita |
IJCNLP | 2 |
| 2011 | picoTrans: an icon-driven user interface for machine translation on mobile devicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved. In our approach we integrate the popular picture-book, in which the user simply points to multiple picture icons representing what they want to say, with a statistical machine translation system that can translate arbitrary word sequences. The simple pointing at pictures paradigm is used as the primary method of user input and the users can use the device as if it were a picture book. The application is then able to generate a complete sentence in the user's native language for what they wish to say from the sequence of picture icons chosen by the user. Once the user is satisfied that the sentence provided by the system adequately represents what they wish to convey, the application can automatically translate the sentence into the language of the other party, who can interpret the intended meaning of the first party by combining evidence from both modes of communication: the picture sequence, and the machine translation. The prototype system we have developed inherits many of the positive features of both approaches, while at the same time mitigating their main weaknesses. The user may combine the pictures in considerably more combinations than is possible with a picture book designed with combinations from only within the same page spread of the book in mind, making the application more expressive than a book. The machine translation system can contribute a detailed and precise translation which is supported by the picture-based mode which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the sentence, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IUI | 4 |
| 2011 | A Comparison Study of Parsers for Patent Machine Translation
Isao Goto, Masao Utiyama, Takashi Onishi, Eiichiro Sumita |
MTSummit | 4 |
| 2011 | A Comparison of Unsupervised Bilingual Term Extraction Methods Using Phrase-Tables
Masamichi Ideue, Kazuhide Yamamoto, Masao Utiyama, Eiichiro Sumita |
MTSummit | 4 |
| 2011 | Searching Translation Memories for Paraphrases
Masao Utiyama, Graham Neubig, Takashi Onishi, Eiichiro Sumita |
MTSummit | 4 |
| 2010 | Community-based Construction of Draft and Final Translation Corpus Through a Translation Hosting Site Minna no Hon'yaku (MNH)
Takeshi Abekawa, Masao Utiyama, Eiichiro Sumita, Kyo Kageura |
LREC | 3 |
| 2009 | The Asian network-based speech-to-speech translation systemabstractThis paper outlines the first Asian network-based speech-to-speech translation system developed by the Asian Speech Translation Advanced Research (A-STAR) consortium. The system was designed to translate common spoken utterances of travel conversations from a certain source language into multiple target languages in order to facilitate multiparty travel conversations between people speaking different Asian languages. Each A-STAR member contributes one or more of the following spoken language technologies: automatic speech recognition, machine translation, and text-to-speech through Web servers. Currently, the system has successfully covered 9 languages-namely, 8 Asian languages (Hindi, Indonesian, Japanese, Korean, Malay, Thai, Vietnamese, Chinese) and additionally, the English language. The system's domain covers about 20,000 travel expressions, including proper nouns that are names of famous places or attractions in Asian countries. In this paper, we discuss the difficulties involved in connecting various different spoken language translation systems through Web servers. We also present speech-translation results on the first A-STAR demo experiments carried out in July 2009. Sakriani Sakti, Noriyuki Kimura, Michael Paul, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
ASRU | 5 |
| 2009 | Bidirectional Phrase-based Statistical Machine Translation
Andrew M. Finch, Eiichiro Sumita |
EMNLP | 2 |
| 2009 | Mining Parallel Texts from Mixed-Language Web Pages
Masao Utiyama, Daisuke Kawahara, Keiji Yasuda, Eiichiro Sumita |
MTSummit | 4 |
| 2008 | Phrase-based Machine Transliteration
Andrew M. Finch, Eiichiro Sumita |
IJCNLP | 2 |
| 2008 | Method of Selecting Training Data to Build a Compact and Efficient Translation Model
Keiji Yasuda, Ruiqiang Zhang, Hirofumi Yamamoto, Eiichiro Sumita |
IJCNLP | 4 |
| 2008 | Chinese Unknown Word Translation by Subword Re-segmentation
Ruiqiang Zhang, Eiichiro Sumita |
IJCNLP | 2 |
| 2008 | Achilles: NiCT/ATR Chinese Morphological Analyzer for the Fourth Sighan Bakeoff
Ruiqiang Zhang, Eiichiro Sumita |
IJCNLP | 2 |
| 2007 | NICT-ATR Speech-to-Speech Translation System
Eiichiro Sumita, Tohru Shimizu, Satoshi Nakamura 0001 |
ACL | 1 |
| 2007 | Boosting Statistical Machine Translation by Lemmatization and Linear Interpolation
Ruiqiang Zhang, Eiichiro Sumita |
ACL | 2 |
| 2007 | Bilingual Cluster Based Models for Statistical Machine Translation
Hirofumi Yamamoto, Eiichiro Sumita |
EMNLP-CoNLL | 2 |
| 2007 | Introducing translation dictionary into phrase-based SMT
Hideo Okuma, Hirofumi Yamamoto, Eiichiro Sumita |
MTSummit | 3 |
| 2007 | The Infinite Markov ModelabstractWe present a nonparametric Bayesian method of estimating variable order Markov processes up to a theoretically infinite order. By extending a stick-breaking prior, which is usually defined on a unit interval, “vertically” to the trees of infinite depth associated with a hierarchical Chinese restaurant process, our model directly infers the hidden orders of Markov dependencies from which each symbol originated. Experiments on character and word sequences in natural language showed that the model has a comparative performance with an exponentially large full-order model, while computationally much efficient in both time and space. We expect that this basic model will also extend to the variable order hierarchical clustering of general data. Daichi Mochihashi, Eiichiro Sumita |
NIPS | 2 |
| 2006 | Using Lexical Dependency and Ontological Knowledge to Improve a Detailed Syntactic and Semantic Tagger of English
Andrew M. Finch, Ezra Black, Young-Sook Hwang, Eiichiro Sumita |
ACL | 4 |
| 2006 | Subword-Based Tagging for Confidence-Dependent Chinese Word Segmentation
Ruiqiang Zhang, Gen-ichiro Kikui, Eiichiro Sumita |
ACL | 3 |
| 2006 | Exploiting Variant Corpora for Machine Translation
Michael Paul, Eiichiro Sumita |
HLT-NAACL | 2 |
| 2006 | Using the Web to Disambiguate Acronyms
Eiichiro Sumita, Fumiaki Sugaya |
HLT-NAACL | 1 |
| 2006 | Word Pronunciation Disambiguation using the Web
Eiichiro Sumita, Fumiaki Sugaya |
HLT-NAACL | 1 |
| 2006 | Subword-based Tagging by Conditional Random Fields for Chinese Word Segmentation
Ruiqiang Zhang, Gen-ichiro Kikui, Eiichiro Sumita |
HLT-NAACL | 3 |
| 2006 | Using multiple edit distances to automatically grade outputs from Machine translation systemsabstractThis paper addresses the challenging problem of automatically evaluating output from machine translation (MT) systems that are subsystems of speech-to-speech MT (SSMT) systems. Conventional automatic MT evaluation methods include BLEU, which MT researchers have frequently used. However, BLEU has two drawbacks in SSMT evaluation. First, BLEU assesses errors lightly at the beginning of translations and heavily in the middle, even though its assessments should be independent of position. Second, BLEU lacks tolerance in accepting colloquial sentences with small errors, although such errors do not prevent us from continuing an SSMT-mediated conversation. In this paper, the authors report a new evaluation method called “g Rader based on Edit Distances (RED)” that automatically grades each MT output by using a decision tree (DT). The DT is learned from training data that are encoded by using multiple edit distances, that is, normal edit distance (ED) defined by insertion, deletion, and replacement, as well as its extensions. The use of multiple edit distances allows more tolerance than either ED or BLEU. Each evaluated MT output is assigned a grade by using the DT. RED and BLEU were compared for the task of evaluating MT systems of varying quality on ATR's Basic Travel Expression Corpus (BTEC). Experimental results show that RED significantly outperforms BLEU. Yasuhiro Akiba, Kenji Imamura, Eiichiro Sumita, Hiromi Nakaiwa, Shun'ichi Yamamoto, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Comparative study on corpora for speech translationabstractThis paper investigates issues in preparing corpora for developing speech-to-speech translation (S2ST). It is impractical to create a broad-coverage parallel corpus only from dialog speech. An alternative approach is to have bilingual experts write conversational-style texts in the target domain, with translations. There is, however, a risk of losing fidelity to the actual utterances. This paper focuses on balancing a tradeoff between these two kinds of corpora through the analysis of two newly developed corpora in the travel domain: a bilingual parallel corpus with 420 K utterances and a collection of in-domain dialogs using actual S2ST systems. We found that the first corpus is effective for covering utterances in the second corpus if complimented with a small number of utterances taken from monolingual dialogs. We also found that characteristics of in-domain utterances become closer to those of the first corpus when more restrictive conditions and instructions to speakers are given. These results suggest the possibility of a bootstrap-style of development of corpora and S2ST systems, where an initial S2ST system is developed with parallel texts, and is then gradually improved with in-domain utterances collected by the system as restrictions are relaxed Gen-ichiro Kikui, Seiichi Yamamoto, Toshiyuki Takezawa, Eiichiro Sumita |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | The ATR multilingual speech-to-speech translation systemabstractIn this paper, we describe the ATR multilingual speech-to-speech translation (S2ST) system, which is mainly focused on translation between English and Asian languages (Japanese and Chinese). There are three main modules of our S2ST system: large-vocabulary continuous speech recognition, machine text-to-text (T2T) translation, and text-to-speech synthesis. All of them are multilingual and are designed using state-of-the-art technologies developed at ATR. A corpus-based statistical machine learning framework forms the basis of our system design. We use a parallel multilingual database consisting of over 600 000 sentences that cover a broad range of travel-related conversations. Recent evaluation of the overall system showed that speech-to-speech translation quality is high, being at the level of a person having a Test of English for International Communication (TOEIC) score of 750 out of the perfect score of 990. Satoshi Nakamura 0001, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, Jinsong Zhang 0001, Hirofumi Yamamoto, Eiichiro Sumita, Seiichi Yamamoto |
IEEE Trans. Speech Audio Process. | 9 |
| 2005 | Acquiring Synonyms from Monolingual Comparable Texts
Mitsuo Shimohata, Eiichiro Sumita |
IJCNLP | 2 |
| 2005 | Practical Approach to Syntax-based Statistical Machine TranslationabstractThis paper presents a practical approach to statistical machine translation (SMT) based on syntactic transfer. Conventionally, phrase-based SMT generates an output sentence by combining phrase (multiword sequence) translation and phrase reordering without syntax. On the other hand, SMT based on tree-to-tree mapping, which involves syntactic information, is theoretical, so its features remain unclear from the viewpoint of a practical system. The SMT proposed in this paper translates phrases with hierarchical reordering based on the bilingual parse tree. In our experiments, the best translation was obtained when both phrases and syntactic information were used for the translation process. Kenji Imamura, Hideo Okuma, Eiichiro Sumita |
MTSummit | 3 |
| 2005 | Example-based machine translation using efficient sentence retrieval based on edit-distanceabstractAn Example-Based Machine Translation (EBMT) system, whose translation example unit is a sentence, can produce an accurate and natural translation if translation examples similar enough to an input sentence are retrieved. Such a system, however, suffers from the problem of narrow coverage. To reduce the problem, a large-scale parallel corpus is required and, therefore, an efficient method is needed to retrieve translation examples from a large-scale corpus. The authors propose an efficient retrieval method for a sentence-wise EBMT using edit-distance. The proposed retrieval method efficiently retrieves the most similar sentences using the measure of edit-distance without omissions. The proposed method employs search-space division, word graphs, and an A* search algorithm. The performance of the EBMT was evaluated through Japanese-to-English translation experiments using a bilingual corpus comprising hundreds of thousands of sentences from a travel conversation domain. The EBMT system achieved a high-quality translation ability by using a large corpus and also achieved efficient processing by using the proposed retrieval method. Takao Doi, Hirofumi Yamamoto, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2004 | Using a Mixture of N-Best Lists from Multiple MT Systems in Rank-Sum-Based Confidence Measure for MT Outputs
Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno |
COLING | 2 |
| 2004 | Splitting Input Sentence for Machine Translation Using Language Model with Sentence Similarity
Takao Doi, Eiichiro Sumita |
COLING | 2 |
| 2004 | Example-based Machine Translation Based on Syntactic Transfer with Statistical Models
Kenji Imamura, Hideo Okuma, Taro Watanabe, Eiichiro Sumita |
COLING | 4 |
| 2004 | Reordering Constraints for Phrase-Based Statistical Machine Translation
Richard Zens, Hermann Ney, Taro Watanabe, Eiichiro Sumita |
COLING | 4 |
| 2004 | Improved spoken language translation using n-best speech recognition hypothesesabstractWe intended to demonstrate the effect of using N-best speech recognition hypotheses for improving speech translation performance. A log-linear model, which integrated features from speech recognition and statistical machine translation, was used to rescore the translation candidates. Model parameters were estimated by optimizing an objectively measurable but subjectively relevant translation quality metric. Experimental results have shown that the proposed N-best approach improved translation quality over the conventional single-best approach. The improvements were confirmed consistently by several automatic translation evaluation metrics. 1. Ruiqiang Zhang, Gen-ichiro Kikui, Hirofumi Yamamoto, Frank K. Soong, Taro Watanabe, Eiichiro Sumita, Wai Kit Lo |
INTERSPEECH | 6 |
| 2004 | Incremental Methods to Select Test Sentences for Evaluating Translation Ability
Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno |
LREC | 2 |
| 2004 | How Does Automatic Machine Translation Evaluation Correlate with Human Scoring as the Number of Reference Translations Increases?
Andrew M. Finch, Yasuhiro Akiba, Eiichiro Sumita |
LREC | 3 |
| 2004 | Building a Paraphrase Corpus for Speech Translation
Mitsuo Shimohata, Eiichiro Sumita, Yuji Matsumoto 0001 |
LREC | 2 |
| 2003 | Feedback Cleaning of Machine Translation Rules Using Automatic EvaluationabstractWhen rules of transfer-based machine translation (MT) are automatically acquired from bilingual corpora, incorrect/redundant rules are generated due to acquisition errors or translation variety in the corpora. As a new countermeasure to this problem, we propose a feedback cleaning method using automatic evaluation of MT quality, which removes incorrect/redundant rules as a way to increase the evaluation score. BLEU is utilized for the automatic evaluation. The hill-climbing algorithm, which involves features of this task, is applied to searching for the optimal combination of rules. Our experiments show that the MT quality improves by 10% in test sentences according to a subjective evaluation. This is considerable improvement over previous methods. Kenji Imamura, Eiichiro Sumita, Yuji Matsumoto 0001 |
ACL | 2 |
| 2003 | Chunk-Based Statistical TranslationabstractThis paper describes an alternative translation model based on a text chunk under the framework of statistical machine translation. The translation model suggested here first performs chunking. Then, each word in a chunk is translated. Finally, translated chunks are reordered. Under this scenario of translation modeling, we have experimented on a broad-coverage Japanese-English traveling corpus and achieved improved performance. Taro Watanabe, Eiichiro Sumita, Hiroshi G. Okuno |
ACL | 2 |
| 2003 | Automatic Construction of Machine Translation Knowledge Using Translation Literalness
Kenji Imamura, Eiichiro Sumita, Yuji Matsumoto 0001 |
EACL | 2 |
| 2003 | A corpus-centered approach to spoken language translation
Eiichiro Sumita, Yasuhiro Akiba, Takao Doi, Andrew M. Finch, Kenji Imamura, Michael Paul, Mitsuo Shimohata, Taro Watanabe |
EACL | 1 |
| 2003 | Creating corpora for speech-to-speech translation
Gen-ichiro Kikui, Eiichiro Sumita, Toshiyuki Takezawa, Seiichi Yamamoto |
INTERSPEECH | 2 |
| 2003 | Experimental comparison of MT evaluation methods: RED vs.BLEUabstractThis paper experimentally compares two automatic evaluators, RED and BLEU, to determine how close the evaluation results of each automatic evaluator are to average evaluation results by human evaluators, following the ATR standard of MT evaluation. This paper gives several cautionary remarks intended to prevent MT developers from drawing misleading conclusions when using the automatic evaluators. In addition, this paper reports a way of using the automatic evaluators so that their results agree with those of human evaluators. Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno |
MTSummit | 2 |
| 2003 | Example-based rough translation for speech-to-speech translationabstractExample-based machine translation (EBMT) is a promising translation method for speech-to-speech translation (S2ST) because of its robustness. However, it has two problems in that the performance degrades when input sentences are long and when the style of the input sentences and that of the example corpus are different. This paper proposes example-based rough translation to overcome these two problems. The rough translation method relies on “meaning-equivalent sentences,” which share the main meaning with an input sentence despite missing some unimportant information. This method facilitates retrieval of meaning-equivalent sentences for long input sentences. The retrieval of meaning-equivalent sentences is based on content words, modality, and tense. This method also provides robustness against the style differences between the input sentence and the example corpus. Mitsuo Shimohata, Eiichiro Sumita, Yuji Matsumoto 0001 |
MTSummit | 2 |
| 2003 | Example-based decoding for statistical machine translationabstractThis paper presents a decoder for statistical machine translation that can take advantage of the example-based machine translation framework. The decoder presented here is based on the greedy approach to the decoding problem, but the search is initiated from a similar translation extracted from a bilingual corpus. The experiments on multilingual translations showed that the proposed method was far superior to a word-by-word generation beam search algorithm. Taro Watanabe, Eiichiro Sumita |
MTSummit | 2 |
| 2003 | Adaptation Using Out-of-Domain Corpus within EBMT
Takao Doi, Eiichiro Sumita, Hirofumi Yamamoto |
HLT-NAACL | 2 |
| 2003 | Automatic Expansion of Equivalent Sentence Set Based on Syntactic Substitution
Kenji Imamura, Yasuhiro Akiba, Eiichiro Sumita |
HLT-NAACL | 3 |
| 2002 | Using Language and Translation Models to Select the Best among Outputs from Multiple MT Systems
Yasuhiro Akiba, Taro Watanabe, Eiichiro Sumita |
COLING | 3 |
| 2002 | Corpus-based Generation of Numeral Classifier using Phrase Alignment
Michael Paul, Eiichiro Sumita, Seiichi Yamamoto |
COLING | 2 |
| 2002 | Bidirectional Decoding for Statistical Machine Translation
Taro Watanabe, Eiichiro Sumita |
COLING | 2 |
| 2002 | Bilingual corpus cleaning focusing on translation literality
Kenji Imamura, Eiichiro Sumita |
INTERSPEECH | 2 |
| 2002 | Reliability measures for translation quality
Eiichiro Sumita, Yasuhiro Akiba, Kenji Imamura |
INTERSPEECH | 1 |
| 2002 | Statistical machine translation decoder based on phraseabstractThis paper describes a decoding algorithm for statistical machine translation based on phrases. In the past, the solution to the decoding problem were inspired from that of speech recognizers, translating each input word into one or more output words generating in left-to-right direction. The algorithm presented here iteratively constructs phrases or chunks of cepts until all the input words are consumed. This behavior resulted in computational complexity higher than those with left-to-right constraints, though the translation accuracy is better from the Japanese-to-English translation experiments. 1. Taro Watanabe, Eiichiro Sumita |
INTERSPEECH | 2 |
| 2002 | Automatic paraphrasing based on parallel corpus for normalization
Mitsuo Shimohata, Eiichiro Sumita |
LREC | 2 |
| 2002 | Toward a Broad-coverage Bilingual Corpus for Speech Translation of Travel Conversations in the Real World
Toshiyuki Takezawa, Eiichiro Sumita, Fumiaki Sugaya, Hirofumi Yamamoto, Seiichi Yamamoto |
LREC | 2 |
| 2002 | Statistical Machine Translation on Paraphrased Corpora
Taro Watanabe, Mitsuo Shimohata, Eiichiro Sumita |
LREC | 3 |
| 2001 | Converting Morphological Information Using Lexicalized and General Conversion
Mitsuo Shimohata, Eiichiro Sumita |
CICLing | 2 |
| 2001 | Using multiple edit distances to automatically rank machine translation outputabstractThis paper addresses the challenging problem of automatically evaluating output from machine translation (MT) systems in order to support the developers of these systems. Conventional approaches to the problem include methods that automatically assign a rank such as A, B, C, or D to MT output according to a single edit distance between this output and a correct translation example. The single edit distance can be differently designed, but changing its design makes assigning a certain rank more accurate, but another rank less accurate. This inhibits improving accuracy of rank assignment. To overcome this obstacle, this paper proposes an automatic ranking method that, by using multiple edit distances, encodes machine-translated sentences with a rank assigned by humans into multi-dimensional vectors from which a classifier of ranks is learned in the form of a decision tree (DT). The proposed method assigns a rank to MT output through the learned DT. The proposed method is evaluated using transcribed texts of real conversations in the travel arrangement domain. Experimental results show that the proposed method is more accurate than the single-edit-distance-based ranking methods, in both closed and open tests. Moreover, the proposed method could estimate MT quality within 3% error in some cases. Yasuhiro Akiba, Kenji Imamura, Eiichiro Sumita |
MTSummit | 3 |
| 2000 | Lexical Transfer Using a Vector-Space ModelabstractBuilding a bilingual dictionary for transfer in a machine translation system is conventionally done by hand and is very time-consuming.In order to overcome this bottleneck, we propose a new mechanism for lexical transfer, which is simple and suitable for learning from bilingual corpora.It exploits a vector-space model developed in information retrieval research.We present a preliminary result from our computational experiment. Eiichiro Sumita |
ACL | 1 |
| 2000 | Multiple decision-tree strategy for input-error robustness: a simulation of tree combinations
Kazuhide Yamamoto, Eiichiro Sumita |
INTERSPEECH | 2 |
| 2000 | Utilization of Coreferences for the Translation of Utterances Containing Anaphoric Expressions
Michael Paul, Eiichiro Sumita |
PRICAI | 2 |
| 2000 | Word Alignment Using a Matrix
Eiichiro Sumita |
PRICAI | 1 |
| 1999 | Error correction translation using text corporaabstractThis paper presents a parametric matching and smoothing method that is applied to a sinusoidal representation and auditory model-based speech analysis/synthesis system. A 2.6kbps speechcoding algorithm is finally derived based on the speech analysis/synthesis system. The synthetic speech is almost same as that of 3.25kbps speech coding algorithm with overlapping and adding method. A linear interpolation method is utilized to smooth the amplitude parameters, and a nonlinear polynomial interpolation method is used to smooth the frequency and phase parameters. The experimental results demonstrate that the parametric matching and smoothing method can reduce the bit-rate with the speech quality unchanged when it is applied to the sinusoidal representation and auditory model-based speech-coding algorithm. Kai Ishikawa, Eiichiro Sumita |
EUROSPEECH | 2 |
| 1999 | Solutions to problems inherent in spoken-language translation: the ATR-MATRIX approachabstractATR has built a multi-language speech translation system called ATR-MATRIX. It consists of a spoken-language translation subsystem, which is the focus of this paper, together with a highly accurate speech recognition subsystem and a high-definition speech synthesis subsystem. This paper gives a road map of solutions to the problems inherent in spoken-language translation. Spoken-language translation systems need to tackle difficult problems such as ungrammaticality. contextual phenomena, speech recognition errors, and the high-speeds required for real-time use. We have made great strides towards solving these problems in recent years. Our approach mainly uses an example-based translation model called TDMT. We have added the use of extra-linguistic information, a decision tree learning mechanism, and methods dealing with recognition errors. Eiichiro Sumita, Setsuo Yamada, Kazuhide Yamamoto, Michael Paul, Hideki Kashioka, Kai Ishikawa, Satoshi Shirai |
MTSummit | 1 |
| 1998 | Example-based error recovery method for speech translation: repairing sub-trees according to the semantic distance
Kai Ishikawa, Eiichiro Sumita, Hitoshi Iida |
ICSLP | 2 |
| 1996 | Spoken-Language Translation Method Using Examples
Hitoshi Iida, Eiichiro Sumita, Osamu Furuse |
COLING | 2 |
| 1994 | The Relationship between Architectures and Example-Retrieval Times
Eiichiro Sumita, Naoya Nisiyama, Hitoshi Iida |
AAAI | 1 |
| 1993 | Example-Based Machine Translation on Massively Parallel Processors
Eiichiro Sumita, Kozo Oi, Osamu Furuse, Hitoshi Iida, Tetsuya Higuchi, Naoto Takahashi, Hiroaki Kitano |
IJCAI | 1 |
| 1991 | Experiments and Prospects of Example-Based Machine TranslationabstractEBMT (Example-Based Machine Translation) is proposed. EBMT retrieves similar examples (pairs of source phrases, sentences, or texts and their translations) from a database of examples, adapting the examples to translate a new input. EBMT has the following features: (1) It is easily upgraded simply by inputting appropriate examples to the database; (2) It assigns a reliability factor to the translation result; (3) It is accelerated effectively by both indexing and parallel computing; (4) It is robust because of best-match reasoning; and (5) It well utilizes translator expertise. A prototype system has been implemented to deal with a difficult translation problem for conventional Rule-Based Machine Translation (RBMT), i.e., translating Japanese noun phrases of the form "N1 no N2" into English. The system has achieved about a 78% success rate on average. This paper explains the basic idea of EBMT, illustrates the experiment in detail, explains the broad applicability of EBMT to several difficult translation problems for RBMT and discusses the advantages of integrating EBMT with RBMT. Eiichiro Sumita, Hitoshi Iida |
ACL | 1 |