EDBT 2026 Demo / reviewers in the wild / expert
Masao Utiyama
dblp:76/5745
· DBLP profile ↗
120ranked-venue papers
12as first author
28since 2021 · last 2025
0000-0003-1111-9245ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 118 · 12 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine TranslationabstractZhi Qu, Yiran Wang, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhi Qu 0001, Yiran Wang 0006, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe |
ACL (1) | 6 |
| 2025 | PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language GenerationabstractThis work introduces PrahokBART, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus quality and addressing the linguistic issues of Khmer, which are ignored in existing multilingual models, by incorporating linguistic components such as word segmentation and normalization. We evaluate PrahokBART on three generative tasks: machine translation, text summarization, and headline generation, where our results demonstrate that it outperforms mBART50, a strong multilingual pre-trained model. Additionally, our analysis provides insights into the impact of each linguistic module and evaluates how effectively our model handles space during text generation, which is crucial for the naturalness of texts in Khmer. Hour Kaing, Raj Dabre, Haiyue Song, Van-Hien Tran, Hideki Tanaka, Masao Utiyama |
COLING | 6 |
| 2025 | Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto |
ICCV | 5 |
| 2024 | To be Continuous, or to be Discrete, Those are Bits of QuestionsabstractRecently, binary representation has been proposed as a novel representation that lies between continuous and discrete representations.It exhibits considerable information-preserving capability when being used to replace continuous input vectors.In this paper, we investigate the feasibility of further introducing it to the output side, aiming to allow models to output binary labels instead.To preserve the structural information on the output side along with label information, we extend the previous contrastive hashing method as structured contrastive hashing.More specifically, we upgrade CKY from label-level to bit-level, define a new similarity function with span marginal probabilities, and introduce a novel contrastive loss function with a carefully designed instance selection strategy.Our model 1 achieves competitive performance on various structured prediction tasks, and demonstrates that binary representation can be considered a novel representation that further bridges the gap between the continuous nature of deep learning and the discrete intrinsic property of natural languages. Yiran Wang 0006, Masao Utiyama |
ACL (1) | 2 |
| 2024 | On Eliciting Syntax from Language Models via HashingabstractUnsupervised parsing, also known as grammar induction, aims to infer syntactic structure from raw text.Recently, binary representation has exhibited remarkable informationpreserving capabilities at both lexicon and syntax levels.In this paper, we explore the possibility of leveraging this capability to deduce parsing trees from raw text, relying solely on the implicitly induced grammars within models.To achieve this, we upgrade the bit-level CKY from zero-order to first-order to encode the lexicon and syntax in a unified binary representation space, switch training from supervised to unsupervised under the contrastive hashing framework, and introduce a novel loss function to impose stronger yet balanced alignment signals.Our model 1 shows competitive performance on various datasets, therefore, we claim that our method is effective and efficient enough to acquire high-quality parsing trees from pre-trained language models at a low cost. Yiran Wang 0006, Masao Utiyama |
EMNLP | 2 |
| 2023 | Language Model Pre-training on True NegativesabstractDiscriminative pre-trained language models (PrLMs) learn to predict original texts from intentionally corrupted ones. Taking the former text as positive and the latter as negative samples, the PrLM can be trained effectively for contextualized representation. However, the training of such a type of PrLMs highly relies on the quality of the automatically constructed samples. Existing PrLMs simply treat all corrupted texts as equal negative without any examination, which actually lets the resulting model inevitably suffer from the false negative issue where training is carried out on pseudo-negative data and leads to less efficiency and less robustness in the resulting PrLMs. In this work, on the basis of defining the false negative issue in discriminative PrLMs that has been ignored for a long time, we design enhanced pre-training methods to counteract false negative predictions and encourage pre-training language models on true negatives by correcting the harmful gradient updates subject to false negative predictions. Experimental results on GLUE and SQuAD benchmarks show that our counter-false-negative pre-training methods indeed bring about better performance together with stronger robustness. Zhuosheng Zhang 0001, Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita |
AAAI | 3 |
| 2023 | Subset Retrieval Nearest Neighbor Machine TranslationabstractHiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita |
ACL (1) | 4 |
| 2023 | 24-bit LanguagesabstractYiran Wang, Taro Watanabe, Masao Utiyama, Yuji Matsumoto. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yiran Wang 0006, Taro Watanabe, Masao Utiyama, Yuji Matsumoto 0001 |
IJCNLP (1) | 3 |
| 2023 | Pivot Translation for Zero-resource Language Pairs Based on a Multilingual Pretrained ModelabstractA multilingual translation model enables a single model to handle multiple languages. However, the translation qualities of unlearned language pairs (i.e., zero-shot translation qualities) are still poor. By contrast, pivot translation translates source texts into target ones via a pivot language such as English, thus enabling machine translation without parallel texts between the source and target languages. In this paper, we perform pivot translation using a multilingual model and compare it with direct translation. We improve the translation quality without using parallel texts of direct translation by fine-tuning the model with machine-translated pseudo-translations. We also discuss what type of parallel texts are suitable for effectively improving the translation quality in multilingual pivot translation. Kenji Imamura, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 2 |
| 2023 | Improving Embedding Transfer for Low-Resource Machine TranslationabstractLow-resource machine translation (LRMT) poses a substantial challenge due to the scarcity of parallel training data. This paper introduces a new method to improve the transfer of the embedding layer from the Parent model to the Child model in LRMT, utilizing trained token embeddings in the Parent model’s high-resource vocabulary. Our approach involves projecting all tokens into a shared semantic space and measuring the semantic similarity between tokens in the low-resource and high-resource languages. These measures are then utilized to initialize token representations in the Child model’s low-resource vocabulary. We evaluated our approach on three benchmark datasets of low-resource language pairs: Myanmar-English, Indonesian-English, and Turkish-English. The experimental results demonstrate that our method outperforms previous methods regarding translation quality. Additionally, our approach is computationally efficient, leading to reduced training time compared to prior works. Van-Hien Tran, Chenchen Ding, Hideki Tanaka, Masao Utiyama |
MTSummit (1) | 4 |
| 2023 | Improving Zero-Shot Dependency Parsing by Unsupervised Learning
Jiannan Mao, Chenchen Ding, Hour Kaing, Hideki Tanaka, Masao Utiyama, Tadahiro Matsumoto |
PACLIC | 5 |
| 2023 | Universal Multimodal Representation for Language UnderstandingabstractRepresentation learning is the foundation of natural language processing (NLP). This work presents new methods to employ visual information as assistant signals to general NLP tasks. For each sentence, we first retrieve a flexible number of images either from a light topic-image lookup table extracted over the existing sentence-image pairs or a shared cross-modal embedding space that is pre-trained on out-of-shelf text-image pairs. Then, the text and images are encoded by a Transformer encoder and convolutional neural network, respectively. The two sequences of representations are further fused by an attention layer for the interaction of the two modalities. In this study, the retrieval process is controllable and flexible. The universal visual representation overcomes the lack of large-scale bilingual sentence-image pairs. Our method can be easily applied to text-only tasks without manually annotated multimodal parallel corpora. We apply the proposed method to a wide range of natural language generation and understanding tasks, including neural machine translation, natural language inference, and semantic similarity. Experimental results show that our method is generally effective for different tasks and languages. Analysis indicates that the visual signals enrich textual representations of content words, provide fine-grained grounding information about the relationship between concepts and events, and potentially conduce to disambiguation. Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Low-resource Multilingual Neural Translation Using Linguistic Feature-based Relevance MechanismsabstractThis article investigates approaches to effectively harness source-side linguistic features for low-resource multilingual neural machine translation (MNMT). Previous works focus on using various features of a word such as lemma, part-of-speech tag, dependency label, and so on, to improve translation quality in a low-resource scenario. However, these studies deal with bilingual translation and do not focus on using features in multilingual training setups. Our work focuses on this particular point and experiments with low-resource multilingual models incorporating source-side linguistic features. Although techniques for integrating features into an NMT model such as concatenation and feature relevance perform quite well in bilingual settings, they do not work well in multilingual settings. To remedy this, we propose the use of dummy features and language indicator features in MNMT models. Experiments are conducted on English to Asian language translation on a multilingual, multi-parallel corpus spanning English and eight Asian languages where for each language pair, the training data size does not exceed 20,000 parallel sentences. After establishing strong bilingual baselines using feature relevance mechanisms and multilingual baselines without any features, we show that our proposed dummy features and language indicator features, in combination with feature relevance mechanisms, yield significant improvements in BLEU points for all language pairs. We then analyze our models from the perspectives of model sizes, the impact of individual linguistic features, validation perplexity computed during training, visualization of the activations of the relevance mechanisms, and exhaustive tuning of hyperparameters. We also report preliminary results for multilingual multi-way models using linguistic features. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Refining History for Future-Aware Neural Machine TranslationabstractNeural machine translation uses a decoder to generate target words auto-regressively by predicting the next target word conditioned on a given source sentence and its previously predicted target words, i.e, its translation history, which suffers from two limitations: 1) the prediction of next word depends heavily on the quality of its history information. Moreover, the discrepancy between training and inference exacerbates this limitation; 2) this left-to-right decoding way cannot make full use of the target-side future information, which leads to the issue of unbalanced outputs. On the one hand, we alleviate the first limitation with a history-refining module, which learns to examine the quality of each history word by assigning it a confidence score. The confidence score is further used as a gate to control the amount of its word embedding flowing to the decoder. On the other hand, we attack the second limitation with a future-foreseeing module, which learns the distribution of future translation at each decoding time step. More importantly, we further propose refining history for future-aware NMT since the two modules can be closely incorporated as they focus on different kinds of context. Experimental results on various translation tasks with different scaled datasets, including WMT English$\leftrightarrow${German, French, Romanian}, show that our proposed approach achieves significant improvements over strong Transformer-based NMT baselines. Xinglin Lyu, Junhui Li 0001, Min Zhang 0005, Chenchen Ding, Hideki Tanaka, Masao Utiyama |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | FeatureBART: Feature Based Sequence-to-Sequence Pre-Training for Low-Resource NMTabstractIn this paper we present FeatureBART, a linguistically motivated sequence-to-sequence monolingual pre-training strategy in which syntactic features such as lemma, part-of-speech and dependency labels are incorporated into the span prediction based pre-training framework (BART). These automatically extracted features are incorporated via approaches such as concatenation and relevance mechanisms, among which the latter is known to be better than the former. When used for low-resource NMT as a downstream task, we show that these feature based models give large improvements in bilingual settings and modest ones in multilingual settings over their counterparts that do not use features. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Eiichiro Sumita |
COLING | 5 |
| 2022 | Effective Graph Context Representation for Document-level Machine TranslationabstractDocument-level neural machine translation (DocNMT) universally encodes several local sentences or the entire document. Thus, DocNMT does not consider the relevance of document-level contextual information, for example, some context (i.e., content words, logical order, and co-occurrence relation) is more effective than another auxiliary context (i.e., functional and auxiliary words). To address this issue, we first utilize the word frequency information to recognize content words in the input document, and then use heuristical relations to summarize content words and sentences as a graph structure without relying on external syntactic knowledge. Furthermore, we apply graph attention networks to this graph structure to learn its feature representation, which allows DocNMT to more effectively capture the document-level context. Experimental results on several widely-used document-level benchmarks demonstrated the effectiveness of the proposed approach. Kehai Chen, Muyun Yang, Masao Utiyama, Eiichiro Sumita, Rui Wang 0015, Min Zhang 0005 |
IJCAI | 3 |
| 2022 | Explicit Alignment Learning for Neural Machine TranslationabstractEven though neural machine translation (NMT) has become the state-of-the-art solution for end-to-end translation, it still suffers from a lack of translation interpretability, which may be conveniently enhanced by explicit alignment learning (EAL), as performed in traditional statistical machine translation (SMT). To provide the benefits of both NMT and SMT, this paper presents a novel model design that enhances NMT with an additional training process for EAL, in addition to the end-to-end translation training. Thus, we propose two approaches an explicit alignment learning approach, in which we further remove the need for the additional alignment model, and perform embedding mixup with the alignment based on encoder--decoder attention weights in the NMT model. We conducted experiments on both small-scale (IWSLT14 De->En and IWSLT13 Fr->En) and large-scale (WMT14 En->De, En->Fr, WMT17 Zh->En) benchmarks. Evaluation results show that our EAL methods significantly outperformed strong baseline methods, which shows the effectiveness of EAL. Further explorations show that the translation improvements are due to a better spatial alignment of the source and target language embeddings. Our method improves translation performance without the need to increase model parameters and training data, which verifies that the idea of incorporating techniques of SMT into NMT is worthwhile. Zuchao Li, Hai Zhao 0001, Fengshun Xiao, Masao Utiyama, Eiichiro Sumita |
IJCAI | 4 |
| 2022 | Text Compression-Aided Transformer EncodingabstractText encoding is one of the most important steps in Natural Language Processing (NLP). It has been done well by the self-attention mechanism in the current state-of-the-art Transformer encoder, which has brought about significant improvements in the performance of many NLP tasks. Though the Transformer encoder may effectively capture general information in its resulting representations, the backbone information, meaning the gist of the input text, is not specifically focused on. In this paper, we propose explicit and implicit text compression approaches to enhance the Transformer encoding and evaluate models using this approach on several typical downstream tasks that rely on the encoding heavily. Our explicit text compression approaches use dedicated models to compress text, while our implicit text compression approach simply adds an additional module to the main model to handle text compression. We propose three ways of integration, namely backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the backbone information into Transformer-based models for various downstream tasks. Our evaluation on benchmark datasets shows that the proposed explicit and implicit text compression approaches improve results in comparison to strong baselines. We therefore conclude, when comparing the encodings to the baseline models, text compression helps the encoders to learn better language representations. Zuchao Li, Zhuosheng Zhang 0001, Hai Zhao 0001, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Integrating Prior Translation Knowledge Into Neural Machine TranslationabstractNeural machine translation (NMT), which is an encoder-decoder joint neural language model with an attention mechanism, has achieved impressive results on various machine translation tasks in the past several years. However, the language model attribute of NMT tends to produce fluent yet sometimes unfaithful translations, which hinders the improvement of translation capacity. In response to this problem, we propose a simple and efficient method to integrate prior translation knowledge into NMT in a universal manner that is compatible with neural networks. Meanwhile, it enables NMT to consider the crossing language translation knowledge from the source-side of the training pipeline of NMT, thereby making full use of the prior translation knowledge to enhance the performance of NMT. The experimental results on two large-scale benchmark translation tasks demonstrated that our approach achieved a significant improvement over a strong baseline. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Which Apple Keeps Which Doctor Away? Colorful Word Representations With Visual OraclesabstractRecent pre-trained language models (PrLMs) offer a new performant method of contextualized word representations by leveraging the sequence-level context for modeling. Although the PrLMs generally provide more effective contextualized word representations than non-contextualized models, they are still subject to a sequence of text contexts without diverse hints from multimodality. This paper thus proposes a visual representation method to explicitly enhance conventional word embedding with multiple-aspect senses from visual guidance. In detail, we build a small-scale word-image dictionary from a multimodal seed dataset where each word corresponds to diverse related images. Experiments on 12 natural language understanding and machine translation tasks further verify the effectiveness and the generalization capability of the proposed approach. Analysis shows that our method with visual guidance pays more attention to content words, improves the representation diversity, and is potentially beneficial for enhancing the accuracy of disambiguation. Zhuosheng Zhang 0001, Haojie Yu, Hai Zhao 0001, Masao Utiyama |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Smoothing Dialogue States for Open Conversational Machine ReadingabstractConversational machine reading (CMR) requires machines to communicate with humans through multi-turn interactions between two salient dialogue states of decision making and question generation processes.In open CMR settings, as the more realistic scenario, the retrieved background knowledge would be noisy, which results in severe challenges in the information transmission.Existing studies commonly train independent or pipeline systems for the two subtasks.However, those methods are trivial by using hard-label decisions to activate question generation, which eventually hinders the model performance.In this work, we propose an effective gating strategy by smoothing the two dialogue states in only one decoder and bridge decision making and question generation to provide a richer dialogue state reference.Experiments on the OR-ShARC dataset show the effectiveness of our method, which achieves new state-of-the-art results. Zhuosheng Zhang 0001, Siru Ouyang, Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita |
EMNLP (1) | 4 |
| 2021 | Unsupervised Neural Machine Translation with Universal GrammarabstractMachine translation usually relies on parallel corpora to provide parallel signals for training.The advent of unsupervised machine translation has brought machine translation away from this reliance, though performance still lags behind traditional supervised machine translation.In unsupervised machine translation, the model seeks symmetric language similarities as a source of weak parallel signal to achieve translation.Chomsky's Universal Grammar theory postulates that grammar is an innate form of knowledge to humans and is governed by universal principles and constraints.Therefore, in this paper, we seek to leverage such shared grammar clues to provide more explicit language parallel signals to enhance the training of unsupervised machine translation models.Through experiments on multiple typical language pairs, we demonstrate the effectiveness of our proposed approaches. Zuchao Li, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP (1) | 2 |
| 2021 | User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical NormalizationabstractShohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shohei Higashiyama, Masao Utiyama, Taro Watanabe, Eiichiro Sumita |
NAACL-HLT | 2 |
| 2021 | Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data ScenariosabstractHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
NAACL-HLT | 4 |
| 2021 | Context-aware positional representation for self-attention networks
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
Neurocomputing | 3 |
| 2021 | Towards Tokenization and Part-of-Speech Tagging for Khmer: Data and DiscussionabstractAs a highly analytic language, Khmer has considerable ambiguities in tokenization and part-of-speech (POS) tagging processing. This topic is investigated in this study. Specifically, a 20,000-sentence Khmer corpus with manual tokenization and POS-tagging annotation is released after a series of work over the last 4 years. This is the largest morphologically annotated Khmer dataset as of 2020, when this article was prepared. Based on the annotated data, experiments were conducted to establish a comprehensive benchmark on the automatic processing of tokenization and POS-tagging for Khmer. Specifically, a support vector machine, a conditional random field (CRF) , a long short-term memory (LSTM) -based recurrent neural network, and an integrated LSTM-CRF model have been investigated and discussed. As a primary conclusion, processing at morpheme-level is satisfactory for the provided data. However, it is intrinsically difficult to identify further grammatical constituents of compounds or phrases because of the complex analytic features of the language. Syntactic annotation and automatic parsing for Khmer will be scheduled in the near future. Hour Kaing, Chenchen Ding, Masao Utiyama, Eiichiro Sumita, Sam Sethserey, Sopheap Seng, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical StudyabstractUnsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines. Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Modeling Future Cost for Neural Machine TranslationabstractExisting neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline. Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Explicit Sentence Compression for Neural Machine TranslationabstractState-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT English-to-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines. Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
AAAI | 4 |
| 2020 | Content Word Aware Neural Machine TranslationabstractNeural machine translation (NMT) encodes the source sentence in a universal way to generate the target sentence word-byword.However, NMT does not consider the importance of word in the sentence meaning, for example, some words (i.e., content words) express more important meaning than others (i.e., function words).To address this limitation, we first utilize word frequency information to distinguish between content and function words in a sentence, and then design a content word-aware NMT to improve translation performance.Empirical results on the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chineseto-English translation tasks show that the proposed methods can significantly improve the performance of Transformer-based NMT. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL | 3 |
| 2020 | A Three-Parameter Rank-Frequency Relation in Natural LanguagesabstractWe present that, the rank-frequency relation in textual data follows f ∝ r -α (r + γ) -β , where f is the token frequency and r is the rank by frequency, with (α, β, γ) as parameters.The formulation is derived based on the empirical observation that d 2 (x+y)/dx 2 is a typical impulse function, where (x, y) = (log r, log f ).The formulation is the power law when β = 0 and the Zipf-Mandelbrot law when α = 0. We illustrate that α is related to the analytic features of syntax and β + γ to those of morphology in natural languages from an investigation of multilingual corpora. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACL | 2 |
| 2020 | Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL | 4 |
| 2020 | Improving Low-Resource NMT through Relevance Based Linguistic Features IncorporationabstractIn this study, linguistic knowledge at different levels are incorporated into the neural machine translation (NMT) framework to improve translation quality for language pairs with extremely limited data.Integrating manually designed or automatically extracted features into the NMT framework is known to be beneficial.However, this study emphasizes that the relevance of the features is crucial to the performance.Specifically, we propose two methods, 1) self relevance and 2) word-based relevance, to improve the representation of features for NMT.Experiments are conducted on translation tasks from English to eight Asian languages, with no more than twenty thousand sentences for training.The proposed methods improve translation quality for all tasks by up to 3.09 BLEU points.Discussions with visualization provide the explainability of the proposed methods where we show that the relevance methods provide weights to features thereby enhancing their impact on low-resource machine translation. Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
COLING | 4 |
| 2020 | Bilingual Subword Segmentation for Neural Machine TranslationabstractThis paper proposed a new subword segmentation method for neural machine translation, "Bilingual Subword Segmentation," which tokenizes sentences to minimize the difference between the number of subword units in a sentence and that of its translation.While existing subword segmentation methods tokenize a sentence without considering its translation, the proposed method tokenizes a sentence by using subword units induced from bilingual sentences; this method could be more favorable to machine translation.Evaluations on WAT Asian Scientific Paper Excerpt Corpus (ASPEC) English-to-Japanese and Japanese-to-English translation tasks and WMT14 English-to-German and German-to-English translation tasks show that our bilingual subword segmentation improves the performance of Transformer neural machine translation (up to +0.81 BLEU). Hiroyuki Deguchi 0002, Masao Utiyama, Akihiro Tamura, Takashi Ninomiya, Eiichiro Sumita |
COLING | 2 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 5 |
| 2020 | Neural Machine Translation with Universal Visual Representation
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
ICLR | 4 |
| 2020 | Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ICLR | 4 |
| 2020 | A Myanmar (Burmese)-English Named Entity Transliteration DictionaryabstractTransliteration is generally a phonetically based transcription across different writing systems. It is a crucial task for various downstream natural language processing applications. For the Myanmar (Burmese) language, robust automatic transliteration for borrowed English words is a challenging task because of the complex Myanmar writing system and the lack of data. In this study, we constructed a Myanmar-English named entity dictionary containing more than eighty thousand transliteration instances. The data have been released under a CC BY-NC-SA license. We evaluated the automatic transliteration performance using statistical and neural network-based approaches based on the prepared data. The neural network model outperformed the statistical model significantly in terms of the BLEU score on the character level. Different units used in the Myanmar script for processing were also compared and discussed. Aye Myat Mon, Chenchen Ding, Hour Kaing, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
LREC | 5 |
| 2020 | Agreement on Target-Bidirectional Recurrent Neural Networks for Sequence-to-Sequence LearningabstractRecurrent neural networks are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a shortcoming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus performance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional RNNs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of either sequence level or non-sequence level metrics. Extensive experiments were performed on three standard sequence-to-sequence transduction tasks: machine transliteration, grapheme-to-phoneme transformation and machine translation. The results show that the proposed approach achieves consistent and substantial improvements, compared to many state-of-the-art systems. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
J. Artif. Intell. Res. | 3 |
| 2020 | Extremely low-resource neural machine translation for Asian languagesabstractAbstract This paper presents a set of effective approaches to handle extremely low-resource language pairs for self-attention based neural machine translation (NMT) focusing on English and four Asian languages. Starting from an initial set of parallel sentences used to train bilingual baseline models, we introduce additional monolingual corpora and data processing techniques to improve translation quality. We describe a series of best practices and empirically validate the methods through an evaluation conducted on eight translation directions, based on state-of-the-art NMT approaches such as hyper-parameter search, data augmentation with forward and backward translation in combination with tags and noise, as well as joint multilingual training. Experiments show that the commonly used default architecture of self-attention NMT models does not reach the best results, validating previous work on the importance of hyper-parameter tuning. Additionally, empirical results indicate the amount of synthetic data required to efficiently increase the parameters of the models leading to the best translation quality measured by automatic metrics. We show that the best NMT models trained on large amount of tagged back-translations outperform three other synthetic data generation approaches. Finally, comparison with statistical machine translation (SMT) indicates that extremely low-resource NMT requires a large amount of synthetic parallel data obtained with back-translation in order to close the performance gap with the preceding SMT approach. Raphaël Rubino, Benjamin Marie, Raj Dabre, Atsushi Fujita, Masao Utiyama, Eiichiro Sumita |
Mach. Transl. | 5 |
| 2020 | Improving neural machine translation through phrase-based soft forced decoding
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 2 |
| 2020 | Towards Burmese (Myanmar) Morphological Analysis: Syllable-based Tokenization and Part-of-speech TaggingabstractThis article presents a comprehensive study on two primary tasks in Burmese (Myanmar) morphological analysis: tokenization and part-of-speech (POS) tagging. Twenty thousand Burmese sentences of newswire are annotated with two-layer tokenization and POS-tagging information, as one component of the Asian Language Treebank Project. The annotated corpus has been released under a CC BY-NC-SA license, and it is the largest open-access database of annotated Burmese when this manuscript was prepared in 2017. Detailed descriptions of the preparation, refinement, and features of the annotated corpus are provided in the first half of the article. Facilitated by the annotated corpus, experiment-based investigations are presented in the second half of the article, wherein the standard sequence-labeling approach of conditional random fields and a long short-term memory (LSTM)-based recurrent neural network (RNN) are applied and discussed. We obtained several general conclusions, covering the effect of joint tokenization and POS-tagging and importance of ensemble from the viewpoint of stabilizing the performance of LSTM-based RNN. This study provides a solid basis for further studies on Burmese processing. Chenchen Ding, Hnin Thu Zar Aye, Win Pa Pa, Khin Thandar Nwet, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2020 | A Burmese (Myanmar) Treebank: Guideline and AnalysisabstractA 20,000-sentence Burmese (Myanmar) treebank on news articles has been released under a CC BY-NC-SA license. Complete phrase structure annotation was developed for each sentence from the morphologically annotated data prepared in previous work of Ding et al. [1]. As the final result of the Burmese component in the Asian Language Treebank Project , this is the first large-scale, open-access treebank for the Burmese language. The annotation details and features of this treebank are presented. Chenchen Ding, Sann Su Su Yee, Win Pa Pa, Khin Mar Soe, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2020 | Towards More Diverse Input Representation for Neural Machine TranslationabstractSource input information plays a very important role in the Transformer-based translation system. In practice, word embedding and positional embedding of each word are added as the input representation. Then self-attention networks are used to encode the global dependencies in the input representation to generate a source representation. However, this processing on the source representation only adopts a single source feature and excludes richer and more diverse features such as recurrence features, local features, and syntactic features, which results in tedious representation and thereby hinders the further translation performance improvement. In this paper, we introduce a simple and efficient method to encode more diverse source features into the input representation simultaneously, and thereby learning an effective source representation by self-attention networks. In particular, the proposed grouped strategy is only applied to the input representation layer, to keep the diversity of translation information and the efficiency of the self-attention networks at the same time. Experimental results show that our approach improves the translation performance over the state-of-the-art baselines of Transformer in regard to WMT14 English-to-German and NIST Chinese-to-English machine translation tasks. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao, Muyun Yang, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Unsupervised Neural Machine Translation With Cross-Lingual Language Representation AgreementabstractUnsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Neural Machine Translation with Reordering EmbeddingsabstractThe reordering model plays an important role in phrase-based statistical machine translation.However, there are few works that exploit the reordering information in neural machine translation.In this paper, we propose a reordering mechanism to learn the reordering embedding of a word based on its contextual information.These reordering embeddings are stacked together with self-attention networks to learn sentence representation for machine translation.The reordering mechanism can be easily integrated into both the encoder and the decoder in the Transformer translation system.Experimental results on WMT'14 English-to-German, NIST Chinese-to-English, and WAT ASPEC Japanese-to-English translation tasks demonstrate that the proposed methods can significantly improve the performance of the Transformer translation system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL (1) | 3 |
| 2019 | Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine TranslationabstractUnsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 4 |
| 2019 | Sentence-Level Agreement for Neural Machine TranslationabstractThe training objective of neural machine translation (NMT) is to minimize the loss between the words in the translated sentences and those in the references. In NMT, there is a natural correspondence between the source sentence and the target sentence. However, this relationship has only been represented using the entire neural network and the training objective is computed in word-level. In this paper, we propose a sentence-level agreement module to directly minimize the difference between the representation of source and target sentence. The proposed agreement module can be integrated into NMT as an additional training objective function and can also be used to enhance the representation of the source sentences. Empirical results on the NIST Chinese-to-English and WMT English-to-German tasks show the proposed agreement module can significantly improve the NMT performance. Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Min Zhang 0005, Tiejun Zhao |
ACL (1) | 4 |
| 2019 | Recurrent Positional Embedding for Neural Machine TranslationabstractKehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Online Sentence Segmentation for Simultaneous Interpretation using Multi-Shifted Recurrent Neural Network
Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 2 |
| 2019 | NOVA: A Feasible and Flexible Annotation System for Joint Tokenization and Part-of-Speech TaggingabstractA feasible and flexible annotation system is designed for joint tokenization and part-of-speech (POS) tagging to annotate those languages without natural definitions of words . This design was motivated by the fact that word separators are not used in many highly analytic East and Southeast Asian languages. Although several of the languages are well-studied, e.g., Chinese and Japanese, many are understudied with low resources, e.g., Burmese (Myanmar) and Khmer. In the first part of the article, the proposed annotation system, named nova, is introduced. nova contains only four basic tags (n, v, a, and o); these tags can be further modified and combined to adapt complex linguistic phenomena in tokenization and POS tagging. In the second part of the article, the feasibility and flexibility of nova is illustrated from the annotation practice on Burmese and Khmer. The relation between nova and two universal POS tagsets is discussed in the final part of the article. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Neural Machine Translation With Sentence-Level Topic ContextabstractTraditional neural machine translation (NMT) methods use the word-level context to predict target language translation while neglecting the sentence-level context, which has been shown to be beneficial for translation prediction in statistical machine translation. This paper represents the sentence-level context as latent topic representations by using a convolution neural network, and designs a topic attention to integrate source sentence-level topic context information into both attention-based and Transformer-based NMT. In particular, our method can improve the performance of NMT by modeling source topics and translations jointly. Experiments on the large-scale LDC Chinese-to-English translation tasks and WMT'14 English-to-German translation tasks show that the proposed approach can achieve significant improvements compared with baseline systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Syntax-Directed Attention for Neural Machine TranslationabstractAttention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
AAAI | 3 |
| 2018 | Forest-Based Neural Machine TranslationabstractTree-based neural machine translation (NMT) approaches, although achieved impressive performance, suffer from a major drawback: they only use the 1best parse tree to direct the translation, which potentially introduces translation mistakes due to parsing errors.For statistical machine translation (SMT), forestbased methods have been proven to be effective for solving this problem, while for NMT this kind of approach has not been attempted.This paper proposes a forest-based NMT method that translates a linearized packed forest under a simple sequence-to-sequence framework (i.e., a forest-to-string NMT model).The BLEU score of the proposed method is higher than that of the string-to-string NMT, treebased NMT, and forest-based SMT systems. Chunpeng Ma, Akihiro Tamura, Masao Utiyama, Tiejun Zhao, Eiichiro Sumita |
ACL (1) | 3 |
| 2018 | Exploring Recombination for Efficient Decoding of Neural Machine TranslationabstractIn Neural Machine Translation (NMT), the decoder can capture the features of the entire prediction history with neural connections and representations.This means that partial hypotheses with different prefixes will be regarded differently no matter how similar they are.However, this might be inefficient since some partial hypotheses can contain only local differences that will not influence future predictions.In this work, we introduce recombination in NMT decoding based on the concept of the "equivalence" of partial hypotheses.Heuristically, we use a simple n-gram suffix based equivalence function and adapt it into beam search decoding.Through experiments on large-scale Chinese-to-English and English-to-Germen translation tasks, we show that the proposed method can obtain similar translation quality with a smaller beam size, making NMT decoding more efficient. Zhisong Zhang, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 3 |
| 2018 | Guiding Neural Machine Translation with Retrieved Translation PiecesabstractJingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, Satoshi Nakamura. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
NAACL-HLT | 2 |
| 2018 | Graph-Based Bilingual Word Embedding for Statistical Machine TranslationabstractBilingual word embedding has been shown to be helpful for Statistical Machine Translation (SMT). However, most existing methods suffer from two obvious drawbacks. First, they only focus on simple contexts such as an entire document or a fixed-sized sliding window to build word embedding and ignore latent useful information from the selected context. Second, the word sense but not the word should be the minimal semantic unit; however, most existing methods still use word representation. To overcome these drawbacks, this article presents a novel Graph-Based Bilingual Word Embedding (GBWE) method that projects bilingual word senses into a multidimensional semantic space. First, a bilingual word co-occurrence graph is constructed using the co-occurrence and pointwise mutual information between the words. Then, maximum complete subgraphs (cliques), which play the role of a minimal unit for bilingual sense representation, are dynamically extracted according to the contextual information. Consequently, correspondence analysis, principal component analyses, and neural networks are used to summarize the clique-word matrix into lower dimensions to build the embedding model. Without contextual information, the proposed GBWE can be applied to lexical translation. In addition, given contextual information, GBWE is able to give a dynamic solution for bilingual word representations, which can be applied to phrase translation and generation. Empirical results show that GBWE can enhance the performance of lexical translation, as well as Chinese/French-to-English and Chinese-to-Japanese phrase-based SMT tasks (IWSLT, NTCIR, NIST, and WAT). Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2018 | A Neural Approach to Source Dependence Based Context Model for Statistical Machine TranslationabstractIn statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods. Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain AdaptationabstractNeural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines. Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Neural Machine Translation with Source Dependency RepresentationabstractSource dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao |
EMNLP | 3 |
| 2017 | Instance Weighting for Neural Machine Translation Domain AdaptationabstractInstance weighting has been widely applied to phrase-based machine translation domain adaptation.However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.In this paper, two instance weighting technologies, i.e., sentence weighting and domain weighting with a dynamic weight learning strategy, are proposed for NMT domain adaptation.Empirical results on the IWSLT English-German/French tasks show that the proposed methods can substantially improve NMT performance by up to 2.7-6.7 BLEU points, outperforming the existing baselines by up to 1.6-3.6BLEU points. Rui Wang 0015, Masao Utiyama, Lemao Liu, Kehai Chen, Eiichiro Sumita |
EMNLP | 2 |
| 2017 | Context-Aware Smoothing for Neural Machine TranslationabstractIn Neural Machine Translation (NMT), each word is represented as a low-dimension, real-value vector for encoding its syntax and semantic information. This means that even if the word is in a different sentence context, it is represented as the fixed vector to learn source representation. Moreover, a large number of Out-Of-Vocabulary (OOV) words, which have different syntax and semantic information, are represented as the same vector representation of “unk”. To alleviate this problem, we propose a novel context-aware smoothing method to dynamically learn a sentence-specific vector for each word (including OOV words) depending on its local context words in a sentence. The learned context-aware representation is integrated into the NMT to improve the translation performance. Empirical results on NIST Chinese-to-English translation task show that the proposed approach achieves 1.78 BLEU improvements on average over a strong attentional NMT, and outperforms some existing systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IJCNLP(1) | 3 |
| 2017 | Improving Neural Machine Translation through Phrase-based Forced DecodingabstractCompared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an existing phrase-based SMT model to compute the phrase-based decoding cost for an NMT output and then using the phrase-based decoding cost to rerank the n-best NMT outputs. The main challenge in implementing this approach is that NMT outputs may not be in the search space of the standard phrase-based decoding algorithm, because the search space of phrase-based SMT is limited by the phrase-based translation rule table. We propose a soft forced decoding algorithm, which can always successfully find a decoding path for any NMT output. We show that using the forced decoding cost to rerank the NMT outputs can successfully improve translation quality on four different language pairs. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
IJCNLP(1) | 2 |
| 2017 | Empirical Study of Dropout Scheme for Neural Machine Translation
Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
MTSummit (1) | 2 |
| 2017 | Translation Quality Estimation Using Only Bilingual CorporaabstractIn computer-aided translation scenarios, quality estimation of machine translation hypotheses plays a critical role. Existing methods for word-level translation quality estimation (TQE) rely on the availability of manually annotated TQE training data obtained via direct annotation or postediting. However, due to the cost of human labor, such data are either limited in size or is only available for few tasks in practice. To avoid the reliance on such annotated TQE data, this paper proposes an approach to train word-level TQE models using bilingual corpora, which are typically used in machine translation training and is relatively easier to access. We formalize the training of our proposed method under the framework of maximum marginal likelihood estimation. To avoid degenerated solutions, we propose a novel regularized training objective whose optimization is achieved by an efficient approximation. Extensive experiments on both written and spoken language datasets empirically show that our approach yields comparable performance to the standard training on annotated data. Lemao Liu, Atsushi Fujita, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence LearningabstractRecurrent neural networks, particularly the long short- term memory networks, are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a fundamental short- coming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus perfor- mance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional LSTMs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of sequence-level losses. Extensive experiments were performed on two standard sequence-to-sequence trans- duction tasks: machine transliteration and grapheme-to- phoneme transformation. The results show that the proposed approach achieves consistent and substantial im- provements, compared to six state-of-the-art systems. In particular, our approach outperforms the best reported error rates by a margin (up to 9% relative gains) on the grapheme-to-phoneme task. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
AAAI | 3 |
| 2016 | A Continuous Space Rule Selection Model for Syntax-based Statistical Machine Translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
ACL (1) | 2 |
| 2016 | Neural Machine Translation with Supervised AttentionabstractThe attention mechanism is appealing for neural machine translation, since it is able to dynamically encode a source sentence by generating a alignment between a target word and source words. Unfortunately, it has been proved to be worse than conventional alignment models in alignment accuracy. In this paper, we analyze and explain this issue from the point view of reordering, and propose a supervised attention which is learned with guidance from conventional alignment models. Experiments on two Chinese-to-English translation tasks show that the supervised attention mechanism yields better alignments leading to substantial gains over the standard attention based NMT. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
COLING | 2 |
| 2016 | Connecting Phrase based Statistical Machine Translation AdaptationabstractAlthough more additional corpora are now available for Statistical Machine Translation (SMT), only the ones which belong to the same or similar domains of the original corpus can indeed enhance SMT performance directly. A series of SMT adaptation methods have been proposed to select these similar-domain data, and most of them focus on sentence selection. In comparison, phrase is a smaller and more fine grained unit for data selection, therefore we propose a straightforward and efficient connecting phrase based adaptation method, which is applied to both bilingual phrase pair and monolingual n-gram adaptation. The proposed method is evaluated on IWSLT/NIST data sets, and the results show that phrase based SMT performances are significantly improved (up to +1.6 in comparison with phrase based SMT baseline system and +0.9 in comparison with existing methods). Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
COLING | 4 |
| 2016 | Assessing Translation Ability through Vocabulary Ability Assessment
Yo Ehara, Yukino Baba, Masao Utiyama, Eiichiro Sumita |
IJCAI | 3 |
| 2016 | A Bilingual Graph-Based Semantic Model for Statistical Machine Translation
Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama |
IJCAI | 5 |
| 2016 | ASPEC: Asian Scientific Paper Excerpt Corpus
Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, Hitoshi Isahara |
LREC | 4 |
| 2016 | Introducing the Asian Language Treebank (ALT)
Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
LREC | 3 |
| 2016 | Agreement on Target-bidirectional Neural Machine TranslationabstractLemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
HLT-NAACL | 2 |
| 2016 | Learning local word reorderings for hierarchical phrase-based statistical machine translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 2 |
| 2016 | Word Segmentation for Burmese (Myanmar)abstractExperiments on various word segmentation approaches for the Burmese language are conducted and discussed in this note. Specifically, dictionary-based, statistical, and machine learning approaches are tested. Experimental results demonstrate that statistical and machine learning approaches perform significantly better than dictionary-based approaches. We believe that this note, based on an annotated corpus of relatively considerable size (containing approximately a half million words), is the first systematic comparison of word segmentation approaches for Burmese. This work aims to discover the properties and proper approaches to Burmese textual processing and to promote further researches on this understudied language. Chenchen Ding, Ye Kyaw Thu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | Converting Continuous-Space Language Models into N-gram Language Models with Efficient Bilingual Pruning for Statistical Machine TranslationabstractThe Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N -grams based Converting (NNGC) method for transforming a Continuous-Space Language Model (CSLM) to a Back-off N -gram Language Model (BNLM). Furthermore, a Bilingual LM Pruning (BLMP) approach is developed for enhancing LMs in SMT decoding and speeding up CSLM converting. The proposed pruning and converting methods can convert a large LM efficiently by working jointly. That is, a LM can be effectively pruned before it is converted from CSLM without sacrificing performance, and further improved if an additional corpus contains out-of-domain information. For different SMT tasks, our experimental results indicate that the proposed NNGC and BLMP methods outperform the existing counterpart approaches significantly in BLEU and computational cost. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | MNH-TT: A Platform to Support Collaborative Translator Training
Masao Utiyama, Kyo Kageura, Martin Thomas, Anthony Hartley |
EAMT | 1 |
| 2015 | Improving fast_align by Reorderingabstractfast align is a simple, fast, and efficient approach for word alignment based on the IBM model 2. fast align performs well for language pairs with relatively similar word orders; however, it does not perform well for language pairs with drastically different word orders.We propose a segmenting-reversing reordering process to solve this problem by alternately applying fast align and reordering source sentences during training.Experimental results with Japanese-English translation demonstrate that the proposed approach improves the performance of fast align significantly without the loss of efficiency.Experiments using other languages are also reported. Chenchen Ding, Masao Utiyama, Eiichiro Sumita |
EMNLP | 2 |
| 2015 | Hierarchical Phrase-based Stream DecodingabstractThis paper proposes a method for hierarchical phrase-based stream decoding.A stream decoder is able to take a continuous stream of tokens as input, and segments this stream into word sequences that are translated and output as a stream of target word sequences.Phrase-based stream decoding techniques have been shown to be effective as a means of simultaneous interpretation.In this paper we transfer the essence of this idea into the framework of hierarchical machine translation.The hierarchical decoding framework organizes the decoding process into a chart; this structure is naturally suited to the process of stream decoding, leading to an efficient stream decoding algorithm that searches a restricted subspace containing only relevant hypotheses.Furthermore, the decoder allows more explicit access to the word re-ordering process that is of critical importance in decoding while interpreting.The decoder was evaluated on TED talk data for English-Spanish and English-Chinese.Our results show that like the phrase-based stream decoder, the hierarchical is capable of approaching the performance of the underlying hierarchical phrase-based machine translation decoder, at useful levels of latency.In addition the hierarchical approach appeared to be robust to the difficulties presented by the more challenging English-Chinese task. Andrew M. Finch, Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
EMNLP | 3 |
| 2015 | Leave-one-out Word Alignment without Garbage Collector EffectsabstractExpectation-maximization algorithms, such as those implemented in GIZA++ pervade the field of unsupervised word alignment.However, these algorithms have a problem of over-fitting, leading to "garbage collector effects," where rare words tend to be erroneously aligned to untranslated words.This paper proposes a leave-one-out expectationmaximization algorithm for unsupervised word alignment to address this problem.The proposed method excludes information derived from the alignment of a sentence pair from the alignment models used to align it.This prevents erroneous alignments within a sentence pair from supporting themselves.Experimental results on Chinese-English and Japanese-English corpora show that the F 1 , precision and recall of alignment were consistently increased by 5.0% -17.2%, and BLEU scores of end-to-end translation were raised by 0.03 -1.30.The proposed method also outperformed l 0 -normalized GIZA++ and Kneser-Ney smoothed GIZA++. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Taro Watanabe, Eiichiro Sumita |
EMNLP | 2 |
| 2015 | A Binarized Neural Network Joint Model for Machine TranslationabstractThe neural network joint model (NNJM), which augments the neural network language model (NNLM) with an m-word source context window, has achieved large gains in machine translation accuracy, but also has problems with high normalization cost when using large vocabularies.Training the NNJM with noise-contrastive estimation (NCE), instead of standard maximum likelihood estimation (MLE), can reduce computation cost.In this paper, we propose an alternative to NCE, the binarized NNJM (BNNJM), which learns a binary classifier that takes both the context and target words as input, and can be efficiently trained using MLE.We compare the BNNJM and NNJM trained by NCE on various translation tasks. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 2 |
| 2015 | Patent claim translation based on sublanguage-specific sentence structure
Masaru Fuji, Atsushi Fujita, Masao Utiyama, Eiichiro Sumita, Yuji Matsumoto 0001 |
MTSummit | 3 |
| 2015 | A Large-scale Study of Statistical Machine Translation Methods for Khmer Language
Ye Kyaw Thu, Vichet Chea, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
PACLIC | 4 |
| 2015 | Preordering using a Target-Language Parser via Cross-Language Syntactic Projection for Statistical Machine TranslationabstractWhen translating between languages with widely different word orders, word reordering can present a major challenge. Although some word reordering methods do not employ source-language syntactic structures, such structures are inherently useful for word reordering. However, high-quality syntactic parsers are not available for many languages. We propose a preordering method using a target-language syntactic parser to process source-language syntactic structures without a source-language syntactic parser. To train our preordering model based on ITG, we produced syntactic constituent structures for source-language training sentences by (1) parsing target-language training sentences, (2) projecting constituent structures of the target-language sentences to the corresponding source-language sentences, (3) selecting parallel sentences with highly synchronized parallel structures, (4) producing probabilistic models for parsing using the projected partial structures and the Pitman-Yor process, and (5) parsing to produce full binary syntactic structures maximally synchronized with the corresponding target-language syntactic structures, using the constraints of the projected partial structures and the probabilistic models. Our ITG-based preordering model is trained using the produced binary syntactic structures and word alignments. The proposed method facilitates the learning of ITG by producing highly synchronized parallel syntactic structures based on cross-language syntactic projection and sentence selection. The preordering model jointly parses input sentences and identifies their reordered structures. Experiments with Japanese--English and Chinese--English patent translation indicate that our method outperforms existing methods, including string-to-tree syntax-based SMT, a preordering method that does not require a parser, and a preordering method that uses a source-language dependency parser. Isao Goto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Bilingual Continuous-Space Language Model Growing for Statistical Machine TranslationabstractLarger n-gram language models (LMs) perform better in statistical machine translation (SMT). However, the existing approaches have two main drawbacks for constructing larger LMs: 1) it is not convenient to obtain larger corpora in the same domain as the bilingual parallel corpora in SMT; 2) most of the previous studies focus on monolingual information from the target corpora only, and redundant n-grams have not been fully utilized in SMT. Nowadays, continuous-space language model (CSLM), especially neural network language model (NNLM), has been shown great improvement in the estimation accuracies of the probabilities for predicting the target words. However, most of these CSLM and NNLM approaches still consider monolingual information only or require additional corpus. In this paper, we propose a novel neural network based bilingual LM growing method. Compared to the existing approaches, the proposed method enables us to use bilingual parallel corpus for LM growing in SMT. The results show that our new method outperforms the existing approaches on both SMT performance and computational efficiency significantly. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine TranslationabstractLanguages that have no explicit word delimiters often have to be segmented for statistical machine translation (SMT).This is commonly performed by automated segmenters trained on manually annotated corpora.However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT.An analysis was performed to test this hypothesis using a manually annotated word alignment (WA) corpus for Chinese-English SMT.An analysis revealed that 74.60% of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA annotations.We formulated an approach based on word splitting with reference to the annotated WA to alleviate these conflicts.Experimental results show that the refined WS reduced word alignment error rate by 6.82% and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine translation (OpenMT) corpora compared to related work. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
EMNLP | 2 |
| 2014 | Neural Network Based Bilingual Language Model Growing for Statistical Machine TranslationabstractSince larger n-gram Language Model (LM) usually performs better in Statistical Machine Translation (SMT), how to construct efficient large LM is an important topic in SMT.However, most of the existing LM growing methods need an extra monolingual corpus, where additional LM adaption technology is necessary.In this paper, we propose a novel neural network based bilingual LM growing method, only using the bilingual parallel corpus in SMT.The results show that our method can improve both the perplexity score for LM evaluation and BLEU score for SMT, and significantly outperforms the existing LM growing methods without extra corpus. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
EMNLP | 4 |
| 2014 | Learning Hierarchical Translation SpansabstractWe propose a simple and effective approach to learn translation spans for the hierarchical phrase-based translation model.Our model evaluates if a source span should be covered by translation rules during decoding, which is integrated into the translation system as soft constraints.Compared to syntactic constraints, our model is directly acquired from an aligned parallel corpus and does not require parsers.Rich source side contextual features and advanced machine learning methods were utilized for this learning task.The proposed approach was evaluated on NTCIR-9 Chinese-English and Japanese-English translation tasks and showed significant improvement over the baseline system. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 2 |
| 2014 | Distortion Model Based on Word Sequence Labeling for Statistical Machine TranslationabstractThis article proposes a new distortion model for phrase-based statistical machine translation. In decoding, a distortion model estimates the source word position to be translated next (subsequent position; SP) given the last translated source word position (current position; CP). We propose a distortion model that can simultaneously consider the word at the CP, the word at an SP candidate, the context of the CP and an SP candidate, relative word order among the SP candidates, and the words between the CP and an SP candidate. These considered elements are called rich context . Our model considers rich context by discriminating label sequences that specify spans from the CP to each SP candidate. It enables our model to learn the effect of relative word order among SP candidates as well as to learn the effect of distances from the training data. In contrast to the learning strategy of existing methods, our learning strategy is that the model learns preference relations among SP candidates in each sentence of the training data. This leaning strategy enables consideration of all of the rich context simultaneously. In our experiments, our model had higher BLUE and RIBES scores for Japanese-English, Chinese-English, and German-English translation compared to the lexical reordering models. Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2013 | Distortion Model Considering Rich Context for Statistical Machine Translation
Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACL (1) | 2 |
| 2013 | An Empirical Study on Word Segmentation for Chinese Machine Translation
Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita, Bao-Liang Lu |
CICLing (2) | 2 |
| 2013 | Converting Continuous-Space Language Models into N-Gram Language Models for Statistical Machine TranslationabstractNeural network language models, or continuous-space language models (CSLMs), have been shown to improve the performance of statistical machine translation (SMT) when they are used for reranking n-best translations.However, CSLMs have not been used in the first pass decoding of SMT, because using CSLMs in decoding takes a lot of time.In contrast, we propose a method for converting CSLMs into back-off n-gram language models (BNLMs) so that we can use converted CSLMs in decoding.We show that they outperform the original BNLMs and are comparable with the traditional use of CSLMs in reranking. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
EMNLP | 2 |
| 2013 | Post-Ordering by Parsing with ITG for Japanese-English Statistical Machine TranslationabstractWord reordering is a difficult task for translation between languages with widely different word orders, such as Japanese and English. A previously proposed post-ordering method for Japanese-to-English translation first translates a Japanese sentence into a sequence of English words in a word order similar to that of Japanese, then reorders the sequence into an English word order. We employed this post-ordering framework and improved upon its reordering method. The existing post-ordering method reorders the sequence of English words via SMT, whereas our method reorders the sequence by (1) parsing the sequence using ITG to obtain syntactic structures which are similar to Japanese syntactic structures, and (2) transferring the obtained syntactic structures into English syntactic structures according to the ITG. The experiments using Japanese-to-English patent translation demonstrated the effectiveness of our method and showed that both the RIBES and BLEU scores were improved over compared methods. Isao Goto, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2011 | A Comparison Study of Parsers for Patent Machine Translation
Isao Goto, Masao Utiyama, Takashi Onishi, Eiichiro Sumita |
MTSummit | 2 |
| 2011 | A Comparison of Unsupervised Bilingual Term Extraction Methods Using Phrase-Tables
Masamichi Ideue, Kazuhide Yamamoto, Masao Utiyama, Eiichiro Sumita |
MTSummit | 3 |
| 2011 | Searching Translation Memories for Paraphrases
Masao Utiyama, Graham Neubig, Takashi Onishi, Eiichiro Sumita |
MTSummit | 1 |
| 2010 | Community-based Construction of Draft and Final Translation Corpus Through a Translation Hosting Site Minna no Hon'yaku (MNH)
Takeshi Abekawa, Masao Utiyama, Eiichiro Sumita, Kyo Kageura |
LREC | 2 |
| 2009 | Incorporating Prior Knowledge into Task Decomposition for Large-Scale Patent Classification
Bao-Liang Lu, Masao Utiyama |
ISNN (2) | 3 |
| 2009 | Mining Parallel Texts from Mixed-Language Web Pages
Masao Utiyama, Daisuke Kawahara, Keiji Yasuda, Eiichiro Sumita |
MTSummit | 1 |
| 2009 | Evaluating effects of machine translation accuracy on cross-lingual patent retrievalabstractWe organized a machine translation (MT) task at the Seventh NTCIR Workshop. Participating groups were requested to machine translate sentences in patent documents and also search topics for retrieving patent documents across languages. We analyzed the relationship between the accuracy of MT and its effects on the retrieval accuracy. Atsushi Fujii, Masao Utiyama, Mikio Yamamoto, Takehito Utsuro |
SIGIR | 2 |
| 2009 | Incorporating prior knowledge into learning by dividing training data
Bao-Liang Lu, Xiaolin Wang 0002, Masao Utiyama |
Frontiers Comput. Sci. China | 3 |
| 2008 | Large-scale patent classification with min-max modular support vector machinesabstractPatent classification is a large-scale, hierarchical, imbalanced, multi-label problem. The number of samples in a real-world patent classification typically exceeds one million, and this number increases every year. An effective patent classifier must be able to deal with this situation. This paper discusses the use of min-max modular support vector machine (M3-SVM) to deal with large-scale patent classification problems. The method includes three steps: decomposing a large-scale and imbalanced patent classification problem into a group of relatively smaller and more balanced two-class subproblems which are independent of each other, learning these subproblems using support vector machines (SVMs) in parallel, and combining all of the trained SVMs according to the minimization and the maximization rules. M3-SVM has two attractive features which are urgently needed to deal with large-scale patent classification problems. First, it can be realized in a massively parallel form. Second, it can be built up incrementally. Results from experiments using the NTCIR-5 patent data set, which contains more than two million patents, have confirmed these two attractive features, and demonstrate that M3-SVM outperforms conventional SVMs in terms of both training time and generalization performance. Xiao-Lei Chu, Jing Li 0137, Bao-Liang Lu, Masao Utiyama, Hitoshi Isahara |
IJCNN | 5 |
| 2008 | Producing a Test Collection for Patent Machine Translation in the Seventh NTCIR Workshop
Atsushi Fujii, Masao Utiyama, Mikio Yamamoto, Takehito Utsuro |
LREC | 2 |
| 2008 | Development of the Japanese WordNet
Hitoshi Isahara, Francis Bond, Kiyotaka Uchimoto, Masao Utiyama, Kyoko Kanzaki |
LREC | 4 |
| 2008 | Application of Resource-based Machine Translation to Real Business Scenes
Hitoshi Isahara, Masao Utiyama, Eiko Yamamoto, Akira Terada, Yasunori Abe |
LREC | 2 |
| 2008 | An empirical comparison of min-max-modular k -NN with different voting methods to large-scale text categorization
Bao-Liang Lu, Masao Utiyama, Hitoshi Isahara |
Soft Comput. | 3 |
| 2007 | A Japanese-English patent parallel corpus
Masao Utiyama, Hitoshi Isahara |
MTSummit | 1 |
| 2007 | A Comparison of Pivot Methods for Phrase-Based Statistical Machine Translation
Masao Utiyama, Hitoshi Isahara |
HLT-NAACL | 1 |
| 2006 | Relevance Feedback Models for Recommendation
Masao Utiyama, Mikio Yamamoto |
EMNLP | 1 |
| 2006 | Getting Deeper Semantics than Berkeley FrameNet with MSFA
Kow Kuroda, Masao Utiyama, Hitoshi Isahara |
LREC | 2 |
| 2005 | Organizing English Reading Materials for Vocabulary Learning
Masao Utiyama, Midori Tanimura, Hitoshi Isahara |
ACL | 1 |
| 2005 | Correction of errors in a verb modality corpus for machine translation with a machine-learning methodabstractIn recent years, various types of tagged corpora have been constructed and much research using tagged corpora has been done. However, tagged corpora contain errors, which impedes the progress of research. Therefore, the correction of errors in corpora is an important research issue. In this study we investigate the correction of such errors, which we call corpus correction. Using machine-learning methods, we applied corpus correction to a verb modality corpus for machine translation. We used the maximum-entropy and decision-list methods as machine-learning methods. We compared several kinds of methods for corpus correction in our experiments, and determined which is most effective by using a statistical test. We obtained several noteworthy findings: (1) Precision was almost the same for both detection and correction, so it is more convenient to do both correction and detection, rather than detection only. (2) In general, the maximum-entropy method worked better than the decision-list method; but the two methods had almost the same precision for the top 50 pieces of extracted data when closed data was used. (3) In terms of precision, the use of closed data was better than the use of open data; however, in terms of the total number of extracted errors, the use of open data was better than the use of closed data. Based on our analysis of these results, we developed a good method for corpus correction. We confirmed the effectiveness of our method by carrying out experiments on machine translation. As corpus-based machine translation continues to be developed, the corpus correction we discuss in this article should prove to be increasingly significant. Masaki Murata, Masao Utiyama, Kiyotaka Uchimoto, Hitoshi Isahara |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2004 | A part-versus-part method for massively parallel training of support vector machinesabstractThis work presents a part-versus-part decomposition method for massively parallel training of multi-class support vector machines (SVMs). By using this method, a massive multi-class classification problem is decomposed into a number of two-class subproblems as small as needed. An important advantage of the part-versus-part method over existing popular pair wise-classification approach is that a large-scale two-class subproblem can be further divided into a number of relatively smaller and balanced two-class subproblems, and fast training of SVMs on massive multi-class classification problems can be easily implemented in a massively parallel way. To demonstrate the effectiveness of the proposed method, we perform simulations on a large-scale text categorization problem. The experimental results show that the proposed method is faster than the existing pairwise-classification approach, better generalization performance can be achieved, and the method scales up to massive, complex multi-class classification problems. Bao-Liang Lu, Kai-An Wang, Masao Utiyama, Hitoshi Isahara |
IJCNN | 3 |
| 2004 | Constructing English Reading Courseware
Masao Utiyama, Midori Tanimura, Hitoshi Isahara |
PACLIC | 1 |
| 2003 | Reliable Measures for Aligning Japanese-English News Articles and SentencesabstractWe have aligned Japanese and English news articles and sentences to make a large parallel corpus. We first used a method based on cross-language information retrieval (CLIR) to align the Japanese and English articles and then used a method based on dynamic programming (DP) matching to align the Japanese and English sentences in these articles. However, the results included many incorrect alignments. To remove these, we propose two measures (scores) that evaluate the validity of alignments. The measure for article alignment uses similarities in sentences aligned by DP matching and that for sentence alignment uses similarities in articles aligned by CLIR. They enhance each other to improve the accuracy of alignment. Using these measures, we have successfully constructed a large-scale article and sentence alignment corpus available to the public. Masao Utiyama, Hitoshi Isahara |
ACL | 1 |
| 2001 | A Statistical Model for Domain-Independent Text SegmentationabstractWe propose a statistical method that finds the maximum-probability segmentation of a given text. This method does not require training data because it estimates probabilities from the given text. Therefore, it can be applied to any text in any domain. An experiment showed that the method is more accurate than or at least as accurate as a state-of-the-art text segmentation system. Masao Utiyama, Hitoshi Isahara |
ACL | 1 |
| 2000 | Multi-Topic Multi-Document Summarization
Masao Utiyama, Kôiti Hasida |
COLING | 1 |
| 2000 | A Statistical Approach to the Processing of Metonymy
Masao Utiyama, Masaki Murata, Hitoshi Isahara |
COLING | 1 |
| 2000 | Self-Organizing Semantic Maps of Japanese Nouns in Terms of Adnominal ConstituentsabstractAs a beginning study on self-organizing a general Japanese semantic map, which will be very useful in natural language processing, particularly in document organization and information retrieval, this paper describes the construction of a semantic map of Japanese nouns mapped according to their adnominal constituents. These maps are not only an important part of a general Japanese semantic map that we aim to construct, but can also be a powerful tool for supporting the analysis of the relation between head nouns and their adnominal constituents, an important issue in studies of Japanese pragmatics. Kyoko Kanzaki, Masaki Murata, Masao Utiyama, Kiyotaka Uchimoto, Hitoshi Isahara |
IJCNN (6) | 4 |