Masaaki Nagata

dblp:16/3520 · DBLP profile ↗
← Back
104ranked-venue papers
14as first author
16since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 95 · 13 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-authorDatabases, data management, data science and information retrieval · 7Computer networks · 1
YearPublicationVenuePosition
2026 NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
abstract
Neologism-aware machine translation 1 aims to translate source sentences containing neologisms into target languages.This field remains underexplored compared with general machine translation (MT).In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary.The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump.The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump.We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologismaware machine translation.Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits "translation difficulty" to further improve the translation quality of translation agents using our search toolkit 2 .
Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)3
2026 Coordinate Structure Extraction for Patent Claims Using Multilingual LLMs
Tsukasa Ishimaru, Takehito Utsuro, Masaaki Nagata
LREC3
2025 Case-Based Decision-Theoretic Decoding with Quality Memories
abstract
Minimum Bayes risk (MBR) decoding is a decision rule of text generation, which selects the hypothesis that maximizes the expected utility and robustly generates higher-quality texts than maximum a posteriori (MAP) decoding.However, it depends on sample texts drawn from the text generation model; thus, it is difficult to find a hypothesis that correctly captures the knowledge or information of out-of-domain.To tackle this issue, we propose case-based decision-theoretic (CBDT) decoding, another method to estimate the expected utility using examples of domain data.CBDT decoding not only generates higher-quality texts than MAP decoding, but also the combination of MBR and CBDT decoding outperformed MBR decoding in seven domain De-En and Ja↔En translation tasks and image captioning tasks on MSCOCO and nocaps datasets.
Hiroyuki Deguchi 0002, Masaaki Nagata
EMNLP2
2025 Patent Claim Translation via Continual Pre-training of Large Language Models with Parallel Data
abstract
Recent advancements in large language models (LLMs) have enabled their application across various domains. However, in the field of patent translation, Transformer encoder-decoder based models remain the standard approach, and the potential of LLMs for translation tasks has not been thoroughly explored. In this study, we conducted patent claim translation using an LLM fine-tuned with parallel data through continual pre-training and supervised fine-tuning, following the methodology proposed by Guo et al. (2024) and Kondo et al. (2024). Comparative evaluation against the Transformer encoder-decoder based translations revealed that the LLM achieved high scores for both BLEU and COMET. This demonstrated improvements in addressing issues such as omissions and repetitions. Nonetheless, hallucination errors, which were not observed in the traditional models, occurred in some cases and negatively affected the translation quality. This study highlights the promise of LLMs for patent translation while identifying the challenges that warrant further investigation.
Haruto Azami, Minato Kondo, Takehito Utsuro, Masaaki Nagata
MTSummit (1)4
2025 Improving Japanese-English Patent Claim Translation with Clause Segmentation Models based on Word Alignment
abstract
In patent documents, patent claims represent a particularly important section as they define the scope of the claims. However, due to the length and unique formatting of these sentences, neural machine translation (NMT) systems are prone to translation errors, such as omissions and repetitions. To address these challenges, this study proposes a translation method that first segments the source sentences into multiple shorter clauses using a clause segmentation model tailored to facilitate translation. These segmented clauses are then translated using a clause translation model specialized for clause-level translation. Finally, the translated clauses are rearranged and edited into the final translation using a reordering and editing model. In addition, this study proposes a method for constructing clause-level parallel corpora required for training the clause segmentation and clause translation models. This method leverages word alignment tools to create clause-level data from sentence-level parallel corpora. Experimental results demonstrate that the proposed method achieves statistically significant improvements in BLEU scores compared to conventional NMT models. Furthermore, for sentences where conventional NMT models exhibit omissions and repetitions, the proposed method effectively suppresses these errors, enabling more accurate translations.
Masato Nishimura, Kosei Buma, Takehito Utsuro, Masaaki Nagata
MTSummit (1)4
2024 JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
abstract
We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO). We also obtained patent family information from the DOCDB, that is a bibliographic database maintained by the European Patent Office (EPO). We extracted approximately 1.4M Japanese-English document pairs, which are translations of each other based on the patent families, and extracted about 350M sentence pairs from the document pairs using a translation-based sentence alignment method whose initial translation model is bootstrapped from a dictionary-based sentence alignment. We experimentally improved the accuracy of the patent translations by 20 bleu points by adding more than 300M sentence pairs obtained from patent applications to 22M sentence pairs obtained from the web.
Masaaki Nagata, Makoto Morishita, Katsuki Chousa, Norihito Yasuda
LREC/COLING1
2024 Argument Mining as a Text-to-Text Generation Task
abstract
Masayuki Kawarada, Tsutomu Hirao, Wataru Uchida, Masaaki Nagata. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Masayuki Kawarada, Tsutomu Hirao, Wataru Uchida, Masaaki Nagata
EACL (1)4
2024 Detector-Corrector: Edit-Based Automatic Post Editing for Human Post Editing
abstract
Post-editing is crucial in the real world because neural machine translation (NMT) sometimes makes errors.Automatic post-editing (APE) attempts to correct the outputs of an MT model for better translation quality.However, many APE models are based on sequence generation, and thus their decisions are harder to interpret for actual users.In this paper, we propose “detector–corrector”, an edit-based post-editing model, which breaks the editing process into two steps, error detection and error correction.The detector model tags each MT output token whether it should be corrected and/or reordered while the corrector model generates corrected words for the spans identified as errors by the detector.Experiments on the WMT’20 English–German and English–Chinese APE tasks showed that our detector–corrector improved the translation edit rate (TER) compared to the previous edit-based model and a black-box sequence-to-sequence APE model, in addition, our model is more explainable because it is based on edit operations.
Hiroyuki Deguchi 0002, Masaaki Nagata, Taro Watanabe
EAMT (1)2
2024 Word Alignment as Preference for Machine Translation
abstract
The problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena.In this work, we mitigate the problem in an LLM-based MT model by guiding it to better word alignment.We first study the correlation between word alignment and the phenomena of hallucination and omission in MT.Then we propose to utilize word alignment as preference to optimize the LLM-based MT model.The preference data are constructed by selecting chosen and rejected translations from multiple MT tools.Subsequently, direct preference optimization is used to optimize the LLM-based model towards the preference signal.Given the absence of evaluators specifically designed for hallucination and omission in MT, we further propose selecting hard instances and utilizing GPT-4 to directly evaluate the performance of the models in mitigating these issues.We verify the rationality of these designed evaluation methods by experiments, followed by extensive results demonstrating the effectiveness of word alignment-based preference optimization to mitigate hallucination and omission.On the other hand, although it shows promise in mitigating hallucination and omission, the overall performance of MT in different language directions remains mixed, with slight increases in BLEU and decreases in COMET.
Qiyu Wu 0001, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka
EMNLP2
2023 WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction
abstract
Most existing word alignment methods rely on manual alignment datasets or parallel corpora, which limits their usefulness.Here, to mitigate the dependence on manual data, we broaden the source of supervision by relaxing the requirement for correct, fully-aligned, and parallel sentences.Specifically, we make noisy, partially aligned, and non-parallel paragraphs.We then use such a large-scale weakly-supervised dataset for word alignment pre-training via span prediction.Extensive experiments with various settings empirically demonstrate that our approach, which is named WSPAlign, is an effective and scalable way to pre-train word aligners without manual data.When fine-tuned on standard benchmarks, WSPAlign has set a new state of the art by improving upon the best supervised baseline by 3.3~6.1 points in F1 and 1.5~6.1 points in AER .Furthermore, WSPAlign also achieves competitive performance compared with the corresponding baselines in few-shot, zero-shot and cross-lingual tests, which demonstrates that WSPAlign is potentially more practical for low-resource languages than existing methods. 1(1) Data Collection and Annotation (2) Pre-training for word alignment Transformer Encoder
Qiyu Wu 0001, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)2
2023 Target Language Monolingual Translation Memory based NMT by Cross-lingual Retrieval of Similar Translations and Reranking
abstract
Retrieve-edit-rerank is a text generation framework composed of three steps: retrieving for sentences using the input sentence as a query, generating multiple output sentence candidates, and selecting the final output sentence from these candidates. This simple approach has outperformed other existing and more complex methods. This paper focuses on the retrieving and the reranking steps. In the retrieving step, we propose retrieving similar target language sentences from a target language monolingual translation memory using language-independent sentence embeddings generated by mSBERT or LaBSE. We demonstrate that this approach significantly outperforms existing methods that use monolingual inter-sentence similarity measures such as edit distance, which is only applicable to a parallel translation memory. In the reranking step, we propose a new reranking score for selecting the best sentences, which considers both the log-likelihood of each candidate and the sentence embeddings based similarity between the input and the candidate. We evaluated the proposed method for English-to-Japanese translation on the ASPEC and English-to-French translation on the EU Bookshop Corpus (EUBC). The proposed method significantly exceeded the baseline in BLEU score, especially observing a 1.4-point improvement in the EUBC dataset over the original Retrieve-Edit-Rerank method.
Takuya Tamura, Takehito Utsuro, Masaaki Nagata
MTSummit (1)4
2023 Leveraging Highly Accurate Word Alignment for Low Resource Translation by Pretrained Multilingual Model
abstract
Recently, there has been a growing interest in pretraining models in the field of natural language processing. As opposed to training models from scratch, pretrained models have been shown to produce superior results in low-resource translation tasks. In this paper, we introduced the use of pretrained seq2seq models for preordering and translation tasks. We utilized manual word alignment data and mBERT-based generated word alignment data for training preordering and compared the effectiveness of various types of mT5 and mBART models for preordering. For the translation task, we chose mBART as our baseline model and evaluated several input manners. Our approach was evaluated on the Asian Language Treebank dataset, consisting of 20,000 parallel data in Japanese, English and Hindi, where Japanese is either on the source or target side. We also used in-house 3,000 parallel data in Chinese and Japanese. The results indicated that mT5-large trained with manual word alignment achieved a preordering performance exceeding 0.9 RIBES score on Ja-En and Ja-Zh pairs. Moreover, our proposed approach significantly outperformed the baseline model in most translation directions of Ja-En, Ja-Zh, and Ja-Hi pairs in at least one of BLEU/COMET scores.
Minato Kondo, Takuya Tamura, Takehito Utsuro, Masaaki Nagata
MTSummit (1)5
2023 Enhanced Retrieve-Edit-Rerank Framework with kNN-MT
Takuya Tamura, Takehito Utsuro, Masaaki Nagata
PACLIC4
2022 JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus
abstract
Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few language pairs, effectively dealing with most language pairs is difficult due to a lack of publicly available parallel corpora. This paper creates a large parallel corpus for English-Japanese, a language pair for which only limited resources are available, compared to such resource-rich languages as English-German. It introduces a new web-based English-Japanese parallel corpus named JParaCrawl v3.0. Our new corpus contains more than 21 million unique parallel sentence pairs, which is more than twice as many as the previous JParaCrawl v2.0 corpus. Through experiments, we empirically show how our new corpus boosts the accuracy of machine translation models on various domains. The JParaCrawl v3.0 corpus will eventually be publicly available online for research purposes.
Makoto Morishita, Katsuki Chousa, Jun Suzuki 0001, Masaaki Nagata
LREC4
2021 Context-aware Neural Machine Translation with Mini-batch Embedding
abstract
It is crucial to provide an inter-sentence context in Neural Machine Translation (NMT) models for higher-quality translation.With the aim of using a simple approach to incorporate inter-sentence information, we propose minibatch embedding (MBE) as a way to represent the features of sentences in a mini-batch.We construct a mini-batch by choosing sentences from the same document, and thus the MBE is expected to have contextual information across sentences.Here, we incorporate MBE in an NMT model, and our experiments show that the proposed method consistently outperforms the translation capabilities of strong baselines and improves writing style or terminology to fit the document's context. 1
Makoto Morishita, Jun Suzuki 0001, Tomoharu Iwata, Masaaki Nagata
EACL4
2021 Improving Neural RST Parsing Model with Silver Agreement Subtrees
abstract
Naoki Kobayashi, Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata
NAACL-HLT5
2020 Top-Down RST Parsing Utilizing Granularity Levels in Documents
abstract
Some downstream NLP tasks exploit discourse dependency trees converted from RST trees. To obtain better discourse dependency trees, we need to improve the accuracy of RST trees at the upper parts of the structures. Thus, we propose a novel neural top-down RST parsing method. Then, we exploit three levels of granularity in a document, paragraphs, sentences and Elementary Discourse Units (EDUs), to parse a document accurately and efficiently. The parsing is done in a top-down manner for each granularity level, by recursively splitting a larger text span into two smaller ones while predicting nuclearity and relation labels for the divided spans. The results on the RST-DT corpus show that our method achieved the state-of-the-art results, 87.0 unlabeled span score, 74.6 nuclearity labeled span score, and the comparable result with the state-of-the-art, 60.0 relation labeled span score. Furthermore, discourse dependency trees converted from our RST trees also achieved the state-of-the-art results, 64.9 unlabeled attachment score and 48.5 labeled attachment score.
Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata
AAAI5
2020 SpanAlign: Sentence Alignment Method based on Cross-Language Span Prediction and ILP
abstract
We propose a novel method of automatic sentence alignment from noisy parallel documents.We first formalize the sentence alignment problem as the independent predictions of spans in the target document from sentences in the source document.We then introduce a total optimization method using integer linear programming to prevent span overlapping and obtain non-monotonic alignments.We implement cross-language span prediction by fine-tuning pre-trained multilingual language models based on BERT architecture and train them using pseudo-labeled data obtained from unsupervised sentence alignment method.While the baseline methods use sentence embeddings and assume monotonic alignment, our method can capture the token-to-token interaction between the tokens of source and target text and handle non-monotonic alignments.In sentence alignment experiments on English-Japanese, our method achieved 70.3 F 1 scores, which are +8.0 points higher than the baseline method.In particular, our method improved by +53.9 F 1 scores for extracting non-parallel sentences.Our method improved the downstream machine translation accuracy by 4.1 BLEU scores when the extracted bilingual sentences are used for fine-tuning a pre-trained Japanese-to-English translation model. 1
Katsuki Chousa, Masaaki Nagata, Masaaki Nishino
COLING2
2020 SODA: Story Oriented Dense Video Captioning Evaluation Framework
Soichiro Fujita, Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata
ECCV (6)5
2020 A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT
abstract
We present a novel supervised word alignment method based on cross-language span prediction.We first formalize a word alignment problem as a collection of independent predictions from a token in the source sentence to a span in the target sentence.Since this step is equivalent to a SQuAD v2.0 style question answering task, we solve it using the multilingual BERT, which is fine-tuned on manually created gold word alignment data.It is nontrivial to obtain accurate alignment from a set of independently predicted spans.We greatly improved the word alignment accuracy by adding to the question the source token's context and symmetrizing two directional predictions.In experiments using five word alignment datasets from among Chinese, Japanese, German, Romanian, French, and English, we show that our proposed method significantly outperformed previous supervised and unsupervised word alignment methods without any bitexts for pretraining.For example, we achieved 86.7 F1 score for the Chinese-English data, which is 13.3 points higher than the previous state-of-the-art supervised method. 1
Masaaki Nagata, Katsuki Chousa, Masaaki Nishino
EMNLP (1)1
2020 JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus
abstract
Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs can benefit from them. We constructed a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited. We constructed the parallel corpus by broadly crawling the web and automatically aligning parallel sentences. Our collected corpus, called JParaCrawl, amassed over 8.7 million sentence pairs. We show how it includes a broader range of domains and how a neural machine translation model trained with it works as a good pre-trained model for fine-tuning specific domains. The pre-training and fine-tuning approaches achieved or surpassed performance comparable to model training from the initial state and reduced the training time. Additionally, we trained the model with an in-domain dataset and JParaCrawl to show how we achieved the best performance with them. JParaCrawl and the pre-trained models are freely available online for research purposes.
Makoto Morishita, Jun Suzuki 0001, Masaaki Nagata
LREC3
2020 A Test Set for Discourse Translation from Japanese to English
abstract
We made a test set for Japanese-to-English discourse translation to evaluate the power of context-aware machine translation. For each discourse phenomenon, we systematically collected examples where the translation of the second sentence depends on the first sentence. Compared with a previous study on test sets for English-to-French discourse translation (CITATION), we needed different approaches to make the data because Japanese has zero pronouns and represents different senses in different characters. We improved the translation accuracy using context-aware neural machine translation, and the improvement mainly reflects the betterment of the translation of zero pronouns.
Masaaki Nagata, Makoto Morishita
LREC1
2019 Character n-Gram Embeddings to Improve RNN Language Models
abstract
This paper proposes a novel Recurrent Neural Network (RNN) language model that takes advantage of character information. We focus on character n-grams based on research in the field of word embedding construction (Wieting et al. 2016). Our proposed method constructs word embeddings from character ngram embeddings and combines them with ordinary word embeddings. We demonstrate that the proposed method achieves the best perplexities on the language modeling datasets: Penn Treebank, WikiText-2, and WikiText-103. Moreover, we conduct experiments on application tasks: machine translation and headline generation. The experimental results indicate that our proposed method also positively affects these tasks
Sho Takase, Jun Suzuki 0001, Masaaki Nagata
AAAI3
2019 Answering while Summarizing: Multi-task Learning for Multi-hop QA with Evidence Extraction
abstract
Kosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Atsushi Otsuka, Itsumi Saito, Hisako Asano, Junji Tomita. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Kosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Atsushi Otsuka, Itsumi Saito, Hisako Asano, Junji Tomita
ACL (1)3
2019 Split or Merge: Which is Better for Unsupervised RST Parsing?
abstract
Naoki Kobayashi, Tsutomu Hirao, Kengo Nakamura, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tsutomu Hirao, Kengo Nakamura 0001, Hidetaka Kamigaito, Manabu Okumura, Masaaki Nagata
EMNLP/IJCNLP (1)6
2019 Generating Natural Anagrams: Towards Language Generation Under Hard Combinatorial Constraints
abstract
Masaaki Nishino, Sho Takase, Tsutomu Hirao, Masaaki Nagata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Masaaki Nishino, Sho Takase, Tsutomu Hirao, Masaaki Nagata
EMNLP/IJCNLP (1)4
2019 ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis
abstract
We propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long speech recording, which is decoded as a confusion network form of an automatic speech recognition (ASR) hypothesis sequence. It selects as many different content words as possible from the speech input that inevitably includes a high level of redundancy (e.g. the repetition of the same word) under a given length constraint. In experiments using a lecture speech corpus, we obtained higher summarization performance in terms of ROUGE scores than with a baseline extractive summarization method. We further conduct experimental analyses to obtain the oracle (upper bound) performance of the summarization methods. The analysis results show that the oracle performance is very high even though the ASR hypotheses include recognition errors. It is significantly higher than the system performance and, in addition, the oracle performance of the compressive method is significantly higher than that of the extractive method. These results confirm that our method is a promising approach.
Atsunori Ogawa, Tsutomu Hirao, Tomohiro Nakatani, Masaaki Nagata
ICASSP4
2019 Selecting Informative Context Sentence by Forced Back-Translation
Ryuichiro Kimura, Shohei Iida, Hongyi Cui, Po-Hsuan Hung, Takehito Utsuro, Masaaki Nagata
MTSummit (1)6
2018 Improving Neural Machine Translation by Incorporating Hierarchical Subword Features
abstract
This paper focuses on subword-based Neural Machine Translation (NMT). We hypothesize that in the NMT model, the appropriate subword units for the following three modules (layers) can differ: (1) the encoder embedding layer, (2) the decoder embedding layer, and (3) the decoder output layer. We find the subword based on Sennrich et al. (2016) has a feature that a large vocabulary is a superset of a small vocabulary and modify the NMT model enables the incorporation of several different subword units in a single embedding layer. We refer these small subword features as hierarchical subword features. To empirically investigate our assumption, we compare the performance of several different subword units and hierarchical subword features for both the encoder and decoder embedding layers. We confirmed that incorporating hierarchical subword features in the encoder consistently improves BLEU scores on the IWSLT evaluation datasets.
Makoto Morishita, Jun Suzuki 0001, Masaaki Nagata
COLING3
2018 Automatic Pyramid Evaluation Exploiting EDU-based Extractive Reference Summaries
abstract
This paper tackles automation of the pyramid method, a reliable manual evaluation framework.To construct a pyramid, we transform human-made reference summaries into extractive reference summaries that consist of Elementary Discourse Units (EDUs) obtained from source documents and then weight every EDU by counting the number of extractive reference summaries that contain the EDU.A summary is scored by the correspondences between EDUs in the summary and those in the pyramid.Experiments on DUC and TAC data sets show that our methods strongly correlate with various manual evaluations.Ref.
Tsutomu Hirao, Hidetaka Kamigaito, Masaaki Nagata
EMNLP3
2018 Direct Output Connection for a High-Rank Language Model
abstract
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers.Our proposed method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. (2018).The proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets.Moreover, we indicate our proposed method contributes to two application tasks: machine translation and headline generation.
Sho Takase, Jun Suzuki 0001, Masaaki Nagata
EMNLP3
2018 Optimizing Network Reliability via Best-First Search over Decision Diagrams
abstract
Communication networks are an essential infrastructure and must be designed carefully to ensure high reliability. Identifying a fully reliable design is, however, computationally very tough since it requires that a reliability evaluation, which is known to be #P-complete, be repeated an exponential number of times. Existing studies, therefore, attempt to avoid exact optimization to reduce the computational burden by applying heuristics. Due to the importance of communication networks and to better assess the accuracy of heuristic approaches, exact optimization remains a key goal. This paper proposes an exact method for two network design problems: reliability maximization under budget constraints and cost minimization with assurance of reliability. Our method employs a common idea to solve these problems, i.e., a best-first search algorithm that runs on decision diagrams. Our method employs just a single binary decision diagram (BDD) to compute the reliability for any solution and is also used as the basis of a novel heuristic function, called the cost-aware BDD heuristic function, as a search guide. Numerical experiments show that our method scales well; it successfully optimizes a network with 189 links. In addition, our method reveals the poor performance of existing heuristic approaches; a well-known existing heuristic method is shown to yield a solution that offers less than half the optimal reliability.
Masaaki Nishino, Takeru Inoue, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
INFOCOM5
2018 Neural Tensor Networks with Diagonal Slice Matrices
abstract
Takahiro Ishihara, Katsuhiko Hayashi, Hitoshi Manabe, Masashi Shimbo, Masaaki Nagata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Takahiro Ishihara, Katsuhiko Hayashi 0001, Hitoshi Manabe, Masashi Shimbo, Masaaki Nagata
NAACL-HLT5
2018 Higher-Order Syntactic Attention Network for Longer Sentence Compression
abstract
Hidetaka Kamigaito, Katsuhiko Hayashi, Tsutomu Hirao, Masaaki Nagata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Tsutomu Hirao, Masaaki Nagata
NAACL-HLT4
2018 Provable Fast Greedy Compressive Summarization with Any Monotone Submodular Function
abstract
Shinsaku Sakaue, Tsutomu Hirao, Masaaki Nishino, Masaaki Nagata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shinsaku Sakaue, Tsutomu Hirao, Masaaki Nishino, Masaaki Nagata
NAACL-HLT4
2018 Reducing Odd Generation from Neural Headline Generation
Shun Kiyono, Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Kentaro Inui, Masaaki Nagata
PACLIC6
2017 Dancing with Decision Diagrams: A Combined Approach to Exact Cover
abstract
Exact cover is the problem of finding subfamilies, S*, of a family of sets, S, over universe U, where S* forms a partition of U. It is a popular NP-hard problem appearing in a wide range of computer science studies. Knuth's algorithm DLX, a backtracking-based depth-first search implemented with the data structure called dancing links, is known as state-of-the-art for finding all exact covers. We propose a method to accelerate DLX. Our method constructs a Zero-suppressed Binary Decision Diagram (ZDD) that represents the set of solutions while running depth-first search in DLX. Constructing ZDDs enables the efficient use of memo cache to speed up the search. Moreover, our method has a virtue that it outputs ZDDs; we can perform several useful operations with them. Experiments confirm that the proposed method is up to several orders of magnitude faster than DLX.
Masaaki Nishino, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
AAAI4
2017 Compiling Graph Substructures into Sentential Decision Diagrams
abstract
The Zero-suppressed Sentential Decision Diagram (ZSDD) is a recentlydiscovered tractable representation of Boolean functions. ZSDD subsumes theZero-suppressed Binary Decision Diagram (ZDD) as a strict subset, andsimilar to ZDD, it can perform several useful operations like model countingand Apply operations. We propose a top-down compilation algorithmfor ZSDD that represents sets of specific graph substructures, e.g.,matchings and simple paths of a graph. We experimentally confirm that theproposed algorithm is faster than other construction methods includingbottom-up methods and top-down methods for ZDDs, and the resulting ZSDDsare smaller than ZDDs representing the same graph substructures. We alsoshow that the size constructed ZSDDs can be bounded by the branch-width of thegraph. This bound is tighter than that of ZDDs.
Masaaki Nishino, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
AAAI4
2017 Learning to Rank for Coordination Detection
Rumeng Li, Hiroyuki Shindo, Katsuhito Sudoh, Masaaki Nagata
CICLing (1)5
2017 Enumeration of Extractive Oracle Summaries
abstract
To analyze the limitations and the future directions of the extractive summarization paradigm, this paper proposes an Integer Linear Programming (ILP) formulation to obtain extractive oracle summaries in terms of ROUGE n .We also propose an algorithm that enumerates all of the oracle summaries for a set of reference summaries to exploit F-measures that evaluate which system summaries contain how many sentences that are extracted as an oracle summary.Our experimental results obtained from Document Understanding Conference (DUC) corpora demonstrated the following: (1) room still exists to improve the performance of extractive summarization; (2) the F-measures derived from the enumerated oracle summaries have significantly stronger correlations with human judgment than those derived from single oracle summaries.
Tsutomu Hirao, Masaaki Nishino, Jun Suzuki 0001, Masaaki Nagata
EACL (1)4
2016 Zero-Suppressed Sentential Decision Diagrams
abstract
The Sentential Decision Diagram (SDD) is a prominent knowledge representation language that subsumes the Ordered Binary Decision Diagram (OBDD) as a strict subset. Like OBDDs, SDDs have canonical forms and support bottom-up operations for combining SDDs, but they are more succinct than OBDDs. In this paper we introduce an SDD variant, called the Zero-suppressed Sentential Decision Diagram (ZSDD). The key idea of ZSDD is to employ new trimming rules for obtaining a canonical form. As a result, ZSDD subsumes the Zero-suppressed Binary Decision Diagram (ZDD) as a strict subset. ZDDs are known for their effectiveness on representing sparse Boolean functions. Likewise, ZSDDs can be more succinct than SDDs when representing sparse Boolean functions. We propose several polytime bottom-up operations over ZSDDs, and a technique for reducing ZSDD size, while maintaining applicability to important queries. We also specify two distinct upper bounds on ZSDD sizes; one is derived from the treewidth of a CNF and the other from the size of a family of sets. Experiments show that ZSDDs are smaller than SDDs or ZDDs for a standard benchmark dataset.
Masaaki Nishino, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
AAAI4
2016 Exploring Text Links for Coherent Multi-Document Summarization
abstract
Summarization aims to represent source documents by a shortened passage. Existing methods focus on the extraction of key information, but often neglect coherence. Hence the generated summaries suffer from a lack of readability. To address this problem, we have developed a graph-based method by exploring the links between text to produce coherent summaries. Our approach involves finding a sequence of sentences that best represent the key information in a coherent way. In contrast to the previous methods that focus only on salience, the proposed method addresses both coherence and informativeness based on textual linkages. We conduct experiments on the DUC2004 summarization task data set. A performance comparison reveals that the summaries generated by the proposed system achieve comparable results in terms of the ROUGE metric, and show improvements in readability by human evaluation.
Masaaki Nishino, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata
COLING5
2016 Neural Headline Generation on Abstract Meaning Representation
abstract
Neural network-based encoder-decoder models are among recent attractive methodologies for tackling natural language generation tasks.This paper investigates the usefulness of structural syntactic and semantic information additionally incorporated in a baseline neural attention-based model.We encode results obtained from an abstract meaning representation (AMR) parser using a modified version of Tree-LSTM.Our proposed attention-based AMR encoder-decoder model improves headline generation benchmarks compared with the baseline neural attention-based model.
Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Tsutomu Hirao, Masaaki Nagata
EMNLP5
2016 Learning Compact Neural Word Embeddings by Parameter Space Sharing
Jun Suzuki 0001, Masaaki Nagata
IJCAI2
2016 Right-truncatable Neural Word Embeddings
abstract
This paper proposes an incremental learning strategy for neural word embedding methods, such as SkipGrams and Global Vectors.Since our method iteratively generates embedding vectors one dimension at a time, obtained vectors equip a unique property.Namely, any right-truncated vector matches the solution of the corresponding lower-dimensional embedding.Therefore, a single embedding vector can manage a wide range of dimensional requirements imposed by many different uses and applications.
Jun Suzuki 0001, Masaaki Nagata
HLT-NAACL2
2016 Empirical comparison of dependency conversions for RST discourse trees
abstract
Two heuristic rules that transform Rhetorical Structure Theory discourse trees into discourse dependency trees (DDTs) have recently been proposed (Hirao et al., 2013;Li et al., 2014), but these rules derive significantly different DDTs because their conversion schemes on multinuclear relations are not identical.This paper reveals the difference among DDT formats with respect to the following questions: (1) How complex are the formats from a dependency graph theoretic point of view?(2) Which formats are analyzed more accurately by automatic parsers?(3) Which are more suitable for text summarization task?Experimental results showed that Hirao's conversion rule produces DDTs that are more useful for text summarization, even though it derives more complex dependency structures.
Katsuhiko Hayashi 0001, Tsutomu Hirao, Masaaki Nagata
SIGDIAL Conference3
2015 BDD-Constrained Search: A Unified Approach to Constrained Shortest Path Problems
abstract
Dynamic programming (DP) is a fundamental tool used to obtain exact, optimal solutions for many combinatorial optimization problems. Among these problems, important ones including the knapsack problems and the computation of edit distances between string pairs can be solved with a kind of DP that corresponds to solving the shortest path problem on a directed acyclic graph (DAG). These problems can be solved efficiently with DP, however, in practical situations, we want to solve the customized problems made by adding logical constraints to the original problems. Developing an algorithm specifically for each combination of a problem and a constraint set is unrealistic. The proposed method, BDD-Constrained Search (BCS), exploits a Binary Decision Diagram (BDD) that represents the logical constraints in combination with the DAG that represents the problem. The BCS runs DP on the DAG while using the BDD to check the equivalence and the validity of intermediate solutions to efficiently solve the problem. The important feature of BCS is that it can be applied to problems with various types of logical constraints in a unified way once we represent the constraints as a BDD. We give a theoretical analysis on the time complexity of BCS and also conduct experiments to compare its performance to that of a state-of-the-art integer linear programming solver.
Masaaki Nishino, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
AAAI4
2015 Enhanced Word Embeddings from a Hierarchical Neural Language Model
abstract
This paper proposes a neural language model to capture the interaction of text units of different levels, i.e.., documents, paragraphs, sentences, words in an hierarchical structure. At each paralleled level, the model incorporates Markov property while each higher-level unit hierarchically influences its containing units. Such an architecture enables the learned word embeddings to encode both global and local information. We evaluate the learned word embeddings and experiments demonstrate the effectiveness of our model.
Katsuhito Sudoh, Masaaki Nagata
CIKM3
2015 Empty Category Detection using Path Features and Distributed Case Frames
abstract
We describe an approach for machine learning-based empty category detection that is based on the phrase structure analysis of Japanese.The problem is formalized as tree node classification, and we find that the path feature, the sequence of node labels from the current node to the root, is highly effective.We also find that the set of dot products between the word embeddings for a verb and those for case particles can be used as a substitution for case frames.Experiments show that the proposed method outperforms the previous state-of the art method by 68.6% to 73.2% in terms of F-measure.
Shunsuke Takeno, Masaaki Nagata, Kazuhide Yamamoto
EMNLP2
2015 A Dynamic Programming Algorithm for Tree Trimming-based Text Summarization
abstract
Masaaki Nishino, Norihito Yasuda, Tsutomu Hirao, Shin-ichi Minato, Masaaki Nagata. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Masaaki Nishino, Norihito Yasuda, Tsutomu Hirao, Shin-ichi Minato, Masaaki Nagata
HLT-NAACL5
2015 Empty Category Detection With Joint Context-Label Embeddings
abstract
This paper presents a novel technique for empty category (EC) detection using distributed word representations.A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space.In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection.Experiments on Chinese Treebank prove the effectiveness of the proposed method.We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB.
Katsuhito Sudoh, Masaaki Nagata
HLT-NAACL3
2015 Rating Entities and Aspects Using a Hierarchical Model
Katsuhito Sudoh, Masaaki Nagata
PAKDD (2)3
2015 Summarizing a Document by Trimming the Discourse Tree
abstract
Recent studies on extractive text summarization formulate it as a combinatorial optimization problem, extracting the optimal subset from a set of the textual units that maximizes an objective function without violating the length constraint. Although these methods successfully improve automatic evaluation scores, they do not consider the discourse structure in the source document. Thus, summaries generated by these methods may lack logical coherence. In previous work, we proposed a method that exploits a discourse tree structure to produce coherent summaries. By transforming a traditional discourse tree, namely a rhetorical structure theory-based discourse tree (RST-DT), into a dependency-based discourse tree (DEP-DT), we formulated the summarization procedure as a Tree Knapsack Problem whose tree corresponds to the DEP-DT. This paper extends the work with a detailed discussion of the approach together with a novel efficient dynamic programming algorithm for solving the Tree Knapsack Problem. Experiments show that our method not only achieved the highest score in both automatic and human evaluation, but also obtained good performance in terms of the linguistic qualities of the summaries.
Tsutomu Hirao, Masaaki Nishino, Yasuhisa Yoshida, Jun Suzuki 0001, Norihito Yasuda, Masaaki Nagata
IEEE ACM Trans. Audio Speech Lang. Process.6
2015 Summarization Based on Task-Oriented Discourse Parsing
abstract
Previous research demonstrates that discourse relations can help generate high-quality summaries. Existing studies usually adopt existing discourse parsers directly with no modifications, hence cannot take full advantage of discourse parsing. This paper describes a new single document summarization system. In contrast to previous work, we train a discourse parser specially for summarization by using summaries. The training data are dynamically changed during the training phase to enable the parser to grab the text units that are important for summaries. A special tree-based summary extraction algorithm is designed to work with the new parser. The proposed system enables us to combine discourse parsing and summarization in a unified scheme. Experiments on both the RST-DT and DUC2001 datasets show the effectiveness of the proposed method.
Yasuhisa Yoshida, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 Fused Feature Representation Discovery for High-Dimensional and Sparse Data
Jun Suzuki 0001, Masaaki Nagata
AAAI2
2014 Dependency-based Discourse Parser for Single-Document Summarization
abstract
The current state-of-the-art single-document summarization method gen-erates a summary by solving a Tree Knapsack Problem (TKP), which is the problem of finding the optimal rooted sub-tree of the dependency-based discourse tree (DEP-DT) of a document. We can obtain a gold DEP-DT by transforming a gold Rhetorical Structure Theory-based discourse tree (RST-DT). However, there is still a large difference between the ROUGE scores of a system with a gold DEP-DT and a system with a DEP-DT obtained from an automatically parsed RST-DT. To improve the ROUGE score, we propose a novel discourse parser that directly generates the DEP-DT. The evaluation results showed that the TKP with our parser outperformed that with the state-of-the-art RST-DT parser, and achieved almost equivalent ROUGE scores to the TKP with the gold DEP-DT. 1
Yasuhisa Yoshida, Jun Suzuki 0001, Tsutomu Hirao, Masaaki Nagata
EMNLP4
2014 Accelerating Graph Adjacency Matrix Multiplications with Adjacency Forest
abstract
We propose a method for accelerating matrix multiplications that are iteratively performed with a sparse adjacency matrix. These operations appear in a wide range of data analyses and data mining situations, which include the computation of Personalized PageRank (PPR) and Nonnegative Matrix Factorization (NMF). We exploit the fact that the intermediate computational results for the matrix multiplication of equivalent partial row vectors of a matrix are the same. Our new data structure, the adjacency forest, uses this property and represents an adjacency matrix as a rooted tree that is made by sharing the common suffixes of the row vectors of the matrix. By exploiting the structure of the tree, we can perform a matrix multiplication while sharing intermediate computational results to reduce the number of required operations. We also show that we can further accelerate computation by dividing a matrix into several sub-matrices and representing the original matrix as a forest. We confirm experimentally that our approach can speed up the computation of Personalized PageRank and NMF up to 300%.
Masaaki Nishino, Norihito Yasuda, Shin-ichi Minato, Masaaki Nagata
SDM4
2013 Text Summarization while Maximizing Multiple Objectives with Lagrangian Relaxation
Masaaki Nishino, Norihito Yasuda, Tsutomu Hirao, Jun Suzuki 0001, Masaaki Nagata
ECIR5
2013 Sub-sentence Extraction Based on Combinatorial Optimization
Norihito Yasuda, Masaaki Nishino, Tsutomu Hirao, Masaaki Nagata
ECIR4
2013 Shift-Reduce Word Reordering for Machine Translation
abstract
This paper presents a novel word reordering model that employs a shift-reduce parser for inversion transduction grammars.Our model uses rich syntax parsing features for word reordering and runs in linear time.We apply it to postordering of phrase-based machine translation (PBMT) for Japanese-to-English patent tasks.Our experimental results show that our method achieves a significant improvement of +3.1 BLEU scores against 30.15BLEU scores of the baseline PBMT system.
Katsuhiko Hayashi 0001, Katsuhito Sudoh, Hajime Tsukada, Jun Suzuki 0001, Masaaki Nagata
EMNLP5
2013 Single-Document Summarization as a Tree Knapsack Problem
abstract
Recent studies on extractive text summarization formulate it as a combinatorial optimization problem such as a Knapsack Problem, a Maximum Coverage Problem or a Budgeted Median Problem.These methods successfully improved summarization quality, but they did not consider the rhetorical relations between the textual units of a source document.Thus, summaries generated by these methods may lack logical coherence.This paper proposes a single document summarization method based on the trimming of a discourse tree.This is a two-fold process.First, we propose rules for transforming a rhetorical structure theorybased discourse tree into a dependency-based discourse tree, which allows us to take a treetrimming approach to summarization.Second, we formulate the problem of trimming a dependency-based discourse tree as a Tree Knapsack Problem, then solve it with integer linear programming (ILP).Evaluation results showed that our method improved ROUGE scores.
Tsutomu Hirao, Yasuhisa Yoshida, Masaaki Nishino, Norihito Yasuda, Masaaki Nagata
EMNLP5
2013 Noise-Aware Character Alignment for Bootstrapping Statistical Machine Transliteration from Bilingual Corpora
abstract
This paper proposes a novel noise-aware character alignment method for bootstrapping statistical machine transliteration from automatically extracted phrase pairs.The model is an extension of a Bayesian many-to-many alignment method for distinguishing nontransliteration (noise) parts in phrase pairs.It worked effectively in the experiments of bootstrapping Japanese-to-English statistical machine transliteration in patent domain using patent bilingual corpora.
Katsuhito Sudoh, Shinsuke Mori, Masaaki Nagata
EMNLP3
2013 Statistical Parsing with Probabilistic Symbol-Refined Tree Substitution Grammars
Hiroyuki Shindo, Yusuke Miyao, Akinori Fujino, Masaaki Nagata
IJCAI4
2013 Two-Stage Pre-ordering for Japanese-to-English Statistical Machine Translation
Sho Hoshino, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata
IJCNLP4
2013 Effects of Parsing Errors on Pre-Reordering Performance for Chinese-to-Japanese SMT
Pascual Martínez-Gómez, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata
PACLIC5
2013 Adaptive semi-supervised learning on labeled and unlabeled data with different distributions
Akinori Fujino, Naonori Ueda, Masaaki Nagata
Knowl. Inf. Syst.3
2013 Syntax-Based Post-Ordering for Efficient Japanese-to-English Translation
abstract
This article proposes a novel reordering method for efficient two-step Japanese-to-English statistical machine translation (SMT) that isolates reordering from SMT and solves it after lexical translation. This reordering problem, called post-ordering , is solved as an SMT problem from Head-Final English (HFE) to English. HFE is syntax-based reordered English that is very successfully used for reordering with English-to-Japanese SMT. The proposed method incorporates its advantage into the reverse direction, Japanese-to-English, and solves the post-ordering problem by accurate syntax-based SMT with target language syntax. Two-step SMT with the proposed post-ordering empirically reduces the decoding time of the accurate but slow syntax-based SMT by its good approximation using intermediate HFE. The proposed method improves the decoding speed of syntax-based SMT decoding by about six times with comparable translation accuracy in Japanese-to-English patent translation experiments.
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata
ACM Trans. Asian Lang. Inf. Process.5
2012 Learning to Translate with Multiple Objectives
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata
ACL (1)5
2012 Bayesian Symbol-Refined Tree Substitution Grammars for Syntactic Parsing
Hiroyuki Shindo, Yusuke Miyao, Akinori Fujino, Masaaki Nagata
ACL (1)4
2011 Transfer Learning for Multiple-Domain Sentiment Analysis - Identifying Domain Dependent/Independent Word Polarity
abstract
Sentiment analysis is the task of determining the attitude (positive or negative) of documents. While the polarity of words in the documents is informative for this task, polarity of some words cannot be determined without domain knowledge. Detecting word polarity thus poses a challenge for multiple-domain sentiment analysis. Previous approaches tackle this problem with transfer learning techniques, but they cannot handle multiple source domains and multiple target domains. This paper proposes a novel Bayesian probabilistic model to handle multiple source and multiple target domains. In this model, each word is associated with three factors: Domain label, domain dependence/independence and word polarity. We derive an efficient algorithm using Gibbs sampling for inferring the parameters of the model, from both labeled and unlabeled texts. Using real data, we demonstrate the effectiveness of our model in a document polarity classification task compared with a method not considering the differences between domains. Moreover our method can also tell whether each word's polarity is domain-dependent or domain-independent. This feature allows us to construct a word polarity dictionary for each domain.
Yasuhisa Yoshida, Tsutomu Hirao, Tomoharu Iwata, Masaaki Nagata, Yuji Matsumoto 0001
AAAI4
2011 Providing Cross-Lingual Editing Assistance to Wikipedia Editors
Ching-man Au Yeung, Kevin Duh, Masaaki Nagata
CICLing (2)3
2011 Generalized Minimum Bayes Risk System Combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata
IJCNLP5
2011 Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction of Second Language Learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, Yuji Matsumoto 0001
IJCNLP3
2011 Distributed Minimum Error Rate Training of SMT using Particle Swarm Optimization
Jun Suzuki 0001, Kevin Duh, Masaaki Nagata
IJCNLP3
2011 Extracting Pre-ordering Rules from Predicate-Argument Structures
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata
IJCNLP5
2011 Post-ordering in Statistical Machine Translation
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata
MTSummit5
2011 Extracting Pre-ordering Rules from Chunk-based Dependency Trees for Japanese-to-English Translation
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata
MTSummit5
2010 A robust semi-supervised classification method for transfer learning
abstract
The transfer learning problem of designing good classifiers with a high generalization ability by using labeled samples whose distribution is different from that of test samples is an important and challenging research issue in the fields of machine learning and data mining. This paper focuses on designing a semi-supervised classifier trained by using unlabeled samples drawn by the same distribution as test samples, and presents a semi-supervised classification method to deal with the transfer learning problem, based on a hybrid discriminative and generative model. Although JESS-CM is one of the most successful semi-supervised classifier design frameworks and has achieved the best published results in NLP tasks, it has an overfitting problem in transfer learning settings that we consider in this paper. We expect the overfitting problem to be mitigated with the proposed method, which utilizes both labeled and unlabeled samples for the discriminative training of classifiers. We also present a refined objective that formalizes the training algorithm and classifier form. Our experimental results for text classification using three typical benchmark test collections confirmed that the proposed method outperformed the JESS-CM framework with most transfer learning settings.
Akinori Fujino, Naonori Ueda, Masaaki Nagata
CIKM3
2010 Enriching Dictionaries with Images from the Internet - Targeting Wikipedia and a Japanese Semantic Lexicon: Lexeed -
Sanae Fujita, Masaaki Nagata
COLING2
2009 BaseNP Supersense Tagging for Japanese Texts
Hirotoshi Taira, Sen Yoshida, Masaaki Nagata
PACLIC3
2009 Utilizing Features of Verbs in Statistical Zero Pronoun Resolution for Japanese Speech
Sen Yoshida, Masaaki Nagata
PACLIC2
2008 A Japanese Predicate Argument Structure Analysis using Decision Lists
Hirotoshi Taira, Sanae Fujita, Masaaki Nagata
EMNLP3
2006 A Clustered Global Phrase Reordering Model for Statistical Machine Translation
abstract
In this paper, we present a novel global reordering model that can be incorporated into standard phrase-based statistical machine translation.Unlike previous local reordering models that emphasize the reordering of adjacent phrase pairs (Tillmann and Zhang, 2005), our model explicitly models the reordering of long distances by directly estimating the parameters from the phrase alignments of bilingual training sentences.In principle, the global phrase reordering model is conditioned on the source and target phrases that are currently being translated, and the previously translated source and target phrases.To cope with sparseness, we use N-best phrase alignments and bilingual phrase clustering, and investigate a variety of combinations of conditioning factors.Through experiments, we show, that the global reordering model significantly improves the translation accuracy of a standard Japanese-English translation task.
Masaaki Nagata, Kuniko Saito, Kazuhide Yamamoto, Kazuteru Ohashi
ACL1
2005 Portable Translator Capable of Recognizing Characters on Signboard and Menu Captured by its Built-in Camera
Hideharu Nakajima, Yoshihiro Matsuo, Masaaki Nagata, Kuniko Saito
ACL3
2004 Efficient Decoding for Statistical Machine Translation with a Fully Expanded WFST Model
Hajime Tsukada, Masaaki Nagata
EMNLP2
2003 Say-as classification for alphabetic words in Japanese texts
Hisako Asano, Masaaki Nagata, Masanobu Abe
INTERSPEECH2
2003 Estimating Japanese word accent from syllable sequence using support vector machine
Hideharu Nakajima, Masaaki Nagata, Hisako Asano, Masanobu Abe
INTERSPEECH2
2003 Improving translation models by applying asymmetric learning
abstract
The statistical Machine Translation Model has two components: a language model and a translation model. This paper describes how to improve the quality of the translation model by using the common word pairs extracted by two asymmetric learning approaches. One set of word pairs is extracted by Viterbi alignment using a translation model, the other set is extracted by Viterbi alignment using another translation model created by reversing the languages. The common word pairs are extracted as the same word pairs in the two sets of word pairs. We conducted experiments using English and Japanese. Our method improves the quality of a original translation model by 5.7%. The experiments also show that the proposed learning method improves the word alignment quality independent of the training domain and the translation model. Moreover, we show that common word pairs are almost as useful as regular dictionary entries for training purposes.
Setsuo Yamada, Masaaki Nagata, Kenji Yamada
MTSummit2
2000 Synchronous Morphological Analysis of Grapheme and Phoneme for Japanese OCR
abstract
We developed a novel language model for Japanese based on grapheme-phoneme tuples, which is one order of magnitude smaller than word-based models. We also developed an alignment algorithm of graphemes and phonemes for both ordinary text and OCR output. We show, by experiment, that the combination of the grapheme-phoneme tuple ngram model and the grapheme-phoneme alignment algorithm significantly improve character recognition accuracy if both grapheme and phoneme representations are given.
Masaaki Nagata
ACL1
1999 A Part of Speech Estimation Method for Japanese Unknown Words using a Statistical Model of Morphology and Context
abstract
We present a statistical model of Japanese unknown words consisting of a set of length and spelling models classified by the character types that constitute a word. The point is quite simple: different character sets should be treated differently and the changes between character types are very important because Japanese script has both ideograms like Chinese (kanji) and phonograms like English (katakana). Both word segmentation accuracy and part of speech tagging accuracy are improved by the proposed model. The model can achieve 96.6% tagging accuracy if unknown words are correctly segmented.
Masaaki Nagata
ACL1
1996 Context-Based Spelling Correction for Japanese OCR
Masaaki Nagata
COLING1
1996 Automatic Extraction of New Words from Japanese Texts using Generalized Forward-Backward Search
Masaaki Nagata
EMNLP1
1996 Automatic acquisition of probabilistic dialogue models
Kenji Kita, Yoshikazu Fukui, Masaaki Nagata, Tsuyoshi Morimoto
ICSLP3
1994 A Stochastic Japanese Morphological Analyzer Using a Forward-DP Backward-A* N-Best Search Algorithm
Masaaki Nagata
COLING1
1994 Efficient chart parsing of speech recognition candidates
abstract
In a spoken language system that uses the N-best interface between the continuous-speech recognition component and the natural language processing component, the role of the language processing component is to filter speech recognition candidates and to determine the semantic interpretation of the input. We propose as efficient chart parsing method of finding the most plausible sentence within the N-best candidate sentences and determining the semantic interpretation of the sentence, while avoiding the re-computation of the common substrings. We use syntactic and semantic constraints described in the unification-based framework and context-sensitive conditional probability CFG preferences to reorder the N-best candidates. In preliminary tests, our parser has been successful in reducing parsing steps by sharing common substrings and in selecting correct sentences first.>
Toshihisa Tashiro, Toshiyuki Takezawa, Tsuyoshi Morimoto, Masaaki Nagata
ICASSP (2)4
1994 A stochastic morphological analyzer for spontaneously spoken languages
Masaaki Nagata
ICSLP1
1994 First steps towards statistical modeling of dialogue to predict the speech act type of the next utterance
Masaaki Nagata, Tsuyoshi Morimoto
Speech Communication1
1993 Spoken Language Translation System
Gen-ichiro Kikui, Mark Seligman, Toshiyuki Takezawa, Masami Suzuki, Kenji Kita, Tsuyoshi Morimoto, Masaaki Nagata, Toshihisa Tashiro, Herbert S. Tropf, Shigeki Sagayama, Jun-ichi Takami, Kazumi Ohkura, Akira Kurematsu
IJCAI7
1993 ATR's speech translation system: ASURA
Tsuyoshi Morimoto, Toshiyuki Takezawa, Fumihiro Yato, Shigeki Sagayama, Toshihisa Tashiro, Masaaki Nagata, Akira Kurematsu
EUROSPEECH6
1992 A Spoken Language Translation System: SL-TRANS2
Tsuyoshi Morimoto, Masami Suzuki, Toshiyuki Takezawa, Gen-ichiro Kikui, Masaaki Nagata, Mutsuko Tomokiyo
COLING5
1992 An Empirical Study on Rule Granularity and Unification Interleaving Toward an Efficient Unification-Based Parsing System
Masaaki Nagata
COLING1
1992 Enhancement of ATR's spoken language translation system: SL-TRANS2
Tsuyoshi Morimoto, Toshiyuki Takezawa, Kazumi Ohkura, Masaaki Nagata, Fumihiro Yato, Shigeki Sagayama, Akira Kurematsu
ICSLP4
1992 Using pragmatics to rule out recognition errors in cooperative task-oriented dialogues
Masaaki Nagata
ICSLP1
1990 HPSG-Based Lattice Parser for Spoken Japanese in a Spoken Language Translation System
Masaaki Nagata, Kiyoshi Kogure
ECAI1