Yuji Matsumoto 0001

dblp:11/4619 · DBLP profile ↗
← Back
194ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0003-4946-9574ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 182 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorTheory of computation · 2 · 1 first-author
YearPublicationVenuePosition
2026 Docora: A System for Interactive Knowledge Extraction and Visualization from Scientific PDFs
abstract
Scientific research articles, typically distributed in PDF format, contain valuable knowledge but remain challenging to convert into structured datasets due to fragmented workflows that separate parsing, annotation, and visualization. Existing annotation platforms operate on plain text, which requires an additional PDF-to-text conversion step before annotation, while PDF parsing tools lack automated annotation suggestions. To bridge this gap, we introduce Docora, a system that unifies PDF parsing, automated annotation assistance, and multi-view visualization into a single interactive platform. Docora enables researchers to configure entity and relation schemas for any domain, automatically generates initial annotations using rule-based, model-based, or LLM-based extractors, and provides synchronized visualizations across PDF, text, and graph views. Users can refine annotations directly on the PDF canvas, ensuring consistency between document layout and structured representations. The system’s source code is publicly available to facilitate further research and development.
Dinh-Truong Do, Hoang-An Trieu, Van-Thuy Phi, Minh Le Nguyen 0001, Yuji Matsumoto 0001
AAAI5
2026 CancerRAGent: Evidence-Linked and Safety-Guided Oncology Question Answering
Trung Vo, An Trieu, Yuji Matsumoto 0001, Minh Le Nguyen 0001
ECIR (4)4
2026 Biomedical concept recognition with error-aware negative-enhanced ranking framework
abstract
MOTIVATION: Mention-agnostic biomedical concept recognition (MA-BCR) requires inferring ontology concepts directly from passages, without relying on explicit mention spans. Prior work has mainly focused on generative and classification-based approaches. Ranking-based methods typically use a retrieve-rerank pipeline, and this paradigm has not been systematically studied for MA-BCR. Consequently, it remains unclear how ranking-based approaches compare with existing paradigms and what types of supervision are most beneficial for ranker training under limited annotation settings. RESULTS: Through a systematic comparison of ranking-, generative-, and classification-based paradigms, we show that a two-stage retrieve-rerank architecture is the most robust and scalable backbone for MA-BCR. Building on this finding, we propose ENR, an error-aware negative-enhanced ranking framework that augments training with false positives collected from heterogeneous recognizers, improving reranking performance without increasing inference-time cost. Experiments on MM-HPO and MM-GO (two datasets derived from MedMentions-ST21pv) demonstrate that ENR substantially outperforms prior approaches. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are available in Github at https://github.com/sl-633/enr-recognizer or in Zenodo at https://doi.org/10.5281/zenodo.20730803.
Noriki Nishida, Fei Cheng 0002, Takehito Utsuro, Yuji Matsumoto 0001
Bioinform.5
2026 Dissecting GraphRAG: A Modular Analysis of Knowledge Structuring for Factoid Question Answering
abstract
Abstract We present a systematic analysis of module-level design choices in GraphRAG, a retrieval-augmented generation framework that integrates structured knowledge graphs into question answering. Focusing on triple extraction, community clustering, and report generation, we evaluate multiple strategies across two knowledge-intensive benchmarks. Our results show that high-quality triple extraction is critical, as the accuracy and coverage of the resulting knowledge graph can become a bottleneck for downstream reasoning. We also find that the granularity of fundamental knowledge units, as determined by community clustering, has a significant impact on downstream performance: Achieving a balance between factual detail and topical coherence within each unit is important to enable precise and comprehensive retrieval and to facilitate effective multi-hop reasoning. In addition, simple template-based reporting outperforms LLM-based summarization in both accuracy and efficiency. These findings provide practical guidance for the structure- aware design of retrieval-augmented systems.
Noriki Nishida, Rumana Ferdous Munne, Narumi Tokunaga, Yuki Yamagata, Fei Cheng 0002, Kouji Kozaki, Yuji Matsumoto 0001
Trans. Assoc. Comput. Linguistics8
2025 Zero-Shot Entailment Learning for Ontology-Based Biomedical Annotation Without Explicit Mentions
abstract
Automatic biomedical annotation is essential for advancing medical research, diagnosis, and treatment. However, it presents significant challenges, especially when entities are not explicitly mentioned in the text, leading to difficulties in extraction of relevant information. These challenges are intensified by unclear terminology, implicit background knowledge, and the lack of labeled training data. Annotating with a specific ontology adds another layer of complexity, as it requires aligning text with a predefined set of concepts and relationships. Manual annotation is time-consuming and expensive, highlighting the need for automated systems to handle large volumes of biomedical data efficiently. In this paper, we propose an entailment-based zero-shot text classification approach to annotate biomedical text passages using the Homeostasis Imbalance Process (HOIP) ontology. Our method reformulates the annotation task as a multi-class, multi-label classification problem and uses natural language inference to classify text into related HOIP processes. Experimental results show promising performance, especially when processes are not explicitly mentioned, highlighting the effectiveness of our approach for ontological annotation of biomedical literature.
Rumana Ferdous Munne, Noriki Nishida, Narumi Tokunaga, Yuki Yamagata, Kouji Kozaki, Yuji Matsumoto 0001
COLING7
2024 Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
abstract
Enhancing user engagement through personalization in conversational agents has gained significance, especially with the advent of large language models that generate fluent responses. Personalized dialogue generation, however, is multifaceted and varies in its definition – ranging from instilling a persona in the agent to capturing users’ explicit and implicit cues. This paper seeks to systemically survey the recent landscape of personalized dialogue generation, including the datasets employed, methodologies developed, and evaluation metrics applied. Covering 22 datasets, we highlight benchmark datasets and newer ones enriched with additional features. We further analyze 17 seminal works from top conferences between 2021-2023 and identify five distinct types of problems. We also shed light on recent progress by LLMs in personalized dialogue generation. Our evaluation section offers a comprehensive summary of assessment facets and metrics utilized in these works. In conclusion, we discuss prevailing challenges and envision prospect directions for future research in personalized dialogue generation.
Yi-Pei Chen 0001, Noriki Nishida, Hideki Nakayama, Yuji Matsumoto 0001
LREC/COLING4
2024 PolyNERE: A Novel Ontology and Corpus for Named Entity Recognition and Relation Extraction in Polymer Science Domain
abstract
Polymers are widely used in diverse fields, and the demand for efficient methods to extract and organize information about them is increasing. An automated approach that utilizes machine learning can accurately extract relevant information from scientific papers, providing a promising solution for automating information extraction using annotated training data. In this paper, we introduce a polymer-relevant ontology featuring crucial entities and relations to enhance information extraction in the polymer science field. Our ontology is customizable to adapt to specific research needs. We present PolyNERE, a high-quality named entity recognition (NER) and relation extraction (RE) corpus comprising 750 polymer abstracts annotated using our ontology. Distinctive features of PolyNERE include multiple entity types, relation categories, support for various NER settings, and the ability to assert entities and relations at different levels. PolyNERE also facilitates reasoning in the RE task through supporting evidence. While our experiments with recent advanced methods achieved promising results, challenges persist in adapting NER and RE from abstracts to full-text paragraphs. This emphasizes the need for robust information extraction systems in the polymer domain, making our corpus a valuable benchmark for future developments.
Van-Thuy Phi, Hiroki Teranishi, Yuji Matsumoto 0001, Hiroyuki Oka, Masashi Ishii
LREC/COLING3
2024 A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages
abstract
User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages.
Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas 0001, Tomohiro Nishiyama, Sebastian Möller 0001, Eiji Aramaki, Yuji Matsumoto 0001, Roland Roller, Pierre Zweigenbaum
LREC/COLING12
2023 24-bit Languages
abstract
Yiran Wang, Taro Watanabe, Masao Utiyama, Yuji Matsumoto. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yiran Wang 0006, Taro Watanabe, Masao Utiyama, Yuji Matsumoto 0001
IJCNLP (1)4
2023 Switching to Discriminative Image Captioning by Relieving a Bottleneck of Reinforcement Learning
abstract
Discriminativeness is a desirable feature of image captions: captions should describe the characteristic details of input images. However, recent high-performing captioning models, which are trained with reinforcement learning (RL), tend to generate overly generic captions despite their high performance in various other criteria. First, we investigate the cause of the unexpectedly low discriminativeness and show that RL has a deeply rooted side effect of limiting the output words to high-frequency words. The limited vocabulary is a severe bottleneck for discriminativeness as it is difficult for a model to describe the details beyond its vocabulary. Then, based on this identification of the bottleneck, we drastically recast discriminative image captioning as a much simpler task of encouraging low-frequency word generation. Hinted by long-tail classification and debiasing methods, we propose methods that easily switch off-the-shelf RL models to discriminativeness-aware models with only a single-epoch fine-tuning on the part of the parameters. Extensive experiments demonstrate that our methods significantly enhance the discriminative-ness of off-the-shelf RL models and even outperform previous discriminativeness-aware methods with much smaller computational costs. Detailed analysis and human evaluation also verify that our methods boost the discriminativeness without sacrificing the overall quality of captions.1
Ukyo Honda, Taro Watanabe, Yuji Matsumoto 0001
WACV3
2022 Unsupervised Lexical Substitution with Decontextualised Embeddings
abstract
We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language models to predict substitutes, our method retrieves substitutes based on the similarity of contextualised and decontextualised word embeddings, i.e. the average contextual representation of a word in multiple contexts. We conduct experiments in English and Italian, and show that our method substantially outperforms strong baselines and establishes a new state-of-the-art without any explicit supervision or fine-tuning. We further show that our method performs particularly well at predicting low-frequency substitutes, and also generates a diverse list of substitute candidates, reducing morphophonetic or morphosyntactic biases induced by article-noun agreement.
Takashi Wada 0001, Timothy Baldwin, Yuji Matsumoto 0001, Jey Han Lau
COLING3
2022 Coordination Generation via Synchronized Text-Infilling
abstract
Generating synthetic data for supervised learning from large-scale pre-trained language models has enhanced performances across several NLP tasks, especially in low-resource scenarios. In particular, many studies of data augmentation employ masked language models to replace words with other words in a sentence. However, most of them are evaluated on sentence classification tasks and cannot immediately be applied to tasks related to the sentence structure. In this paper, we propose a simple yet effective approach to generating sentences with a coordinate structure in which the boundaries of its conjuncts are explicitly specified. For a given span in a sentence, our method embeds a mask with a coordinating conjunction in two ways (”X and [mask]”, ”[mask] and X”) and forces masked language models to fill the two blanks with an identical text. To achieve this, we introduce decoding methods for BERT and T5 models with the constraint that predictions for different masks are synchronized. Furthermore, we develop a training framework that effectively selects synthetic examples for the supervised coordination disambiguation task. We demonstrate that our method produces promising coordination instances that provide gains for the task in low-resource settings.
Hiroki Teranishi, Yuji Matsumoto 0001
COLING2
2022 Global Entity Disambiguation with BERT
abstract
Ikuya Yamada, Koki Washio, Hiroyuki Shindo, Yuji Matsumoto. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ikuya Yamada, Koki Washio, Hiroyuki Shindo, Yuji Matsumoto 0001
NAACL-HLT4
2022 Out-of-Domain Discourse Dependency Parsing via Bootstrapping: An Empirical Analysis on Its Effectiveness and Limitation
abstract
Abstract Discourse parsing has been studied for decades. However, it still remains challenging to utilize discourse parsing for real-world applications because the parsing accuracy degrades significantly on out-of-domain text. In this paper, we report and discuss the effectiveness and limitations of bootstrapping methods for adapting modern BERT-based discourse dependency parsers to out-of-domain text without relying on additional human supervision. Specifically, we investigate self-training, co-training, tri-training, and asymmetric tri-training of graph-based and transition-based discourse dependency parsing models, as well as confidence measures and sample selection criteria in two adaptation scenarios: monologue adaptation between scientific disciplines and dialogue genre adaptation. We also release COVID-19 Discourse Dependency Treebank (COVID19-DTB), a new manually annotated resource for discourse dependency parsing of biomedical paper abstracts. The experimental results show that bootstrapping is significantly and consistently effective for unsupervised domain adaptation of discourse dependency parsing, but the low coverage of accurately predicted pseudo labels is a bottleneck for further improvement. We show that active learning can mitigate this limitation.
Noriki Nishida, Yuji Matsumoto 0001
Trans. Assoc. Comput. Linguistics2
2022 Characterization of Pulmonary Nodules in Computed Tomography Images Based on Pseudo-Labeling Using Radiology Reports
abstract
A computer-aided diagnosis (CAD) system that characterizes nodules in medical images can help radiologists determine its malignancy. Preparing large volumes of labeled data for CAD systems, however, requires advanced medical knowledge. This makes it extremely difficult to develop such systems, despite their growing demand. In this paper, we propose a new training method to build an image classifier for characterization of nodules utilizing pseudo-labels, i.e., image labels automatically retrieved from radiology reports. A radiology report is a type of record in which radiologists present a summary of lesion characteristics and diagnosis. Labeling radiology reports is much easier than labeling radiology images, and can be done without high expertise. Using several thousand labeled reports, we constructed a hierarchical attention network-based text classifier to assign pseudo-labels of the characteristics of pulmonary nodules with high accuracy (macro F1-score of 0.941). Experimental results show that the image classifier trained with the pseudo-labels can achieve almost the same performance as the one trained with the labels annotated by radiologists: AUC 0.848 for the model trained with the pseudo-labels on 3,000 computed tomography (CT) images and 0.847 for the model trained with the manual labels on 800 CT images.
Yohei Momoki, Akimichi Ichinose, Yutaro Shigeto, Ukyo Honda, Keigo Nakamura, Yuji Matsumoto 0001
IEEE Trans. Circuits Syst. Video Technol.6
2021 Nested Named Entity Recognition via Explicitly Excluding the Influence of the Best Path
abstract
Yiran Wang, Hiroyuki Shindo, Yuji Matsumoto, Taro Watanabe. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yiran Wang 0006, Hiroyuki Shindo, Yuji Matsumoto 0001, Taro Watanabe
ACL/IJCNLP (1)3
2021 Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image Captioning
abstract
Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto, Taro Watanabe, Yuji Matsumoto. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Taro Watanabe, Yuji Matsumoto 0001
EACL5
2021 Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts
abstract
Textual emotion detection is a challenge in computational linguistics and affective computing study as it involves the discovery of all associated emotions expressed within a given piece of text. It becomes an even more difficult problem when applied to conversation transcripts, as we need to model the spoken utterances between speakers, keeping in mind the context of the entire conversation. In this paper, we propose a semisupervised multilabel method of predicting emotions from conversation transcripts. The corpus contains conversational quotes extracted from movies. A small number of them are annotated, while the rest are used for unsupervised training. We use the word2vec word-embedding method to build an emotion lexicon from the corpus and to embed the utterances into vector representations. A deep-learning autoencoder is then used to discover the underlying structure of the unsupervised data. We fine-tune the learned model on labeled training data, and measure its performance on a test set. The experiment result suggests that the method is effective and is only slightly behind human annotators.
Duc Anh Phan, Yuji Matsumoto 0001, Hiroyuki Shindo
IEEE Trans. Affect. Comput.2
2020 Coordination Boundary Identification without Labeled Data for Compound Terms Disambiguation
abstract
Yuya Sawada, Takashi Wada, Takayoshi Shibahara, Hiroki Teranishi, Shuhei Kondo, Hiroyuki Shindo, Taro Watanabe, Yuji Matsumoto. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Yuya Sawada, Takashi Wada 0001, Takayoshi Shibahara, Hiroki Teranishi, Shuhei Kondo, Hiroyuki Shindo, Taro Watanabe, Yuji Matsumoto 0001
COLING8
2020 LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention
abstract
Entity representations are useful in natural language tasks involving entities.In this paper, we propose new pretrained contextualized representations of words and entities based on the bidirectional transformer (Vaswani et al., 2017).The proposed model treats words and entities in a given text as independent tokens, and outputs contextualized representations of them.Our model is trained using a new pretraining task based on the masked language model of BERT (Devlin et al., 2019).The task involves predicting randomly masked words and entities in a large entity-annotated corpus retrieved from Wikipedia.We also propose an entity-aware self-attention mechanism that is an extension of the self-attention mechanism of the transformer, and considers the types of tokens (words or entities) when computing attention scores.The proposed model achieves impressive empirical performance on a wide range of entity-related tasks.In particular, it obtains state-of-the-art results on five well-known datasets: Open Entity (entity typing), TACRED (relation classification), CoNLL-2003 (named entity recognition), ReCoRD (cloze-style question answering), and SQuAD 1.1 (extractive question answering).Our source code and pretrained representations are available at https: //github.com/studio-ousia/luke.
Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 0001, Yuji Matsumoto 0001
EMNLP (1)5
2019 Scientific Article Search System Based on Discourse Facet Representation
abstract
We present a browser-based scientific article search system with graphical visualization. This system is based on triples of distributed representations of articles, each triple representing a scientific discourse facet (Objective, Method, or Result) using both text and citation information. Because each facet of an article is encoded as a separate vector, the similarity between articles can be measured by considering the articles not only in their entirety but also on a facet-by-facet basis. Our system provides three search options: a similarity ranking search, a citation graph with facet-labeled edges, and a scatter plot visualization with facets as the axes.
Yuta Kobayashi, Hiroyuki Shindo, Yuji Matsumoto 0001
AAAI3
2019 Stochastic Tokenization with a Language Model for Neural Text Classification
abstract
For unsegmented languages such as Japanese and Chinese, tokenization of a sentence has a significant impact on the performance of text classification. Sentences are usually segmented with words or subwords by a morphological analyzer or byte pair encoding and then encoded with word (or subword) representations for neural networks. However, segmentation is potentially ambiguous, and it is unclear whether the segmented tokens achieve the best performance for the target task. In this paper, we propose a method to simultaneously learn tokenization and text classification to address these problems. Our model incorporates a language model for unsupervised tokenization into a text classifier and then trains both models simultaneously. To make the model robust against infrequent tokens, we sampled segmentation for each sentence stochastically during training, which resulted in improved performance of text classification. We conducted experiments on sentiment analysis as a text classification task and show that our method achieves better performance than previous methods.
Tatsuya Hiraoka, Hiroyuki Shindo, Yuji Matsumoto 0001
ACL (1)3
2019 Unsupervised Multilingual Word Embedding with Limited Resources using Neural Language Models
abstract
Recently, a variety of unsupervised methods have been proposed that map pre-trained word embeddings of different languages into the same space without any parallel data.These methods aim to find a linear transformation based on the assumption that monolingual word embeddings are approximately isomorphic between languages.However, it has been demonstrated that this assumption holds true only on specific conditions, and with limited resources, the performance of these methods decreases drastically.To overcome this problem, we propose a new unsupervised multilingual embedding method that does not rely on such assumption and performs well under resource-poor scenarios, namely when only a small amount of monolingual data (i.e., 50k sentences) are available, or when the domains of monolingual data are different across languages.Our proposed model, which we call 'Multilingual Neural Language Models', shares some of the network parameters among multiple languages, and encodes sentences of multiple languages into the same space.The model jointly learns word embeddings of different languages in the same space, and generates multilingual embeddings without any parallel data or pre-training.Our experiments on word alignment tasks have demonstrated that, on the low-resource condition, our model substantially outperforms existing unsupervised and even supervised methods trained with 500 bilingual pairs of words.Our model also outperforms unsupervised methods given different-domain corpora across languages.Our code is publicly available 1 .
Takashi Wada 0001, Tomoharu Iwata, Yuji Matsumoto 0001
ACL (1)3
2019 ATAR: Aspect-Based Temporal Analog Retrieval System for Document Archives
abstract
In recent years, we have witnessed a rapid increase of text content stored in digital archives such as newspaper archives or web archives. With the passage of time, it is however difficult to effectively perform search within such collections due to vocabulary and context change. In this paper, we present a system that helps to find analogical terms across temporal text collections by applying non-linear transformation. We implement two approaches for analog retrieval where one of them allows users to also input an aspect term specifying particular perspective of a query. The current prototype system permits temporal analog search across two different time periods based on New York Times Annotated Corpus.
Adam Jatowt, Sourav S. Bhowmick, Yuji Matsumoto 0001
WSDM4
2018 Dynamic Feature Selection with Attention in Incremental Parsing
abstract
One main challenge for incremental transition-based parsers, when future inputs are invisible, is to extract good features from a limited local context. In this work, we present a simple technique to maximally utilize the local features with an attention mechanism, which works as context- dependent dynamic feature selection. Our model learns, for example, which tokens should a parser focus on, to decide the next action. Our multilingual experiment shows its effectiveness across many languages. We also present an experiment with augmented test dataset and demon- strate it helps to understand the model’s behavior on locally ambiguous points.
Ryosuke Kohita, Hiroshi Noji, Yuji Matsumoto 0001
COLING3
2018 A Span Selection Model for Semantic Role Labeling
abstract
We present a simple and accurate span-based model for semantic role labeling (SRL).Our model directly takes into account all possible argument spans and scores them for each label.At decoding time, we greedily select higher scoring labeled spans.One advantage of our model is to allow us to design and use spanlevel features, that are difficult to use in tokenbased BIO tagging approaches.Experimental results demonstrate that our ensemble model achieves the state-of-the-art results, 87.4 F1 and 87.0 F1 on the CoNLL-2005 and 2012 datasets, respectively.
Hiroki Ouchi, Hiroyuki Shindo, Yuji Matsumoto 0001
EMNLP3
2018 Interpretable Adversarial Perturbation in Input Embedding Space for Text
abstract
Following great success in the image processing field, the idea of adversarial training has been applied to tasks in the natural language processing (NLP) field. One promising approach directly applies adversarial training developed in the image processing field to the input word embedding space instead of the discrete input space of texts. However, this approach abandons such interpretability as generating adversarial texts to significantly improve the performance of NLP tasks. This paper restores interpretability to such methods by restricting the directions of perturbations toward the existing words in the input embedding space. As a result, we can straightforwardly reconstruct each input with perturbations to an actual text by considering the perturbations to be the replacement of words in the sentence while maintaining or even improving the task performance.
Motoki Sato, Jun Suzuki 0001, Hiroyuki Shindo, Yuji Matsumoto 0001
IJCAI4
2018 Universal Dependencies Version 2 for Japanese
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto 0001, Mai Omura, Yugo Murawaki
LREC7
2018 A Parallel Corpus of Arabic-Japanese News Articles
Go Inoue, Nizar Habash, Yuji Matsumoto 0001, Hiroyuki Aoyama
LREC3
2018 Construction of Large-scale English Verbal Multiword Expression Annotated Corpus
Akihiko Kato, Hiroyuki Shindo, Yuji Matsumoto 0001
LREC3
2018 EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text
Duc Anh Phan, Yuji Matsumoto 0001
LREC2
2018 PDFAnno: a Web-based Linguistic Annotation Tool for PDF Documents
Hiroyuki Shindo, Yohei Munesada, Yuji Matsumoto 0001
LREC3
2018 Sudachi: a Japanese Tokenizer for Business
Kazuma Takaoka, Sorami Hisamoto, Noriko Kawahara, Miho Sakamoto, Yoshitaka Uchida, Yuji Matsumoto 0001
LREC6
2018 Chemical Compounds Knowledge Visualization with Natural Language Processing and Linked Data
Kazunari Tanaka, Tomoya Iwakura, Yusuke Koyanagi, Noriko Ikeda, Hiroyuki Shindo, Yuji Matsumoto 0001
LREC6
2018 Automatic Error Correction on Japanese Functional Expressions Using Character-based Neural Machine Translation
Fei Cheng 0002, Yiran Wang 0006, Hiroyuki Shindo, Yuji Matsumoto 0001
PACLIC5
2018 Reduction of Parameter Redundancy in Biaffine Classifiers with Symmetric and Circulant Weight Matrices
Tomoki Matsuno, Katsuhiko Hayashi 0001, Takahiro Ishihara, Hitoshi Manabe, Yuji Matsumoto 0001
PACLIC5
2017 Neural Modeling of Multi-Predicate Interactions for Japanese Predicate Argument Structure Analysis
abstract
The performance of Japanese predicate argument structure (PAS) analysis has improved in recent years thanks to the joint modeling of interactions between multiple predicates.However, this approach relies heavily on syntactic information predicted by parsers, and suffers from error propagation.To remedy this problem, we introduce a model that uses grid-type recurrent neural networks.The proposed model automatically induces features sensitive to multi-predicate interactions from the word sequence information of a sentence.Experiments on the NAIST Text Corpus demonstrate that without syntactic information, our model outperforms previous syntax-dependent models.
Hiroki Ouchi, Hiroyuki Shindo, Yuji Matsumoto 0001
ACL (1)3
2017 A* CCG Parsing with a Supertag and Dependency Factored Model
abstract
We propose a new A* CCG parsing model in which the probability of a tree is decomposed into factors of CCG categories and its syntactic dependencies both defined on bi-directional LSTMs.Our factored model allows the precomputation of all probabilities and runs very efficiently, while modeling sentence structures explicitly via dependencies.Our model achieves the stateof-the-art results on English and Japanese CCG parsing. 1
Masashi Yoshikawa, Hiroshi Noji, Yuji Matsumoto 0001
ACL (1)3
2017 Joint Prediction of Morphosyntactic Categories for Fine-Grained Arabic Part-of-Speech Tagging Exploiting Tag Dictionary Information
abstract
Part-of-speech (POS) tagging for morphologically rich languages such as Arabic is a challenging problem because of their enormous tag sets.One reason for this is that in the tagging scheme for such languages, a complete POS tag is formed by combining tags from multiple tag sets defined for each morphosyntactic category.Previous approaches in Arabic POS tagging applied one model for each morphosyntactic tagging task, without utilizing shared information between the tasks.In this paper, we propose an approach that utilizes this information by jointly modeling multiple morphosyntactic tagging tasks with a multi-task learning framework.We also propose a method of incorporating tag dictionary information into our neural models by combining word representations with representations of the sets of possible tags.Our experiments showed that the joint model with tag dictionary information results in an accuracy of 91.38% on the Penn Arabic Treebank data set, with an absolute improvement of 2.11% over the current state-of-the-art tagger. 1
Go Inoue, Hiroyuki Shindo, Yuji Matsumoto 0001
CoNLL3
2017 Knowledge Transfer for Out-of-Knowledge-Base Entities : A Graph Neural Network Approach
abstract
Knowledge base completion (KBC) aims to predict missing information in a knowledge base. In this paper, we address the out-of-knowledge-base (OOKB) entity problem in KBC: how to answer queries concerning test entities not observed at training time. Existing embedding-based KBC models assume that all test entities are available at training time, making it unclear how to obtain embeddings for new entities without costly retraining. To solve the OOKB entity problem without retraining, we use graph neural networks (Graph-NNs) to compute the embeddings of OOKB entities, exploiting the limited auxiliary knowledge provided at test time. The experimental results show the effectiveness of our proposed model in the OOKB setting. Additionally, in the standard KBC setting in which OOKB entities are not involved, our model achieves state-of-the-art performance on the WordNet dataset.
Takuo Hamaguchi, Hidekazu Oiwa, Masashi Shimbo, Yuji Matsumoto 0001
IJCAI4
2017 Improving Sequence to Sequence Neural Machine Translation by Utilizing Syntactic Dependency Information
abstract
Sequence to Sequence Neural Machine Translation has achieved significant performance in recent years. Yet, there are some existing issues that Neural Machine Translation still does not solve completely. Two of them are translation for long sentences and the “over-translation”. To address these two problems, we propose an approach that utilize more grammatical information such as syntactic dependencies, so that the output can be generated based on more abundant information. In our approach, syntactic dependencies is employed in decoding. In addition, the output of the model is presented not as a simple sequence of tokens but as a linearized tree construction. In order to assess the performance, we construct model based on an attention mechanism encoder-decoder model in which the source language is input to the encoder as a sequence and the decoder generates the target language as a linearized dependency tree structure. Experiments on the Europarl-v7 dataset of French-to-English translation demonstrate that our proposed method improves BLEU scores by 1.57 and 2.40 on datasets consisting of sentences with up to 50 and 80 tokens, respectively. Furthermore, the proposed method also solved the two existing problems, ineffective translation for long sentences and over-translation in Neural Machine Translation.
An Nguyen Le, Ander Martinez, Akifumi Yoshimoto, Yuji Matsumoto 0001
IJCNLP(1)4
2017 Coordination Boundary Identification with Similarity and Replaceability
abstract
We propose a neural network model for coordination boundary detection. Our method relies on the two common properties - similarity and replaceability in conjuncts - in order to detect both similar pairs of conjuncts and dissimilar pairs of conjuncts. The model improves identification of clause-level coordination using bidirectional RNNs incorporating two properties as features. We show that our model outperforms the existing state-of-the-art methods on the coordination annotated Penn Treebank and Genia corpus without any syntactic information from parsers.
Hiroki Teranishi, Hiroyuki Shindo, Yuji Matsumoto 0001
IJCNLP(1)3
2017 Sentence Complexity Estimation for Chinese-speaking Learners of Japanese
Yuji Matsumoto 0001
PACLIC2
2017 A Fast and Easy Regression Technique for k-NN Classification Without Using Negative Pairs
Yutaro Shigeto, Masashi Shimbo, Yuji Matsumoto 0001
PAKDD (1)3
2016 Non-Linear Similarity Learning for Compositionality
abstract
Many NLP applications rely on the existence ofsimilarity measures over text data.Although word vector space modelsprovide good similarity measures between words,phrasal and sentential similarities derived from compositionof individual words remain as a difficult problem.In this paper, we propose a new method of ofnon-linear similarity learning for semantic compositionality.In this method, word representations are learnedthrough the similarity learning of sentencesin a high-dimensional space with kernel functions.On the task of predicting the semantic similarity oftwo sentences (SemEval 2014, Task 1),our method outperforms linear baselines,feature engineering approaches,recursive neural networks,and achieve competitive results with long short-term memory models.
Masashi Tsubaki, Kevin Duh, Masashi Shimbo, Yuji Matsumoto 0001
AAAI4
2016 Modelling the Usage of Discourse Connectives as Rational Speech Acts
abstract
Discourse relations can either be implicit or explicitly expressed by markers, such as 'therefore' and 'but'.How a speaker makes this choice is a question that is not well understood.We propose a psycholinguistic model that predicts whether a speaker will produce an explicit marker given the discourse relation s/he wishes to express.Based on the framework of the Rational Speech Acts model, we quantify the utility of producing a marker based on the information-theoretic measure of surprisal, the cost of production, and a bias to maintain uniform information density throughout the utterance.Experiments based on the Penn Discourse Treebank show that our approach outperforms stateof-the-art approaches, while giving an explanatory account of the speaker's choice.1 'Speakers' and 'listeners' are interchangeably used with 'authors' and 'readers' in this article
Frances Yung, Kevin Duh, Taku Komura, Yuji Matsumoto 0001
CoNLL4
2016 Joint Transition-based Dependency Parsing and Disfluency Detection for Automatic Speech Recognition Texts
abstract
Joint dependency parsing with disfluency detection is an important task in speech language processing.Recent methods show high performance for this task, although most authors make the unrealistic assumption that input texts are transcribed by human annotators.In real-world applications, the input text is typically the output of an automatic speech recognition (ASR) system, which implies that the text contains not only disfluency noises but also recognition errors from the ASR system.In this work, we propose a parsing method that handles both disfluency and ASR errors using an incremental shift-reduce algorithm with several novel features suited to ASR output texts.Because the gold dependency information is usually annotated only on transcribed texts, we also introduce an alignment-based method for transferring the gold dependency annotation to the ASR output texts to construct training data for our parser.We conducted an experiment on the Switchboard corpus and show that our method outperforms conventional methods in terms of dependency parsing and disfluency detection.
Masashi Yoshikawa, Hiroyuki Shindo, Yuji Matsumoto 0001
EMNLP3
2016 Construction of an English Dependency Corpus incorporating Compound Function Words
Akihiko Kato, Hiroyuki Shindo, Yuji Matsumoto 0001
LREC3
2016 Universal Dependencies for Japanese
Takaaki Tanaka, Yusuke Miyao, Masayuki Asahara, Sumire Uematsu, Hiroshi Kanayama, Shinsuke Mori, Yuji Matsumoto 0001
LREC7
2016 Discriminative Reranking for Grammatical Error Correction with Statistical Machine Translation
abstract
Research on grammatical error correction has received considerable attention. For dealing with all types of errors, grammatical error correction methods that employ statistical machine translation (SMT) have been proposed in recent years. An SMT system generates candidates with scores for all candidates and selects the sentence with the highest score as the correction result. However, the 1-best result of an SMT system is not always the best result. Thus, we propose a reranking approach for grammatical error correction. The reranking approach is used to re-score N-best results of the SMT and reorder the results. Our experiments show that our reranking system using parts of speech and syntactic features improves performance and achieves state-of-theart quality, with an F0.5 score of 40.0.
Tomoya Mizumoto, Yuji Matsumoto 0001
HLT-NAACL2
2016 Multiple Emotions Detection in Conversation Transcripts
Duc Anh Phan, Hiroyuki Shindo, Yuji Matsumoto 0001
PACLIC3
2016 Integrating Word Embedding Offsets into the Espresso System for Part-Whole Relation Extraction
Van-Thuy Phi, Yuji Matsumoto 0001
PACLIC2
2016 A Generalized Framework for Hierarchical Word Sequence Language Model
Xiaoyi Wu, Kevin Duh, Yuji Matsumoto 0001
PACLIC3
2016 Transition-Based Dependency Parsing Exploiting Supertags
abstract
Lexical information, including surface word form and part-of-speech (POS) information, plays a crucial role when predicting ambiguous dependency relationships in dependency parsing. However, for resolving dependency ambiguities, surface word information may be too sparse, while POS information may be too coarse. Supertags, which are lexical templates that represent rich syntactic information, have been shown to provide effective features at an intermediate level on the coarse-to-fine scale. In this work, we present a supertag design framework that allows us to instantiate various supertag sets based on the dependency structures. Using this framework, we instantiate various supertag sets and utilize them as features in transition-based dependency parsing systems. Performing experiments on the Penn Treebank and Universal Dependencies data sets, we show that our supertags are effective for transition-based parsers in multilingual parsing as well as English parsing. The comparison of the results of the different supertag sets shows that it is crucial to incorporate the head directionality, head labels, and dependent possession information in supertags to improve the parser performance.
Hiroki Ouchi, Kevin Duh, Hiroyuki Shindo, Yuji Matsumoto 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Joint Case Argument Identification for Japanese Predicate Argument Structure Analysis
abstract
Hiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Hiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto 0001
ACL (1)4
2015 Patent claim translation based on sublanguage-specific sentence structure
Masaru Fuji, Atsushi Fujita, Masao Utiyama, Eiichiro Sumita, Yuji Matsumoto 0001
MTSummit5
2015 An Efficient Annotation for Phrasal Verbs using Dependency Information
Masayuki Komai, Hiroyuki Shindo, Yuji Matsumoto 0001
PACLIC3
2015 An Improved Hierarchical Word Sequence Language Model Using Directional Information
Xiaoyi Wu, Yuji Matsumoto 0001
PACLIC2
2015 Ridge Regression, Hubness, and Zero-Shot Learning
Yutaro Shigeto, Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Yuji Matsumoto 0001
ECML/PKDD (1)5
2015 A Hybrid Ranking Approach to Chinese Spelling Check
abstract
We propose a novel framework for Chinese Spelling Check (CSC), which is an automatic algorithm to detect and correct Chinese spelling errors. Our framework contains two key components: candidate generation and candidate ranking . Our framework differs from previous research, such as Statistical Machine Translation (SMT) based model or Language Model (LM) based model, in that we use both SMT and LM models as components of our framework for generating the correction candidates, in order to obtain maximum recall; to improve the precision, we further employ a Support Vector Machines (SVM) classifier to rank the candidates generated by the SMT and the LM. Experiments show that our framework outperforms other systems, which adopted the same or similar resources as ours in the SIGHAN 7 shared task; even comparing with the state-of-the-art systems, which used more resources, such as a considerable large dictionary, an idiom dictionary and other semantic information, our framework still obtains competitive results. Furthermore, to address the resource scarceness problem for training the SMT model, we generate around 2 million artificial training sentences using the Chinese character confusion sets, which include a set of Chinese characters with similar shapes and similar pronunciations, provided by the SIGHAN 7 shared task.
Xiaodong Liu 0003, Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2015 Multilingual Topic Models for Bilingual Dictionary Extraction
abstract
A machine-readable bilingual dictionary plays a crucial role in many natural language processing tasks, such as statistical machine translation and cross-language information retrieval. In this article, we propose a framework for extracting a bilingual dictionary from comparable corpora by exploiting a novel combination of topic modeling and word aligners such as the IBM models. Using a multilingual topic model, we first convert a comparable document -aligned corpus into a parallel topic -aligned corpus. This novel topic-aligned corpus is similar in structure to the sentence -aligned corpus frequently employed in statistical machine translation and allows us to extract a bilingual dictionary using a word alignment model. The main advantages of our framework is that (1) no seed dictionary is necessary for bootstrapping the process, and (2) multilingual comparable corpora in more than two languages can also be exploited. In our experiments on a large-scale Wikipedia dataset, we demonstrate that our approach can extract higher precision dictionaries compared to previous approaches and that our method improves further as we add more languages to the dataset.
Xiaodong Liu 0003, Kevin Duh, Yuji Matsumoto 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2014 Improving Dependency Parsers with Supertags
abstract
Transition-based dependency parsing systems can utilize rich feature representations.However, in practice, features are generally limited to combinations of lexical tokens and part-of-speech tags.In this paper, we investigate richer features based on supertags, which represent lexical templates extracted from dependency structure annotated corpus.First, we develop two types of supertags that encode information about head position and dependency relations in different levels of granularity.Then, we propose a transition-based dependency parser that incorporates the predictions from a CRF-based supertagger as new features.On standard English Penn Treebank corpus, we show that our supertag features achieve parsing improvements of 1.3% in unlabeled attachment, 2.07% root attachment, and 3.94% in complete tree accuracy.
Hiroki Ouchi, Kevin Duh, Yuji Matsumoto 0001
EACL3
2014 Analysis and Prediction of Unalignable Words in Parallel Text
abstract
Professional human translators usually do not employ the concept of word alignments, producing translations 'sense-forsense' instead of 'word-for-word'.This suggests that unalignable words may be prevalent in the parallel text used for machine translation (MT).We analyze this phenomenon in-depth for Chinese-English translation.We further propose a simple and effective method to improve automatic word alignment by pre-removing unalignable words, and show improvements on hierarchical MT systems in both translation directions.
Frances Yung, Kevin Duh, Yuji Matsumoto 0001
EACL3
2014 Parsing Chinese Synthetic Words with a Character-based Dependency Model
Fei Cheng 0002, Kevin Duh, Yuji Matsumoto 0001
LREC3
2014 Collocation or Free Combination? ― Applying Machine Translation Techniques to identify collocations in Japanese
Lis Pereira, Elga Strafella, Yuji Matsumoto 0001
LREC3
2014 A Hierarchical Word Sequence Language Model
Xiaoyi Wu, Yuji Matsumoto 0001
PACLIC2
2013 Topic Models + Word Alignment = A Flexible Framework for Extracting Bilingual Dictionary from Comparable Corpus
Xiaodong Liu 0003, Kevin Duh, Yuji Matsumoto 0001
CoNLL3
2013 Modeling and Learning Semantic Co-Compositionality through Prototype Projections and Neural Networks
abstract
We present a novel vector space model for semantic co-compositionality.Inspired by Generative Lexicon Theory (Pustejovsky, 1995), our goal is a compositional model where both predicate and argument are allowed to modify each others' meaning representations while generating the overall semantics.This readily addresses some major challenges with current vector space models, notably the polysemy issue and the use of one representation per word type.We implement cocompositionality using prototype projections on predicates/arguments and show that this is effective in adapting their word representations.We further cast the model as a neural network and propose an unsupervised algorithm to jointly train word representations with co-compositionality.The model achieves the best result to date (ρ = 0.47) on the semantic similarity task of transitive verbs (Grefenstette and Sadrzadeh, 2011).
Masashi Tsubaki, Kevin Duh, Masashi Shimbo, Yuji Matsumoto 0001
EMNLP4
2013 What Information is Helpful for Dependency Based Semantic Role Labeling
Yanyan Luo, Kevin Duh, Yuji Matsumoto 0001
IJCNLP3
2013 Towards Automatic Error Type Classification of Japanese Language Learners' Writings
Hiromi Oyama, Mamoru Komachi, Yuji Matsumoto 0001
PACLIC3
2013 Efficient Stacked Dependency Parsing by Forest Reranking
abstract
This paper proposes a discriminative forest reranking algorithm for dependency parsing that can be seen as a form of efficient stacked parsing. A dynamic programming shift-reduce parser produces a packed derivation forest which is then scored by a discriminative reranker, using the 1-best tree output by the shift-reduce parser as guide features in addition to third-order graph-based features. To improve efficiency and accuracy, this paper also proposes a novel shift-reduce parser that eliminates the spurious ambiguity of arc-standard transition systems. Testing on the English Penn Treebank data, forest reranking gave a state-of-the-art unlabeled dependency accuracy of 93.12.
Katsuhiko Hayashi 0001, Shuhei Kondo, Yuji Matsumoto 0001
Trans. Assoc. Comput. Linguistics3
2012 Investigating the Effectiveness of Laplacian-Based Kernels in Hub Reduction
abstract
A “hub” is an object closely surrounded by, or very similar to, many other objects in the dataset. Recent studies by Radovanovi´c et al. indicate that in high dimensional spaces, hubs almost always emerge, and objects close to the data centroid tend to become hubs. In this paper, we show that the family of kernels based on the graph Laplacian makes all objects in the dataset equally similar to the centroid, and thus they are expected to make less hubs when used as a similarity measure. We investigate this hypothesis using both synthetic and real-world data. It turns out that these kernels suppress hubs in some cases but not always, and the results seem to be affected by the size of the data—a factor not discussed previously. However, for the datasets in which hubs are indeed reduced by the Laplacian-based kernels, these kernels work well in ranking and classification tasks. This result suggests that the amount of hubs, which can be readily computed in an unsupervised fashion, can be a yardstick of whether Laplacian-based kernels work effectively for a given data.
Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Yuji Matsumoto 0001, Marco Saerens
AAAI4
2012 Head-driven Transition-based Parsing with Top-down Prediction
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
ACL (1)4
2012 Walk-based Computation of Contextual Word Similarity
Kazuo Hara, Ikumi Suzuki, Masashi Shimbo, Yuji Matsumoto 0001
COLING4
2012 Joint English Spelling Error Correction and POS Tagging for Language Learners Writing
Keisuke Sakaguchi, Tomoya Mizumoto, Mamoru Komachi, Yuji Matsumoto 0001
COLING4
2012 UniDic for Early Middle Japanese: a Dictionary for Morphological Analysis of Classical Japanese
Toshinobu Ogiso, Mamoru Komachi, Yasuharu Den, Yuji Matsumoto 0001
LREC4
2012 Things between Lexicon and Grammar
Yuji Matsumoto 0001
PACLIC1
2012 Mining Rules for Rewriting States in a Transition-Based Dependency Parser
Akihiro Inokuchi, Ayumu Yamaoka, Takashi Washio, Yuji Matsumoto 0001, Masayuki Asahara, Masakazu Iwatate, Hideto Kazawa
PRICAI4
2011 Transfer Learning for Multiple-Domain Sentiment Analysis - Identifying Domain Dependent/Independent Word Polarity
abstract
Sentiment analysis is the task of determining the attitude (positive or negative) of documents. While the polarity of words in the documents is informative for this task, polarity of some words cannot be determined without domain knowledge. Detecting word polarity thus poses a challenge for multiple-domain sentiment analysis. Previous approaches tackle this problem with transfer learning techniques, but they cannot handle multiple source domains and multiple target domains. This paper proposes a novel Bayesian probabilistic model to handle multiple source and multiple target domains. In this model, each word is associated with three factors: Domain label, domain dependence/independence and word polarity. We derive an efficient algorithm using Gibbs sampling for inferring the parameters of the model, from both labeled and unlabeled texts. Using real data, we demonstrate the effectiveness of our model in a document polarity classification task compared with a method not considering the differences between domains. Moreover our method can also tell whether each word's polarity is domain-dependent or domain-independent. This feature allows us to construct a word polarity dictionary for each domain.
Yasuhisa Yoshida, Tsutomu Hirao, Tomoharu Iwata, Masaaki Nagata, Yuji Matsumoto 0001
AAAI5
2011 Co-related Verb Argument Selectional Preferences
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001
CICLing (1)3
2011 Using the Mutual k-Nearest Neighbor Graphs for Semi-supervised Classification on Natural Language Data
Kohei Ozaki, Masashi Shimbo, Mamoru Komachi, Yuji Matsumoto 0001
CoNLL4
2011 Multilayer Sequence Labeling
Ai Azuma, Yuji Matsumoto 0001
EMNLP2
2011 Third-order Variational Reranking on Packed-Shared Dependency Forests
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
EMNLP4
2011 Japanese Predicate Argument Structure Analysis Exploiting Argument Position and Type
Yuta Hayashibe, Mamoru Komachi, Yuji Matsumoto 0001
IJCNLP3
2011 Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction of Second Language Learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, Yuji Matsumoto 0001
IJCNLP4
2011 Automatic Labeling of Voiced Consonants for Morphological Analysis of Modern Japanese Literature
Teruaki Oka, Mamoru Komachi, Toshinobu Ogiso, Yuji Matsumoto 0001
IJCNLP4
2011 Jointly Extracting Japanese Predicate-Argument Relation with Markov Logic
Katsumasa Yoshikawa, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2011 Dependency-based Analysis for Tagalog Sentences
Erlyn Manguilimotan, Yuji Matsumoto 0001
PACLIC2
2010 Annotating Event Mentions in Text with Modality, Focus, and Source Information
Suguru Matsuyoshi, Megumi Eguchi, Chitose Sao, Koji Murakami, Kentaro Inui, Yuji Matsumoto 0001
LREC6
2009 Coordinate Structure Analysis with Global Structural Constraints and Alignment-Based Local Features
Kazuo Hara, Masashi Shimbo, Hideharu Okuma, Yuji Matsumoto 0001
ACL/IJCNLP4
2009 Capturing Salience with a Trainable Cache Model for Zero-anaphora Resolution
Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001
ACL/IJCNLP3
2009 Jointly Identifying Temporal Relations with Markov Logic
Katsumasa Yoshikawa, Sebastian Riedel 0001, Masayuki Asahara, Yuji Matsumoto 0001
ACL/IJCNLP4
2009 Learning Co-relations of Plausible Verb Arguments with a WSM and a Distributional Thesaurus
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001
CIARP3
2009 Interpolated PLSI for Learning Plausible Verb Arguments
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001
PACLIC3
2009 Factors Affecting Part-of-Speech Tagging for Tagalog
Erlyn Manguilimotan, Yuji Matsumoto 0001
PACLIC2
2009 A Generalization of Forward-Backward Algorithm
Ai Azuma, Yuji Matsumoto 0001
ECML/PKDD (1)2
2009 On the properties of von Neumann kernels for link analysis
Masashi Shimbo, Takahiko Ito, Daichi Mochihashi, Yuji Matsumoto 0001
Mach. Learn.4
2008 Two-Phased Event Relation Acquisition: Coupling the Relation-Oriented and Argument-Oriented Approaches
Shuya Abe, Kentaro Inui, Yuji Matsumoto 0001
COLING3
2008 Japanese Dependency Parsing Using a Tournament Model
Masakazu Iwatate, Masayuki Asahara, Yuji Matsumoto 0001
COLING3
2008 Emotion Classification Using Massive Examples Extracted from the Web
Ryoko Tokuhisa, Kentaro Inui, Yuji Matsumoto 0001
COLING3
2008 Training Conditional Random Fields Using Incomplete Annotations
Yuta Tsuboi, Hisashi Kashima, Shinsuke Mori, Hiroki Oda, Yuji Matsumoto 0001
COLING5
2008 A Pipeline Approach for Syntactic and Semantic Dependency Parsing
Yotaro Watanabe, Masakazu Iwatate, Masayuki Asahara, Yuji Matsumoto 0001
CoNLL4
2008 Graph-based Analysis of Semantic Drift in Espresso-like Bootstrapping Algorithms
Mamoru Komachi, Taku Kudo, Masashi Shimbo, Yuji Matsumoto 0001
EMNLP4
2008 Learning by switching generation and reasoning methods - acquisition of meta-knowledge for switching with reinforcement learning
abstract
When we generate knowledge, we initially have no knowledge and acquire it by observing data one by one. We memorize the raw data when the number of observed data is small and generate general knowledge when it becomes large. To simulate this learning process, we proposed a learning model with switching several knowledge representation and reasoning methods. In this model, the time when to switch is decided with the fixed rules. These rules are considered to be meta-knowledge because they control the learning process. In this paper, we propose a method acquiring the meta-knowledge for deciding the time of switching knowledge representation or reasoning method. For learning of the meta-knowledge, the correct answers can not to be given but just the evaluation of the learning process. We use Q-learning, therefore, a method of reinforcement learning. In the simulation, we apply the method to the iris plant data to acquire the meta-knowledge. The system with the acquired meta-knowledge has smaller number of rules than the old method for the similar rate correctly classified.
Masahiro Tomaru, Motohide Umano, Yuji Matsumoto 0001, Kazuhisa Seta
FUZZ-IEEE3
2008 Acquiring Event Relation Knowledge by Learning Cooccurrence Patterns and Fertilizing Cooccurrence Samples with Verbal Nouns
Shuya Abe, Kentaro Inui, Yuji Matsumoto 0001
IJCNLP3
2008 Generic Text Summarization Using Probabilistic Latent Semantic Indexing
Harendra Bhandari, Masashi Shimbo, Takahiko Ito, Yuji Matsumoto 0001
IJCNLP4
2008 Use of Event Types for Temporal Relation Identification in Chinese Text
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2008 Analyzing Chinese Synthetic Words with Tree-based Information and a Survey on Chinese Morphologically Derived Words
Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2008 Japanese-Spanish Thesaurus Construction Using English as a Pivot
Jessica C. Ramírez, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2008 Visualization with Voronoi tessellation and moving output units in Self-Organizing map of the real-number system
abstract
The Self-Organizing map (SOM) proposed by T. Kohonen is a method to produce a low-dimensional representation from high-dimensional input data automatically, where output units are restrictedly placed on grid points. We propose real-number SOM (RSOM), where output units are freely placed on the real-number coordinates plane and visualized as a Voronoi diagram. RSOM is a natural extension of the conventional SOM because Voronoi tessellation for the output units on the square grid generates square regions on the output plane, the same as the conventional SOM. We propose two methods of moving with preserving topology of the input data and several visualization method such as minimum spanning tree, variable boundary width and spherical RSOM. We illustrate moving methods decrease errors in results of simulation.
Yuji Matsumoto 0001, Motohide Umano, Masahiro Inuiguchi
IJCNN1
2008 Large Scale Corpus Analysis and Recent Applications
Yuji Matsumoto 0001
PRICAI1
2007 Extracting Aspect-Evaluation and Aspect-Of Relations in Opinion Mining
Nozomi Kobayashi, Kentaro Inui, Yuji Matsumoto 0001
EMNLP-CoNLL3
2007 A Graph-Based Approach to Named Entity Categorization in Wikipedia Using Conditional Random Fields
Yotaro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
EMNLP-CoNLL3
2007 Constructing a Temporal Relation Tagged Corpus of Chinese Based on Dependency Structure Analysis
abstract
This paper describes an annotation guideline for a temporal relation tagged corpus. Our goal is to construct a machine learnable model that automatically analyzes temporal events and relations between events. Since analyzing all combinations of events is inefficient, we examine use of dependency structure analysis to efficiently recognize meaningful temporal relations. We survey a small tagged data set to investigate the coverage of our method. Although the coverage of our methods is about 49%, we find that the dependency structure appears useful for reducing manual efforts in constructing a tagged corpus with temporal relations.
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
TIME3
2007 Zero-anaphora resolution by learning rich syntactic pattern features
abstract
We approach the zero-anaphora resolution problem by decomposing it into intrasentential and intersentential zero-anaphora resolution tasks. For the former task, syntactic patterns of zeropronouns and their antecedents are useful clues. Taking Japanese as a target language, we empirically demonstrate that incorporating rich syntactic pattern features in a state-of-the-art learning-based anaphora resolution model dramatically improves the accuracy of intrasentential zero-anaphora, which consequently improves the overall performance of zero-anaphora resolution.
Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001
ACM Trans. Asian Lang. Inf. Process.3
2006 Exploiting Syntactic Patterns as Clues in Zero-Anaphora Resolution
abstract
We approach the zero-anaphora resolution problem by decomposing it into intra-sentential and inter-sentential zero-anaphora resolution. For the former problem, syntactic patterns of the appearance of zero-pronouns and their antecedents are useful clues. Taking Japanese as a target language, we empirically demonstrate that incorporating rich syntactic pattern features in a state-of-the-art learning-based anaphora resolution model dramatically improves the accuracy of intra-sentential zero-anaphora, which consequently improves the overall performance of zero-anaphora resolution.
Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001
ACL3
2006 Guessing Parts-of-Speech of Unknown Words Using Global Information
abstract
In this paper, we present a method for guessing POS tags of unknown words using local and global information. Although many existing methods use only local information (i.e. limited window size or intra-sentential features), global information (extra-sentential features) provides valuable clues for predicting POS tags of unknown words. We propose a probabilistic model for POS guessing of unknown words using global information as well as local information, and estimate its parameters using Gibbs sampling. We also attempt to apply the model to semi-supervised learning, and conduct experiments on multiple corpora.
Tetsuji Nakagawa, Yuji Matsumoto 0001
ACL2
2006 Multi-lingual Dependency Parsing at NAIST
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
CoNLL3
2006 Learning by Switching Knowledge Representations-Limiting the Number of Stored Data
abstract
When we solve a problem, we initially have no knowledge and we memorize the raw data with observing data. Finally we have general knowledge for solving the problem. To simulate this learning process, we proposed a learning method with switching different levels of knowledge representations, reconstructing knowledge and switching reasoning methods. In the system, all given data are stored to generate new knowledge, but it is different from the one of our human's knowledge acquisition, in which we just memorize a limit number of data. Therefore, we limit it and when the number of stored data exceeds specified size, the system throws away the oldest data. In the simulation, we apply the method to the data set whose classes are changed periodically, and get a better result than the old method.
Yuji Matsumoto 0001, Motohide Umano, Masahiro Tomaru, Kazuhisa Seta
FUZZ-IEEE1
2006 Augmenting a Semantic Verb Lexicon with a Large Scale Collection of Example Sentences
Kentaro Inui, Toru Hirano, Ryu Iida, Atsushi Fujita, Yuji Matsumoto 0001
LREC5
2006 An Annotated Corpus Management Tool: ChaKi
Yuji Matsumoto 0001, Masayuki Asahara, Kiyota Hashimoto, Yukio Tono, Akira Ohtani, Toshio Morita
LREC1
2006 The Construction of a Dictionary for a Two-layer Chinese Morphological Analyzer
Chooi-Ling Goh, Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
PACLIC5
2006 Exploring Multiple Communities with Kernel-Based Link Analysis
Takahiko Ito, Masashi Shimbo, Daichi Mochihashi, Yuji Matsumoto 0001
PKDD4
2005 Exploiting Lexical Conceptual Structure for Paraphrase Generation
Atsushi Fujita, Kentaro Inui, Yuji Matsumoto 0001
IJCNLP3
2005 Building a Japanese-Chinese Dictionary Using Kanji/Hanzi Conversion
Chooi-Ling Goh, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2005 Automatic Extraction of Fixed Multiword Expressions
Campbell Hore, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2005 Application of kernels to link analysis
abstract
The application of kernel methods to link analysis is explored. In particular, Kandola et al.'s Neumann kernels are shown to subsume not only the co-citation and bibliographic coupling relatedness but also Kleinberg's HITS importance. These popular measures of relatedness and importance correspond to the Neumann kernels at the extremes of their parameter range, and hence these kernels can be interpreted as defining a spectrum of link analysis measures intermediate between co-citation/bibliographic coupling and HITS. We also show that the kernels based on the graph Laplacian, including the regularized Laplacian and diffusion kernels, provide relatedness measures that overcome some limitations of co-citation relatedness. The property of these kernel-based link analysis measures is examined with a network of bibliographic citations. Practical issues in applying these methods to real data are discussed, and possible solutions are proposed.
Takahiko Ito, Masashi Shimbo, Taku Kudo, Yuji Matsumoto 0001
KDD4
2005 Context as Filtering
abstract
Long-distance language modeling is important not only in speech recognition and machine translation, but also in high-dimensional discrete sequence modeling in general. However, the problem of context length has almost been neglected so far and a nave bag-of-words history has been i employed in natural language processing. In contrast, in this paper we view topic shifts within a text as a latent stochastic process to give an explicit probabilistic generative model that has partial exchangeability. We propose an online inference algorithm using particle filters to recognize topic shifts to employ the most appropriate length of context automatically. Experiments on the BNC corpus showed consistent improvement over previous methods involving no chronological order.
Daichi Mochihashi, Yuji Matsumoto 0001
NIPS2
2005 Anaphora resolution by antecedent identification followed by anaphoricity determination
abstract
We propose a machine learning-based approach to noun-phrase anaphora resolution that combines the advantages of previous learning-based models while overcoming their drawbacks. Our anaphora resolution process reverses the order of the steps in the classification-then-search model proposed by Ng and Cardie [2002b], inheriting all the advantages of that model. We conducted experiments on resolving noun-phrase anaphora in Japanese. The results show that with the selection-then-classification-based modifications, our proposed model outperforms earlier learning-based approaches.
Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001
ACM Trans. Asian Lang. Inf. Process.3
2005 Acquiring causal knowledge from text using the connective marker tame
abstract
In this paper, we deal with automatic knowledge acquisition from text, specifically the acquisition of causal relations . A causal relation is the relation existing between two events such that one event causes (or enables) the other event, such as “hard rain causes flooding” or “taking a train requires buying a ticket.” In previous work these relations have been classified into several types based on a variety of points of view. In this work, we consider four types of causal relations--- cause , effect , precond(ition) and means ---mainly based on agents' volitionality, as proposed in the research field of discourse understanding. The idea behind knowledge acquisition is to use resultative connective markers, such as “because,” “but,” and “if” as linguistic cues. However, there is no guarantee that a given connective marker always signals the same type of causal relation. Therefore, we need to create a computational model that is able to classify samples according to the causal relation. To examine how accurately we can automatically acquire causal knowledge, we attempted an experiment using Japanese newspaper articles, focusing on the resultative connective “tame.” By using machine-learning techniques, we achieved 80% recall with over 95% precision for the cause , precond , and means relations, and 30% recall with 90% precision for the effect relation. Furthermore, the classification results suggest that one can expect to acquire over 27,000 instances of causal relations from 1 year of Japanese newspaper articles.
Takashi Inui, Kentaro Inui, Yuji Matsumoto 0001
ACM Trans. Asian Lang. Inf. Process.3
2005 Introduction to the special issue: Recent advances in information processing and access for Japanese
abstract
No abstract available.
Tetsuya Sakai, Yuji Matsumoto 0001
ACM Trans. Asian Lang. Inf. Process.2
2004 Japanese Unknown Word Identification by Character-based Chunking
Masayuki Asahara, Yuji Matsumoto 0001
COLING2
2004 Trajectory Based Word Sense Disambiguation
Yuji Matsumoto 0001
COLING2
2004 Modeling Category Structures with a Kernel Function
Hiroya Takamura, Yuji Matsumoto 0001, Hiroyasu Yamada
CoNLL2
2004 A Boosting Algorithm for Classification of Semi-Structured Text
Taku Kudo, Yuji Matsumoto 0001
EMNLP2
2004 Applying Conditional Random Fields to Japanese Morphological Analysis
Taku Kudo, Kaoru Yamamoto, Yuji Matsumoto 0001
EMNLP3
2004 Deterministic Dependency Structure Analyzer for Chinese
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP3
2004 Detection of Incorrect Case Assignments in Paraphrase Generation
Atsushi Fujita, Kentaro Inui, Yuji Matsumoto 0001
IJCNLP3
2004 Practical Translation Pattern Acquisition from Combined Language Resources
Mihoko Kitamura, Yuji Matsumoto 0001
IJCNLP2
2004 Collecting Evaluative Expressions for Opinion Extraction
Nozomi Kobayashi, Kentaro Inui, Yuji Matsumoto 0001, Kenji Tateishi, Toshikazu Fukushima
IJCNLP3
2004 Improving Word Sense Disambiguation by Pseudo-samples
Yuji Matsumoto 0001
IJCNLP2
2004 Building a Paraphrase Corpus for Speech Translation
Mitsuo Shimohata, Eiichiro Sumita, Yuji Matsumoto 0001
LREC3
2004 An Application of Boosting to Graph Classification
abstract
This paper presents an application of Boosting for classifying labeled graphs, general structures for modeling a number of real-world data, such as chemical compounds, natural language texts, and bio sequences. The proposal consists of i) decision stumps that use subgraph as features, and ii) a Boosting algorithm in which subgraph-based decision stumps are used as weak learners. We also discuss the relation between our al- gorithm and SVMs with convolution kernels. Two experiments using natural language data and chemical compounds show that our method achieves comparable or even better performance than SVMs with convo- lution kernels as well as improves the testing efficiency. 1 Introduction Most machine learning (ML) algorithms assume that given instances are represented in numerical vectors. However, much real-world data is not represented as numerical vectors, but as more complicated structures, such as sequences, trees, or graphs. Examples include biological sequences (e.g., DNA and RNA), chemical compounds, natural language texts, and semi-structured data (e.g., XML and HTML documents). Kernel methods, such as support vector machines (SVMs) [11], provide an elegant solution to handling such structured data. In this approach, instances are implicitly mapped into a high-dimensional space, where information about their similarities (inner-products) is only used for constructing a hyperplane for classification. Recently, a number of kernels have been proposed for such structured data, such as sequences [7], trees [2, 5], and graphs [6]. Most are based on the idea that a feature vector is implicitly composed of the counts of substructures (e.g., subsequences, subtrees, subpaths, or subgraphs). Although kernel methods show remarkable performance, their implicit definitions of fea- ture space make it difficult to know what kind of features (substructures) are relevant or which features are used in classifications. To use ML algorithms for data mining or as knowledge discovery tools, they must output a list of relevant features (substructures). This information may be useful not only for a detailed analysis of individual data but for the hu- man decision-making process. In this paper, we present a new machine learning algorithm for classifying labeled graphs that has the following characteristics: 1) It performs learning and classification using the Figure 1: Labeled connected graphs and subgraph relation structural information of a given graph. 2) It uses a set of all subgraphs (bag-of-subgraphs) as a feature set without any constraints, which is essentially the same idea as a convolution kernel [4]. 3) Even though the size of the candidate feature set becomes quite large, it automatically selects a compact and relevant feature set based on Boosting. 2 Classifier for Graphs We first assume that an instance is represented in a labeled graph. The focused problem can be formalized as a general problem called the graph classification problem. The graph classification problem is to induce a mapping f (x) : X {1}, from given training examples T = { xi, yi }L i=1, where xi X is a labeled graph and yi {1} is a class label associated with the training data. We here focus on the problem of binary classifica- tion. The important characteristic is that input example xi is represented not as a numerical feature vector but as a labeled graph. 2.1 Preliminaries In this paper we focus on undirected, labeled, and connected graphs, since we can easily extend our algorithm to directed or unlabeled graphs with minor modifications. Let us in- troduce a labeled connected graph (or simply a labeled graph), its definitions and notations. Definition 1 Labeled Connected Graph A labeled graph is represented in a 4-tuple G = (V, E, L, l), where V is a set of vertices, E V V is a set of edges, L is a set of labels, and l : V E L is a mapping that assigns labels to the vertices and the edges. A labeled connected graph is a labeled graph such that there is a path between any pair of verticies. Definition 2 Subgraph Let G = (V , E , L , l ) and G = (V, E, L, l) be labeled connected graphs. G matches G, or G is a subgraph of G (G G) if the following conditions are satisfied: (1) V V , (2) E E, (3) L L, and (4) l = l. If G is a subgraph of G, then G is a supergraph of G . Figure 1 shows an example of a labeled graph and its subgraph and non-subgraph. 2.2 Decision Stumps Decision stumps are simple classifiers in which the final decision is made by a single hy- pothesis or feature. Boostexter [10] uses word-based decision stumps for text classification. To classify graphs, we define the subgraph-based decision stumps as follows. Definition 3 Decision Stumps for Graphs Let t and x be labeled graphs and y be a class label (y {1}). A decision stump classifier for graphs is given by y t x h t,y (x) def = -y otherwise. The parameter for classification is a tuple t, y , hereafter referred to as a rule of decision stumps. The decision stumps are trained to find a rule ^ t, ^ y that minimizes the error rate for the given training data T = { xi, yi }L i=1: L L ^ 1 1 t, ^ y = argmin I(yi = h t,y (xi)) = argmin (1 - yih t,y (xi)), (1) tF ,y{1} L tF ,y{1} 2L i=1 i=1 where F is a set of candidate graphs or a feature set (i.e., F = L {t|t x i=1 i}) and I () is the indicator function. The gain function for a rule t, y is defined as L gain( t, y ) def = yih t,y (xi). (2) i=1 Using the gain, the search problem (1) becomes equivalent to the problem: ^ t, ^ y = argmaxtF,y{1} gain( t, y ). In this paper, we use gain instead of error rate for clarity. 2.3 Applying Boosting The decision stump classifiers are too inaccurate to be applied to real applications, since the final decision relies on the existence of a single graph. However, accuracies can be boosted by the Boosting algorithm [3, 10]. Boosting repeatedly calls a given weak learner and finally produces a hypothesis f , which is a linear combination of K hypotheses produced by the weak learners, i,e.: f (x) = sgn( K (x)). A weak learner is built k=1 k h tk,yk at each iteration k with different distributions or weights d(k) = (d(k), . . . , d(k)) on the i L training data, where L d(k) = 1, d(k) 0. The weights are calculated to concentrate i=1 i i more on hard examples than easy examples. To use decision stumps as the weak learner of Boosting, we redefine the gain function (2) as: L gain( t, y ) def = yidih t,y (xi). (3) i=1 In this paper, we use the AdaBoost algorithm, the original and the best known algorithm among many variants of Boosting. However, it is trivial to fit our decision stumps to other boosting algorithms, such as Arc-GV [1] and Boosting with soft margins [8]. 3 Efficient Computation In this section, we introduce an efficient and practical algorithm to find the optimal rule ^ t, ^ y from given training data. This problem is formally defined as follows. Problem 1 Find Optimal Rule Let T = { x1, y1, d1 , . . . , xL, yL, dL } be training data where xi is a labeled graph, yi {1} is a class label associated with xi and di ( L d i=1 i = 1, di 0) is a normal- ized weight assigned to xi. Given T , find the optimal rule ^ t, ^ y that maximizes the gain, i.e., ^ t, ^ y = argmax {t|t x tF ,y{1} diyih t,y , where F = L i=1 i}. The most naive and exhaustive method in which we first enumerate all subgraphs F and then calculate the gains for all subgraphs is usually impractical, since the number of sub- graphs is exponential to its size. We thus adopt an alternative strategy to avoid such ex- haustive enumerations. The method to find the optimal rule is modeled as a variant of branch-and-bound algorithm and will be summarized as the following strategies: 1) Define Figure 2: Example of DFS Code Tree for a graph a canonical search space in which a whole set of subgraphs can be enumerated. 2) Find the optimal rule by traversing this search space. 3) Prune the search space by proposing a criteria for the upper bound of the gain. We will describe these steps more precisely in the next subsections. 3.1 Efficient Enumeration of Graphs Yan et al. proposed an efficient depth-first search algorithm to enumerate all subgraphs from a given graph [12]. The key idea of their algorithm is a DFS (depth first search) code, a lexicographic order to the sequence of edges. The search tree given by the DFS code is called a DFS Code Tree. Leaving the details to [12], the order of the DFS code is defined by the lexicographic order of labels as well as the topology of graphs. Figure 2 illustrates an example of a DFS Code Tree. Each node in this tree is represented in a 5-tuple [i, j, vi, eij, vj], where eij, vi and vj are the labels of i-j edge, i-th vertex, and j-th vertex respectively. By performing a pre-order search of the DFS Code Tree, we can obtain all the subgraphs of a graph in order of their DFS code. However, one cannot avoid isomorphic enumerations even giving pre-order traverse, since one graph can have several DFS codes in a DFS Code Tree. So, canonical DFS code (minimum DFS code) is defined as its first code in the pre-order search of the DFS Code Tree. Yan et al. show that two graphs G and G are isomorphic if and only if minimum DFS codes for the two graphs min(G) and min(G ) are the same. We can thus ignore non-minimum DFS codes in subgraph enumerations. In other words, in depth-first traverse, we can prune a node with DFS code c, if c is not minimum. The isomorphic graph represented in minimum code has already been enumerated in the depth-first traverse. For example, in Figure 2, if G1 is identical to G0, G0 has been discovered before the node for G1 is reached. This property allows us to avoid an explicit isomorphic test of the two graphs. 3.2 Upper bound of gain DFS Code Tree defines a canonical search space in which one can enumerate all subgraphs from a given set of graphs. We consider an upper bound of the gain that allows pruning of subspace in this canonical search space. The following lemma gives a convenient method of computing a tight upper bound on gain( t , y ) for any supergraph t of t. Lemma 1 Upper bound of the gain: (t) For any t t and y {1}, the gain of t , y is bounded by (t) (i.e., gain( t y ) (t)), where (t) is given by L L def (t) = max 2 di - yi di, 2 di + yi di . {i|yi=+1,txi} i=1 {i|yi=-1,txi} i=1 Proof 1 L L gain( t , y ) = diyih t ,y (xi) = diyi y (2I(t xi) - 1), i=1 i=1 where I() is the indicator function. If we focus on the case y = +1, then L L gain( t , +1 ) = 2 yidi - yi di 2 di - yi di {i|t xi} i=1 {i|yi=+1,t xi} i=1 L 2 di - yi di, {i|yi=+1,txi} i=1 since |{i|yi = +1, t xi}| |{i|yi = +1, t xi}| for any t t. Similarly, L gain( t , -1 ) 2 di + yi di. {i|yi=-1,txi} i=1 Thus, for any t t and y {1}, gain( t , y ) (t). 2 We can efficiently prune the DFS Code Tree using the upper bound of gain u(t). During pre-order traverse in a DFS Code Tree, we always maintain the temporally suboptimal gain among all the gains calculated previously. If (t) < , the gain of any supergraph t t is no greater than , and therefore we can safely prune the search space spanned from the subgraph t. If (t) , then we cannot prune this space since a supergraph t t might exist such that gain(t ) . 3.3 Efficient Computation in Boosting At each Boosting iteration, the suboptimal value is reset to 0. However, if we can calcu- late a tighter upper bound in advance, the search space can be pruned more effectively. For this purpose, a cache is used to maintain all rules found in the previous iterations. Subop- timal value is calculated by selecting one rule from the cache that maximizes the gain of the current distribution. This idea is based on our observation that a rule in the cache tends to be reused as the number of Boosting iterations increases. Furthermore, we also maintain the search space built by a DFS Code Tree as long as memory allows. This cache reduces duplicated constructions of a DFS Code Tree at each Boosting iteration. 4 Connection to Convolution Kernel Recent studies [1, 9, 8] have shown that both Boosting and SVMs [11] work according to similar strategies: constructing an optimal hypothesis that maximizes the smallest margin between positive and negative examples. The difference between the two algorithms is the metric of margin; the margin of Boosting is measured in l1-norm, while that of SVMs is measured in l2-norm. We describe how maximum margin properties are translated in the two algorithms. AdaBoost and Arc-GV asymptotically solve the following linear program, [1, 9, 8], J max ; s.t. yi wjhj(xi) , ||w||1 = 1 (4) wIRJ ,IR+ j=1 where J is the number of hypotheses. Note that in the case of decision stumps for graphs, J = |{1} F | = 2|F |. SVMs, on the other hand, solve the following quadratic optimization problem [11]: 1 max ; s.t. yi (w (xi)) , ||w||2 = 1. (5) wIRJ ,IR+ 1For simplicity, we omit the bias term (b) and the extension of Soft Margin. The function (x) maps the original input example x into a J -dimensional feature vector (i.e., (x) IRJ ). The l2-norm margin gives the separating hyperplane expressed by dot- products in feature space. The feature space in SVMs is thus expressed implicitly by using a Marcer kernel function, which is a generalized dot-product between two objects, (i.e., K(x1, x2) = (x1) (x2)). The best known kernel for modeling structured data is a convolution kernel [4] (e.g., string kernel [7] and tree kernel [2, 5]), which argues that a feature vector is implicitly composed of the counts of substructures. 2 The implicit mapping defined by the convolution kernel is given as: (x) = (#(t1 x), . . . , #(t|F| x)), where tj F and #(u) is the cardinality of u. Noticing that a decision stump can be expressed as h t,y (x) = y (2I(t x) - 1), we see that the constraints or feature space of Boosting with substructure-based decision stumps are essentially the same as those of SVMs with the convolution kernel 3. The critical difference is the definition of margin: Boosting uses l1-norm, and SVMs use l2-norm. The difference between them can be explained by sparseness. It is well known that the solution or separating hyperplane of SVMs is expressed in a linear combination of training examples using coefficients , (i.e., w = L i=1 i(xi)) [11]. Maximizing l2-norm margin gives a sparse solution in the example space, (i.e., most of i becomes 0). Examples having non-zero coefficients are called support vectors that form the final solution. Boosting, in contrast, performs the computation explicitly in feature space. The concept behind Boosting is that only a few hypotheses are needed to express the final solution. l1-norm margin realizes such a property [8]. Boosting thus finds a sparse solution in the feature space. The accuracies of these two methods depend on the given training data. However, we argue that Boosting has the following practical advantages. First, sparse hypotheses allow the construction of an efficient classification algorithm. The complexity of SVMs with tree kernel is O(l|n1||n2|), where n1 and n2 are trees, and l is the number of support vectors, which is too heavy to be applied to real applications. Boosting, in contrast, performs faster since the complexity depends only on a small number of decision stumps. Second, sparse hypotheses are useful in practice as they provide "transparent" models with which we can analyze how the model performs or what kind of features are useful. It is difficult to give such analysis with kernel methods since they define feature space implicitly. 5 Experiments and Discussion To evaluate our algorithm, we employed two experiments using two real-world data. (1) Cellphone review classification (REV) The goal of this task is to classify reviews for cellphones as positive or negative. 5,741 sen- tences were collected from an Web-BBS discussion about cellphones in which users were directed to submit positive reviews separately from negative reviews. Each sentence is rep- resented in a word-based dependency tree using a Japanese dependency parser CaboCha4. (2) Toxicology prediction of chemical compounds (PTC) The task is to classify chemical compounds by carcinogenicity. We used the PTC data set5 consisting of 417 compounds with 4 types of test animals: male mouse (MM), female 2Strictly speaking, graph kernel [6] is not a convolution kernel because it is not based on the count of subgraphs, but on random walks in a graph. 3The difference between decision stumps and the convolution kernels is that the former uses a binary feature denoting the existence (or absence) of each substructure, whereas the latter uses the cardinality of each substructure. However, it makes little difference since a given graph is often sparse and the cardinality of substructures will be approximated by their existence. 4http://chasen.naist.jp/~ taku/software/cabocha/ 5http://www.predictive-toxicology.org/ptc/ Table 1: Classification F-scores of the REV and PTC tasks REV PTC MM FM MR FR Boosting BOL-based Decision Stumps 76.6 47.0 52.9 42.7 26.9 Subgraph-based Decision Stumps 79.0 48.9 52.5 55.1 48.5 SVMs BOL Kernel 77.2 40.9 39.9 43.9 21.8 Tree/Graph Kernel 79.4 42.3 34.1 53.2 25.9 mouse (FM), male rat (MR) and female rat (FR). Each compound is assigned one of the following labels: {EE,IS,E,CE,SE,P,NE,N}. We here assume that CE,SE, and P are "posi- tive" and that NE and NN are "negative", which is exactly the same setting as [6]. We thus have four binary classifiers (MM/FM/MR/FR) in this data set. We compared the performance of our Boosting algorithm and support vector machines with tree kernel [2, 5] (for REV) and graph kernel [6] (for PTC) according to their F-score in 5-fold cross validation. Table 1 summarizes the best results of REV and PCT task, varying the hyperparameters of Boosting and SVMs (e.g., maximum iteration of Boosting, soft margin parameter of SVMs, and termination probability of random walks in graph kernel [6]). We also show the results with bag-of-label (BOL) features as a baseline. In most tasks and categories, ML algorithms with structural features outperform the baseline systems (BOL). These re- sults support our first intuition that structural features are important for the classification of structured data, such as natural language texts and chemical compounds. Comparing our Boosting algorithm with SVMs using tree kernel, no significant difference can be found the REV data set. However, in the PTC task, our method outperforms SVMs using graph kernel on the categories MM, FM, and FR at a statistically significant level. Furthermore, the number of active features (subgraphs) used in Boosting is much smaller than those of SVMs. With our methods, about 1800 and 50 features (subgraphs) are used in the REV and PTC tasks respectively, while the potential number of features is quite large. Even giving all subgraphs as feature candidates, Boosting selects a small and highly relevant subset of features. Figure 3 show an example of extracted support features (subgraphs) in the REV and PTC task respectively. In the REV task, features reflecting the domain knowledge (cellphone reviews) are extracted: 1) "want to use " positive, 2) "hard to use" negative, 3) "recharging time is short" positive, 4) "recharging time is long" negative. These features are interesting because we cannot determine the correct label (positive/negative) only using such bag-of-label features as "charging," "short," or "long." In the PTC task, similar structures show different behavior. For instance, Trihalomethanes (TTHMs), well- known carcinogenic substances (e.g., chloroform, bromodichloromethane, and chlorodi- bromomethane), contain the common substructure H-C-Cl (Fig. 3(a)). However, TTHMs do not contain the similar but different structure H-C(C)-Cl (Fig. 3(b)). Such structural information is useful for analyzing how the system classifies the input data in a category and what kind of features are used in the classification. We cannot examine such analysis in kernel methods, since they define their feature space implicitly. The reason why graph kernel shows poor performance on the PTC data set is that it cannot identify subtle difference between two graphs because it is based on a random walks in a graph. For example, kernel dot-product between the similar but different structures 3(c) and 3(d) becomes quite large, although they show different behavior. To classify chemical compounds by their functions, the system must be capable of capturing subtle differences among given graphs. The testing speed of our Boosting algorithm is also much faster than SVMs with tree/graph Figure 3: Support features and their weights kernels. In the REV task, the speed of Boosting and SVMs are 0.135 sec./1,149 instances and 57.91 sec./1,149 instances respectively6. Our method is significantly faster than SVMs with tree/graph kernels without a discernible loss of accuracy.
Taku Kudo, Eisaku Maeda, Yuji Matsumoto 0001
NIPS3
2004 Pruning False Unknown Words to Improve Chinese Word Segmentation
Chooi-Ling Goh, Masayuki Asahara, Yuji Matsumoto 0001
PACLIC3
2004 Machine Learning based NLP : Experiences and Supporting Tools
Yuji Matsumoto 0001
PACLIC1
2004 Japanese Subjects and Information Structure : A Constraint-based Approach
Akira Ohtani, Yuji Matsumoto 0001
PACLIC2
2004 Use of morphological analysis in protein name recognition
Kaoru Yamamoto, Taku Kudo, Akihiko Konagaya, Yuji Matsumoto 0001
J. Biomed. Informatics4
2003 Feedback Cleaning of Machine Translation Rules Using Automatic Evaluation
abstract
When rules of transfer-based machine translation (MT) are automatically acquired from bilingual corpora, incorrect/redundant rules are generated due to acquisition errors or translation variety in the corpora. As a new countermeasure to this problem, we propose a feedback cleaning method using automatic evaluation of MT quality, which removes incorrect/redundant rules as a way to increase the evaluation score. BLEU is utilized for the automatic evaluation. The hill-climbing algorithm, which involves features of this task, is applied to searching for the optimal combination of rules. Our experiments show that the MT quality improves by 10% in test sentences according to a subjective evaluation. This is considerable improvement over previous methods.
Kenji Imamura, Eiichiro Sumita, Yuji Matsumoto 0001
ACL3
2003 Fast Methods for Kernel-Based Text Analysis
abstract
Kernel-based learning (e.g., Support Vector Machines) has been successfully applied to many hard problems in Natural Language Processing (NLP). In NLP, although feature combinations are crucial to improving performance, they are heuristically selected. Kernel methods change this situation. The merit of the kernel methods is that effective feature combination is implicitly expanded without loss of generality and increasing the computational costs. Kernel-based text analysis shows an excellent performance in terms in accuracy; however, these methods are usually too slow to apply to large-scale text analysis. In this paper, we extend a Basket Mining algorithm to convert a kernel-based classifier into a simple and fast linear classifier. Experimental results on English BaseNP Chunking, Japanese Word Segmentation and Japanese Dependency Parsing show that our new classifiers are about 30 to 300 times faster than the standard kernel-based classifiers.
Taku Kudo, Yuji Matsumoto 0001
ACL2
2003 What Kinds and Amounts of Causal Knowledge Can Be Acquired from Text by Using Connective Markers as Clues?
Takashi Inui, Kentaro Inui, Yuji Matsumoto 0001
Discovery Science3
2003 Automatic Construction of Machine Translation Knowledge Using Translation Literalness
Kenji Imamura, Eiichiro Sumita, Yuji Matsumoto 0001
EACL3
2003 Example-based rough translation for speech-to-speech translation
abstract
Example-based machine translation (EBMT) is a promising translation method for speech-to-speech translation (S2ST) because of its robustness. However, it has two problems in that the performance degrades when input sentences are long and when the style of the input sentences and that of the example corpus are different. This paper proposes example-based rough translation to overcome these two problems. The rough translation method relies on “meaning-equivalent sentences,” which share the main meaning with an input sentence despite missing some unimportant information. This method facilitates retrieval of meaning-equivalent sentences for long input sentences. The retrieval of meaning-equivalent sentences is based on content words, modality, and tense. This method also provides robustness against the style differences between the input sentence and the example corpus.
Mitsuo Shimohata, Eiichiro Sumita, Yuji Matsumoto 0001
MTSummit3
2003 Japanese Named Entity Extraction with Redundant Morphological Analysis
Masayuki Asahara, Yuji Matsumoto 0001
HLT-NAACL2
2003 The diversity-based approach to open-domain text summarization
Tadashi Nomoto, Yuji Matsumoto 0001
Inf. Process. Manag.2
2002 Revision Learning and its Application to Part-of-Speech Tagging
abstract
This paper presents a revision learning method that achieves high performance with small computational cost by combining a model with high generalization capacity and a model with small computational cost. This method uses a high capacity model to revise the output of a small cost model. We apply this method to English part-of-speech tagging and Japanese morphological analysis, and show that the method performs well.
Tetsuji Nakagawa, Taku Kudo, Yuji Matsumoto 0001
ACL3
2002 Supervised Ranking in Open-Domain Text Summarization
abstract
The paper proposes and empirically motivates an integration of supervised learning with unsupervised learning to deal with human biases in summarization. In particular, we explore the use of probabilistic decision tree within the clustering framework to account for the variation as well as regularity in human created summaries. The corpus of human created extracts is created from a newspaper corpus and used as a test set. We build probabilistic decision trees of different flavors and integrate each of them with the clustering framework. Experiments with the corpus demonstrate that the mixture of the two paradigms generally gives a significant boost in performance compared to cases where either of the two is considered alone.
Tadashi Nomoto, Yuji Matsumoto 0001
ACL2
2002 Extracting Important Sentences with Support Vector Machines
Tsutomu Hirao, Hideki Isozaki, Eisaku Maeda, Yuji Matsumoto 0001
COLING4
2002 Detecting Errors in Corpora Using Support Vector Machines
Tetsuji Nakagawa, Yuji Matsumoto 0001
COLING2
2002 Japanese Dependency Analysis using Cascaded Chunking
Taku Kudo, Yuji Matsumoto 0001
CoNLL2
2002 Two-dimensional Clustering for Text Categorization
Hiroya Takamura, Yuji Matsumoto 0001
CoNLL2
2002 Learning by switching generation and reasoning methods in several knowledge representations towards the simulation of human learning process
abstract
When we solve a problem, we firstly have no knowledge and gradually acquire some piece of knowledge by observing new data, and at last arrive at complete knowledge for solving the problem. We have a simple form of specific knowledge in the first stage and a complex form of a general one in the final stage. To simulate this kind of learning mechanism, we must combine several kinds of learning methods in several stages. We proposed a method of not only reconstructing rules and switching reasoning methods in each knowledge representation but also switching rule generation methods in several knowledge representation. We simulated the method by applying to the iris classification problem.
Motohide Umano, Yuji Matsumoto 0001, Yushi Uno, Kazuhisa Seta
FUZZ-IEEE2
2002 Use of XML and Relational Databases for Consistent Development and Maintenance of Lexicons and Annotated Corpora
Masayuki Asahara, Ryuichi Yoneda, Akiko Yamashita, Yasuharu Den, Yuji Matsumoto 0001
LREC5
2002 Modeling (in)variability of human judgments for text summarization
abstract
The paper proposes and empirically motivates an integration of supervised learning with unsupervised learning to deal with human biases in summarization. In particular, we explore the use of probabilistic decision tree within the clustering framework to account for the variation as well as regularity in human created summaries.
Tadashi Nomoto, Yuji Matsumoto 0001
SIGIR2
2001 Feature Space Restructuring for SVMs with Application to Text Categorization
Hiroya Takamura, Yuji Matsumoto 0001
EMNLP2
2001 An Experimental Comparison of Supervised and Unsupervised Approaches to Text Summarization
abstract
The paper presents a direct comparison of supervised and unsupervised approaches to text summarization. As a representative supervised method, we use the C4.5 decision tree algorithm, extended with the minimum description length principle (MDL), and compare it against several unsupervised methods. It is found that a particular unsupervised method based on an extension of the K-means clustering algorithm, performs equal to and in some cases superior to the decision tree based method.
Tadashi Nomoto, Yuji Matsumoto 0001
ICDM2
2001 Chunking with Support Vector Machines
Taku Kudo, Yuji Matsumoto 0001
NAACL2
2001 An HPSG Account of the Hierarchical Clause Formation in Japanese : HPSG-Based Japanese Grammar for Practical Parsing
Takashi Miyata, Akira Otani, Yuji Matsumoto 0001
PACLIC3
2001 A New Approach to Unsupervised Text Summarization
abstract
The paper presents a novel approach to unsupervised text summarization. The novelty lies in exploiting the diversity of concepts in text for summarization, which has not received much attention in the summarization literature. A diversity-based approach here is a principled generalization of Maximal Marginal Relevance criterion by Carbonell and Goldstein \cite{carbonell-goldstein98}.
Tadashi Nomoto, Yuji Matsumoto 0001
SIGIR2
2000 Extended Models and Tools for High-performance Part-of-speech
Masayuki Asahara, Yuji Matsumoto 0001
COLING2
2000 Acquisition of Phrase-level Bilingual Correspondence using Dependency Structure
Kaoru Yamamoto, Yuji Matsumoto 0001
COLING2
2000 Japanese Dependency Structure Analysis Based on Support Vector Machines
abstract
This paper presents a method of Japanese dependency structure analysis based on Sup-port Vector Machines (SVMs).Conventional parsing techniques based on Machine Learning framework, such as Decision Trees and Maximum Entropy Models, have difficulty in selecting useful features as well as finding appropriate combination of selected features.On the other hand, it is well-known that SVMs achieve high generalization performance even with input data of very high dimensional feature space.Furthermore, by introducing the Kernel principle, SVMs can carry out the training in high-dimensional • spaces with a smaller computational cost independent of their dimensionality.We apply SVMs to Japanese dependency structure identification problem.Experimental results on Kyoto University corpus show that our system achieves the accuracy of 89.09% even with small training data (7958 sentences).
Taku Kudo, Yuji Matsumoto 0001
EMNLP2
2000 Comparing the Minimum Description Length Principle and Boosting in the Automatic Analysis of Discourse
Tadashi Nomoto, Yuji Matsumoto 0001
ICML2
2000 Using Machine Learning Methods to Improve Quality of Tagged Corpora and Learning Models
Yuji Matsumoto 0001, Tatsuo Yamashita
LREC1
1999 Learning Discourse Relations with Active Data Selection
Tadashi Nomoto, Yuji Matsumoto 0001
EMNLP2
1998 Japanese Dependency Structure Analysis based on Lexicalized Statistics
Masakazu Fujio, Yuji Matsumoto 0001
EMNLP2
1998 Minimum detection error training for acoustic signal monitoring
abstract
In this paper we propose a novel approach to the detection of acoustic irregular signals using minimum detection error (MDE) training. The MDE training is based on the generalized probabilistic descent method, which was originally developed as a general concept for a discriminative pattern recognizer design. We demonstrate its fundamental utility by experiments in which several acoustic events are detected in a noisy environment.
Hideyuki Watanabe, Yuji Matsumoto 0001, Shigeru Katagiri
ICASSP2
1998 A fast method for statistical grammar induction
Wide R. Hogenhout, Yuji Matsumoto 0001
Nat. Lang. Eng.2
1997 Mistake-Driven Mixture of Hierarchical Tag Context Trees
abstract
This paper proposes a mistake-driven mixture method for learning a tag model. The method iteratively performs two procedures: 1. constructing a tag model based on the current data distribution and 2. updating the distribution by focusing on data that are not well predicted by the constructed model. The final tag model is constructed by mixing all the models according to their performance. To well reflect the data distribution, we represent each tag model as a hierarchical tag (i.e., NTT <proper noun
Masahiko Haruno, Yuji Matsumoto 0001
ACL2
1997 Automatic Extraction of Aspectual Information from a Monolingual Corpus
abstract
This paper describes an approach to extract the aspectual information of Japanese verb phrases from a monoligual corpus. We classify verbs into six categories by means of the aspectual features which are defined on the basis of the possibility of co-occurrence with aspectual forms and adverbs. A unique category could be identified for 96% of the target verbs. To evaluate the result of the experiment, we examined the meaning of -teiru which is one of the most fundamental aspectual markers in Japanese, and obtained the correct recognition score of 71% for the 200 sentences.
Akira Oishi, Yuji Matsumoto 0001
ACL2
1997 A Preliminary Study of Word Clustering Based on Syntactic Behavior
Wide R. Hogenhout, Yuji Matsumoto 0001
CoNLL2
1996 Towards a More Careful Evaluation of Broad Coverage Parsing Systems
Wide R. Hogenhout, Yuji Matsumoto 0001
COLING2
1996 Reversible delayed lexical choice in a bidirectional framework
Graham Wilcock, Yuji Matsumoto 0001
COLING2
1996 A Proposal of Korean Conjugation System and its Application to Morphological Analysis
Yoshitaka Hirano, Yuji Matsumoto 0001
PACLIC2
1996 Fast Statistical Grammar Induction
Wide R. Hogenhout, Yuji Matsumoto 0001
PACLIC2
1995 Integration of Syntactic, Semantic and Contextual Information in Processing Grammatically Ill-Formed Inputs
Osamu Imaichi, Yuji Matsumoto 0001
IJCAI2
1995 HMM Parameter Learning for Japanese Morphological Analyzer
Koichi Takeuchi, Yuji Matsumoto 0001
PACLIC2
1994 Bilingual Text, Matching using Bilingual Dictionary and Statistics
Takehito Utsuro, Hiroshi Ikeda, Masaya Yamane, Yuji Matsumoto 0001, Makoto Nagao
COLING4
1993 Bidirectional Chart Generation of Natural Language Texts
Masahiko Haruno, Yasuharu Den, Yuji Matsumoto 0001, Makoto Nagao
AAAI3
1993 Sructural Matching of Parallel Texts
abstract
This paper describes a method for finding structural matching between parallel sentences of two languages, (such as Japanese and English). Parallel sentences are analyzed based on unification grammars, and structural matching is performed by making use of a similarity measure of word pairs in the two languages. Syntactic ambiguities are resolved simultaneously in the matching process. The results serve as a useful source for extracting linguistic and lexical knowledge.
Yuji Matsumoto 0001, Hiroyuki Ishimoto, Takehito Utsuro
ACL1
1993 Verbal Case Frame Acquisition from Bilingual Corpora
Takehito Utsuro, Yuji Matsumoto 0001, Makoto Nagao
IJCAI2
1992 Lexical Knowledge Acquisition from Bilingual Corpora
Takehito Utsuro, Yuji Matsumoto 0001, Makoto Nagao
COLING2
1987 A Parsing System Based on Logic Programming
Yuji Matsumoto 0001, Ryôichi Sugimura
IJCAI1
1986 A Parallel Parsing System for Natural Language Analysis
Yuji Matsumoto 0001
ICLP1
1982 Prolog Interpreter Based on Concurrent Programming
Koichi Furukawa, Katsumi Nitta, Yuji Matsumoto 0001
ICLP3