Yoshimasa Tsuruoka

dblp:18/3787 · DBLP profile ↗
← Back
64ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0002-0707-1077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
abstract
Neologism-aware machine translation 1 aims to translate source sentences containing neologisms into target languages.This field remains underexplored compared with general machine translation (MT).In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary.The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump.The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump.We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologismaware machine translation.Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits "translation difficulty" to further improve the translation quality of translation agents using our search toolkit 2 .
Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)4
2024 Musical Scene Detection in Comics: Comparing Perception of Humans and GPT-4
abstract
This paper aims to detect musical scenes for improving background music generation for comics. Musical scenes can be defined as scenes where background music is enhancing the media experience. A significant barrier in detecting such scenes for comics is the absence of a ground truth to compare against. This makes our work one of the first in this field to address this task. Hence, we first analyse musical scenes through the dialogues in anime films adapted from Japanese comics (manga). Through our analysis we discover that, apart from the existing literature on dimensions of change in scene segmentation, musical scenes are also triggered by emotions. We then build an internal dataset to engineer prompts for GPT-4 to recognise musical scenes. The results of the generated outcomes are evaluated against the internal dataset as well as human evaluation. We are able to prove that prompt engineering can objectively improve the musical scene detection. Our results indicate that GPT-4 has 62.5% agreement rate with human evaluators. The implications of this research on background music generation are relevant for various media. This paper invites further investigation on this topic.
Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka
IEEE Big Data4
2024 MJ-DLVAT: A Deep Learning Value Assessment Technique for Mahjong
abstract
Abstract-In games with stochastic outcomes, evaluating agent performance from limited data is challenging. Results of Monte Carlo sampling do not provide a reliable indicator due to the significant variance. The difficulty of evaluating agents is particularly prominent in mahjong, an incomplete information game with a huge state space. For example, Suphx, which outperformed humans in mahjong, played 5, 760 games against humans in online mahjong to evaluate its performance, which took as long as four months. In this study, we propose MJ-DLVAT, a Deep Learning Value Assessment Technique for Mahjong, which is an evaluation method for mahjong players that provides an unbiased estimate of average ranking with reduced variance. MJ-DLVAT introduces three techniques to manage the extensive game tree and board information in mahjong: splitting the game into subgames, dealing with the variance caused by drawn tiles, dealt tiles and hidden-dora, and introducing neural networks. We created a dataset using online mahjong records and trained a neural network-based value function from scratch. We evaluated MJ-DLVAT on the online mahjong records. We confirmed that the average estimated rankings are unbiased estimators of average ranking and the variance of the estimated ranking is $45.5 \%$ smaller than that of the average ranking. As a result, the number of games required to correctly evaluate a player’s ability is reduced by $45.5 \%$.
Takuya Ogami, Katsutoshi Amano, Yoshimasa Tsuruoka
CoG3
2024 Leveraging Multi-lingual Positive Instances in Contrastive Learning to Improve Sentence Embedding
abstract
Learning multilingual sentence embeddings is a fundamental task in natural language processing.Recent trends in learning both monolingual and multilingual sentence embeddings are mainly based on contrastive learning (CL) among an anchor, one positive, and multiple negative instances.In this work, we argue that leveraging multiple positives should be considered for multilingual sentence embeddings because (1) positives in a diverse set of languages can benefit cross-lingual learning, and (2) transitive similarity across multiple positives can provide reliable structural information for learning.In order to investigate the impact of multiple positives in CL, we propose a novel approach, named MPCL, to effectively utilize multiple positive instances to improve the learning of multilingual sentence embeddings.Experimental results on various backbone models and downstream tasks demonstrate that MPCL leads to better retrieval, semantic similarity, and classification performance compared to conventional CL.We also observe that in unseen languages, sentence embedding models trained on multiple positives show better cross-lingual transfer performance than models trained on a single positive instance.
Kaiyan Zhao, Qiyu Wu 0001, Xin-Qiang Cai, Yoshimasa Tsuruoka
EACL (1)4
2024 Word Alignment as Preference for Machine Translation
abstract
The problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena.In this work, we mitigate the problem in an LLM-based MT model by guiding it to better word alignment.We first study the correlation between word alignment and the phenomena of hallucination and omission in MT.Then we propose to utilize word alignment as preference to optimize the LLM-based MT model.The preference data are constructed by selecting chosen and rejected translations from multiple MT tools.Subsequently, direct preference optimization is used to optimize the LLM-based model towards the preference signal.Given the absence of evaluators specifically designed for hallucination and omission in MT, we further propose selecting hard instances and utilizing GPT-4 to directly evaluate the performance of the models in mitigating these issues.We verify the rationality of these designed evaluation methods by experiments, followed by extensive results demonstrating the effectiveness of word alignment-based preference optimization to mitigate hallucination and omission.On the other hand, although it shows promise in mitigating hallucination and omission, the overall performance of MT in different language directions remains mixed, with slight increases in BLEU and decreases in COMET.
Qiyu Wu 0001, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka
EMNLP4
2024 Interpretable Imitation Learning with Symbolic Rewards
abstract
Sample inefficiency of deep reinforcement learning methods is a major obstacle for their use in real-world tasks as they naturally feature sparse rewards. In fact, this from-scratch approach is often impractical in environments where extreme negative outcomes are possible. Recent advances in imitation learning have improved sample efficiency by leveraging expert demonstrations. Most work along this line of research employs neural network-based approaches to recover an expert cost function. However, the complexity and lack of transparency make neural networks difficult to trust and deploy in the real world. In contrast, we present a method for extracting interpretable symbolic reward functions from expert data, which offers several advantages. First, the learned reward function can be parsed by a human to understand, verify and predict the behavior of the agent. Second, the reward function can be improved and modified by an expert. Finally, the structure of the reward function can be leveraged to extract explanations that encode richer domain knowledge than standard scalar rewards. To this end, we use an autoregressive recurrent neural network that generates hierarchical symbolic rewards represented by simple symbolic trees. The recurrent neural network is trained via risk-seeking policy gradients. We test our method in MuJoCo environments as well as a chemical plant simulator. We show that the discovered rewards can significantly accelerate the training process and achieve similar or better performance than neural network-based algorithms.
Nicolas Bougie, Takashi Onishi, Yoshimasa Tsuruoka
ACM Trans. Intell. Syst. Technol.3
2023 WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction
abstract
Most existing word alignment methods rely on manual alignment datasets or parallel corpora, which limits their usefulness.Here, to mitigate the dependence on manual data, we broaden the source of supervision by relaxing the requirement for correct, fully-aligned, and parallel sentences.Specifically, we make noisy, partially aligned, and non-parallel paragraphs.We then use such a large-scale weakly-supervised dataset for word alignment pre-training via span prediction.Extensive experiments with various settings empirically demonstrate that our approach, which is named WSPAlign, is an effective and scalable way to pre-train word aligners without manual data.When fine-tuned on standard benchmarks, WSPAlign has set a new state of the art by improving upon the best supervised baseline by 3.3~6.1 points in F1 and 1.5~6.1 points in AER .Furthermore, WSPAlign also achieves competitive performance compared with the corresponding baselines in few-shot, zero-shot and cross-lingual tests, which demonstrates that WSPAlign is potentially more practical for low-resource languages than existing methods. 1(1) Data Collection and Annotation (2) Pre-training for word alignment Transformer Encoder
Qiyu Wu 0001, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)3
2022 Pretraining with Artificial Language: Studying Transferable Knowledge in Language Models
abstract
We investigate what kind of structural knowledge learned in neural network encoders is transferable to processing natural language.We design artificial languages with structural properties that mimic natural language, pretrain encoders on the data, and see how much performance the encoder exhibits on downstream tasks in natural language.Our experimental results show that pretraining with an artificial language with a nesting dependency structure provides some knowledge transferable to natural language.A follow-up probing analysis indicates that its success in the transfer is related to the amount of encoded contextual information and what is transferred is the knowledge of position-aware context dependence of language.Our results provide insights into how neural network encoders process human languages and the source of cross-lingual transferability of recent multilingual language models.
Ryokan Ri, Yoshimasa Tsuruoka
ACL (1)2
2022 mLUKE: The Power of Entity Representations in Multilingual Pretrained Language Models
abstract
Recent studies have shown that multilingual pretrained language models can be effectively improved with cross-lingual alignment information from Wikipedia entities.However, existing methods only exploit entity information in pretraining and do not explicitly use entities in downstream tasks.In this study, we explore the effectiveness of leveraging entity representations for downstream cross-lingual tasks.We train a multilingual language model with 24 languages with entity representations and show the model consistently outperforms word-based pretrained models in various crosslingual transfer tasks.We also analyze the model and the key insight is that incorporating entity representations into the input allows us to extract more language-agnostic features.We also evaluate the model with a multilingual cloze prompt task with the mLAMA dataset.We show that entity-based prompt elicits correct factual knowledge more likely than using only word representations.Our source code and pretrained models are available at https: //github.com/studio-ousia/luke.
Ryokan Ri, Ikuya Yamada, Yoshimasa Tsuruoka
ACL (1)3
2022 A Multilingual Bag-of-Entities Model for Zero-Shot Cross-Lingual Text Classification
abstract
We present a multilingual bag-of-entities model that effectively boosts the performance of zeroshot cross-lingual text classification by extending a multilingual pre-trained language model (e.g., M-BERT).It leverages the multilingual nature of Wikidata: entities in multiple languages representing the same concept are defined with a unique identifier.This enables entities described in multiple languages to be represented using shared embeddings.A model trained on entity features in a resource-rich language can thus be directly applied to other languages.Our experimental results on crosslingual topic classification (using the MLDoc and TED-CLDC datasets) and entity typing (using the SHINRA2020-ML dataset) show that the proposed model consistently outperforms state-of-the-art models.
Sosuke Nishikawa, Ikuya Yamada, Yoshimasa Tsuruoka, Isao Echizen
CoNLL3
2022 Dropout Q-Functions for Doubly Efficient Reinforcement Learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, Yoshimasa Tsuruoka
ICLR5
2022 EASE: Entity-Aware Contrastive Learning of Sentence Embedding
abstract
Sosuke Nishikawa, Ryokan Ri, Ikuya Yamada, Yoshimasa Tsuruoka, Isao Echizen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Sosuke Nishikawa, Ryokan Ri, Ikuya Yamada, Yoshimasa Tsuruoka, Isao Echizen
NAACL-HLT4
2021 Meta-Model-Based Meta-Policy Optimization
abstract
Model-based meta-reinforcement learning (RL) methods have recently been shown to be a promising approach to improving the sample efficiency of RL in multi-task settings. However, the theoretical understanding of those methods is yet to be established, and there is currently no theoretical guarantee of their performance in a real-world environment. In this paper, we analyze the performance guarantee of model-based meta-RL methods by extending the theorems proposed by Janner et al. (2019). On the basis of our theoretical results, we propose Meta-Model-Based Meta-Policy Optimization (M3PO), a model-based meta-RL method with a performance guarantee. We demonstrate that M3PO outperforms existing meta-RL methods in continuous-control benchmarks.
Takuya Hiraoka, Takahisa Imagawa, Voot Tangkaratt, Takayuki Osa, Takashi Onishi, Yoshimasa Tsuruoka
ACML6
2021 Modeling Target-side Inflection in Placeholder Translation
abstract
Placeholder translation systems enable the users to specify how a specific phrase is translated in the output sentence. The system is trained to output special placeholder tokens and the user-specified term is injected into the output through the context-free replacement of the placeholder token. However and this approach could result in ungrammatical sentences because it is often the case that the specified term needs to be inflected according to the context of the output and which is unknown before the translation. To address this problem and we propose a novel method of placeholder translation that can inflect specified terms according to the grammatical construction of the output sentence. We extend the seq2seq architecture with a character-level decoder that takes the lemma of a user-specified term and the words generated from the word-level decoder to output a correct inflected form of the lemma. We evaluate our approach with a Japanese-to-English translation task in the scientific writing domain and and show our model can incorporate specified terms in a correct form more successfully than other comparable models.
Ryokan Ri, Toshiaki Nakazawa, Yoshimasa Tsuruoka
MTSummit (1)3
2020 Revisiting the Context Window for Cross-lingual Word Embeddings
abstract
Existing approaches to mapping-based crosslingual word embeddings are based on the assumption that the source and target embedding spaces are structurally similar.The structures of embedding spaces largely depend on the cooccurrence statistics of each word, which the choice of context window determines.Despite this obvious connection between the context window and mapping-based cross-lingual embeddings, their relationship has been underexplored in prior work.In this work, we provide a thorough evaluation, in various languages, domains, and tasks, of bilingual embeddings trained with different context windows.The highlight of our findings is that increasing the size of both the source and target window sizes improves the performance of bilingual lexicon induction, especially the performance on frequent nouns.
Ryokan Ri, Yoshimasa Tsuruoka
ACL2
2020 Parallelizing and optimizing neural Encoder-Decoder models without padding on multi-core architecture
Yuchen Qiao, Kazuma Hashimoto, Akiko Eriguchi, Haixia Wang 0001, Dongsheng Wang 0002, Yoshimasa Tsuruoka, Kenjiro Taura
Future Gener. Comput. Syst.6
2019 Monte Carlo Tree Search with Variable Simulation Periods for Continuously Running Tasks
abstract
Monte Carlo Tree Search (MCTS) is widely used for planning in domains where the potential actions can be represented as a tree of sequential decisions. To efficiently select an action, MCTS usually needs to perform many simulations to build a reliable tree representation of the decision space. As such, a bottleneck to MCTS arises when enough simulations cannot be performed between action selections. This is particularly highlighted in continuously running tasks, for which the time available to perform simulations between actions tends to be limited due to the environment's state constantly changing. In this paper, we present an approach that extends the time available for Monte Carlo simulations when allowed. Our approach is to effectively balance the prospect of selecting the right action with the time that can be spared to perform MCTS simulations before the next action selection. For that, we considered the simulation time as a decision variable to be selected alongside an action. We extended the Hierarchical Optimistic Optimization applied to Tree (HOOT) method to adapt our approach to environments with a continuous decision space. We evaluated our approach on tasks with a continuous decision space using OpenAI gym's Pendulum and Continuous Mountain Car environments and on those with discrete action space using the arcade learning environment (ALE) platform. The evaluation results show that, with variable simulation times, the proposed approach outperforms the conventional MCTS in the evaluated continuous decision space tasks and improves the performance of MCTS in most of the ALE tasks.
Seydou Ba, Takuya Hiraoka, Takashi Onishi, Toru Nakata, Yoshimasa Tsuruoka
ICTAI5
2019 Learning Robust Options by Conditional Value at Risk Optimization
abstract
Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters, these methods only consider either the worst case or the average (ordinary) case for learning options. This limited consideration of the cases often produces options that do not work well in the unconsidered case. In this paper, we propose a conditional value at risk (CVaR)-based method to learn options that work well in both the average and worst cases. We extend the CVaR-based policy gradient method proposed by Chow and Ghavamzadeh (2014) to deal with robust Markov decision processes and then apply the extended method to learning robust options. We conduct experiments to evaluate our method in multi-joint robot control tasks (HopperIceBlock, Half-Cheetah, and Walker2D). Experimental results show that our method produces options that 1) give better worst-case performance than the options learned only to minimize the average-case loss, and 2) give better average-case performance than the options learned only to minimize the worst-case loss.
Takuya Hiraoka, Takahisa Imagawa, Tatsuya Mori 0001, Takashi Onishi, Yoshimasa Tsuruoka
NeurIPS5
2019 Incorporating Source-Side Phrase Structures into Neural Machine Translation
abstract
Neural machine translation (NMT) has shown great success as a new alternative to the traditional Statistical Machine Translation model in multiple languages. Early NMT models are based on sequence-to-sequence learning that encodes a sequence of source words into a vector space and generates another sequence of target words from the vector. In those NMT models, sentences are simply treated as sequences of words without any internal structure. In this article, we focus on the role of the syntactic structure of source sentences and propose a novel end-to-end syntactic NMT model, which we call a tree-to-sequence NMT model, extending a sequence-to-sequence model with the source-side phrase structure. Our proposed model has an attention mechanism that enables the decoder to generate a translated word while softly aligning it with phrases as well as words of the source sentence. We have empirically compared the proposed model with sequence-to-sequence models in various settings on Chinese-to-Japanese and English-to-Japanese translation tasks. Our experimental results suggest that the use of syntactic structure can be beneficial when the training data set is small, but is not as effective as using a bi-directional encoder. As the size of training data set increases, the benefits of using a syntactic tree tends to diminish.
Akiko Eriguchi, Kazuma Hashimoto, Yoshimasa Tsuruoka
Comput. Linguistics3
2017 Neural Machine Translation with Source-Side Latent Graph Parsing
abstract
This paper presents a novel neural machine translation model which jointly learns translation and source-side latent graph representations of sentences.Unlike existing pipelined approaches using syntactic parsers, our end-to-end model learns a latent graph parser as part of the encoder of an attention-based neural machine translation model, and thus the parser is optimized according to the translation objective.In experiments, we first show that our model compares favorably with state-of-the-art sequential and pipelined syntax-based NMT models.We also show that the performance of our model can be further improved by pretraining it with a small amount of treebank annotations.Our final ensemble model significantly outperforms the previous best models on the standard Englishto-Japanese translation dataset.
Kazuma Hashimoto, Yoshimasa Tsuruoka
EMNLP2
2017 A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
abstract
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks.Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in a single model.We introduce a joint many-task model together with a strategy for successively growing its depth to solve increasingly complex tasks.Higher layers include shortcut connections to lower-level task predictions to reflect linguistic hierarchies.We use a simple regularization term to allow for optimizing all model weights to improve one task's loss without exhibiting catastrophic interference of the other tasks.Our single end-to-end model obtains state-of-the-art or competitive results on five different tasks from tagging, parsing, relatedness, and entailment tasks.
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, Richard Socher
EMNLP3
2017 Game State Retrieval with Keyword Queries
abstract
There are many databases of game records available online. In order to retrieve a game state from such a database, users usually need to specify the target state in a domain-specific language, which may be difficult to learn for novice users. In this work, we propose a search system that allows users to retrieve game states from a game record database by using keywords. In our approach, we first train a neural network model for symbol grounding using a small number of pairs of a game state and a commentary on it. We then apply it to all the states in the database to associate each of them with characteristic terms and their scores. The enhanced database thus enables users to search for a state using keywords. To evaluate the performance of the proposed method, we conducted experiments of game state retrieval using game records of Shogi (Japanese chess) with commentaries. The results demonstrate that our approach gives significantly better results than full-text search and an LSTM language model.
Atsushi Ushiku, Shinsuke Mori, Hirotaka Kameko, Yoshimasa Tsuruoka
SIGIR4
2016 Tree-to-Sequence Attentional Neural Machine Translation
abstract
Most of the existing Neural Machine Translation (NMT) models focus on the conversion of sequential data and do not directly use syntactic information.We propose a novel end-to-end syntactic NMT model, extending a sequenceto-sequence model with the source-side phrase structure.Our model has an attention mechanism that enables the decoder to generate a translated word while softly aligning it with phrases as well as words of the source sentence.Experimental results on the WAT'15 Englishto-Japanese dataset demonstrate that our proposed model considerably outperforms sequence-to-sequence attentional NMT models and compares favorably with the state-of-the-art tree-to-string SMT system.
Akiko Eriguchi, Kazuma Hashimoto, Yoshimasa Tsuruoka
ACL (1)3
2016 Adaptive Joint Learning of Compositional and Non-Compositional Phrase Embeddings
abstract
We present a novel method for jointly learning compositional and noncompositional phrase embeddings by adaptively weighting both types of embeddings using a compositionality scoring function.The scoring function is used to quantify the level of compositionality of each phrase, and the parameters of the function are jointly optimized with the objective for learning phrase embeddings.In experiments, we apply the adaptive joint learning method to the task of learning embeddings of transitive verb phrases, and show that the compositionality scores have strong correlation with human ratings for verb-object compositionality, substantially outperforming the previous state of the art.Moreover, our embeddings improve upon the previous best model on a transitive verb disambiguation task.We also show that a simple ensemble technique further improves the results for both tasks.
Kazuma Hashimoto, Yoshimasa Tsuruoka
ACL (1)2
2016 A Japanese Chess Commentary Corpus
Shinsuke Mori, Atsushi Ushiku, Tetsuro Sasada, Hirotaka Kameko, Yoshimasa Tsuruoka
LREC6
2016 Modification of improved upper confidence bounds for regulating exploration in Monte-Carlo tree search
Yun-Ching Liu, Yoshimasa Tsuruoka
Theor. Comput. Sci.2
2015 Task-Oriented Learning of Word Embeddings for Semantic Relation Classification
abstract
We present a novel learning method for word embeddings designed for relation classification.Our word embeddings are trained by predicting words between noun pairs using lexical relation-specific features on a large unlabeled corpus.This allows us to explicitly incorporate relationspecific information into the word embeddings.The learned word embeddings are then used to construct feature vectors for a relation classification model.On a wellestablished semantic relation classification task, our method significantly outperforms a baseline based on a previously introduced word embedding method, and compares favorably to previous state-of-the-art models that use syntactic information or manually constructed external resources.
Kazuma Hashimoto, Pontus Stenetorp, Makoto Miwa, Yoshimasa Tsuruoka
CoNLL4
2015 Can Symbol Grounding Improve Low-Level NLP? Word Segmentation as a Case Study
abstract
We propose a novel framework for improving a word segmenter using information acquired from symbol grounding.We generate a term dictionary in three steps: generating a pseudo-stochastically segmented corpus, building a symbol grounding model to enumerate word candidates, and filtering them according to the grounding scores.We applied our method to game records of Japanese chess with commentaries.The experimental results show that the accuracy of a word segmenter can be improved by incorporating the generated dictionary.
Hirotaka Kameko, Shinsuke Mori, Yoshimasa Tsuruoka
EMNLP3
2015 Wide-coverage relation extraction from MEDLINE using deep syntax
abstract
BACKGROUND: Relation extraction is a fundamental technology in biomedical text mining. Most of the previous studies on relation extraction from biomedical literature have focused on specific or predefined types of relations, which inherently limits the types of the extracted relations. With the aim of fully leveraging the knowledge described in the literature, we address much broader types of semantic relations using a single extraction framework. RESULTS: Our system, which we name PASMED, extracts diverse types of binary relations from biomedical literature using deep syntactic patterns. Our experimental results demonstrate that it achieves a level of recall considerably higher than the state of the art, while maintaining reasonable precision. We have then applied PASMED to the whole MEDLINE corpus and extracted more than 137 million semantic relations. The extracted relations provide a quantitative understanding of what kinds of semantic relations are actually described in MEDLINE and can be ultimately extracted by (possibly type-specific) relation extraction systems. CONCLUSION: PASMED extracts a large number of relations that have previously been missed by existing text mining systems. The entire collection of the relations extracted from MEDLINE is publicly available in machine-readable form, so that it can serve as a potential knowledge base for high-level text-mining applications.
Nhung T. H. Nguyen 0001, Makoto Miwa, Yoshimasa Tsuruoka, Takashi Chikayama, Satoshi Tojo
BMC Bioinform.3
2015 Identifying synonymy between relational phrases using word embeddings
Nhung T. H. Nguyen 0001, Makoto Miwa, Yoshimasa Tsuruoka, Satoshi Tojo
J. Biomed. Informatics3
2014 Jointly Learning Word Representations and Composition Functions Using Predicate-Argument Structures
abstract
We introduce a novel compositional lan-guage model that works on Predicate-Argument Structures (PASs). Our model jointly learns word representations and their composition functions using bag-of-words and dependency-based con-texts. Unlike previous word-sequence-based models, our PAS-based model com-poses arguments into predicates by using the category information from the PAS. This enables our model to capture long-range dependencies between words and to better handle constructs such as verb-object and subject-verb-object relations. We verify this experimentally using two phrase similarity datasets and achieve re-sults comparable to or higher than the pre-vious best results. Our system achieves these results without the need for pre-trained word vectors and using a much smaller training corpus; despite this, for the subject-verb-object dataset our model improves upon the state of the art by as much as 10 % in relative performance. 1
Kazuma Hashimoto, Pontus Stenetorp, Makoto Miwa, Yoshimasa Tsuruoka
EMNLP4
2013 Simple Customization of Recursive Neural Networks for Semantic Relation Classification
abstract
In this paper, we present a recursive neural network (RNN) model that works on a syntactic tree.Our model differs from previous RNN models in that the model allows for an explicit weighting of important phrases for the target task.We also propose to average parameters in training.Our experimental results on semantic relation classification show that both phrase categories and task-specific weighting significantly improve the prediction accuracy of the model.We also show that averaging the model parameters is effective in stabilizing the learning and improves generalization capacity.The proposed model marks scores competitive with state-of-the-art RNN-based models.
Kazuma Hashimoto, Makoto Miwa, Yoshimasa Tsuruoka, Takashi Chikayama
EMNLP3
2013 A System-Design Outline of the Distributed-Shogi-System Akara 2010
abstract
This paper describes Akara 2010, the distributed shogi system that has defeated a professional shogi player in a public game for the first time in history. The system employs a novel design to build a high-performance computer shogi player for standard tournament conditions. The design enhances the performance of the entire system by means of distributed computing. To utilize a large number of computers, a majority-voting method using four existing programs is combined with a distributed-search method. Although the performance of the entire system could not be tested, the majority-voting component increased the winning percentage from 62% to 73%, and the distributed-search component increased it from 50% to 70% or more.
Kunihito Hoki, Tomoyuki Kaneko, Daisaku Yokoyama, Takuya Obata, Hiroshi Yamashita, Yoshimasa Tsuruoka, Takeshi Ito
SNPD6
2013 Probabilistic Chinese word segmentation with non-local information and stochastic training
Xu Sun 0001, Takuya Matsuzaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii
Inf. Process. Manag.4
2012 Bitext Dependency Parsing With Auto-Generated Bilingual Treebank
abstract
This paper proposes a method to improve the accuracy of bilingual texts (bitexts) dependency parsing by using an auto-generated bilingual treebank created with the help of statistical machine translation (SMT) systems. Previous bitext parsing methods use human-annotated bilingual treebanks that are costly and troublesome to obtain. In the proposed method, we use an auto-generated bilingual treebank to train the parsing models. First, an SMT system is used to translate a monolingual treebank into the target language; then, a monolingual parser for the target language is used to parse the translated sentences. Since the auto-translated sentences and auto-parsed trees in the auto-generated bilingual treebank are far from perfect, the bilingual constraints are not sufficiently reliable. To overcome this problem, we propose a method to verify the reliability of the constraints using a large amount of target monolingual and bilingual unannotated data. Finally, we design a set of effective bilingual features for parsing models on the basis of the verified constraints. We conduct the experiments using a standard test data. The experimental results show that our bitext parser significantly outperforms monolingual parsers. Moreover, our method is still able to provide improvement when we use a larger monolingual treebank containing over 50 000 sentences. We also test the proposed method with different SMT systems and the results show that our method is very robust to the noise. In particular, the proposed method can be used in a purely monolingual setting with the help of SMT. That is, it does not need the human translation of the test set as previous methods do.
Wenliang Chen, Jun'ichi Kazama, Min Zhang 0005, Yoshimasa Tsuruoka, Yiou Wang, Kentaro Torisawa, Haizhou Li 0001
IEEE Trans. Speech Audio Process.4
2011 Learning with Lookahead: Can History-Based Models Rival Globally Optimized Models?
Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Kazama
CoNLL1
2011 SMT Helps Bitext Dependency Parsing
Wenliang Chen, Jun'ichi Kazama, Min Zhang 0005, Yoshimasa Tsuruoka, Yiou Wang, Kentaro Torisawa, Haizhou Li 0001
EMNLP4
2011 Improving Chinese Word Segmentation and POS Tagging with Semi-supervised Methods Using Large Auto-Analyzed Data
Yiou Wang, Jun'ichi Kazama, Yoshimasa Tsuruoka, Wenliang Chen, Kentaro Torisawa
IJCNLP3
2011 AGRA: analysis of gene ranking algorithms
abstract
UNLABELLED: Often, the most informative genes have to be selected from different gene sets and several computer gene ranking algorithms have been developed to cope with the problem. To help researchers decide which algorithm to use, we developed the analysis of gene ranking algorithms (AGRA) system that offers a novel technique for comparing ranked lists of genes. The most important feature of AGRA is that no previous knowledge of gene ranking algorithms is needed for their comparison. Using the text mining system finding-associated concepts with text analysis. AGRA defines what we call biomedical concept space (BCS) for each gene list and offers a comparison of the gene lists in six different BCS categories. The uploaded gene lists can be compared using two different methods. In the first method, the overlap between each pair of two gene lists of BCSs is calculated. The second method offers a text field where a specific biomedical concept can be entered. AGRA searches for this concept in each gene lists' BCS, highlights the rank of the concept and offers a visual representation of concepts ranked above and below it. AVAILABILITY AND IMPLEMENTATION: Available at http://agra.fzv.uni-mb.si/, implemented in Java and running on the Glassfish server. CONTACT: [email protected].
Simon Kocbek, Rune Sætre, Gregor Stiglic, Jin-Dong Kim, Igor Pernek, Yoshimasa Tsuruoka, Peter Kokol, Sophia Ananiadou, Jun'ichi Tsujii
Bioinform.6
2011 Discovering and visualizing indirect associations between biomedical concepts
abstract
MOTIVATION: Discovering useful associations between biomedical concepts has been one of the main goals in biomedical text-mining, and understanding their biomedical contexts is crucial in the discovery process. Hence, we need a text-mining system that helps users explore various types of (possibly hidden) associations in an easy and comprehensible manner. RESULTS: This article describes FACTA+, a real-time text-mining system for finding and visualizing indirect associations between biomedical concepts from MEDLINE abstracts. The system can be used as a text search engine like PubMed with additional features to help users discover and visualize indirect associations between important biomedical concepts such as genes, diseases and chemical compounds. FACTA+ inherits all functionality from its predecessor, FACTA, and extends it by incorporating three new features: (i) detecting biomolecular events in text using a machine learning model, (ii) discovering hidden associations using co-occurrence statistics between concepts, and (iii) visualizing associations to improve the interpretability of the output. To the best of our knowledge, FACTA+ is the first real-time web application that offers the functionality of finding concepts involving biomolecular events and visualizing indirect associations of concepts with both their categories and importance. AVAILABILITY: FACTA+ is available as a web application at http://refine1-nactem.mc.man.ac.uk/facta/, and its visualizer is available at http://refine1-nactem.mc.man.ac.uk/facta-visualizer/. CONTACT: [email protected].
Yoshimasa Tsuruoka, Makoto Miwa, Kaisei Hamamoto, Jun'ichi Tsujii, Sophia Ananiadou
Bioinform.1
2010 PathText: a text mining integrator for biological pathway visualizations
abstract
MOTIVATION: Metabolic and signaling pathways are an increasingly important part of organizing knowledge in systems biology. They serve to integrate collective interpretations of facts scattered throughout literature. Biologists construct a pathway by reading a large number of articles and interpreting them as a consistent network, but most of the models constructed currently lack direct links to those articles. Biologists who want to check the original articles have to spend substantial amounts of time to collect relevant articles and identify the sections relevant to the pathway. Furthermore, with the scientific literature expanding by several thousand papers per week, keeping a model relevant requires a continuous curation effort. In this article, we present a system designed to integrate a pathway visualizer, text mining systems and annotation tools into a seamless environment. This will enable biologists to freely move between parts of a pathway and relevant sections of articles, as well as identify relevant papers from large text bases. The system, PathText, is developed by Systems Biology Institute, Okinawa Institute of Science and Technology, National Centre for Text Mining (University of Manchester) and the University of Tokyo, and is being used by groups of biologists from these locations.
Brian Kemper, Takuya Matsuzaki, Yukiko Matsuoka, Yoshimasa Tsuruoka, Hiroaki Kitano, Sophia Ananiadou, Jun'ichi Tsujii
Bioinform.4
2009 Stochastic Gradient Descent Training for L1-regularized Log-linear Models with Cumulative Penalty
Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou
ACL/IJCNLP1
2009 Fast Full Parsing by Linear-Chain Conditional Random Fields
Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou
EACL1
2009 A Discriminative Latent Variable Chinese Segmenter with Hybrid Word/Character Information
Xu Sun 0001, Takuya Matsuzaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii
HLT-NAACL4
2008 Modeling Latent-Dynamic in Shallow Parsing: A Latent Conditional Model with Imrpoved Inference
Xu Sun 0001, Louis-Philippe Morency, Daisuke Okanohara, Yoshimasa Tsuruoka, Jun'ichi Tsujii
COLING4
2008 A Discriminative Candidate Generator for String Transformations
Naoaki Okazaki, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii
EMNLP2
2008 Towards Data and Goal Oriented Analysis: Tool Inter-operability and Combinatorial Comparison
Yoshinobu Kano, Ngan L. T. Nguyen, Rune Sætre, Kazuhiro Yoshida, Keiichiro Fukamachi, Yusuke Miyao, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii
IJCNLP7
2008 Connecting Text Mining and Pathways using the PathText Resource
Rune Sætre, Brian Kemper, Kanae Oda, Naoaki Okazaki, Yukiko Matsuoka, Norihiro Kikuchi, Hiroaki Kitano, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii
LREC8
2008 Kleio: a knowledge-enriched information retrieval system for biology
abstract
Kleio is an advanced information retrieval (IR) system developed at the UK National Centre for Text Mining (NaCTeM)1. The system offers textual and metadata searches across MEDLINE and provides enhanced searching functionality by leveraging terminology management technologies.
Chikashi Nobata, Philip Cotter, Naoaki Okazaki, Brian Rea, Yutaka Sasaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou
SIGIR6
2008 FACTA: a text search engine for finding associated biomedical concepts
abstract
UNLABELLED: FACTA is a text search engine for MEDLINE abstracts, which is designed particularly to help users browse biomedical concepts (e.g. genes/proteins, diseases, enzymes and chemical compounds) appearing in the documents retrieved by the query. The concepts are presented to the user in a tabular format and ranked based on the co-occurrence statistics. Unlike existing systems that provide similar functionality, FACTA pre-indexes not only the words but also the concepts mentioned in the documents, which enables the user to issue a flexible query (e.g. free keywords or Boolean combinations of keywords/concepts) and receive the results immediately even when the number of the documents that match the query is very large. The user can also view snippets from MEDLINE to get textual evidence of associations between the query terms and the concepts. The concept IDs and their names/synonyms for building the indexes were collected from several biomedical databases and thesauri, such as UniProt, BioThesaurus, UMLS, KEGG and DrugBank. AVAILABILITY: The system is available at http://www.nactem.ac.uk/software/facta/
Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou
Bioinform.1
2008 How to make the most of NE dictionaries in statistical NER
abstract
BACKGROUND: When term ambiguity and variability are very high, dictionary-based Named Entity Recognition (NER) is not an ideal solution even though large-scale terminological resources are available. Many researches on statistical NER have tried to cope with these problems. However, it is not straightforward how to exploit existing and additional Named Entity (NE) dictionaries in statistical NER. Presumably, addition of NEs to an NE dictionary leads to better performance. However, in reality, the retraining of NER models is required to achieve this. We chose protein name recognition as a case study because it most suffers the problems related to heavy term variation and ambiguity. METHODS: We have established a novel way to improve the NER performance by adding NEs to an NE dictionary without retraining. In our approach, first, known NEs are identified in parallel with Part-of-Speech (POS) tagging based on a general word dictionary and an NE dictionary. Then, statistical NER is trained on the POS/PROTEIN tagger outputs with correct NE labels attached. RESULTS: We evaluated performance of our NER on the standard JNLPBA-2004 data set. The F-score on the test set has been improved from 73.14 to 73.78 after adding protein names appearing in the training data to the POS tagger dictionary without any model retraining. The performance further increased to 78.72 after enriching the tagging dictionary with test set protein names. CONCLUSION: Our approach has demonstrated high performance in protein name recognition, which indicates how to make the most of known NEs in statistical NER.
Yutaka Sasaki, Yoshimasa Tsuruoka, John McNaught, Sophia Ananiadou
BMC Bioinform.2
2008 Normalizing biomedical terms by minimizing ambiguity and variability
abstract
BACKGROUND: One of the difficulties in mapping biomedical named entities, e.g. genes, proteins, chemicals and diseases, to their concept identifiers stems from the potential variability of the terms. Soft string matching is a possible solution to the problem, but its inherent heavy computational cost discourages its use when the dictionaries are large or when real time processing is required. A less computationally demanding approach is to normalize the terms by using heuristic rules, which enables us to look up a dictionary in a constant time regardless of its size. The development of good heuristic rules, however, requires extensive knowledge of the terminology in question and thus is the bottleneck of the normalization approach. RESULTS: We present a novel framework for discovering a list of normalization rules from a dictionary in a fully automated manner. The rules are discovered in such a way that they minimize the ambiguity and variability of the terms in the dictionary. We evaluated our algorithm using two large dictionaries: a human gene/protein name dictionary built from BioThesaurus and a disease name dictionary built from UMLS. CONCLUSIONS: The experimental results showed that automatically discovered rules can perform comparably to carefully crafted heuristic rules in term mapping tasks, and the computational overhead of rule application is small enough that a very fast implementation is possible. This work will help improve the performance of term-concept mapping tasks in biomedical information extraction especially when good normalization heuristics for the target terminology are not fully known.
Yoshimasa Tsuruoka, John McNaught, Sophia Ananiadou
BMC Bioinform.1
2008 Accelerating the annotation of sparse named entities by dynamic sentence selection
abstract
BACKGROUND: Previous studies of named entity recognition have shown that a reasonable level of recognition accuracy can be achieved by using machine learning models such as conditional random fields or support vector machines. However, the lack of training data (i.e. annotated corpora) makes it difficult for machine learning-based named entity recognizers to be used in building practical information extraction systems. RESULTS: This paper presents an active learning-like framework for reducing the human effort required to create named entity annotations in a corpus. In this framework, the annotation work is performed as an iterative and interactive process between the human annotator and a probabilistic named entity tagger. Unlike active learning, our framework aims to annotate all occurrences of the target named entities in the given corpus, so that the resulting annotations are free from the sampling bias which is inevitable in active learning approaches. CONCLUSION: We evaluate our framework by simulating the annotation process using two named entity corpora and show that our approach can reduce the number of sentences which need to be examined by the human annotator. The cost reduction achieved by the framework could be drastic when the target named entities are sparse.
Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou
BMC Bioinform.1
2007 Ambiguous Part-of-Speech Tagging for Improving Accuracy and Domain Portability of Syntactic Parsers
Kazuhiro Yoshida, Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Tsujii
IJCAI2
2007 Learning string similarity measures for gene/protein name dictionary look-up using logistic regression
abstract
MOTIVATION: One of the bottlenecks of biomedical data integration is variation of terms. Exact string matching often fails to associate a name with its biological concept, i.e. ID or accession number in the database, due to seemingly small differences of names. Soft string matching potentially enables us to find the relevant ID by considering the similarity between the names. However, the accuracy of soft matching highly depends on the similarity measure employed. RESULTS: We used logistic regression for learning a string similarity measure from a dictionary. Experiments using several large-scale gene/protein name dictionaries showed that the logistic regression-based similarity measure outperforms existing similarity measures in dictionary look-up tasks. AVAILABILITY: A dictionary look-up system using the similarity measures described in this article is available at http://text0.mib.man.ac.uk/software/mldic/.
Yoshimasa Tsuruoka, John McNaught, Jun'ichi Tsujii, Sophia Ananiadou
Bioinform.1
2006 Semantic Retrieval for the Accurate Identification of Relational Concepts in Massive Textbases
abstract
This paper introduces a novel framework for the accurate retrieval of relational concepts from huge texts. Prior to retrieval, all sentences are annotated with predicate argument structures and ontological identifiers by applying a deep parser and a term recognizer. During the run time, user requests are converted into queries of region algebra on these annotations. Structural matching with pre-computed semantic annotations establishes the accurate and efficient retrieval of relational concepts. This framework was applied to a text retrieval system for MEDLINE. Experiments on the retrieval of biomedical correlations revealed that the cost is sufficiently small for real-time applications and that the retrieval precision is significantly improved.
Yusuke Miyao, Tomoko Ohta, Katsuya Masuda, Yoshimasa Tsuruoka, Kazuhiro Yoshida, Takashi Ninomiya, Jun'ichi Tsujii
ACL4
2006 An Intelligent Search Engine and GUI-based Efficient MEDLINE Search Tool Based on Deep Syntactic Parsing
abstract
We present a practical HPSG parser for English, an intelligent search engine to retrieve MEDLINE abstracts that represent biomedical events and an efficient MED-LINE search tool helping users to find information about biomedical entities such as genes, proteins, and the interactions between them.
Tomoko Ohta, Yusuke Miyao, Takashi Ninomiya, Yoshimasa Tsuruoka, Akane Yakushiji, Katsuya Masuda, Jumpei Takeuchi, Kazuhiro Yoshida, Tadayoshi Hara, Jin-Dong Kim, Yuka Tateisi, Jun'ichi Tsujii
ACL4
2006 Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition
abstract
This paper presents techniques to apply semi-CRFs to Named Entity Recognition tasks with a tractable computational cost. Our framework can handle an NER task that has long named entities and many labels which increase the computational cost. To reduce the computational cost, we propose two techniques: the first is the use of feature forests, which enables us to pack feature-equivalent states, and the second is the introduction of a filtering process which significantly reduces the number of candidate states. This framework allows us to use a rich set of features extracted from the chunk-based representation that can capture informative characteristics of entities. We also introduce a simple trick to transfer information about distant entities by embedding label information into non-entity labels. Experimental results show that our model achieves an F-score of 71.48% on the JNLPBA 2004 shared task without using any external resources or post-processing techniques.
Daisuke Okanohara, Yusuke Miyao, Yoshimasa Tsuruoka, Jun'ichi Tsujii
ACL3
2006 Extremely Lexicalized Models for Accurate and Fast HPSG Parsing
Takashi Ninomiya, Takuya Matsuzaki, Yoshimasa Tsuruoka, Yusuke Miyao, Jun'ichi Tsujii
EMNLP3
2006 Automatic recognition of topic-classified relations between prostate cancer and genes using MEDLINE abstracts
abstract
BACKGROUND: Automatic recognition of relations between a specific disease term and its relevant genes or protein terms is an important practice of bioinformatics. Considering the utility of the results of this approach, we identified prostate cancer and gene terms with the ID tags of public biomedical databases. Moreover, considering that genetics experts will use our results, we classified them based on six topics that can be used to analyze the type of prostate cancers, genes, and their relations. METHODS: We developed a maximum entropy-based named entity recognizer and a relation recognizer and applied them to a corpus-based approach. We collected prostate cancer-related abstracts from MEDLINE, and constructed an annotated corpus of gene and prostate cancer relations based on six topics by biologists. We used it to train the maximum entropy-based named entity recognizer and relation recognizer. RESULTS: Topic-classified relation recognition achieved 92.1% precision for the relation (an increase of 11.0% from that obtained in a baseline experiment). For all topics, the precision was between 67.6 and 88.1%. CONCLUSION: A series of experimental results revealed two important findings: a carefully designed relation recognition system using named entity recognition can improve the performance of relation recognition, and topic-classified relation recognition can be effectively addressed through a corpus-based approach using manual annotation and machine learning techniques.
Hong-Woo Chun, Yoshimasa Tsuruoka, Jin-Dong Kim, Rie Shiba, Naoki Nagata, Teruyoshi Hishiki, Jun'ichi Tsujii
BMC Bioinform.2
2004 Iterative CKY Parsing for Probabilistic Context-Free Grammars
Yoshimasa Tsuruoka, Jun'ichi Tsujii
IJCNLP1
2004 Improving the performance of dictionary-based approaches in protein name recognition
Yoshimasa Tsuruoka, Jun'ichi Tsujii
J. Biomed. Informatics1
2003 Training a Naive Bayes Classifier via the EM Algorithm with a Class Distribution Constraint
Yoshimasa Tsuruoka, Jun'ichi Tsujii
CoNLL1
2003 Probabilistic term variant generator for biomedical terms
abstract
This paper presents an algorithm to generate possible variants for biomedical terms. The algorithm gives each variant its generation probability representing its plausibility, which is potentially useful for query and dictionary expansions. The probabilistic rules for generating variants are automatically learned from raw texts using an existing abbreviation extraction technique. Our method, therefore, requires no linguistic knowledge or labor-intensive natural language resource. We conducted an experiment using 83,142 MEDLINE abstracts for rule induction and 18,930 abstracts for testing. The results indicate that our method will significantly increase the number of retrieved documents for long biomedical terms.
Yoshimasa Tsuruoka, Jun'ichi Tsujii
SIGIR1