Takenobu Tokunaga

dblp:35/4614 · DBLP profile ↗
← Back
75ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-1399-9517ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 9 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author
YearPublicationVenuePosition
2025 Evaluation of LLM-Generated Distractors of Multiple-Choice Questions for the Japanese National Nursing Examination
Yûsei Kido, Hiroaki Yamada 0002, Takenobu Tokunaga, Rika Kimura, Yuriko Miura, Yumi Sakyo, Naoko Hayashi
CSEDU (1)3
2024 Analyzing Interpretability of Summarization Model with Eye-gaze Information
abstract
Interpretation methods provide saliency scores indicating the importance of input words for neural summarization models. Prior work has analyzed models by comparing them to human behavior, often using eye-gaze as a proxy for human attention in reading tasks such as classification. This paper presents a framework to analyze the model behavior in summarization by comparing it to human summarization behavior using eye-gaze data. We examine two research questions: RQ1) whether model saliency conforms to human gaze during summarization and RQ2) how model saliency and human gaze affect summarization performance. For RQ1, we measure conformity by calculating the correlation between model saliency and human fixation counts. For RQ2, we conduct ablation experiments removing words/sentences considered important by models or humans. Experiments on two datasets with human eye-gaze during summarization partially confirm that model saliency aligns with human gaze (RQ1). However, ablation experiments show that removing highly-attended words/sentences from the human gaze does not significantly degrade performance compared with the removal by the model saliency (RQ2).
Fariz Ikhwantri, Hiroaki Yamada 0002, Takenobu Tokunaga
LREC/COLING3
2024 Automatic Question Generation for the Japanese National Nursing Examination Using Large Language Models
Yûsei Kido, Hiroaki Yamada 0002, Takenobu Tokunaga, Rika Kimura, Yuriko Miura, Yumi Sakyo, Naoko Hayashi
CSEDU (1)3
2023 Nearest Neighbor Search for Summarization of Japanese Judgment Documents
abstract
With the increasing demand for summarizing Japanese judgment documents, the automatic generation of high-quality summaries by large language models (LLMs) is expected. We propose a method to select exemplars using the nearest neighbor search for the one-shot learning method. The experiments showed our method outperforms baseline methods.
Akito Shimbo, Yuta Sugawara, Hiroaki Yamada 0002, Takenobu Tokunaga
JURIX4
2023 Improving logical flow in English-as-a-foreign-language learner essays by reordering sentences
abstract
Argumentation is ubiquitous in everyday discourse, and it is a skill that can be learned. In our society, it is also one that must be learned: education systems all over the world agree on the importance of argumentation skills. However, writing effective argumentation is difficult, and even more so if it has to be expressed in a foreign language. Existing artificial intelligence systems for language learning can help learners: they can provide objective feedback (e.g., concerning grammar and spelling), as well as providing learners with opportunities to identify errors and subsequently improve their texts. Even so, systems aiming at higher discourse-level skills, such as persuasiveness and content organisation, are still limited. In this article, we propose the novel task of sentence reordering for improving the logical flow of argumentative essays. To train such a computational system, we present a new corpus called ICNALE-AS2R, containing essays written by English-as-foreign-language learners from various Asian countries, that have been annotated with argumentative structure and sentence reordering. We also propose a novel method to automatically reorder sentences in imperfect essays, which is based on argumentative structure analysis. Given an input essay and its corresponding argumentative structure, we cast the reordering task as a traversal problem. Our sentence reordering system first determines the pairwise ordering relation between pairs of sentences that are connected by argumentative relations. In the second step, the system traverses the argumentative structure that has been augmented with pairwise ordering information, in order to generate the final output text. Empirical evaluation shows that in the task of reconstructing the final reordered essays in the dataset, our reordering system achieves .926 and .879 in longest common subsequence ratio and Kendall's Tau metrics, respectively. The system is also able to perform the reordering operation selectively, that is, it reorders sentences when necessary and retains the original input order when it is already optimal.
Jan Wira Gotama Putra, Simone Teufel, Takenobu Tokunaga
Artif. Intell.3
2023 Looking deep in the eyes: Investigating interpretation methods for neural models on reading tasks using human eye-movement behaviour
abstract
This paper provides the first broad overview of the relation between different interpretation methods and human eye-movement behaviour across different tasks and architectures. The interpretation methods of neural networks provide the information the machine considers important, while the human eye-gaze has been believed to be a proxy of the human cognitive process. Thus, comparing them explains machine behaviour in terms of human behaviour, leading to improvement in machine performance through minimising their difference. We consider three types of natural language processing (NLP) tasks: sentiment analysis, relation classification and question answering, and four interpretation methods based on: simple gradient, integrated gradient, input-perturbation and attention, and three architectures: LSTM, CNN and Transformer. We leverage two corpora annotated with eye-gaze information: the Zuco dataset and the MQA-RC dataset. This research sets up two research questions. First, we investigate whether the saliency (importance) of input-words conform with those from human eye-gaze features. To this end, we compute a saliency distance (SD) between input words (by an interpretation method) and an eye-gaze feature. SD is defined as the KL-divergence between the saliency distribution over input words and an eye-gaze feature. We found that the SD scores vary depending on the combinations of tasks, interpretation methods and architectures. Second, we investigate whether the models with good saliency conformity to human eye-gaze behaviour have better prediction performances. To this end, we propose a novel evaluation device called “SD-performance curve” (SDPC) which represents the cumulative model performance against the SD scores. SDPC enables us to analyse the underlying phenomena that were overlooked using only the macroscopic metrics, such as average SD scores and rank correlations, that are typically used in the past studies. We observe that the impact of good saliency conformity between humans and machines on task performance varies among the combinations of tasks, interpretation methods and architectures. Our findings should be considered when introducing eye-gaze information for model training to improve the model performance.
Fariz Ikhwantri, Jan Wira Gotama Putra, Hiroaki Yamada 0002, Takenobu Tokunaga
Inf. Process. Manag.4
2022 Vocabulary Volume: A New Metric for Assessing Vocabulary Knowledge
Dolça Tellols, Takenobu Tokunaga, Hikaru Yokono
CSEDU (2)2
2022 Annotation Study of Japanese Judgments on Tort for Legal Judgment Prediction with Rationales
abstract
This paper describes a comprehensive annotation study on Japanese judgment documents in civil cases. We aim to build an annotated corpus designed for Legal Judgment Prediction (LJP), especially for torts. Our annotation scheme contains annotations of whether tort is accepted by judges as well as its corresponding rationales for explainability purpose. Our annotation scheme extracts decisions and rationales at character-level. Moreover, the scheme can capture the explicit causal relation between judge’s decisions and their corresponding rationales, allowing multiple decisions in a document. To obtain high-quality annotation, we developed an annotation scheme with legal experts, and confirmed its reliability by agreement studies with Krippendorff’s alpha metric. The result of the annotation study suggests the proposed annotation scheme can produce a dataset of Japanese LJP at reasonable reliability.
Hiroaki Yamada 0002, Takenobu Tokunaga, Ryutaro Ohara, Keisuke Takeshita, Mihoko Sumida
LREC2
2022 Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis
abstract
In this paper, we approach summary evaluation from an applied linguistics (AL) point of view. We provide computational tools to AL researchers to simplify the process of Idea Unit (IU) segmentation. The IU is a segmentation unit that can identify chunks of information. These chunks can be compared across documents to measure the content overlap between a summary and its source text. We propose a full revision of the annotation guidelines to allow machine implementation. The new guideline also improves the inter-annotator agreement, rising from 0.547 to 0.785 (Cohen’s Kappa). We release L2WS 2021, a IU gold standard corpus composed of 40 manually annotated student summaries. We propose IUExtract; i.e. the first automatic segmentation algorithm based on the IU. The algorithm was tested over the L2WS 2021 corpus. Our results are promising, achieving a precision of 0.789 and a recall of 0.844. We tested an existing approach to IU alignment via word embeddings with the state of the art model SBERT. The recorded precision for the top 1 aligned pair of IUs was 0.375. We deemed this result insufficient for effective automatic alignment. We propose “SAT”, an online tool to facilitate the collection of alignment gold standards for future training.
Marcello Gecchele, Hiroaki Yamada 0002, Takenobu Tokunaga, Yasuyo Sawaki, Mika Ishizuka
LREC3
2022 Annotating argumentative structure in English-as-a-Foreign-Language learner essays
abstract
Abstract Argument mining (AM) aims to explain how individual argumentative discourse units (e.g. sentences or clauses) relate to each other and what roles they play in the overall argumentation. The automatic recognition of argumentative structure is attractive as it benefits various downstream tasks, such as text assessment, text generation, text improvement, and summarization. Existing studies focused on analyzing well-written texts provided by proficient authors. However, most English speakers in the world are non-native, and their texts are often poorly structured, particularly if they are still in the learning phase. Yet, there is no specific prior study on argumentative structure in non-native texts. In this article, we present the first corpus containing argumentative structure annotation for English-as-a-foreign-language (EFL) essays, together with a specially designed annotation scheme. The annotated corpus resulting from this work is called “ICNALE-AS” and contains 434 essays written by EFL learners from various Asian countries. The corpus presented here is particularly useful for the education domain. On the basis of the analysis of argumentation-related problems in EFL essays, educators can formulate ways to improve them so that they more closely resemble native-level productions. Our argument annotation scheme is demonstrably stable, achieving good inter-annotator agreement and near-perfect intra-annotator agreement. We also propose a set of novel document-level agreement metrics that are able to quantify structural agreement from various argumentation aspects, thus providing a more holistic analysis of the quality of the argumentative structure annotation. The metrics are evaluated in a crowd-sourced meta-evaluation experiment, achieving moderate to good correlation with human judgments.
Jan Wira Gotama Putra, Simone Teufel, Takenobu Tokunaga
Nat. Lang. Eng.3
2020 Effective Use of Target-side Context for Neural Machine Translation
abstract
In this paper, we deal with two problems in Japanese-English machine translation of news articles. The first problem is the quality of parallel corpora. Neural machine translation (NMT) systems suffer degraded performance when trained with noisy data. Because there is no clean Japanese-English parallel data for news articles, we build a novel parallel news corpus consisting of Japanese news articles translated into English in a content-equivalent manner. This is the first content-equivalent Japanese-English news corpus translated specifically for training NMT systems. The second problem involves the domain-adaptation technique. NMT systems suffer degraded performance when trained with mixed data having different features, such as noisy data and clean data. Though the existing methods try to overcome this problem by using tags for distinguishing the differences between corpora, it is not sufficient. We thus extend a domain-adaptation method using multi-tags to train an NMT model effectively with the clean corpus and existing parallel news corpora with some types of noise. Experimental results show that our corpus increases the translation quality, and that our domain-adaptation method is more effective for learning with the multiple types of corpora than existing domain-adaptation methods are.
Hideya Mino, Hitoshi Ito, Isao Goto, Ichiro Yamada, Takenobu Tokunaga
COLING5
2020 Content-Equivalent Translated Parallel News Corpus and Extension of Domain Adaptation for NMT
abstract
In this paper, we deal with two problems in Japanese-English machine translation of news articles. The first problem is the quality of parallel corpora. Neural machine translation (NMT) systems suffer degraded performance when trained with noisy data. Because there is no clean Japanese-English parallel data for news articles, we build a novel parallel news corpus consisting of Japanese news articles translated into English in a content-equivalent manner. This is the first content-equivalent Japanese-English news corpus translated specifically for training NMT systems. The second problem involves the domain-adaptation technique. NMT systems suffer degraded performance when trained with mixed data having different features, such as noisy data and clean data. Though the existing methods try to overcome this problem by using tags for distinguishing the differences between corpora, it is not sufficient. We thus extend a domain-adaptation method using multi-tags to train an NMT model effectively with the clean corpus and existing parallel news corpora with some types of noise. Experimental results show that our corpus increases the translation quality, and that our domain-adaptation method is more effective for learning with the multiple types of corpora than existing domain-adaptation methods are.
Hideya Mino, Hideki Tanaka, Hitoshi Ito, Isao Goto, Ichiro Yamada, Takenobu Tokunaga
LREC6
2020 Gamification Platform for Collecting Task-oriented Dialogue Data
abstract
Demand for massive language resources is increasing as the data-driven approach has established a leading position in Natural Language Processing. However, creating dialogue corpora is still a difficult task due to the complexity of the human dialogue structure and the diversity of dialogue topics. Though crowdsourcing is majorly used to assemble such data, it presents problems such as less-motivated workers. We propose a platform for collecting task-oriented situated dialogue data by using gamification. Combining a video game with data collection benefits such as motivating workers and cost reduction. Our platform enables data collectors to create their original video game in which they can collect dialogue data of various types of tasks by using the logging function of the platform. Also, the platform provides the annotation function that enables players to annotate their own utterances. The annotation can be gamified aswell. We aim at high-quality annotation by introducing such self-annotation method. We implemented a prototype of the proposed platform and conducted a preliminary evaluation to obtain promising results in terms of both dialogue data collection and self-annotation.
Haruna Ogawa, Hitoshi Nishikawa, Takenobu Tokunaga, Hikaru Yokono
LREC3
2020 TIARA: A Tool for Annotating Discourse Relations and Sentence Reordering
abstract
This paper introduces TIARA, a new publicly available web-based annotation tool for discourse relations and sentence reordering. Annotation tasks such as these, which are based on relations between large textual objects, are inherently hard to visualise without either cluttering the display and/or confusing the annotators. TIARA deals with the visual complexity during the annotation process by systematically simplifying the layout, and by offering interactive visualisation, including coloured links, indentation, and dual-view. TIARA’s text view allows annotators to focus on the analysis of logical sequencing between sentences. A separate tree view allows them to review their analysis in terms of the overall discourse structure. The dual-view gives it an edge over other discourse annotation tools and makes it particularly attractive as an educational tool (e.g., for teaching students how to argue more effectively). As it is based on standard web technologies and can be easily customised to other annotation schemes, it can be easily used by anybody. Apart from the project it was originally designed for, in which hundreds of texts were annotated by three annotators, TIARA has already been adopted by a second discourse annotation study, which uses it in the teaching of argumentation.
Jan Wira Gotama Putra, Simone Teufel, Kana Matsumura, Takenobu Tokunaga
LREC4
2019 Dialogue Systems for the Assessment of Language Learners' Productive Vocabulary
abstract
This paper proposes to use dialogue systems to assess language learners' productive vocabulary. We introduce a new task where dialogue systems try to induce learners to use specific words during a natural conversation to assess their productive vocabulary. To investigate the feasibility of the dialogue systems that are capable of this task, we performed two kinds of experiments.
Dolça Tellols, Hitoshi Nishikawa, Takenobu Tokunaga
HAI3
2019 Neural Network Based Rhetorical Status Classification for Japanese Judgment Documents
abstract
We address the legal text understanding task, and in particular we treat Japanese judgment documents in civil law. Rhetorical status classification (RSC) is the task of classifying sentences according to the rhetorical functions they fulfil; it is an important preprocessing step for our overall goal of legal summarisation. We present several improvements over our previous RSC classifier, which was based on CRF. The first is a BiLSTM-CRF based model which improves performance significantly over previous baselines. The BiLSTM-CRF architecture is able to additionally take the context in terms of neighbouring sentences into account. The second improvement is the inclusion of section heading information, which resulted in the overall best classifier. Explicit structure in the text, such as headings, is an information source which is likely to be important to legal professionals during the reading phase; this makes the automatic exploitation of such information attractive.We also considerably extended the size of our annotated corpus of judgment documents.
Hiroaki Yamada 0002, Simone Teufel, Takenobu Tokunaga
JURIX3
2018 Interpretation of Implicit Conditions in Database Search Dialogues
abstract
Targeting the database search dialogue, we propose to utilise information in the user utterances that do not directly mention the database (DB) field of the backend database system but are useful for constructing database queries. We call this kind of information implicit conditions. Interpreting the implicit conditions enables the dialogue system more natural and efficient in communicating with humans. We formalised the interpretation of the implicit conditions as classifying user utterances into the related DB field while identifying the evidence for that classification at the same time. Introducing this new task is one of the contributions of this paper. We implemented two models for this task: an SVM-based model and an RCNN-based model. Through the evaluation using a corpus of simulated dialogues between a real estate agent and a customer, we found that the SVM-based model showed better performance than the RCNN-based model.
Shun-ya Fukunaga, Hitoshi Nishikawa, Takenobu Tokunaga, Hikaru Yokono, Tetsuro Takahashi
COLING3
2018 Analysis of Implicit Conditions in Database Search Dialogues
Shun-ya Fukunaga, Hitoshi Nishikawa, Takenobu Tokunaga, Hikaru Yokono, Tetsuro Takahashi
LREC3
2018 Effectiveness of Domain Adaptation in Japanese Predicate-Argument Structure Analysis
Mizuki Sango, Hitoshi Nishikawa, Takenobu Tokunaga
PACLIC3
2018 Neural Japanese Zero Anaphora Resolution using Smoothed Large-scale Case Frames with Word Embedding
Souta Yamashiro, Hitoshi Nishikawa, Takenobu Tokunaga
PACLIC3
2017 Automatic Generation of English Reference Question by Utilising Nonrestrictive Relative Clause
Arief Yudha Satria, Takenobu Tokunaga
CSEDU (1)2
2016 Parameter estimation of Japanese predicate argument structure analysis model using eye gaze information
abstract
In this paper, we propose utilising eye gaze information for estimating parameters of a Japanese predicate argument structure (PAS) analysis model. We employ not only linguistic information in the text, but also the information of annotator eye gaze during their annotation process. We hypothesise that annotator’s frequent looks at certain candidates imply their plausibility of being the argument of the predicate. Based on this hypothesis, we consider annotator eye gaze for estimating the model parameters of the PAS analysis. The evaluation experiment showed that introducing eye gaze information increased the accuracy of the PAS analysis by 0.05 compared with the conventional methods.
Ryosuke Maki, Hitoshi Nishikawa, Takenobu Tokunaga
COLING3
2016 Item Difficulty Analysis of English Vocabulary Questions
Yuni Susanti, Hitoshi Nishikawa, Takenobu Tokunaga, Hiroyuki Obari
CSEDU (1)3
2016 Solving the AL Chicken-and-Egg Corpus and Model Problem: Model-free Active Learning for Phenomena-driven Corpus Construction
Dain Kaplan, Neil Rubens, Simone Teufel, Takenobu Tokunaga
LREC4
2015 Automatic Generation of English Vocabulary Tests
Yuni Susanti, Ryu Iida, Takenobu Tokunaga
CSEDU (1)3
2015 Incrementally Tracking Reference in Human/Human Dialogue Using Linguistic and Extra-Linguistic Information
abstract
Casey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Casey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen
HLT-NAACL3
2014 Collecting Pairs of Word Senses and Their Context Sentences for Generating English Vocabulary Tests
Yuni Susanti, Ryu Iida, Takenobu Tokunaga
ICCE3
2014 Building a Corpus of Manually Revised Texts from Discourse Perspective
Ryu Iida, Takenobu Tokunaga
LREC2
2013 Empirical investigation on spatial templates for a diagonal spatial term
Takenobu Tokunaga, Toshiaki Watatani, Ryu Iida, Asuka Terai
CogSci1
2012 Effects of Document Clustering in Modeling Wikipedia-style Term Descriptions
Atsushi Fujii, Yuya Fujii, Takenobu Tokunaga
LREC3
2012 The REX corpora: A collection of multimodal corpora of referring expressions in collaborative problem solving dialogues
Takenobu Tokunaga, Ryu Iida, Asuka Terai, Naoko Kuriyama
LREC1
2012 A Unified Probabilistic Approach to Referring Expressions
Kotaro Funakoshi, Mikio Nakano, Takenobu Tokunaga, Ryu Iida
SIGDIAL Conference3
2011 Multi-modal Reference Resolution in Situated Dialogue by Integrating Linguistic and Extra-Linguistic Clues
Ryu Iida, Masaaki Yasuhara, Takenobu Tokunaga
IJCNLP3
2010 Incorporating Extra-Linguistic Information into Reference Resolution in Collaborative Task Dialogue
Ryu Iida, Syumpei Kobayashi, Takenobu Tokunaga
ACL3
2010 Towards an Extrinsic Evaluation of Referring Expressions in Situated Dialogs
Philipp Spanger, Ryu Iida, Takenobu Tokunaga, Asuka Terai, Naoko Kuriyama
INLG3
2010 Annotation Process Management Revisited
Dain Kaplan, Ryu Iida, Takenobu Tokunaga
LREC3
2009 AdjScales: Differentiating between Similar Adjectives for Language Learners
Vera Sheinman, Takenobu Tokunaga
CSEDU (1)2
2009 Hozumi Tanaka
Timothy Baldwin, Takenobu Tokunaga, Jun'ichi Tsujii
Comput. Linguistics2
2008 Constructing Taxonomy of Numerative Classifiers for Asian Languages
Kiyoaki Shirai, Takenobu Tokunaga, Chu-Ren Huang, Shu-Kai Hsieh, Tzu-Yi Kuo, Virach Sornlertlamvanich, Thatsanee Charoenporn
IJCNLP2
2008 Adapting International Standard for Asian Language Technologies
Takenobu Tokunaga, Dain Kaplan, Chu-Ren Huang, Shu-Kai Hsieh, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Kiyoaki Shirai, Virach Sornlertlamvanich, Thatsanee Charoenporn, Yingju Xia
LREC1
2006 Efficient Sentence Retrieval Based on Syntactic Structure
Hiroshi Ichikawa, Keita Hakoda, Taiichi Hashimoto, Takenobu Tokunaga
ACL4
2006 Infrastructure for Standardization of Asian Language Resources
Takenobu Tokunaga, Virach Sornlertlamvanich, Thatsanee Charoenporn, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Chu-Ren Huang, Yingju Xia, Hao Yu 0005, Laurent Prévot 0001, Kiyoaki Shirai
ACL1
2006 Identifying Repair Targets in Action Control Dialogue
Kotaro Funakoshi, Takenobu Tokunaga
EACL2
2006 Group-Based Generation of Referring Expressions
Kotaro Funakoshi, Satoru Watanabe, Takenobu Tokunaga
INLG3
2006 A new approach to syntactic annotation
Masaki Noguchi, Hiroshi Ichikawa, Taiichi Hashimoto, Takenobu Tokunaga
LREC4
2005 Understanding Referring Expressions Involving Perceptual Grouping
abstract
This paper deals with understanding referring expressions involving perceptual grouping. The ability to use referring expressions is important for conversational agents aimed at real-world interaction. We conducted a psychological experiment to collect referring expressions involving perceptual grouping. A set of methods to identify referents based on the collected data are presented. We were able to identify 78.8% of the referents in the collected expressions
Kotaro Funakoshi, Satoru Watanabe, Takenobu Tokunaga, Naoko Kuriyama
CW3
2004 Generation of Relative Referring Expressions based on Perceptual Grouping
Kotaro Funakoshi, Satoru Watanabe, Naoko Kuriyama, Takenobu Tokunaga
COLING4
2004 Generating Referring Expressions Using Perceptual Groups
Kotaro Funakoshi, Satoru Watanabe, Naoko Kuriyama, Takenobu Tokunaga
INLG4
2004 Classification of Japanese Spatial Nouns
Takenobu Tokunaga, Tomofumi Koyama, Suguru Saito, Masayuki Nakajima 0001
LREC1
2004 Retrieving Annotated Corpora for Corpus Annotation
Kyôsuke Yoshida, Taiichi Hashimoto, Takenobu Tokunaga, Hozumi Tanaka
LREC3
2004 Automatic expansion of abbreviations by using context and character information
Akira Terada, Takenobu Tokunaga, Hozumi Tanaka
Inf. Process. Manag.2
2002 Processing Japanese Self-correction in Speech Dialog Systems
Kotaro Funakoshi, Takenobu Tokunaga, Hozumi Tanaka
COLING2
2002 Enhanced Japanese Electronic Dictionary Look-up
Timothy Baldwin, Slaven Bilac, Ryo Okumura, Takenobu Tokunaga, Hozumi Tanaka
LREC4
2002 Constructing a lexicon of action
Takenobu Tokunaga, Manabu Okumura, Suguru Saito, Hozumi Tanaka
LREC1
2002 Selecting effective index terms using a decision tree
abstract
This paper explores the effectiveness of index terms more complex than the single words used in conventional information retrieval systems. Retrieval is done in two phases: in the first, a conventional retrieval method (the Okapi system) is used; in the second, complex index terms such as syntactic relations and single words with part-of-speech information are introduced to rerank the results of the first phase. We evaluated the effectiveness of the different types of index terms through experiments using the TREC-7 test collection and 50 queries. The retrieval effectiveness was improved for 32 out of 50 queries. Based on this investigation, we then introduce a method to select effective index terms by using a decision tree. Further experiments with the same test collection showed that retrieval effectiveness was improved in 25 of the 50 queries.
Takenobu Tokunaga, Kenji Kimura, Hironori Ogibayashi, Hozumi Tanaka
Nat. Lang. Eng.1
2001 Decision lists for determining adjective dependency in Japanese
abstract
In Japanese constructions of the form [N1 no Adj N2], the adjective Adj modifies either N1 or N2. Determing the semantic dependencies of adjective in such phrase is an important task for machine translation. This paper describes a method for determining the adjective dependency in such constructions using decision lists, and inducing decision lists from training contexts with correct semantic dependencies and without. Based on evaluation, our method is able to determine adjective dependency with an precision of about 94%. We further analyze rules in the induced decision lists and examine effective features to determine the semantic dependencies of adjectives.
Taiichi Hashimoto, Kosuke Nishidate, Kiyoaki Shirai, Takenobu Tokunaga, Hozumi Tanaka
MTSummit4
2000 Semi-automatic Construction of a Tree-annotated Corpus Using an Iterative Learning Statistical Language Model
Kiyoaki Shirai, Hozumi Tanaka, Takenobu Tokunaga
LREC3
2000 Improving information retrieval system performance by combining different text-mining techniques
Rila Mandala, Takenobu Tokunaga, Hozumi Tanaka
Intell. Data Anal.2
2000 Query expansion using heterogeneous thesauri
Rila Mandala, Takenobu Tokunaga, Hozumi Tanaka
Inf. Process. Manag.2
1999 Complementing WordNet with Roget's and Corpus-based Thesauri for Information Retrieval
Rila Mandala, Takenobu Tokunaga, Hozumi Tanaka
EACL2
1999 Combining General Hand-Made and Automatically Constructed Thesauri for Query Expansion in Information Retrieval
Rila Mandala, Takenobu Tokunaga, Hozumi Tanaka
IJCAI2
1999 A Case Based Approach to the Generation of Musical Expression
Taizan Suzuki, Takenobu Tokunaga, Hozumi Tanaka
IJCAI2
1999 Sharing syntactic structures
abstract
Bracketed corpora are a very useful resource for natural language processing, but hard to build efficiently, leading to quantitative insufficiency for practical use. Disparities in morphological information, such as word segmentation and part-of-speech tag sets, are also troublesome. An application specific to a particular corpus often cannot be applied to another corpus. In this paper, we sketch out a method to build a corpus that has a fixed syntactic structure but varying morphological annotation based on the different tag set schemes utilized. Our system uses a two layered grammar, one layer of which is made up of replaceable tag-set-dependent rules while the other has no such tag set dependency. The input sentences of our system are bracketed corresponding to structural information of corpus. The parser can work using any tag set and grammar, and using the same input bracketing, we obtain corpus that shares partial syntactic structure.
Masahiro Ueki, Takenobu Tokunaga, Hozumi Tanaka
MTSummit2
1999 Combining Multiple Evidence from Different Types of Thesaurus for Query Expansion
abstract
Automatic query expansion has been known to be the most important method in overcoming the word mismatch problem in information retrieval. Thesauri have long been used by many researchers as a tool for query expansion. However only one type of thesaurus has generally been used. In this paper we analyze the characteristics of different thesaurus types and propose a method to combine them for query expansion. Experiments using the TREC collection proved the effectiveness of our method over those using one type of thesaurus.
Rila Mandala, Takenobu Tokunaga, Hozumi Tanaka
SIGIR2
1998 An Empirical Evaluation on Statistical Parsing of Japanese Sentences Using Lexical Association Statistics
Kiyoaki Shirai, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka
EMNLP3
1998 The RWC text databases
Kôiti Hasida, Hitoshi Isahara, Takenobu Tokunaga, Minako Hashimoto, Shiho Ogino, Wakako Kashino, Jun Toyoura, Hironobu Takahashi
LREC3
1998 Lessons from BMIR-J2: A Test Collection for Japanese IR Systems
Tsuyoshi Kitani, Yasushi Ogawa, Tetsuya Ishikawa, Haruo Kimoto, Ikuo Keshi, Jun Toyoura, Toshikazu Fukushima, Kunio Matsui, Yoshihiro Ueda, Tetsuya Sakai, Takenobu Tokunaga, Hiroshi Tsuruoka, Hidekazu Nakawatase, Teru Agata
SIGIR11
1998 Selective Sampling for Example-based Word Sense Disambiguation
Atsushi Fujii, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka
Comput. Linguistics3
1996 To what extent does case contribute to verb sense disambiguation?
Atsushi Fujii, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka
COLING3
1995 Hierarchical Bayesian Clustering for Automatic Text Classification
Makoto Iwayama, Takenobu Tokunaga
IJCAI2
1995 Automatic Thesaurus Construction based on Grammatical Relations
Takenobu Tokunaga, Makoto Iwayama, Hozumi Tanaka
IJCAI1
1995 Cluster-Based Text Categorization: A Comparison of Category Search Strategies
abstract
Article Cluster-based text categorization: a comparison of category search strategies Share on Authors: Makoto Iwayama Advanced Research Laboratory, Hitachi Ltd., Hatoyama, Saitama 350-03, Japan Advanced Research Laboratory, Hitachi Ltd., Hatoyama, Saitama 350-03, JapanView Profile , Takenobu Tokunaga Department of Computer Science, Tokyo Institute of Technology, 2-12-1, Ookayama, Meguro, Tokyo 152, Japan Department of Computer Science, Tokyo Institute of Technology, 2-12-1, Ookayama, Meguro, Tokyo 152, JapanView Profile Authors Info & Claims SIGIR '95: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrievalJuly 1995 Pages 273–280https://doi.org/10.1145/215206.215371Online:01 July 1995Publication History 63citation1,299DownloadsMetricsTotal Citations63Total Downloads1,299Last 12 Months38Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Makoto Iwayama, Takenobu Tokunaga
SIGIR2
1994 Analysis of Japanese Compound Nouns using Collocational Information
Yosiyuki Kobayasi, Takenobu Tokunaga, Hozumi Tanaka
COLING2
1990 A Method of Calculating the Measure of Salience in Understanding Metaphors
Makoto Iwayama, Takenobu Tokunaga, Hozumi Tanaka
AAAI2
1988 LangLab: a natural language analysis system
Takenobu Tokunaga, Makoto Iwayama, Hozumi Tanaka, Tadashi Kamiwaki
COLING1