VLDB 2026 Research / reviewers in the wild / expert
Takehito Utsuro
dblp:52/4891
· DBLP profile ↗
74ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0003-4072-1833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 58 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 9Databases, data management, data science and information retrieval · 7 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLMs Understand Punchlines? LLMs' Narrative Understanding Evaluation with Short-shorts
Jiashi Cheng, Takehito Utsuro |
LREC | 2 |
| 2026 | Coordinate Structure Extraction for Patent Claims Using Multilingual LLMs
Tsukasa Ishimaru, Takehito Utsuro, Masaaki Nagata |
LREC | 2 |
| 2026 | Biomedical concept recognition with error-aware negative-enhanced ranking frameworkabstractMOTIVATION: Mention-agnostic biomedical concept recognition (MA-BCR) requires inferring ontology concepts directly from passages, without relying on explicit mention spans. Prior work has mainly focused on generative and classification-based approaches. Ranking-based methods typically use a retrieve-rerank pipeline, and this paradigm has not been systematically studied for MA-BCR. Consequently, it remains unclear how ranking-based approaches compare with existing paradigms and what types of supervision are most beneficial for ranker training under limited annotation settings. RESULTS: Through a systematic comparison of ranking-, generative-, and classification-based paradigms, we show that a two-stage retrieve-rerank architecture is the most robust and scalable backbone for MA-BCR. Building on this finding, we propose ENR, an error-aware negative-enhanced ranking framework that augments training with false positives collected from heterogeneous recognizers, improving reranking performance without increasing inference-time cost. Experiments on MM-HPO and MM-GO (two datasets derived from MedMentions-ST21pv) demonstrate that ENR substantially outperforms prior approaches. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are available in Github at https://github.com/sl-633/enr-recognizer or in Zenodo at https://doi.org/10.5281/zenodo.20730803. Noriki Nishida, Fei Cheng 0002, Takehito Utsuro, Yuji Matsumoto 0001 |
Bioinform. | 4 |
| 2025 | Patent Claim Translation via Continual Pre-training of Large Language Models with Parallel DataabstractRecent advancements in large language models (LLMs) have enabled their application across various domains. However, in the field of patent translation, Transformer encoder-decoder based models remain the standard approach, and the potential of LLMs for translation tasks has not been thoroughly explored. In this study, we conducted patent claim translation using an LLM fine-tuned with parallel data through continual pre-training and supervised fine-tuning, following the methodology proposed by Guo et al. (2024) and Kondo et al. (2024). Comparative evaluation against the Transformer encoder-decoder based translations revealed that the LLM achieved high scores for both BLEU and COMET. This demonstrated improvements in addressing issues such as omissions and repetitions. Nonetheless, hallucination errors, which were not observed in the traditional models, occurred in some cases and negatively affected the translation quality. This study highlights the promise of LLMs for patent translation while identifying the challenges that warrant further investigation. Haruto Azami, Minato Kondo, Takehito Utsuro, Masaaki Nagata |
MTSummit (1) | 3 |
| 2025 | Improving Japanese-English Patent Claim Translation with Clause Segmentation Models based on Word AlignmentabstractIn patent documents, patent claims represent a particularly important section as they define the scope of the claims. However, due to the length and unique formatting of these sentences, neural machine translation (NMT) systems are prone to translation errors, such as omissions and repetitions. To address these challenges, this study proposes a translation method that first segments the source sentences into multiple shorter clauses using a clause segmentation model tailored to facilitate translation. These segmented clauses are then translated using a clause translation model specialized for clause-level translation. Finally, the translated clauses are rearranged and edited into the final translation using a reordering and editing model. In addition, this study proposes a method for constructing clause-level parallel corpora required for training the clause segmentation and clause translation models. This method leverages word alignment tools to create clause-level data from sentence-level parallel corpora. Experimental results demonstrate that the proposed method achieves statistically significant improvements in BLEU scores compared to conventional NMT models. Furthermore, for sentences where conventional NMT models exhibit omissions and repetitions, the proposed method effectively suppresses these errors, enabling more accurate translations. Masato Nishimura, Kosei Buma, Takehito Utsuro, Masaaki Nagata |
MTSummit (1) | 3 |
| 2024 | CoT based Few-Shot Learning of Negative Comments FilteringabstractIn this paper, we filter negative comments on videos using large language models (LLMs). In comment filtering using LLMs such as GPT-4 (zero-shot learning), over filtering occurs frequently. This over filtering can be improved by using few-shot learning based on chain-of-thought reasoning. Few-shot learning is a strategy to improve the performance of LLMs by presenting a small number of examples to LLMs. Chain-of-thought (CoT) [18] is one of the reasoning strategy which allows LLMs to engage in more logical reasoning by generating the process of inference. In this paper, we propose and evaluate an effective method of few-shot learning based on CoT in the task of negative comment filtering. The results showed that over filtering can be effectively reduced by selecting appropriate few-shot comments from the clusters created by GPT-4, and providing them for the prompt as demonstrations of CoT. Takashi Mitadera, Takehito Utsuro |
IEEE Big Data | 2 |
| 2024 | Emotion Classification of Lyrics through Summarization by Large Language ModelsabstractWe propose a method that utilizes a large language model in the task of lyrics emotion classification. We especially employ GPT-4o, which is expected to deliver high performance to conduct emotion classification of lyrics, where we propose a few-shot prompt of GPT-4o for classification into 6 classes. We developed a dataset of 181 lyrics categorized into six classifications to evaluate GPT-4O’s performance. This dataset maintains a reasonable level of validity as it was carefully curated by the first author, then independently reclassified by six annotators, with final classifications determined through majority voting. Performing detailed classification, we achieved over 75% classification performance for the total accuracy. This accuracy was achieved through performing extractive summarization of the lyrics into optimal number of characters. In other words, rather than feeding the collected lyrics directly into GPT-4o, we implemented a two-stage process where GPT-4o first automatically summarizes the lyrics before they are used. Using a confusion matrix, we analyzed the tendencies of classification errors and showed the possibility of improving the total accuracy. Sho Miyakawa, Takehito Utsuro |
IEEE Big Data | 2 |
| 2024 | Emoji Prediction of Japanese X Posts by LLMs
Yijie Hua, Takehito Utsuro |
PACLIC | 2 |
| 2023 | Target Language Monolingual Translation Memory based NMT by Cross-lingual Retrieval of Similar Translations and RerankingabstractRetrieve-edit-rerank is a text generation framework composed of three steps: retrieving for sentences using the input sentence as a query, generating multiple output sentence candidates, and selecting the final output sentence from these candidates. This simple approach has outperformed other existing and more complex methods. This paper focuses on the retrieving and the reranking steps. In the retrieving step, we propose retrieving similar target language sentences from a target language monolingual translation memory using language-independent sentence embeddings generated by mSBERT or LaBSE. We demonstrate that this approach significantly outperforms existing methods that use monolingual inter-sentence similarity measures such as edit distance, which is only applicable to a parallel translation memory. In the reranking step, we propose a new reranking score for selecting the best sentences, which considers both the log-likelihood of each candidate and the sentence embeddings based similarity between the input and the candidate. We evaluated the proposed method for English-to-Japanese translation on the ASPEC and English-to-French translation on the EU Bookshop Corpus (EUBC). The proposed method significantly exceeded the baseline in BLEU score, especially observing a 1.4-point improvement in the EUBC dataset over the original Retrieve-Edit-Rerank method. Takuya Tamura, Takehito Utsuro, Masaaki Nagata |
MTSummit (1) | 3 |
| 2023 | Leveraging Highly Accurate Word Alignment for Low Resource Translation by Pretrained Multilingual ModelabstractRecently, there has been a growing interest in pretraining models in the field of natural language processing. As opposed to training models from scratch, pretrained models have been shown to produce superior results in low-resource translation tasks. In this paper, we introduced the use of pretrained seq2seq models for preordering and translation tasks. We utilized manual word alignment data and mBERT-based generated word alignment data for training preordering and compared the effectiveness of various types of mT5 and mBART models for preordering. For the translation task, we chose mBART as our baseline model and evaluated several input manners. Our approach was evaluated on the Asian Language Treebank dataset, consisting of 20,000 parallel data in Japanese, English and Hindi, where Japanese is either on the source or target side. We also used in-house 3,000 parallel data in Chinese and Japanese. The results indicated that mT5-large trained with manual word alignment achieved a preordering performance exceeding 0.9 RIBES score on Ja-En and Ja-Zh pairs. Moreover, our proposed approach significantly outperformed the baseline model in most translation directions of Ja-En, Ja-Zh, and Ja-Hi pairs in at least one of BLEU/COMET scores. Minato Kondo, Takuya Tamura, Takehito Utsuro, Masaaki Nagata |
MTSummit (1) | 4 |
| 2023 | Enhanced Retrieve-Edit-Rerank Framework with kNN-MT
Takuya Tamura, Takehito Utsuro, Masaaki Nagata |
PACLIC | 3 |
| 2023 | Large Scale Evaluation of End-to-End Pipeline of Speaker to Dialogue Attribution in Japanese Novels
Yuki Zenimoto, Shinzan Komata, Takehito Utsuro |
PACLIC | 3 |
| 2022 | Tweet Review Mining focusing on Celebrities by Machine Reading Comprehension based on BERT
Yuta Nozaki, Kotoe Sugawara, Yuki Zenimoto, Takehito Utsuro |
PACLIC | 4 |
| 2022 | Speaker Identification of Quotes in Japanese Novels based on Gender Classification Model by BERT
Yuki Zenimoto, Takehito Utsuro |
PACLIC | 2 |
| 2022 | Developing and Evaluating a Dataset for How-to Tip Machine Reading at Scale
Fuzhu Zhu, Shuting Bai, Tingxuan Li, Takehito Utsuro |
PACLIC | 4 |
| 2021 | Language and Speaker-Independent Feature Transformation for End-to-End Multilingual Speech Recognition
Tomoaki Hayakawa, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki |
Interspeech | 4 |
| 2021 | Voice Activity Detection for Live Speech of Baseball Game Based on Tandem Connection with Speech/Noise Separation Model
Yuto Nonaka, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki |
Interspeech | 4 |
| 2021 | Evaluating a How-to Tip Machine Comprehension Model with QA Examples collected from a Community QA Site
Tingxuan Li, Shuting Bai, Takehito Utsuro, Fuzhu Zhu |
PACLIC | 3 |
| 2020 | Automatic Fluency Evaluation of Spontaneous Speech Using Disfluency-Based FeaturesabstractThis paper describes an automatic fluency evaluation of spontaneous speech. Although we regularly observe a variety of different disfluencies in spontaneous speech, we focus on two types of phenomena, i.e., filled pauses and word fragments. This paper aims to reveal that these two types of disfluencies have effects on speech fluency evaluation differently. To this end, we conduct a series of SVM classification experiments on the Japanese spontaneous speech corpus. The experimental results show that the features derived from word fragments are effective in evaluating disfluent speech especially when combined with prosodic features such as speech rate and pauses/silence, while the features from filled pauses are not effective in evaluating fluency. Huaijin Deng, Youchao Lin, Takehito Utsuro, Akio Kobayashi, Hiromitsu Nishizaki, Junichi Hoshino |
ICASSP | 3 |
| 2020 | Integrating Disfluency-based and Prosodic Features with Acoustics in Automatic Fluency Evaluation of Spontaneous SpeechabstractThis paper describes an automatic fluency evaluation of spontaneous speech. In the task of automatic fluency evaluation, we integrate diverse features of acoustics, prosody, and disfluency-based ones. Then, we attempt to reveal the contribution of each of those diverse features to the task of automatic fluency evaluation. Although a variety of different disfluencies are observed regularly in spontaneous speech, we focus on two types of phenomena, i.e., filled pauses and word fragments. The experimental results demonstrate that the disfluency-based features derived from word fragments and filled pauses are effective relative to evaluating fluent/disfluent speech, especially when combined with prosodic features, e.g., such as speech rate and pauses/silence. Next, we employed an LSTM based framework in order to integrate the disfluency-based and prosodic features with time sequential acoustic features. The experimental evaluation results of those integrated diverse features indicate that time sequential acoustic features contribute to improving the model with disfluency-based and prosodic features when detecting fluent speech, but not when detecting disfluent speech. Furthermore, when detecting disfluent speech, the model without time sequential acoustic features performs best even without word fragments features, but only with filled pauses and prosodic features. Huaijin Deng, Youchao Lin, Takehito Utsuro, Akio Kobayashi, Hiromitsu Nishizaki, Junichi Hoshino |
LREC | 3 |
| 2020 | Text Mining of Evidence on Infants' Developmental Stages for Developmental Order Acquisition from Picture Book Reviews
Miho Kasamatsu, Takehito Utsuro, Yu Saito, Yumiko Ishikawa |
PACLIC | 2 |
| 2019 | Selecting Informative Context Sentence by Forced Back-Translation
Ryuichiro Kimura, Shohei Iida, Hongyi Cui, Po-Hsuan Hung, Takehito Utsuro, Masaaki Nagata |
MTSummit (1) | 5 |
| 2018 | Identifying Tips Web Sites of a Specific Query based on Search Engine Suggests and the Topic DistributionabstractThis paper proposes techniques of automatically discovering tips Web sites from a large collection of Web pages using a topic model and support vector machine (SVM). Tips refer to practical knowledge or expertise that is used to help accomplish certain tasks in a particular field. We designed several approaches of extracting features with respect to domain names based on their distribution among Web pages and candidate tips Web sites. In addition, search engine suggests, the query keywords used to fetch Web pages from the search engine are also considered to present patterns that can be potential features. It was discovered from our dataset that domain names of tips Web sites (Web sites containing tips on a certain specific theme) are more likely dispersed among topics and Web pages. These domain names also tend to correspond to a larger number of search engine suggests. This paper verifies such observed patterns by training an SVM using those extracted features. Evaluation is performed in precision and recall to measure correctness of classifying whether or not a domain name belongs to a tips Web site. Yohei Ohkawa, Shuto Kawabata, Wenbin Niu, Youchao Lin, Takehito Utsuro, Yasuhide Kawada |
IEEE BigData | 6 |
| 2018 | Learning to Identify Rush Strategies in StarCraft
Teguh Budianto, Hyunwoo Oh, Takehito Utsuro |
ICEC | 3 |
| 2017 | Constraint-Based Modelling as a Tutoring Framework for Japanese Honorifics
Zachary T. Chung, Takehito Utsuro, Ma. Mercedes T. Rodrigo |
AIED | 2 |
| 2017 | Movie Summarization Based on Alignment of Plot and ShotsabstractThis paper proposes a method of assisting movie summarization using plotinformation. A plot of a movie available at Wikipedia contains a majorstory of the movie. From such a plot of a movie, we extract severalimportant sentences as the content of summary. For summarizing movie, the key work is finding the best alignment between sentences of plot andshots which are segmented from a movie. There are two cues used tomeasure the similarity between a sentence and a shot. One is based oncharacter appearing in both sentence and shot, another is based on wordsmatching. Then an alignment method based on dynamic programming is applyto optimize the alignment. Finally an experiment on movie "RomanHoliday" and "Alice in the wonderland" show the effectiveness of this method. Xueshan Li, Takehito Utsuro, Hiroshi Uehara |
AINA | 2 |
| 2017 | Identifying Rush Strategies Employed in StarCraft II Using Support Vector Machines
Teguh Budianto, Hyunwoo Oh, Zi Long, Takehito Utsuro |
ICEC | 5 |
| 2017 | Mining Preferences on Identifying Werewolf Players from Werewolf Game Logs
Yuki Hatori, Youchao Lin, Takehito Utsuro |
ICEC | 4 |
| 2017 | Generating the Expression of the Move of Go by Classifier Learning
Natsumi Mori, Takehito Utsuro |
ICEC | 2 |
| 2017 | Deep Photo Rally: Let's Gather Conversational Pictures
Kazuki Ookawara, Hayaki Kawata, Masafumi Muta, Soh Masuko, Takehito Utsuro, Junichi Hoshino |
ICEC | 5 |
| 2017 | Neural Machine Translation Model with a Large Vocabulary Selected by Branching Entropy
Zi Long, Ryuichiro Kimura, Takehito Utsuro, Tomoharu Mitsuhashi, Mikio Yamamoto |
MTSummit (1) | 3 |
| 2017 | Clustering search engine suggests by integrating a topic model and word embeddingsabstractThe background of this paper is the issue of how to overview the knowledge of a given query keyword. Especially, we focus on concerns of those who search for Web pages with a given query keyword. The Web search information needs of a given query keyword is collected through search engine suggests. Given a query keyword, we collect up to around 1,000 suggests, while many of them are redundant. We cluster redundant search engine suggests based on a topic model. However, one limitation of the topic model based clustering of search engine suggests is that the granularity of the topics, i.e., the clusters of search engine suggests, is too coarse. In order to overcome the problem of the coarse-grained clusters of search engine suggests, this paper further applies the word embedding technique to the Web pages used during the training of the topic model, in addition to the text data of the whole Japanese version of Wikipedia. Then, we examine the word embedding based similarity between search engines suggests and further classify search engine suggests within a single topic into finer-grained subtopics based on the similarity of word embeddings. Evaluation results prove that the proposed approach performs well in the task of subtopic clustering of search engine suggests. Tian Nie, Youchao Lin, Takehito Utsuro, Yasuhide Kawada |
SNPD | 5 |
| 2016 | Generating a werewolf game log digest of inferring each player's roleabstractWhile playing the communication game “Are You a Werewolf”, a player always guesses other players' roles through discussions, based on one's own role and other players' crucial utterances. The underlying goal of this paper is to construct an agent that can analyze the participating players' utterances and play the werewolf game as if it is a human. For the first step of this underlying goal, given a specific player participating in the wolf game, this paper studies how to generate a digest of inferring other players' roles from the viewpoint of the given specific player. In this inference process, we regard the werewolf game rules as well as certain common sense as inference rules. Then, we develop a set of inference rules and apply them to infer the participating players' roles from a real werewolf game log. Youchao Lin, Mizuho Baba, Takehito Utsuro |
ICIS | 3 |
| 2016 | Utilizing texts of picture book reviews for extracting children's behavioral characteristics in language acquisitionabstractPointing behavior in childhood is typical developmental sign having strong correlation to his or her language development. This paper focuses on the pointing behavior accompanied by utterance during picture book reading. With this respect, we make use of picture books' review data amounting to approximately 320 thousand, and analyze the reviews reflecting the pointing behavior with children's utterance. The results show patterns of the pointing with utterance change corresponding to children's developmental stages. Also, one of the pointing patterns is found to have strong relationship with a certain type of picture books. Hiroshi Uehara, Mizuho Baba, Takehito Utsuro |
ICIS | 3 |
| 2016 | Analyzing Time Series Changes of Correlation between Market Share and Concerns on Companies measured through Search Engine Suggests
Takakazu Imada, Yusuke Inoue 0001, Syunya Doi, Tian Nie, Takehito Utsuro, Yasuhide Kawada |
LREC | 7 |
| 2015 | Two-step spoken term detection using SVM classifier trained with pre-indexed keywords based on ASR result
Kentaro Domoto, Takehito Utsuro, Naoki Sawada, Hiromitsu Nishizaki |
INTERSPEECH | 2 |
| 2015 | Detecting an Infant's Developmental Reactions in Reviews on Picture Books
Hiroshi Uehara, Mizuho Baba, Takehito Utsuro |
PACLIC | 3 |
| 2013 | Time Series Topic Modeling and Bursty Topic Detection of Correlated News and Twitter
Daichi Koike, Yusuke Takahashi, Takehito Utsuro, Masaharu Yoshioka, Noriko Kando |
IJCNLP | 3 |
| 2012 | Detecting Japanese Compound Functional Expressions using Canonical/Derivational Relation
Takafumi Suzuki, Yusuke Abe, Itsuki Toyota, Takehito Utsuro, Suguru Matsuyoshi, Masatoshi Tsuchiya |
LREC | 4 |
| 2012 | Cross-Lingual Topic Alignment in Time Series Japanese / Chinese News
Shuo Hu, Yusuke Takahashi, Liyi Zheng, Takehito Utsuro, Masaharu Yoshioka, Noriko Kando, Tomohiro Fukuhara, Hiroshi Nakagawa, Yoji Kiyota |
PACLIC | 4 |
| 2011 | Semi-Automatic Identification of Bilingual Synonymous Technical Terms from Phrase Tables and Parallel Patent Sentences
Takehito Utsuro, Mikio Yamamoto |
PACLIC | 2 |
| 2010 | Utilizing Semantic Equivalence Classes of Japanese Functional Expressions in Translation Rule Acquisition from Parallel Patent Sentences
Taiji Nagasaka, Ran Shimanouchi, Akiko Sakamoto, Takafumi Suzuki, Yohei Morishita, Takehito Utsuro, Suguru Matsuyoshi |
LREC | 6 |
| 2009 | Visualizing Cross-Lingual/Cross-Cultural Differences in Concerns in Multilingual Blogs
Hiroyuki Nakasaki, Mariko Kawaba, Sayuri Yamazaki, Takehito Utsuro, Tomohiro Fukuhara |
ICWSM | 4 |
| 2009 | Towards Conceptual Indexing of the Blogosphere through Wikipedia Topic Hierarchy
Mariko Kawaba, Daisuke Yokomoto, Hiroyuki Nakasaki, Takehito Utsuro, Tomohiro Fukuhara |
PACLIC | 4 |
| 2009 | Identifying and Utilizing the Class of Monosemous Japanese Functional Expressions in Machine Translation
Akiko Sakamoto, Taiji Nagasaka, Takehito Utsuro, Suguru Matsuyoshi |
PACLIC | 3 |
| 2009 | Evaluating effects of machine translation accuracy on cross-lingual patent retrievalabstractWe organized a machine translation (MT) task at the Seventh NTCIR Workshop. Participating groups were requested to machine translate sentences in patent documents and also search topics for retrieving patent documents across languages. We analyzed the relationship between the accuracy of MT and its effects on the retrieval accuracy. Atsushi Fujii, Masao Utiyama, Mikio Yamamoto, Takehito Utsuro |
SIGIR | 4 |
| 2008 | Cross-Lingual Blog Analysis based on Multilingual Blog Distillation from Multilingual Wikipedia Entries
Mariko Kawaba, Hiroyuki Nakasaki, Takehito Utsuro, Tomohiro Fukuhara |
ICWSM | 3 |
| 2008 | Collecting and Analyzing Japanese Splogs based on Characteristics of Keywords
Yuuki Sato, Takehito Utsuro, Tomohiro Fukuhara, Yasuhide Kawada, Yoshiaki Murakami, Hiroshi Nakagawa, Noriko Kando |
ICWSM | 2 |
| 2008 | Producing a Test Collection for Patent Machine Translation in the Seventh NTCIR Workshop
Atsushi Fujii, Masao Utiyama, Mikio Yamamoto, Takehito Utsuro |
LREC | 4 |
| 2006 | Japanese Idiom Recognition: Drawing a Line between Literal and Idiomatic Meanings
Chikara Hashimoto, Satoshi Sato, Takehito Utsuro |
ACL | 3 |
| 2006 | Compiling French-Japanese Terminologies from the Web
Xavier Robitaille, Yasuhiro Sasaki, Masatsugu Tonoike, Satoshi Sato, Takehito Utsuro |
EACL | 5 |
| 2006 | Adjective-to-Verb Paraphrasing in Japanese Based on Lexical Constraints of Verbs
Atsushi Fujita, Naruaki Masuno, Satoshi Sato, Takehito Utsuro |
INLG | 4 |
| 2004 | Integrating Cross-Lingually Relevant News Articles and Monolingual Web Documents in Bilingual Lexicon Acquisition
Takehito Utsuro, Kohei Hino, Mitsuhiro Kida, Seiichi Nakagawa, Satoshi Sato |
COLING | 1 |
| 2004 | Keyword recognition and extraction by multiple-LVCSRs with 60, 000 words in speech-driven WEB retrieval taskabstractThis paper presents speech-driven Web retrieval models which accepts spoken search topics (queries) in the NTCIR-3 Web retrieval task. We experimentally evaluate the techniques of combining outputs of multiple LVCSR models with a language model(LM) with a 60,000 vocabulary size in recognition of spoken queries. As model combination techniques, we use the SVM learning. We show that the techniques of multiple LVCSR model combination can achieve improvement both in speech recognition and retrieval accuracies in speech-driven text retrieval. Comparing with the retrieval accuracies when a LM with a 20,000/60,000 vocabulary size is used in LVCSRs, the LM that has larger size of the vocabulary improves also retrieval accuracies. Masahiko Matsushita, Hiromitsu Nishizaki, Seiichi Nakagawa, Takehito Utsuro |
INTERSPEECH | 4 |
| 2004 | Unsupervised speaker adaptation using high confidence portion recognition results by multiple recognition systems
Tomohiro Watanabe, Hiromitsu Nishizaki, Takehito Utsuro, Seiichi Nakagawa |
INTERSPEECH | 3 |
| 2003 | Effect of Cross-Language IR in Bilingual Lexicon Acquisition from Comparable Corpora
Takehito Utsuro, Takashi Horiuchi, Takeshi Hamamoto, Kohei Hino, Takeaki Nakayama |
EACL | 1 |
| 2003 | Confidence of agreement among multiple LVCSR models and model combination by SVMabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, we have proposed features for confidence measures that are extracted from outputs of more than one LVCSR models. For further analysis of the proposed confidence measure, this paper examines the correlation between each word's confidence and the word's features such as its part-of-speech and syllable length. We then apply SVM learning technique to the task of combining outputs of multiple LVCSR models, where, as features of SVM learning, information such as the pairs of the models which output the hypothesized word are useful for improving the word recognition rate. Experimental results show that the combination results achieve a relative word error reduction of up to 72 % against the best performing single model and that of up to 36 % against ROVER. Takehito Utsuro, Yasuhiro Kodama, Tomohiro Watanabe, Hiromitsu Nishizaki, Seiichi Nakagawa |
ICASSP (1) | 1 |
| 2003 | Evaluating multiple LVCSR model combination in NTCIR-3 speech-driven web retrieval taskabstractThis paper studies speech-driven Web retrieval models which accepts spoken search topics (queries) in the NTCIR-3 Web retrieval task. The major focus of this paper is on improving speech recognition accuracy of spoken queries and then improving retrieval accuracy in speech-driven Web retrieval. We experimentally evaluate the techniques of combining outputs of multiple LVCSR models in recognition of spoken queries. As model combination techniques, we compare the SVM learning technique and conventional voting schemes such as ROVER. We show that the techniques of multiple LVCSR model combination can achieve improvement both in speech recognition and retrieval accuracies in speech-driven text retrieval. We also show that model combination by SVM learning outperforms conventional voting schemes both in speech recognition and retrieval accuracies. Masahiko Matsushita, Hiromitsu Nishizaki, Takehito Utsuro, Yasuhiro Kodama, Seiichi Nakagawa |
INTERSPEECH | 3 |
| 2002 | Combining Outputs of Multiple Japanese Named Entity Chunkers by StackingabstractIn this paper, we propose a method for learning a classifier which combines outputs of more than one Japanese named entity extractors. The proposed combination method belongs to the family of stacked generalizers, which is in principle a technique of combining outputs of several classifiers at the first stage by learning a second stage classifier to combine those outputs at the first stage. Individual models to be combined are based on maximum entropy models, one of which always considers surrounding contexts of a fixed length, while the other considers those of variable lengths according to the number of constituent morphemes of named entities. As an algorithm for learning the second stage classifier, we employ a decision list learning method. Experimental evaluation shows that the proposed method achieves improvement over the best known results with Japanese named entity extractors based on maximum entropy models. Takehito Utsuro, Manabu Sassano, Kiyotaka Uchimoto |
EMNLP | 1 |
| 2002 | A confidence measure based on agreement among multiple LVCSR models - correlation between pair of acoustic models and confidenceabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, this paper studies features for confidence measures that are extracted from outputs of more than one LVCSR models. More specifically, this paper experimentally evaluates the agreement among the outputs of multiple Japanese LVCSR models, with respect to whether it is effective as an estimate of confidence for each hypothesized word. The results of experimental evaluation show that the agreement between the outputs with two LVCSR models with different decoders and acoustic models can achieve quite reliable confidence. Furthermore, among various features of acoustic models based on Gaussian mixture HMMs, it is concluded that ones such as whether or not to have short pause models, as well as different units in HMMs (e.g., triphone model or syllable model) are the most effective in achieving highly reliable confidence. Takehito Utsuro, Tetsuji Harada, Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2002 | A Web-based English Abstract Writing Tool Using a Tagged E-J Parallel Corpus
Masumi Narita, Kazuya Kurokawa, Takehito Utsuro |
LREC | 3 |
| 2001 | Experimental evaluation on confidence of agreement among multiple Japanese LVCSR modelsabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, this paper studies features for confidence measures that are extracted from outputs of more than one LVCSR models. More specifically, this paper experimentally evaluates the agreement among the outputs of multiple Japanese LVCSR models, with respect to whether it is effective as an estimate of confidence for each hypothesized word. The results of experimental evaluation show that the agreement between the outputs with two acoustic models which have different units in HMMs, such as phonemes and syllables, can achieve quite reliable confidence. 1. Yasuhiro Kodama, Takehito Utsuro, Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2000 | Named Entity Chunking Techniques in Supervised Learning for Japanese Named Entity Recognition
Manabu Sassano, Takehito Utsuro |
COLING | 2 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 11 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 7 |
| 2000 | Learning Preference of Dependency between Japanese Subordinate Clauses and its Evaluation in Parsing
Takehito Utsuro |
LREC | 1 |
| 2000 | Minimally Supervised Japanese Named Entity Recognition: Resources and Evaluation
Takehito Utsuro, Manabu Sassano |
LREC | 1 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 8 |
| 1996 | Sense Classification of Verbal Polysemy based-on Bilingual Class/Class Association
Takehito Utsuro |
COLING | 1 |
| 1994 | Bilingual Text, Matching using Bilingual Dictionary and Statistics
Takehito Utsuro, Hiroshi Ikeda, Masaya Yamane, Yuji Matsumoto 0001, Makoto Nagao |
COLING | 1 |
| 1994 | Thesaurus-based Efficient Example Retrieval by Generating Retrieval Queries from Similarities
Takehito Utsuro, Kiyotaka Uchimoto, Mitsutaka Matsumoto, Makoto Nagao |
COLING | 1 |
| 1993 | Sructural Matching of Parallel TextsabstractThis paper describes a method for finding structural matching between parallel sentences of two languages, (such as Japanese and English). Parallel sentences are analyzed based on unification grammars, and structural matching is performed by making use of a similarity measure of word pairs in the two languages. Syntactic ambiguities are resolved simultaneously in the matching process. The results serve as a useful source for extracting linguistic and lexical knowledge. Yuji Matsumoto 0001, Hiroyuki Ishimoto, Takehito Utsuro |
ACL | 3 |
| 1993 | Verbal Case Frame Acquisition from Bilingual Corpora
Takehito Utsuro, Yuji Matsumoto 0001, Makoto Nagao |
IJCAI | 1 |
| 1992 | Lexical Knowledge Acquisition from Bilingual Corpora
Takehito Utsuro, Yuji Matsumoto 0001, Makoto Nagao |
COLING | 1 |