Kazuhiro Seki

dblp:68/4147 · DBLP profile ↗
← Back
29ranked-venue papers
15as first author
5since 2021 · last 2025
0000-0002-1967-4334ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Speech-Scenario Generation Based on the Philosophy of a Prominent Leader Within a Small Community
Tetsuya Kitahata, Kazuhiro Seki, Akiyo Nadamoto
DEXA (1)2
2025 Two-Stage Fine-Tuning for Dialogue Generation with Small Community Prominent Leaders' Philosophies
Tetsuya Kitahata, Kazuhiro Seki, Akiyo Nadamoto
iiWAS2
2023 Topic-Sentiment Analysis of Central Bank Press Conferences: BOJ Case Study
abstract
This paper reports our ongoing work on topic-sentiment analysis of central bank press conferences. Central banks rely on communication to implement monetary policy and convey their objectives to the public. Apart from the policies themselves, the sentiment or tone conveyed during a governor’s press conference could also influence both immediate economic indicators and long-term economic trajectories. This study analyzes the Japanese texts from the press conferences of two contrasting Bank of Japan (BOJ) governors: Shirakawa and Kuroda. By employing topic and sentiment analysis techniques, we define a topic-sentiment distribution representing the sentiment of governors’ statements per topic. Our findings highlight distinct divergences in tone between the two governors in relation to salient economic topics. This divergence provides crucial insights into the nuanced relationship between central bank communication styles and resultant economic outcomes.
Kazuhiro Seki, Masahiko Shibamoto, Takashi Kamihigashi
IEEE Big Data1
2022 Turning News Texts into Business Sentiment
Kazuhiro Seki
ECIR (2)1
2022 News-based business sentiment and its properties as an economic index
abstract
This paper presents an approach to measuring business sentiment based on textual data. Business sentiment has been measured by traditional surveys, which are costly and time-consuming to conduct. To address the issues, we take advantage of daily newspaper articles and adopt a self-attention-based model to define a business sentiment index, named S-APIR, where outlier detection models are investigated to properly handle various genres of news articles. Moreover, we propose a simple approach to temporally analyzing how much any given event contributed to the predicted business sentiment index. To demonstrate the validity of the proposed approach, an extensive analysis is carried out on 12 years’ worth of newspaper articles. The analysis shows that the S-APIR index is strongly and positively correlated with established survey-based index (up to correlation coefficient r=0.937) and that the outlier detection is effective especially for a general newspaper. Also, S-APIR is compared with a variety of economic indices, revealing the properties of S-APIR that it reflects the trend of the macroeconomy as well as the economic outlook and sentiment of economic agents. Moreover, to illustrate how S-APIR could benefit economists and policymakers, several events are analyzed with respect to their impacts on business sentiment over time.
Kazuhiro Seki, Yusuke Ikuta, Yoichi Matsubayashi
Inf. Process. Manag.1
2019 Evaluating Interactive Clustering for Biomedical Information Retrieval
Michael Segundo Ortiz, Kazuhiro Seki, Javed Mostafa
AMIA2
2019 Effectiveness and Efficiency for Document Clustering in Biomedicine
abstract
It is crucial for biomedical information retrieval, or clinical decision support in particular, to discover relevant biomedical/clinical information buried in scientific publications. At present, typical search interface is based on keywords as queries and returns a ranked list of documents, which is suited for finding simple factoids but not ideal for more complex information needs required for clinical decision support. A search interface deemed more suitable for this kind of tasks is cluster-based browsing, where retrieved documents are topically grouped for more intuitive exploration. To adopt this model, however, one needs to consider not only the effectiveness but also the efficiency of clustering framework as clustering is a computationally costly operation. As a first step toward a cluster-based browsing information exploration, this paper empirically studies representative feature selection/extraction methods and clustering algorithms for their effectiveness and efficiency.
Kazuhiro Seki, Michael Segundo Ortiz, Javed Mostafa
BIBM1
2018 Exploring Neural Translation Models for Cross-Lingual Text Similarity
abstract
This paper explores a neural network-based approach to computing similarity of two texts written in different languages. Such similarity can be useful for a variety of applications including cross-lingual information retrieval and cross-lingual text classification. To compute similarity, we focus on neural machine translation models and examine the utility of their intermediate states. Through experiments on an English-Japanese translation corpus, it is demonstrated that the intermediate states of input texts are indeed beneficial for computing cross-lingual text similarity, outperforming other approaches including a strong machine translation-based baseline.
Kazuhiro Seki
CIKM1
2014 Time-Aware Latent Concept Expansion for Microblog Search
Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara
ICWSM2
2014 Predicting Stock Market Trends by Recurrent Deep Neural Networks
Akira Yoshihara, Kazuki Fujikawa, Kazuhiro Seki, Kuniaki Uehara
PRICAI3
2013 Agglomerative co-clustering for synonymous phrases based on common effects and influences
abstract
This paper proposes an approach to clustering synonymous noun phrases focusing on two types of predicate argument relations extracted from potentially big textual data. One is associated with common effects, the other with common influences. Based on the context represented by those relations, a matrix is constructed with rows being noun phrases and columns being a pair of a noun phrase and a verb phrase. Following the distribution hypothesis often adopted in the literature, it is assumed that rows (i.e., noun phrases) with similar distributions share similar meanings. Due to the inherent sparsity of the matrix, however, two strategies are taken to group noun phrases having similar distributions. One strategy is to simply use a large-scale corpus, which however results in an even larger matrix. To handle the large matrix, a parallel distributed programming model, MapReduce, is employed. The other is to adopt hierarchical agglomerative co-clustering and approximates its computation in a way suited to the MapReduce programming model. The proposed approach is evaluated based on a series of experiments in terms of the validity of our underlying assumptions, processing time, quality of the resulting clusters, and effect of parallelization.
Koji Kumanami, Kazuhiro Seki, Kuniaki Uehara
IEEE BigData2
2013 Improving pseudo-relevance feedback via tweet selection
abstract
Query expansion methods using pseudo-relevance feedback have been shown effective for microblog search because they can solve vocabulary mismatch problems often seen in searching short documents such as Twitter messages (tweets), which are limited to 140 characters. Pseudo-relevance feedback assumes that the top ranked documents in the initial search results are relevant and that they contain topic-related words appropriate for relevance feedback. However, those assumptions do not always hold in reality because the initial search results often contain many irrelevant documents. In such a case, only a few of the suggested expansion words may be useful with many others being useless or even harmful. To overcome the limitation of pseudo-relevance feedback for microblog search, we propose a novel query expansion method based on two-stage relevance feedback that models search interests by manual tweet selection and integration of lexical and temporal evidence into its relevance model. Our experiments using a corpus of microblog data (the Tweets2011 corpus) demonstrate that the proposed two-stage relevance feedback approaches considerably improve search result relevance over almost all topics.
Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara
CIKM2
2013 Combining Recency and Topic-Dependent Temporal Variation for Microblog Search
Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara
ECIR2
2013 Supervised Hypothesis Discovery Using Syllogistic Patterns in the Biomedical Literature
Kazuhiro Seki, Kuniaki Uehara
IJCAI1
2013 Block coordinate descent algorithms for large-scale sparse multiclass classification
Mathieu Blondel, Kazuhiro Seki, Kuniaki Uehara
Mach. Learn.2
2013 A shape-based similarity measure for time series data with ensemble learning
Tetsuya Nakamura, Keishi Taki, Hiroki Nomiya, Kazuhiro Seki, Kuniaki Uehara
Pattern Anal. Appl.4
2012 Parallel distributed trajectory pattern mining using MapReduce
abstract
This paper proposes a new approach to trajectory pattern mining, which attempts to discover frequent movement patterns from the trajectories of moving objects. For dealing with a large volume of trajectory data, traditional approaches quantize them by a grid with a fixed resolution. However, an appropriate resolution often varies across different areas of trajectories. Simply increasing the resolution cannot capture broad patterns and consumes unnecessarily large computational resources. To solve the problem, we propose a hierarchical grid-based approach with quadtree search. The approach initially searches for frequent patterns with a coarse grid and drills down into a finer grid level to discover more minute patterns. The algorithm is naturally parallelized and implemented in the MapReduce programming model to accelerate the computation. Our evaluative experiments on real-word data show the effectiveness of our approach in mining complex patterns with lower computational cost than the previous work.
Ryota Jinno, Kazuhiro Seki, Kuniaki Uehara
CloudCom2
2011 Application of Semantic Kernels to Literature-Based Gene Function Annotation
Mathieu Blondel, Kazuhiro Seki, Kuniaki Uehara
Discovery Science2
2011 Tackling class imbalance and data scarcity in literature-based gene function annotation
abstract
In recent years, a number of machine learning approaches to literature-based gene function annotation have been proposed. However, due to issues such as lack of labeled data, class imbalance and computational cost, they have usually been unable to surpass simpler approaches based on string-matching. In this paper, we propose a principled machine learning approach based on kernel classifiers.
Mathieu Blondel, Kazuhiro Seki, Kuniaki Uehara
SIGIR2
2011 Opinionated document retrieval using subjective triggers
abstract
This article proposes a novel application of a statistical language model to opinionated document retrieval targeting weblogs (blogs). In particular, we explore the use of the trigger model—originally developed for incorporating distant word dependencies—in order to model the characteristics of personal opinions that cannot be properly modeled by standard n-grams. Our primary assumption is that there are two constituents to form a subjective opinion. One is the subject of the opinion or the object that the opinion is about, and the other is a subjective expression; the former is regarded as a triggering word and the latter as a triggered word. We automatically identify those subjective trigger patterns to build a language model from a corpus of product customer reviews. Experimental results on the Text Retrieval Conference Blog track test collections show that, when used for reranking initial search results, our proposed model significantly improves opinionated document retrieval. In addition, we report on an experiment on dynamic adaptation of the model to a given query, which is found effective for most of the difficult queries categorized under politics and organizations. We also demonstrate that, without any modification to the proposed model itself, it can be effectively applied to polarized opinion retrieval.
Kazuhiro Seki, Kuniaki Uehara
J. Assoc. Inf. Sci. Technol.1
2010 Unsupervised Learning of Stroke Tagger for Online Kanji Handwriting Recognition
abstract
Traditionally, HMM-based approaches to online Kanji handwriting recognition have relied on a hand-made dictionary, mapping characters to primitives such as strokes or substrokes. We present an unsupervised way to learn a stroke tagger from data, which we eventually use to automatically generate such a dictionary. In addition to not requiring a prior hand-made dictionary, our approach can improve the recognition accuracy by exploiting unlabeled data when the amount of labeled data is limited.
Mathieu Blondel, Kazuhiro Seki, Kuniaki Uehara
ICPR2
2009 Gene Functional Annotation with Dynamic Hierarchical Classification Guided by Orthologs
Kazuhiro Seki, Yoshihiro Kino, Kuniaki Uehara
Discovery Science1
2009 Adaptive subjective triggers for opinionated document retrieval
abstract
This paper proposes a novel application of a statistical language model to opinionated document retrieval targeting weblogs (blogs). In particular, we explore the use of the trigger model---originally developed for incorporating distant word dependencies---in order to model the characteristics of personal opinions that cannot be properly modeled by standard n-grams. Our primary assumption is that there are two constituents to form a subjective opinion. One is the subject of the opinion or the object that the opinion is about, and the other is a subjective expression; the former is regarded as a triggering word and the latter as a triggered word. We automatically identify those subjective trigger patterns to build a language model from a corpus of product customer reviews. Experimental results on the TREC Blog Track test collections show that, when used for reranking initial search results, our proposed model significantly improves opinionated document retrieval by over 20% in MAP. In addition, we report on an experiment on dynamic adaptation of the model to a given query, which is found effective for most of difficult queries categorized under politics and organizations.
Kazuhiro Seki, Kuniaki Uehara
WSDM1
2008 Generating diverse katakana variants based on phonemic mapping
abstract
In Japanese, it is quite common for the same word to be written in several different ways. This is especially true for katakana words which are typically used for transliterating foreign languages. This ambiguity becomes critical for automatic processing such as information retrieval (IR). To tackle this problem, we propose a simple but effective approach to generating katakana variants by considering phonemic representation of the original language for a given word. The proposed approach is evaluated through an assessment of the variants it generates. Also, the impact of the generated variants on IR is studied in comparison to an existing approach using katakana rewriting rules.
Kazuhiro Seki, Hiroyuki Hattori, Kuniaki Uehara
SIGIR1
2008 Gene ontology annotation as text categorization: An empirical study
Kazuhiro Seki, Javed Mostafa
Inf. Process. Manag.1
2007 Literature-Based Discovery by an Enhanced Information Retrieval Model
Kazuhiro Seki, Javed Mostafa
Discovery Science1
2005 An application of text categorization methods to gene ontology annotation
abstract
This paper describes an application of IR and text categorization methods to a highly practical problem in biomedicine, specifically, Gene Ontology (GO) annotation. GO annotation is a major activity in most model organism database projects and annotates gene functions using a controlled vocabulary. As a first step toward automatic GO annotation, we aim to assign GO domain codes given a specific gene and an article in which the gene appears, which is one of the task challenges at the TREC 2004 Genomics Track. We approached the task with careful consideration of the specialized terminology and paid special attention to dealing with various forms of gene synonyms, so as to exhaustively locate the occurrences of the target gene. We extracted the words around the gene occurrences and used them to represent the gene for GO domain code annotation. As a classifier, we adopted a variant of k-Nearest Neighbor (kNN) with supervised term weighting schemes to improve the performance, making our method among the top-performing systems in the TREC official evaluation. Moreover, it is demonstrated that our proposed framework is successfully applied to another task of the Genomics Track, showing comparable results to the best performing system.
Kazuhiro Seki, Javed Mostafa
SIGIR1
2005 A hybrid approach to protein name identification in biomedical texts
Kazuhiro Seki, Javed Mostafa
Inf. Process. Manag.1
2002 A Probabilistic Method for Analyzing Japanese Anaphora Integrating Zero Pronoun Detection and Resolution
Kazuhiro Seki, Atsushi Fujii, Tetsuya Ishikawa
COLING1