VLDB 2026 Research / reviewers in the wild / expert
Kazuo Hara
dblp:86/5964
· DBLP profile ↗
20ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-4699-6136ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hubness Change Point DetectionabstractThis study proposes a new change detection method that leverages hubness. Hubness is a phenomenon that occurs in high-dimensional spaces, where certain special data points, known as hub data, tend to be closer to other data points. Hubness is known to degrade the accuracy of methods based on nearest neighbor search. Therefore, many studies in the past have focused on reducing hubness to improve accuracy. In contrast, this study utilizes hubness to detect changes. Specifically, if there is no change, suppressing the hubness occurring in the two datasets obtained by dividing the time series data will result in a uniform data distribution. However, if there is a change, even if we try to reduce the hubness in the two datasets obtained by dividing the time series data before and after the change, the hubness will not be reduced, and the data distribution will not become uniform. We use this finding to detect changes. Experiments with synthetic data show that the proposed method achieves accuracy comparable to or exceeding that of existing methods. Additionally, the proposed method achieves good accuracy with real-world data from hydraulic systems and gas sensors, along with excellent runtime performance. Ikumi Suzuki, Kazuo Hara, Eiji Murakami |
AAAI | 2 |
| 2021 | Finding Experiential Stress in Tweets by Utilizing both Explicit and Implicit Stress DataabstractStress has become a major social issue. In Japan, many people have been tweeting or confessing their complaints and worries on microblogging sites such as Twitter. However, it is difficult to estimate how many people are truly suffering from stress. This is because some people express their complaints and worries without using the word "stress." Moreover, even if the word "stress" is explicitly written, it does not necessarily represent the author's experience of stress, which we call experiential stress.In this study, we built classifiers that find not only the experiential stress appeared in tweets containing the word "stress," but also that appeared in all tweets related to stress. To build such classifiers, first, we collected two datasets from the web in two different ways. One is a dataset of tweets that explicitly contain the word "stress," and the other is a dataset by collecting tweets with the hashtag "#stress." The former one (we call explicit stress data) is easy to collect on a large scale, but the latter one (we call implicit stress data) is not. We then employed the pre-training and fine-tuning framework of BERT. We used the explicit stress data for pre-training and implicit stress data for fine-tuning. We showed that a classifier built in this way can find experiential stress with higher accuracy than a classifier built using only implicit stress data. Shizuku Iida, Kazuo Hara |
IEEE BigData | 2 |
| 2021 | Bayesian Optimization With an Auxiliary Classifier for the Development of Polymer MaterialsabstractRecently, Bayesian optimization has become commonly used in material development. However, in the development of polymer materials, a problem occurs wherein polymers do not always get formed. To address this, we incorporated an auxiliary classifier into the flow of Bayesian optimization. The results of a preliminary experiment show the potential of this approach. Tomoya Sasaki, Arisa Nakamura, Jun-Ichi Harasawa, Kazuo Hara, Ikumi Suzuki, Tatsuhiro Takahashi |
IEEE BigData | 4 |
| 2021 | Robust Method to Convert HIRAGANA Sequences into Japanese TextabstractWe apply an attention-based sequence-to-sequence model for the Japanese HIRAGANA-KANJI conversion task in spontaneous speech transcripts. Experimental results indicate that short HIRAGANA sequences containing speech-specific errors can be converted into error-free HIRAGANA-KANJI mixed Japanese text. Toshiki Yamaguchi, Kazuo Hara, Ikumi Suzuki |
IEEE BigData | 2 |
| 2021 | Impact of Duplicating Small Training Data on GANs
Yuki Eizuka, Kazuo Hara, Ikumi Suzuki |
DATA | 2 |
| 2021 | Semantic Entanglement on Verb Negation
Yuto Kikuchi, Kazuo Hara, Ikumi Suzuki |
DATA | 2 |
| 2017 | Centered kNN Graph for Semi-Supervised LearningabstractGraph construction is an important process in graph-based semi-supervised learning. Presently, the mutual kNN graph is the most preferred as it reduces hub nodes which can be a cause of failure during the process of label propagation. However, the mutual kNN graph, which is usually very sparse, suffers from over sparsification problem. That is, although the number of edges connecting nodes that have different labels decreases in the mutual kNN graph, the number of edges connecting nodes that have the same labels also reduces. In addition, over sparsification can produce a disconnected graph, which is not desirable for label propagation. So we present a new graph construction method, the centered kNN graph, which not only reduces hub nodes but also avoids the over sparsification problem. Ikumi Suzuki, Kazuo Hara |
SIGIR | 2 |
| 2016 | Flattening the Density Gradient for Eliminating Spatial Centrality to Reduce HubnessabstractSpatial centrality, whereby samples closer to the center of a dataset tend to be closer to all other samples, is regarded as one source of hubness. Hubness is well known to degrade k-nearest-neighbor (k-NN) classification. Spatial centrality can be removed by centering, i.e., shifting the origin to the global center of the dataset, in cases where inner product similarity is used. However, when Euclidean distance is used, centering has no effect on spatial centrality because the distance between the samples is the same before and after centering. As described in this paper, we propose a solution for the hubness problem when Euclidean distance is considered. We provide a theoretical explanation to demonstrate how the solution eliminates spatial centrality and reduces hubness. We then present some discussion of the reason the proposed solution works, from a viewpoint of density gradient, which is regarded as the origin of spatial centrality and hubness. We demonstrate that the solution corresponds to flattening the density gradient. Using real-world datasets, we demonstrate that the proposed method improves k-NN classification performance and outperforms an existing hub-reduction method. Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
AAAI | 1 |
| 2015 | Localized Centering: Reducing Hubness in Large-Sample DataabstractHubness has been recently identified as a problematic phenomenon occurring in high-dimensional space. In this paper, we address a different type of hubness that occurs when the number of samples is large. We investigate the difference between the hubness in high-dimensional data and the one in large-sample data. One finding is that centering, which is known to reduce the former, does not work for the latter. We then propose a new hub-reduction method, called localized centering. It is an extension of centering, yet works effectively for both types of hubness. Using real-world datasets consisting of a large number of documents, we demonstrate that the proposed method improves the accuracy of k-nearest neighbor classification. Kazuo Hara, Ikumi Suzuki, Masashi Shimbo, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
AAAI | 1 |
| 2015 | Ridge Regression, Hubness, and Zero-Shot Learning
Yutaro Shigeto, Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Yuji Matsumoto 0001 |
ECML/PKDD (1) | 3 |
| 2015 | Reducing Hubness: A Cause of Vulnerability in Recommender SystemsabstractIt is known that memory-based collaborative filtering systems are vulnerable to shilling attacks. In this paper, we demonstrate that hubness, which occurs in high dimensional data, is exploited by the attacks. Hence we explore methods for reducing hubness in user-response data to make these systems robust against attacks. Using the MovieLens dataset, we empirically show that the two methods for reducing hubness by transforming a similarity matrix(i) centering and (ii) conversion to a commute time kernel-can thwart attacks without degrading the recommendation performance. Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu |
SIGIR | 1 |
| 2015 | Reducing Hubness for Kernel Regression
Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
SISAP | 1 |
| 2013 | Centering Similarity Measures to Reduce HubsabstractThe performance of nearest neighbor methods is degraded by the presence of hubs, i.e., objects in the dataset that are similar to many other objects.In this paper, we show that the classical method of centering, the transformation that shifts the origin of the space to the data centroid, provides an effective way to reduce hubs.We show analytically why hubs emerge and why they are suppressed by centering, under a simple probabilistic model of data.To further reduce hubs, we also move the origin more aggressively towards hubs, through weighted centering.Our experimental results show that (weighted) centering is effective for natural language data; it improves the performance of the k-nearest neighbor classifiers considerably in word sense disambiguation and document classification tasks. Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Marco Saerens, Kenji Fukumizu |
EMNLP | 2 |
| 2012 | Investigating the Effectiveness of Laplacian-Based Kernels in Hub ReductionabstractA “hub” is an object closely surrounded by, or very similar to, many other objects in the dataset. Recent studies by Radovanovi´c et al. indicate that in high dimensional spaces, hubs almost always emerge, and objects close to the data centroid tend to become hubs. In this paper, we show that the family of kernels based on the graph Laplacian makes all objects in the dataset equally similar to the centroid, and thus they are expected to make less hubs when used as a similarity measure. We investigate this hypothesis using both synthetic and real-world data. It turns out that these kernels suppress hubs in some cases but not always, and the results seem to be affected by the size of the data—a factor not discussed previously. However, for the datasets in which hubs are indeed reduced by the Laplacian-based kernels, these kernels work well in ranking and classification tasks. This result suggests that the amount of hubs, which can be readily computed in an unsupervised fashion, can be a yardstick of whether Laplacian-based kernels work effectively for a given data. Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Yuji Matsumoto 0001, Marco Saerens |
AAAI | 2 |
| 2012 | Walk-based Computation of Contextual Word Similarity
Kazuo Hara, Ikumi Suzuki, Masashi Shimbo, Yuji Matsumoto 0001 |
COLING | 1 |
| 2011 | Mining personal experiences and opinions from Web documentsabstractThis paper proposes a new UGC-oriented language technology application, which we call experience mining. Experience mining aims at automatically collecting instances of personal experiences as well as opinions from vast amounts of user generated cont Shuya Abe, Kentaro Inui, Kazuo Hara, Hiraku Morita, Chitose Sao, Megumi Eguchi, Asuka Sumida, Koji Murakami, Suguru Matsuyoshi |
Web Intell. Agent Syst. | 3 |
| 2010 | Unsupervised WSD by Finding the Predominant Sense Using Context as a Dynamic Thesaurus
Javier Tejada-Cárcamo, Hiram Calvo, Alexander F. Gelbukh, Kazuo Hara |
J. Comput. Sci. Technol. | 4 |
| 2009 | Coordinate Structure Analysis with Global Structural Constraints and Alignment-Based Local Features
Kazuo Hara, Masashi Shimbo, Hideharu Okuma, Yuji Matsumoto 0001 |
ACL/IJCNLP | 1 |
| 2008 | Experience Mining: Building a Large-Scale Database of Personal Experiences and Opinions from Web DocumentsabstractThis paper proposes a new UGC-oriented language technology application, which we call experience mining. Experience mining aims at automatically collecting instances of personal experiences as well as opinions from an explosive number of user generated contents (UGCs) such as Weblog and forum posts and storing them in an experience database with semantically rich indices. After arguing the technical issues of this new task, we focus on the central problem, factuality analysis, among others and propose a machine learning-based solution as well as the task definition itself. Our empirical evaluation indicates that our factuality analysis task is sufficiently well-defined to achieve a high inter-annotator agreement and our factorial CRF-based model considerably outperforms the baseline. We also present an application system, which currently stores over 50M experience instances extracted from 150M Japanese blog posts with semantic indices and is scheduled to start serving as an experience search engine for unrestricted users in October. Kentaro Inui, Shuya Abe, Kazuo Hara, Hiraku Morita, Chitose Sao, Megumi Eguchi, Asuka Sumida, Koji Murakami, Suguru Matsuyoshi |
Web Intelligence | 3 |
| 2007 | A Discriminative Learning Model for Coordinate Conjunctions
Masashi Shimbo, Kazuo Hara |
EMNLP-CoNLL | 2 |