EDBT 2026 Demo / reviewers in the wild / expert
Kazuo Hara
dblp:86/5964
· DBLP profile ↗
10ranked-venue papers in the field
2as first author
5since 2021 · last 2021
0000-0002-4699-6136ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 3Information Retrieval & Web Search · 2 (1 first)Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Finding Experiential Stress in Tweets by Utilizing both Explicit and Implicit Stress DataabstractStress has become a major social issue. In Japan, many people have been tweeting or confessing their complaints and worries on microblogging sites such as Twitter. However, it is difficult to estimate how many people are truly suffering from stress. This is because some people express their complaints and worries without using the word "stress." Moreover, even if the word "stress" is explicitly written, it does not necessarily represent the author's experience of stress, which we call experiential stress.In this study, we built classifiers that find not only the experiential stress appeared in tweets containing the word "stress," but also that appeared in all tweets related to stress. To build such classifiers, first, we collected two datasets from the web in two different ways. One is a dataset of tweets that explicitly contain the word "stress," and the other is a dataset by collecting tweets with the hashtag "#stress." The former one (we call explicit stress data) is easy to collect on a large scale, but the latter one (we call implicit stress data) is not. We then employed the pre-training and fine-tuning framework of BERT. We used the explicit stress data for pre-training and implicit stress data for fine-tuning. We showed that a classifier built in this way can find experiential stress with higher accuracy than a classifier built using only implicit stress data. Shizuku Iida, Kazuo Hara |
IEEE BigData | 2 |
| 2021 | Bayesian Optimization With an Auxiliary Classifier for the Development of Polymer MaterialsabstractRecently, Bayesian optimization has become commonly used in material development. However, in the development of polymer materials, a problem occurs wherein polymers do not always get formed. To address this, we incorporated an auxiliary classifier into the flow of Bayesian optimization. The results of a preliminary experiment show the potential of this approach. Tomoya Sasaki, Arisa Nakamura, Jun-Ichi Harasawa, Kazuo Hara, Ikumi Suzuki, Tatsuhiro Takahashi |
IEEE BigData | 4 |
| 2021 | Robust Method to Convert HIRAGANA Sequences into Japanese TextabstractWe apply an attention-based sequence-to-sequence model for the Japanese HIRAGANA-KANJI conversion task in spontaneous speech transcripts. Experimental results indicate that short HIRAGANA sequences containing speech-specific errors can be converted into error-free HIRAGANA-KANJI mixed Japanese text. Toshiki Yamaguchi, Kazuo Hara, Ikumi Suzuki |
IEEE BigData | 2 |
| 2021 | Impact of Duplicating Small Training Data on GANs
Yuki Eizuka, Kazuo Hara, Ikumi Suzuki |
DATA | 2 |
| 2021 | Semantic Entanglement on Verb Negation
Yuto Kikuchi, Kazuo Hara, Ikumi Suzuki |
DATA | 2 |
| 2017 | Centered kNN Graph for Semi-Supervised LearningabstractGraph construction is an important process in graph-based semi-supervised learning. Presently, the mutual kNN graph is the most preferred as it reduces hub nodes which can be a cause of failure during the process of label propagation. However, the mutual kNN graph, which is usually very sparse, suffers from over sparsification problem. That is, although the number of edges connecting nodes that have different labels decreases in the mutual kNN graph, the number of edges connecting nodes that have the same labels also reduces. In addition, over sparsification can produce a disconnected graph, which is not desirable for label propagation. So we present a new graph construction method, the centered kNN graph, which not only reduces hub nodes but also avoids the over sparsification problem. Ikumi Suzuki, Kazuo Hara |
SIGIR | 2 |
| 2015 | Ridge Regression, Hubness, and Zero-Shot Learning
Yutaro Shigeto, Ikumi Suzuki, Kazuo Hara, Masashi Shimbo, Yuji Matsumoto 0001 |
ECML/PKDD (1) | 3 |
| 2015 | Reducing Hubness: A Cause of Vulnerability in Recommender SystemsabstractIt is known that memory-based collaborative filtering systems are vulnerable to shilling attacks. In this paper, we demonstrate that hubness, which occurs in high dimensional data, is exploited by the attacks. Hence we explore methods for reducing hubness in user-response data to make these systems robust against attacks. Using the MovieLens dataset, we empirically show that the two methods for reducing hubness by transforming a similarity matrix(i) centering and (ii) conversion to a commute time kernel-can thwart attacks without degrading the recommendation performance. Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu |
SIGIR | 1 |
| 2015 | Reducing Hubness for Kernel Regression
Kazuo Hara, Ikumi Suzuki, Kei Kobayashi, Kenji Fukumizu, Milos Radovanovic 0001 |
SISAP | 1 |
| 2008 | Experience Mining: Building a Large-Scale Database of Personal Experiences and Opinions from Web DocumentsabstractThis paper proposes a new UGC-oriented language technology application, which we call experience mining. Experience mining aims at automatically collecting instances of personal experiences as well as opinions from an explosive number of user generated contents (UGCs) such as Weblog and forum posts and storing them in an experience database with semantically rich indices. After arguing the technical issues of this new task, we focus on the central problem, factuality analysis, among others and propose a machine learning-based solution as well as the task definition itself. Our empirical evaluation indicates that our factuality analysis task is sufficiently well-defined to achieve a high inter-annotator agreement and our factorial CRF-based model considerably outperforms the baseline. We also present an application system, which currently stores over 50M experience instances extracted from 150M Japanese blog posts with semantic indices and is scheduled to start serving as an experience search engine for unrestricted users in October. Kentaro Inui, Shuya Abe, Kazuo Hara, Hiraku Morita, Chitose Sao, Megumi Eguchi, Asuka Sumida, Koji Murakami, Suguru Matsuyoshi |
Web Intelligence | 3 |