VLDB 2026 Research / reviewers in the wild / expert
Hiromi Arai
dblp:85/7066
· DBLP profile ↗
15ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-5162-3812ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 80% Generative modeling · 12% Probabilistic and Bayesian machine learning · 4% | |
| Human-computer interaction and pervasive computing
3 papers |
Collaborative and social computing · 67% Human-AI interaction · 33% | |
| Network and information security
4 papers |
Web and mobile security · 56% Privacy and data protection · 44% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.1 | 3 | 2024 | Characterizing the risk of fairwashing · NeurIPS 2021 Fairwashing: the risk of rationalization · ICML 2019 Fair Machine Guidance to Enhance Fair Decision Making in Biased People · CHI 2024 |
Machine learning › Trustworthy machine learning › fairness
fairwashing |
0.9 | 2 | 2021 | Characterizing the risk of fairwashing · NeurIPS 2021 Fairwashing: the risk of rationalization · ICML 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 2 | 2021 | Characterizing the risk of fairwashing · NeurIPS 2021 Fairwashing: the risk of rationalization · ICML 2019 |
Collaborative and social computing › misinformation
misinformation intervention |
0.9 | 1 | 2025 | Beyond Click to Cognition: Effective Interventions for Promoting Examination of False Beliefs in Misinformation · CHI 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | Will Large-scale Generative Models Corrupt Future Datasets? · ICCV 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.7 | 1 | 2023 | Will Large-scale Generative Models Corrupt Future Datasets? · ICCV 2023 |
Collaborative and social computing
misinformation |
0.7 | 1 | 2023 | Who Does Not Benefit from Fact-checking Websites?: A Psychological Characteristic Predicts the Selective Avoidance of Clicking Uncongenial Facts · CHI 2023 |
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation |
0.5 | 1 | 2021 | Characterizing the risk of fairwashing · NeurIPS 2021 |
Web and mobile security
misinformation |
0.5 | 2 | 2025 | Beyond Click to Cognition: Effective Interventions for Promoting Examination of False Beliefs in Misinformation · CHI 2025 Who Does Not Benefit from Fact-checking Websites?: A Psychological Characteristic Predicts the Selective Avoidance of Clicking Uncongenial Facts · CHI 2023 |
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
model-agnostic explanation |
0.4 | 1 | 2019 | Fairwashing: the risk of rationalization · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › posterior inference
gibbs posterior |
0.2 | 1 | 2016 | Differential Privacy without Sensitivity · NIPS 2016 |
Privacy and data protection
differential privacy |
0.2 | 1 | 2016 | Differential Privacy without Sensitivity · NIPS 2016 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2023 | Will Large-scale Generative Models Corrupt Future Datasets? · ICCV 2023 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.1 | 1 | 2019 | Fairwashing: the risk of rationalization · ICML 2019 |
Privacy and data protection
privacy-preserving data analysis |
0.1 | 1 | 2010 | Online Prediction with Privacy · ICML 2010 |
Approximation and online algorithms
online algorithms |
0.1 | 1 | 2010 | Online Prediction with Privacy · ICML 2010 |
Approximation and online algorithms › online learning
online prediction |
0.1 | 1 | 2010 | Online Prediction with Privacy · ICML 2010 |
Bioinformatics and computational biology › structural biology
NMR spectroscopy |
0.1 | 1 | 2009 | A new modeling method in feature construction for the HSQC spectra screening problem · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
intervention design · 1.7fairness-aware machine learning · 1.5between-subjects experiment · 1.5psychological characteristic analysis · 1.3preregistered experiment · 0.7pre-registered experiment · 0.7large-scale generative models · 0.7fidelity-unfairness trade-off analysis · 0.5exponential mechanism · 0.5explanation manipulation · 0.5convex lipschitz loss · 0.5bayesian posterior · 0.5rule list enumeration · 0.4regularization · 0.4differential privacy · 0.2random coil peak model · 0.1machine learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Click to Cognition: Effective Interventions for Promoting Examination of False Beliefs in Misinformation
Yuko Tanaka, Hiromi Arai, Miwa Inuzuka, Yoichi Takahashi, Minao Kukita, Ryuta Iseki, Kentaro Inui |
CHI | 2 |
| 2024 | Fair Machine Guidance to Enhance Fair Decision Making in Biased PeopleabstractTeaching unbiased decision-making is crucial for addressing biased decision-making in daily life. Although both raising awareness of personal biases and providing guidance on unbiased decision-making are essential, the latter topics remains under-researched. In this study, we developed and evaluated an AI system aimed at educating individuals on making unbiased decisions using fairness-aware machine learning. In a between-subjects experimental design, 99 participants who were prone to bias performed personal assessment tasks. They were divided into two groups: a) those who received AI guidance for fair decision-making before the task and b) those who received no such guidance but were informed of their biases. The results suggest that although several participants doubted the fairness of the AI system, fair machine guidance prompted them to reassess their views regarding fairness, reflect on their biases, and modify their decision-making criteria. Our findings provide insights into the design of AI systems for guiding fair decision-making in humans. Mingzhe Yang, Hiromi Arai, Naomi Yamashita, Yukino Baba |
CHI | 2 |
| 2023 | Who Does Not Benefit from Fact-checking Websites?: A Psychological Characteristic Predicts the Selective Avoidance of Clicking Uncongenial FactsabstractFact-checking messages are shared or ignored subjectively. Users tend to seek like-minded information and ignore information that conflicts with their preexisting beliefs, leaving like-minded misinformation uncontrolled on the Internet. To understand the factors that distract fact-checking engagement, we investigated the psychological characteristics associated with users’ selective avoidance of clicking uncongenial facts. In a pre-registered experiment, we measured participants’ (N = 506) preexisting beliefs about COVID-19-related news stimuli. We then examined whether they clicked on fact-checking links to false news that they believed to be accurate. We proposed an index that divided participants into fact-avoidance and fact-exposure groups using a mathematical baseline. The results indicated that 43% of participants selectively avoided clicking on uncongenial facts, keeping 93% of their false beliefs intact. Reflexiveness is the psychological characteristic that predicts selective avoidance. We discuss susceptibility to click bias that prevents users from utilizing fact-checking websites and the implications for future design. Yuko Tanaka, Miwa Inuzuka, Hiromi Arai, Yoichi Takahashi, Minao Kukita, Kentaro Inui |
CHI | 3 |
| 2023 | Fight Bias with Bias? Two Interventions for Mitigating the Selective Avoidance of Clicking Uncongenial Facts
Yuko Tanaka, Hiromi Arai, Miwa Inuzuka, Yoichi Takahashi, Minao Kukita, Kentaro Inui |
CogSci | 2 |
| 2023 | Will Large-scale Generative Models Corrupt Future Datasets?abstractRecently proposed large-scale text-to-image generative models such as DALL•E 2 [47], Midjourney [42], and StableDiffusion [51] can generate high-quality and realistic images from users’ prompts. Not limited to the research community, ordinary Internet users enjoy these generative models, and consequently, a tremendous amount of generated images have been shared on the Internet. Meanwhile, today’s success of deep learning in the computer vision field owes a lot to images collected from the Internet. These trends lead us to a research question: "will such generated images impact the quality of future datasets and the performance of computer vision models positively or negatively?" This paper empirically answers this question by simulating contamination. Namely, we generate ImageNet-scale and COCO-scale datasets using a state-of-the-art generative model and evaluate models trained with "contaminated" datasets on various tasks, including image classification and image generation. Throughout experiments, we conclude that generated images negatively affect downstream performance, while the significance depends on tasks and the amount of generated images. The generated datasets and the codes for experiments will be publicly released for future research. Generated datasets and source codes are available from https://github.com/moskomule/dataset-contamination. Ryuichiro Hataya, Han Bao 0002, Hiromi Arai |
ICCV | 3 |
| 2023 | Designing a Location Trace Anonymization ContestabstractFor a better understanding of anonymization methods for location traces, we have designed and held a location trace anonymization contest that deals with a long trace (400 events per user) and fine-grained locations (1024 regions). In our contest, each team anonymizes her original traces, and then the other teams perform privacy attacks against the anonymized traces. In other words, both defense and attack compete together, which is close to what happens in real life. Prior to our contest, we show that re-identification alone is insufficient as a privacy risk and that trace inference should be added as an additional risk. Specifically, we show an example of anonymization that is perfectly secure against re-identification and is not secure against trace inference. Based on this, our contest evaluates both the re-identification risk and trace inference risk and analyzes their relationship. Through our contest, we show several findings in a situation where both defense and attack compete together. In particular, we show that an anonymization method secure against trace inference is also secure against re-identification under the presence of appropriate pseudonymization. We also report defense and attack algorithms that won first place, and analyze the utility of anonymized traces submitted by teams in various applications such as POI recommendation and geo-data analysis. Takao Murakami, Hiromi Arai, Koki Hamada, Takuma Hatano, Makoto Iguchi, Hiroaki Kikuchi, Atsushi Kuromasa, Hiroshi Nakagawa, Yuichi Nakamura 0004, Kenshiro Nishiyama, Ryo Nojima, Hidenobu Oguri, Chiemi Watanabe, Akira Yamada 0001, Takayasu Yamaguchi, Yuji Yamaoka |
Proc. Priv. Enhancing Technol. | 2 |
| 2021 | Characterizing the risk of fairwashingabstractFairwashing refers to the risk that an unfair black-box model can be explained by a fairer model through post-hoc explanation manipulation. In this paper, we investigate the capability of fairwashing attacks by analyzing their fidelity-unfairness trade-offs. In particular, we show that fairwashed explanation models can generalize beyond the suing group (i.e., data points that are being explained), meaning that a fairwashed explainer can be used to rationalize subsequent unfair decisions of a black-box model. We also demonstrate that fairwashing attacks can transfer across black-box models, meaning that other black-box models can perform fairwashing without explicitly using their predictions. This generalization and transferability of fairwashing attacks imply that their detection will be difficult in practice. Finally, we propose an approach to quantify the risk of fairwashing, which is based on the computation of the range of the unfairness of high-fidelity explainers. Ulrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi Hara 0001 |
NeurIPS | 2 |
| 2019 | Fairwashing: the risk of rationalizationabstractBlack-box explanation is the problem of explaining how a machine learning model – whose internal logic is hidden to the auditor and generally complex – produces its outcomes. Current approaches for solving this problem include model explanation, outcome explanation as well as model inspection. While these techniques can be beneficial by providing interpretability, they can be used in a negative manner to perform fairwashing, which we define as promoting the false perception that a machine learning model respects some ethical values. In particular, we demonstrate that it is possible to systematically rationalize decisions taken by an unfair black-box model using the model explanation as well as the outcome explanation approaches with a given fairness metric. Our solution, LaundryML, is based on a regularized rule list enumeration algorithm whose objective is to search for fair rule lists approximating an unfair black-box model. We empirically evaluate our rationalization technique on black-box models trained on real-world datasets and show that one can obtain rule lists with high fidelity to the black-box model while being considerably less unfair at the same time. Ulrich Aïvodji, Hiromi Arai, Olivier Fortineau, Sébastien Gambs, Satoshi Hara 0001, Alain Tapp |
ICML | 2 |
| 2016 | Differential Privacy without SensitivityabstractThe exponential mechanism is a general method to construct a randomized estimator that satisfies $(\varepsilon, 0)$-differential privacy. Recently, Wang et al. showed that the Gibbs posterior, which is a data-dependent probability distribution that contains the Bayesian posterior, is essentially equivalent to the exponential mechanism under certain boundedness conditions on the loss function. While the exponential mechanism provides a way to build an $(\varepsilon, 0)$-differential private algorithm, it requires boundedness of the loss function, which is quite stringent for some learning problems. In this paper, we focus on $(\varepsilon, \delta)$-differential privacy of Gibbs posteriors with convex and Lipschitz loss functions. Our result extends the classical exponential mechanism, allowing the loss functions to have an unbounded sensitivity. Kentaro Minami, Hiromi Arai, Issei Sato, Hiroshi Nakagawa |
NIPS | 2 |
| 2015 | Privacy-preserving search for chemical compound databasesabstractBACKGROUND: Searching for similar compounds in a database is the most important process for in-silico drug screening. Since a query compound is an important starting point for the new drug, a query holder, who is afraid of the query being monitored by the database server, usually downloads all the records in the database and uses them in a closed network. However, a serious dilemma arises when the database holder also wants to output no information except for the search results, and such a dilemma prevents the use of many important data resources. RESULTS: In order to overcome this dilemma, we developed a novel cryptographic protocol that enables database searching while keeping both the query holder's privacy and database holder's privacy. Generally, the application of cryptographic techniques to practical problems is difficult because versatile techniques are computationally expensive while computationally inexpensive techniques can perform only trivial computation tasks. In this study, our protocol is successfully built only from an additive-homomorphic cryptosystem, which allows only addition performed on encrypted values but is computationally efficient compared with versatile techniques such as general purpose multi-party computation. In an experiment searching ChEMBL, which consists of more than 1,200,000 compounds, the proposed method was 36,900 times faster in CPU time and 12,000 times as efficient in communication size compared with general purpose multi-party computation. CONCLUSION: We proposed a novel privacy-preserving protocol for searching chemical compound databases. The proposed method, easily scaling for large-scale databases, may help to accelerate drug discovery research by making full use of unused but valuable data that includes sensitive information. Kana Shimizu, Koji Nuida, Hiromi Arai, Shigeo Mitsunari, Nuttapong Attrapadung, Michiaki Hamada, Koji Tsuda, Takatsugu Hirokawa, Jun Sakuma, Goichiro Hanaoka, Kiyoshi Asai |
BMC Bioinform. | 3 |
| 2014 | Preserving worker privacy in crowdsourcing
Hiroshi Kajino, Hiromi Arai, Hisashi Kashima |
Data Min. Knowl. Discov. | 2 |
| 2011 | Privacy Preserving Semi-supervised Learning for Labeled Graphs
Hiromi Arai, Jun Sakuma |
ECML/PKDD (1) | 1 |
| 2010 | An Accurate Prediction Method for Protein Structural Class from Signal Patterns of NMR Spectra in the Absence of Chemical Shift AssignmentsabstractThe structural class information about a protein is important to understand its biological properties. NMR is one of the most powerful tools to obtain structural information of proteins in atomic resolution. However, an analysis of protein three-dimensional structure from NMR spectra usually requires laborious chemical shift assignment. We developed a new method for predicting the protein structural class directly from the NMR spectra without any chemical shift assignment. The results show that our method outperforms the methods using current secondary structure prediction. Hiromi Arai, Naoya Tochio, Tsuyoshi Kato, Takanori Kigawa, Masayuki Yamamura |
BIBE | 1 |
| 2010 | Online Prediction with Privacy
Jun Sakuma, Hiromi Arai |
ICML | 2 |
| 2009 | A new modeling method in feature construction for the HSQC spectra screening problemabstractMOTIVATION: Large-scale biological analyses produce huge amounts of data. As a consequence, automation in the data analysis process is needed. Sample screening problems in NMR high-throughput protein structure analysis are the typical examples. Especially, screening by protein (1)H-(15)N heteronuclear single quantum coherence (HSQC) spectra must be done quantitatively by a human expert. One popular solution for this problem is data mining. Machine learning methods can automatically extract rules and achieve high accuracy in prediction when a good quality training dataset is prepared. However, they tend to be a black box and the learned machines suffer the risk of overfitting to the dataset. RESULTS: We propose a model which evaluates HSQC spectra for feature construction. The model calculates similarity between the measured chemical shifts and those of a random coil peak model. We applied our feature construction method for the machine learning discrimination of folded protein HSQC spectra from unfolded ones, and compared our model-based features with those of conventional sequence-based features and image recognition features. The results revealed that our method has sufficient discrimination power and less overfits on training data, as compared to the other methods. In addition, our method succeeded reduction of input data complexity towards further investigation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hiromi Arai, Satoru Watanabe, Takanori Kigawa, Masayuki Yamamura |
Bioinform. | 1 |