Yukino Baba

dblp:41/8393 · DBLP profile ↗
← Back
53ranked-venue papers
11as first author
10since 2021 · last 2026
0000-0001-5310-9841ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 9 first-author · 3 since 2021Databases, data management, data science and information retrieval · 18 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-author
YearPublicationVenuePosition
2026 More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
abstract
Just as people improve decision-making by consulting diverse human advisors, they can now also consult with multiple AI systems. Prior work on group decision-making shows that advice aggregation creates pressure to conform, leading to overreliance. However, the conditions under which multi-AI consultation improves or undermines human decision-making remain unclear. We conducted experiments with three tasks in which participants received advice from panels of AIs. We varied panel size, within-panel consensus, and the human-likeness of presentation. Accuracy improved for small panels relative to a single AI; larger panels yielded no gains. The level of within-panel consensus affected participants’ reliance on AI advice: High consensus fostered overreliance; a single dissent reduced pressure to conform; wide disagreement created confusion and undermined appropriate reliance. Human-like presentations increased perceived usefulness and agency in certain tasks, without raising conformity pressure. These findings yield design implications for presenting multi-AI advice that preserve accuracy while mitigating conformity.
Yuta Tsuchiya, Yukino Baba
CHI2
2025 Do Expressions Change Decisions? Exploring the Impact of AI's Explanation Tone on Decision-Making
Ayano Okoso, Mingzhe Yang, Yukino Baba
CHI3
2025 Self Iterative Label Refinement via Robust Unlabeled Learning
abstract
Recent advances in large language models (LLMs) have yielded impressive performance on various tasks, yet they often depend on high-quality feedback that can be costly. Self-refinement methods attempt to leverage LLMs' internal evaluation mechanisms with minimal human supervision; however, these approaches frequently suffer from inherent biases and overconfidence, especially in domains where the models lack sufficient internal knowledge, resulting in performance degradation. As an initial step toward enhancing self-refinement for broader applications, we introduce an iterative refinement pipeline that employs the Unlabeled-Unlabeled learning framework to improve LLM-generated pseudo-labels for classification tasks. By exploiting two unlabeled datasets with differing positive class ratios, our approach iteratively denoises and refines the initial pseudo-labels, thereby mitigating the adverse effects of internal biases with minimal human supervision. Evaluations on diverse datasets, including low-resource language corpora, patent classifications, and protein structure categorizations, demonstrate that our method consistently outperforms both initial LLM's classification performance and the self-refinement approaches by cutting-edge models (e.g., GPT-4o and DeepSeek-R1). Moreover, we experimentally confirm that our refined classifier facilitates effective post-training alignment for safety in LLMs and demonstrate successful self-refinement in generative tasks as well. Our code is available at https://github.com/HikaruAsano/self-iterative-label-refinement.
Hikaru Asano, Tadashi Kozuno, Yukino Baba
NeurIPS3
2025 Travel itinerary recommendation using interaction-based augmented data
abstract
Itinerary planning is complicated for travelers because the traveling content, including places to visit, acceptable times, and distances, can be diverse. Travel recommender systems (TRSs) recommend the most relevant itineraries for a traveler. In this paper, we propose an interactive framework that allows users to edit itineraries directly on a map to suit their preferences better. The proposed framework collects feedback data by recording user itinerary modifications, infers positive and negative preferences, and fine-tunes the recommender models with our ranking-based loss function. This interaction-based data augmentation approach addresses data sparsity issues due to personalization by capturing a variety of travel item combinations. In our experiments, we evaluate multiple combinations of models and itinerary generation methods to show the effectiveness of integrating interaction data into TRSs. Our experimental evaluations demonstrate that our interactive TRS can provide itineraries that align with users’ preferences more in terms of point-set-wise and rank-wise accuracy; the integration consistently improves the accuracy for all combinations of the components, and particularly, the improvement is large for small backbone models.
Keisuke Otaki, Yukino Baba
Expert Syst. Appl.2
2025 Impact of Tone-Aware Explanations in Recommender Systems
abstract
In recommender systems, explanations are essential for supporting users’ decision-making processes. While many studies have focused on explanation content or user interface, the expression of textual explanations has been largely overlooked. The expression refers to textual styles such as formal or humorous, which we call tone in this article. Although tone contributes to smooth human communication, its impact on users’ perceptions of recommender systems remains largely unexplored. In particular, it is unclear whether the perceived effects of explanation tone differ by domain or user attributes. Therefore, we investigate the effects of explanation tones through two online user studies considering domains and user attributes. In the first study with 470 participants, we generated datasets using a large language model to create fictional items and explanations with six tones across three domains: movies, hotels, and home products. The participants evaluated two explanations for an item, each presented in a different tone, and rated 10 metrics. In the second study with 103 participants, we used a real-world dataset from the hotel domain and incorporated a simple personalized recommender system to examine effects of tone in a more realistic setting. The results revealed that the perceived effects of tones differ by domain and are significantly influenced by user attributes such as age and personality traits. Our findings suggest that appropriately adjusting the tone of explanations according to domains and user attributes can enhance the perceived effects of recommender systems.
Ayano Okoso, Keisuke Otaki, Satoshi Koide, Yukino Baba
Trans. Recomm. Syst.4
2024 Fair Machine Guidance to Enhance Fair Decision Making in Biased People
abstract
Teaching unbiased decision-making is crucial for addressing biased decision-making in daily life. Although both raising awareness of personal biases and providing guidance on unbiased decision-making are essential, the latter topics remains under-researched. In this study, we developed and evaluated an AI system aimed at educating individuals on making unbiased decisions using fairness-aware machine learning. In a between-subjects experimental design, 99 participants who were prone to bias performed personal assessment tasks. They were divided into two groups: a) those who received AI guidance for fair decision-making before the task and b) those who received no such guidance but were informed of their biases. The results suggest that although several participants doubted the fairness of the AI system, fair machine guidance prompted them to reassess their views regarding fairness, reflect on their biases, and modify their decision-making criteria. Our findings provide insights into the design of AI systems for guiding fair decision-making in humans.
Mingzhe Yang, Hiromi Arai, Naomi Yamashita, Yukino Baba
CHI4
2024 SwipeGANSpace: Swipe-to-Compare Image Generation via Efficient Latent Space Exploration
abstract
Generating preferred images using generative adversarial networks (GANs) is challenging owing to the high-dimensional nature of latent space. In this study, we propose a novel approach that uses simple user-swipe interactions to generate preferred images for users. To effectively explore the latent space with only swipe interactions, we apply principal component analysis to the latent space of the StyleGAN, creating meaningful subspaces. We use a multi-armed bandit algorithm to decide the dimensions to explore, focusing on the preferences of the user. Experiments show that our method is more efficient in generating preferred images than the baseline methods. Furthermore, changes in preferred images during image generation or the display of entirely different image styles were observed to provide new inspirations, subsequently altering user preferences. This highlights the dynamic nature of user preferences, which our proposed approach recognizes and enhances.
Yuto Nakashima 0002, Mingzhe Yang, Yukino Baba
IUI3
2024 Toward Tone-Aware Explanations in Recommender Systems
abstract
In recommender systems, the presentation of explanations plays a crucial role in supporting users’ decision-making processes. Although numerous existing studies have focused on the effects (e.g., transparency) of explanation content, explanation expression is largely overlooked. Tone, such as formal and humorous, is directly linked to expressiveness and is an important element in human communication. However, studies on the impact of tone on explanations within the context of recommender systems are insufficient. Therefore, this study investigates the tonal effects of explanations through an online user study. We focus on a hotel domain and six types of tones. The collected data analysis reveals that the tone of explanations influences the perceived effects, such as trust and effectiveness, of recommender systems. Our findings suggest that the tone of explanations can enhance user experience in recommender systems.
Ayano Okoso, Keisuke Otaki, Satoshi Koide, Yukino Baba
UMAP4
2021 A practical and universal framework for generating publicly available medical notes of authentic quality via the power of crowds
abstract
Medical notes written by doctors in hospitals or clinics are information-rich. However, in many countries or cultures, few people have access to them for educational and research purposes, even once anonymized. This is because their contents, including patients’ disease information, are sensitive and require confidentiality. Therefore, publicly available pseudo-medical notes are needed. Authentic pseudo-medical notes must meet two requirements: (1) medical consistency, and (2) informal descriptions and specific sub-language; however, these are empirical knowledge, even for medical doctors, and are not clarified specifically. We combat this by harnessing the power of crowds. We propose a human-in-the-loop framework for generating publicly available professional medical notes utilizing human cognitive traits with a small dataset. The practical and universal framework has three steps. In Step 1, crowd workers imitated actual notes. In Step 2, crowds and algorithms collaboratively identified notes’ characteristics based on comparisons between actual and dummy notes. In Step 3, the texts generated in Step 1 that exhibited the characteristics from Step 2 were evaluated as authentic medical notes that met all requirements. We demonstrated this framework with a total of 1,662 crowds’ power. All data were preprocessed to protect patients’ privacy before the experiments. The crowds’ generated 9,756 notes were evaluated as the most realistic compared to dummy medical notes written by doctors. These crowd-generated medical notes, which are the largest publicly available dataset of Japanese medical notes, are published. This study was the first challenge for the crowds to solve the medical expert-level task.
Rina Kagawa, Yukino Baba, Hideo Tsurushima
IEEE BigData2
2021 Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception
abstract
We propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in which humans accept the naturalness regardless of whether the data are real or not. A Human-GAN was proposed to model the human-acceptable distribution. A DNN-based generator is trained using a human-based discriminator, i.e., humans’ perceptual evaluations, instead of the GAN’s DNN-based discriminator. However, the HumanGAN cannot represent conditional distributions. This paper proposes the HumanACGAN, a theoretical extension of the HumanGAN, to deal with conditional human-acceptable distributions. Our HumanACGAN trains a DNN-based conditional generator by regarding humans as not only a discriminator but also an auxiliary classifier. The generator is trained by deceiving the human-based discriminator that scores the unconditioned naturalness and the human-based classifier that scores the class-conditioned perceptual acceptability. The training can be executed using the backpropagation algorithm involving humans’ perceptual evaluations. Our experimental results in phoneme perception demonstrate that our HumanACGAN can successfully train this conditional generator.
Yota Ueda, Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP5
2020 Stress Prediction from Head Motion
abstract
The measurement of cognitive stress has huge potential for advertising optimization (e.g., neuromarketing), optimization of recommendation systems, and applications in the fields of human-computer interaction and affective computing. Many studies have addressed stress prediction based on machine learning from the features measured by sensors attached to a subject's body. Meanwhile, as virtual reality (VR) and augmented reality (AR) have increased in popularity, head motion data from users watching VR/AR contents have become ubiquitous. In addition, stress prediction from head motion can be valuable because it does not rely on skin condition (e.g., sweat, tattoo, and cosmetics) and is not detrimental to usability. However, the effectiveness of stress prediction based on head motion data is not well understood. In this study, we propose a method to predict stress from head motion and verify the performance of this method in multiple test situations.
Hitoshi Kusano, Yuji Horiguchi, Yukino Baba, Hisashi Kashima
DSAA3
2020 CrowDEA: Multi-View Idea Prioritization with Crowds
abstract
Given a set of ideas collected from crowds with regard to an open-ended question, how can we organize and prioritize them in order to determine the preferred ones based on preference comparisons by crowd evaluators? As there are diverse latent criteria for the value of an idea, multiple ideas can be considered as “the best”. In addition, evaluators can have different preference criteria, and their comparison results often disagree. In this paper, we propose an analysis method for obtaining a subset of ideas, which we call frontier ideas, that are the best in terms of at least one latent evaluation criterion. We propose an approach, called CrowDEA, which estimates the embeddings of the ideas in the multiple-criteria preference space, the best viewpoint for each idea, and preference criterion for each evaluator, to obtain a set of frontier ideas. Experimental results using real datasets containing numerous ideas or designs demonstrate that the proposed approach can effectively prioritize ideas from multiple viewpoints, thereby detecting frontier ideas. The embeddings of ideas learned by the proposed approach provide a visualization that facilitates observation of the frontier ideas. In addition, the proposed approach prioritizes ideas from a wider variety of viewpoints, whereas the baselines tend to use to the same viewpoints; it can also handle various viewpoints and prioritize ideas in situations where only a limited number of evaluators or labels are available.
Yukino Baba, Jiyi Li, Hisashi Kashima
HCOMP1
2020 Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling
abstract
We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent the outside of a real-data distribution. In the case of speech perception, humans can recognize not only human voices but also processed (i.e., a non-existent human) voices as human voice. Such a human-acceptable distribution is typically wider than a real-data one and cannot be modeled by the basic GAN. To model the human-acceptable distribution, we formulate a backpropagation-based generator training algorithm by regarding human perception as a black-boxed discriminator. The training efficiently iterates generator training by using a computer and discrimination by human. We evaluate our HumanGAN in speech naturalness modeling and demonstrate that it can represent a human-acceptable distribution that is wider than a real-data distribution.
Kazuki Fujii, Yuki Saito 0001, Shinnosuke Takamichi, Yukino Baba, Hiroshi Saruwatari
ICASSP4
2020 Performance as a Constraint: An Improved Wisdom of Crowds Using Performance Regularization
abstract
Quality assurance is one of the most important problems in crowdsourcing and human computation, and it has been extensively studied from various aspects. Typical approaches for quality assurance include unsupervised approaches such as introducing task redundancy (i.e., asking the same question to multiple workers and aggregating their answers) and supervised approaches such as using worker performance on past tasks or injecting qualification questions into tasks in order to estimate the worker performance. In this paper, we propose to utilize the worker performance as a global constraint for inferring the true answers. The existing semi-supervised approaches do not consider such use of qualification questions. We also propose to utilize the constraint as a regularizer combined with existing statistical aggregation methods. The experiments using heterogeneous multiple-choice questions demonstrate that the performance constraint not only has the power to estimate the ground truths when used by itself, but also boosts the existing aggregation methods when used as a regularizer.
Jiyi Li, Yasushi Kawase, Yukino Baba, Hisashi Kashima
IJCAI3
2020 Dual graph convolutional neural network for predicting chemical networks
abstract
BACKGROUND: Predicting of chemical compounds is one of the fundamental tasks in bioinformatics and chemoinformatics, because it contributes to various applications in metabolic engineering and drug discovery. The recent rapid growth of the amount of available data has enabled applications of computational approaches such as statistical modeling and machine learning method. Both a set of chemical interactions and chemical compound structures are represented as graphs, and various graph-based approaches including graph convolutional neural networks have been successfully applied to chemical network prediction. However, there was no efficient method that can consider the two different types of graphs in an end-to-end manner. RESULTS: We give a new formulation of the chemical network prediction problem as a link prediction problem in a graph of graphs (GoG) which can represent the hierarchical structure consisting of compound graphs and an inter-compound graph. We propose a new graph convolutional neural network architecture called dual graph convolutional network that learns compound representations from both the compound graphs and the inter-compound network in an end-to-end manner. CONCLUSIONS: Experiments using four chemical networks with different sparsity levels and degree distributions shows that our dual graph convolution approach achieves high prediction performance in relatively dense networks, while the performance becomes inferior on extremely-sparse networks.
Shonosuke Harada, Hirotaka Akita, Masashi Tsubaki, Yukino Baba, Ichigaku Takigawa, Yoshihiro Yamanishi, Hisashi Kashima
BMC Bioinform.4
2020 Synthetic accessibility assessment using auxiliary responses
Shun Ito, Yukino Baba, Tetsu Isomura, Hisashi Kashima
Expert Syst. Appl.2
2019 Active Learning Strategies for Hierarchical Labeling Microtasks
abstract
This paper reports the result of a preliminary experiment on active learning strategies for the hierarchical labeling microtasks. A typical example of hierarchical labeling microtask consists of a set of labeling tasks for partitions of a large image; starting from the whole image, the workers choose to give a label or divide it into smaller ones. This paper shows the result of an experiment to compare several strategies for active learning in the setting. The result suggests that the difference in the strategies affects the performance in the early stage.
Kousuke Uo, Masaki Kobayashi, Masaki Matsubara, Yukino Baba, Atsuyuki Morishima
IEEE BigData4
2019 Probabilistic Modeling of Peer Correction and Peer Assessment
Takeru Sunahase, Yukino Baba, Hisashi Kashima
EDM2
2019 Interdependence Model for Multi-label Classification
Kosuke Yoshimura, Tomoaki Iwase, Yukino Baba, Hisashi Kashima
ICANN (4)3
2019 Crownn: Human-in-the-loop Network with Crowd-generated Inputs
abstract
Input features are indispensable for almost all machine learning methods; however, their definitions themselves are sometimes too abstract to extract automatically. Human-in-the-loop machine learning is a promising solution to such cases where humans extract the feature values for machine learning models. We use crowdsourcing for feature value extraction and consider a problem to aggregate the feature values to improve machine learning classifiers. We propose a novel neural network model called CROWNN, a neural network with crowd-generated inputs with the worker convolution layer, that learns both the capabilities of human feature extractors and the weights of a neural network classifier by applying the idea of the convolution neural network to feature aggregation. Our experiments using four datasets show the proposed method outperforms the baseline method using unsupervised aggregation methods in some datasets. We also show the robustness of the proposed model against the existence of spam workers, especially when they are malicious workers who intentionally flip the feature values.
Yusuke Sakata, Yukino Baba, Hisashi Kashima
ICASSP2
2019 Large-scale Driver Identification Using Automobile Driving Data
abstract
We address a large-scale driver identification problem, which aims to predict the driver of a vehicle from various types of data, such as speed and acceleration information, that are collected during driving by using GPS sensors equipped with smart phones. While existing studies consider at most a few hundreds of drivers, we target a huge number of drivers up to 10,000 drivers. The results of our experiments show that our method identifies drivers more precisely than baseline methods. We also show that location features are quite effective in the large scale driver identification, and speed and acceleration features also contribute to driver identification.
Daiki Tanaka, Yukino Baba, Hisashi Kashima, Yuta Okubo
SMC2
2018 Data Analysis Competition Platform for Educational Purposes: Lessons Learned and Future Challenges
abstract
Data analysis education plays an important role in accelerating the efficient use of data analysis technologies in various domains. Not only the knowledge of statistics and machine learning, but also practical skills of deploying machine learning and data analysis techniques, are required for conducting data analysis projects in the real world. Data analysis competitions, such as Kaggle, have been considered as an efficient system for learning such skills by addressing real data analysis problems. However, current data analysis competitions are not designed for educational purposes and it is not well studied how data analysis competition platforms should be designed for enhancing educational effectiveness. To answer this research question, we built, and subsequently operated an educational data analysis competition platform called University of Big Data for several years. In this paper, we present our approaches for supporting and motivating learners and the results of our case studies. We found that providing a tutorial article is beneficial for encouraging active participation of learners, and a leaderboard system allowing an unlimited number of submissions can motivate the efforts of learners. We further discuss future directions of educational data analysis competitions.
Yukino Baba, Tomoumi Takase, Kyohei Atarashi, Satoshi Oyama, Hisashi Kashima
AAAI1
2018 Predictive Modeling of Learning Continuation in Preschool Education Using Temporal Patterns of Development Tests
abstract
Learning analytics applies data analysis techniques to learning data in order to support students’ learning processes and to improve the quality of education. Despite the increasing attention to learning analytics for higher education, it has not been fully addressed in primary and preschool education. In this research, we apply learning analytics to preschool education to predict the continuation of learning of preschool children. Based on our hypothesis that temporal patterns in the assessment scores of development tests are effective features for prediction, we extract the temporal patterns using time-series clustering, and use them as the features of prediction models. The experimental results using a real preschool education dataset show that the use of the temporal patterns improves the predictive accuracy of future continuation of study.
Junpei Naito, Yukino Baba, Hisashi Kashima, Takenori Takaki, Takuya Funo
AAAI2
2018 AdaFlock: Adaptive Feature Discovery for Human-in-the-loop Predictive Modeling
abstract
Feature engineering is the key to successful application of machine learning algorithms to real-world data. The discovery of informative features often requires domain knowledge or human inspiration, and data scientists expend a certain amount of effort into exploring feature spaces. Crowdsourcing is considered a promising approach for allowing many people to be involved in feature engineering; however, there is a demand for a sophisticated strategy that enables us to acquire good features at a reasonable crowdsourcing cost. In this paper, we present a novel algorithm called AdaFlock to efficiently obtain informative features through crowdsourcing. AdaFlock is inspired by AdaBoost, which iteratively trains classifiers by increasing the weights of samples misclassified by previous classifiers. AdaFlock iteratively generates informative features; at each iteration of AdaFlock, crowdsourcing workers are shown samples selected according to the classification errors of the current classifiers and are asked to generate new features that are helpful for correctly classifying the given examples. The results of our experiments conducted using real datasets indicate that AdaFlock successfully discovers informative features with fewer iterations and achieves high classification accuracy.
Ryusuke Takahama, Yukino Baba, Nobuyuki Shimizu, Sumio Fujita, Hisashi Kashima
AAAI2
2018 Incorporating Worker Similarity for Label Aggregation in Crowdsourcing
Jiyi Li, Yukino Baba, Hisashi Kashima
ICANN (2)2
2018 BayesGrad: Explaining Predictions of Graph Convolutional Networks
Hirotaka Akita, Kosuke Nakago, Tomoki Komatsu, Yohei Sugawara, Shin-ichi Maeda, Yukino Baba, Hisashi Kashima
ICONIP (5)6
2018 Statistical Quality Control for Human Computation and Crowdsourcing
abstract
Human computation is a method for solving difficult problems by combining humans and computers. Quality control is a critical issue in human computation because it relies on a large number of participants (i.e., crowds) and there is an uncertainty about their reliability. A solution for this issue is to leverage the power of the "wisdom of crowds"; for example, we can aggregate the outputs of multiple participants or ask a participant to check the output of another participant to improve its quality. In this paper, we review several statistical approaches for controlling the quality of outputs from crowds.
Yukino Baba
IJCAI1
2018 Simultaneous Clustering and Ranking from Pairwise Comparisons
abstract
When people make decisions with a number of ideas, designs, or other kinds of objects, one attempt is probably to organize them into several groups of objects and to prioritize them according to some preference. The grouping task is referred to as clustering and the prioritizing task is called as ranking. These tasks are often outsourced with the help of human judgments in the form of pairwise comparisons. Two objects are compared on whether they are similar in the clustering problem, while the object of higher priority is determined in the ranking problem. Our research question in this paper is whether the pairwise comparisons for clustering also help ranking (and vice versa). Instead of solving the two tasks separately, we propose a unified formulation to bridge the two types of pairwise comparisons. Our formulation simultaneously estimates the object embeddings and the preference criterion vector. The experiments using real datasets support our hypothesis; our approach can generate better neighbor and preference estimation results than the approaches that only focus on a single type of pairwise comparisons.
Jiyi Li, Yukino Baba, Hisashi Kashima
IJCAI2
2017 Predicting Fuel Consumption and Flight Delays for Low-Cost Airlines
Yuji Horiguchi, Yukino Baba, Hisashi Kashima, Masahito Suzuki, Hiroki Kayahara, Jun Maeno
AAAI2
2017 Pairwise HITS: Quality Estimation from Pairwise Comparisons in Creator-Evaluator Crowdsourcing Process
abstract
A common technique for improving the quality of crowdsourcing results is to assign a same task to multiple workers redundantly, and then to aggregate the results to obtain a higher-quality result; however, this technique is not applicable to complex tasks such as article writing since there is no obvious way to aggregate the results. Instead, we can use a two-stage procedure consisting of a creation stage and an evaluation stage, where we first ask workers to create artifacts, and then ask other workers to evaluate the artifacts to estimate their quality. In this study, we propose a novel quality estimation method for the two-stage procedure where pairwise comparison results for pairs of artifacts are collected at the evaluation stage. Our method is based on an extension of Kleinberg's HITS algorithm to pairwise comparison, which takes into account the ability of evaluators as well as the ability of creators. Experiments using actual crowdsourcing tasks show that our methods outperform baseline methods especially when the number of evaluators per artifact is small.
Takeru Sunahase, Yukino Baba, Hisashi Kashima
AAAI2
2017 Hyper Questions: Unsupervised Targeting of a Few Experts in Crowdsourcing
abstract
Quality control is one of the major problems in crowdsourcing. One of the primary approaches to rectify this issue is to assign the same task to different workers and then aggregate their answers to obtain a reliable answer. In addition to simple aggregation approaches such as majority voting, various sophisticated probabilistic models have been proposed. However, given that most of the existing methods operate by strengthening the opinions of the majority, these models often fail when the tasks require highly specialized knowledge and the ability of a large majority of the workers is inadequate. In this paper, we focus on an important class of answer aggregation problems in which majority voting fails and propose the concept of hyper questions to devise effective aggregation methods. A hyper question is a set of single questions, and our key idea is that experts are more likely to provide correct answers to all of the single questions included in a hyper question than non-experts. Thus, experts are more likely to reach consensus on the hyper questions than non-experts, which strengthen their influences. We incorporate the concept of hyper questions into existing answer aggregation methods. The results of our experiments conducted using both synthetic datasets and real datasets demonstrate that our simple and easily usable approach works effectively in cases where only a few experts are available.
Jiyi Li, Yukino Baba, Hisashi Kashima
CIKM2
2017 Atomic Distance Kernel for Material Property Prediction
Hirotaka Akita, Yukino Baba, Hisashi Kashima, Atsuto Seko
ICONIP (1)2
2017 Quality Control for Crowdsourced Multi-label Classification Using RAkEL
Kosuke Yoshimura, Yukino Baba, Hisashi Kashima
ICONIP (1)2
2017 A Generalized Model for Multidimensional Intransitivity
Jiuding Duan, Jiyi Li, Yukino Baba, Hisashi Kashima
PAKDD (2)3
2017 Distributed Multi-task Learning for Sensor Network
Jiyi Li, Tomohiro Arai, Yukino Baba, Hisashi Kashima, Shotaro Miwa
ECML/PKDD (2)3
2016 Learning to Enumerate
Patrick Jörger, Yukino Baba, Hisashi Kashima
ICANN (1)2
2016 Assessing Translation Ability through Vocabulary Ability Assessment
Yo Ehara, Yukino Baba, Masao Utiyama, Eiichiro Sumita
IJCAI2
2016 Participation recommendation system for crowdsourcing contests
Yukino Baba, Kei Kinoshita, Hisashi Kashima
Expert Syst. Appl.1
2016 Quality control of crowdsourced classification using hierarchical class structures
Naoki Otani, Yukino Baba, Hisashi Kashima
Expert Syst. Appl.2
2015 From one star to three stars: Upgrading legacy open data using crowdsourcing
abstract
Despite recent open data initiatives in many countries, a significant percentage of the data provided is in non-machine-readable formats like image format rather than in a machine-readable electronic format, thereby restricting their usability. This paper describes the first unified framework for converting legacy open data in image format into a machine-readable and reusable format by using crowdsourcing. Crowd workers are asked not only to extract data from an image of a chart but also to reproduce the chart objects in spreadsheets. The properties of the reconstructed chart objects give their data structures including series names and values, which are useful for automatic processing of data by computer. Since results produced by crowdsourcing inherently contain errors, a quality control mechanism was developed that improves the accuracy of extracted tables by aggregating tables created by different workers for the same chart image and by utilizing the data structures obtained from the reproduced chart objects. Experimental results demonstrated that the proposed framework and mechanism are effective.
Satoshi Oyama, Yukino Baba, Ikki Ohmukai, Hiroaki Dokoshi, Hisashi Kashima
DSAA2
2015 Quality Control for Crowdsourced Hierarchical Classification
abstract
Repeated labeling is a widely adopted quality control method in crowdsourcing. This method is based on selecting one reliable label from multiple labels collected by workers because a single label from only one worker has a wide variance of accuracy. Hierarchical classification, where each class has a hierarchical relationship, is a typical task in crowdsourcing. However, direct applications of existing methods designed for multi-class classification have the disadvantage of discriminating among a large number of classes. In this paper, we propose a label aggregation method for hierarchical classification tasks. Our method takes the hierarchical structure into account to handle a large number of classes and estimate worker abilities more precisely. Our method is inspired by the steps model based on item response theory, which models responses of examinees to sequentially dependent questions. We considered hierarchical classification to be a question consisting of a sequence of subquestions and built a worker response model for hierarchical classification. We conducted experiments using real crowdsourced hierarchical classification tasks and demonstrated the benefit of incorporating a hierarchical structure to improve the label aggregation accuracy.
Naoki Otani, Yukino Baba, Hisashi Kashima
ICDM2
2015 Predictive Approaches for Low-Cost Preventive Medicine Program in Developing Countries
abstract
Non-communicable diseases (NCDs) are no longer just a problem for high-income countries, but they are also a problem that affects developing countries. Preventive medicine is definitely the key to combat NCDs; however, the cost of preventive programs is a critical issue affecting the popularization of these medicine programs in developing countries. In this study, we investigate predictive modeling for providing a low-cost preventive medicine program. In our two-year-long field study in Bangladesh, we collected the health checkup results of 15,075 subjects, the data of 6,607 prescriptions, and the follow-up examination results of 2,109 subjects. We address three prediction problems, namely subject risk prediction, drug recommendation, and future risk prediction, by using machine learning techniques; our multiple-classifier approach successfully reduced the costs of health checkups, a multi-task learning method provided accurate recommendation for specific types of drugs, and an active learning method achieved an efficient assignment of healthcare workers for the follow-up care of subjects.
Yukino Baba, Hisashi Kashima, Yasunobu Nohara, Eiko Kai, Partha Pratim Ghosh, Rafiqul Islam Maruf, Ashir Ahmed, Masahiro Kuroda, Sozo Inoue, Tatsuo Hiramatsu, Michio Kimura, Shuji Shimizu, Kunihisa Kobayashi, Koji Tsuda, Masashi Sugiyama, Mathieu Blondel, Naonori Ueda, Masaru Kitsuregawa, Naoki Nakashima
KDD1
2015 Quality Control for Crowdsourced POI Collection
Shunsuke Kajimura, Yukino Baba, Hiroshi Kajino, Hisashi Kashima
PAKDD (2)2
2014 Crowdsourced data analytics: A case study of a predictive modeling competition
abstract
Predictive modeling competitions provide a new data mining approach that leverages crowds of data scientists to examine a wide variety of predictive models and build the best performance model. Competition hosts, who provide their own dataset and specify the problem to be solved, are not only able to obtain the best model from among those submitted but also to aggregate the submitted models to obtain one that outperforms the rest. In this paper, we report the results of a study conducted on CrowdSolving, a platform for predictive modeling competitions in Japan. We hosted a competition on a link prediction task and observed that (i) the prediction performance of the winner significantly outperformed that of a state-of-the-art method, (ii) the aggregated model constructed from all submitted models further improved the final performance, and (iii) the performance of the aggregated model built only from early submissions nevertheless overtook the final performance of the winner. Our results show the power of crowds for predictive modeling, not only in the quality of the obtained model, but also in its speed to achieve it. Furthermore, they demonstrate the possibilities of combining human insights and machine learning in data analytics.
Yukino Baba, Nozomi Nori, Shigeru Saito, Hisashi Kashima
DSAA1
2014 Crowdsourced Data Analytics: A Case Study of a Predictive Modeling Competition
abstract
Predictive modeling competitions provide a new data mining approach that leverages crowds of data scientists to examine a wide variety of predictive models and build the best performance model. In this paper, we report the results of a study conducted on CrowdSolving, a platform for predictive modeling competitions in Japan. We hosted a competition on a link prediction task and observed that (i) the prediction performance of the winner significantly outperformed that of a state-of-the-art method, (ii) the aggregated model constructed from all submitted models further improved the final performance, and (iii) the performance of the aggregated model built only from early submissions nevertheless overtook the final performance of the winner.
Yukino Baba, Nozomi Nori, Shigeru Saito, Hisashi Kashima
HCOMP1
2014 Quality Control for Crowdsourced Enumeration Tasks
abstract
Quality control is one of the central issues in crowdsourcing research. In this paper, we consider a quality control problem of crowdsourced enumeration tasks that request workers to enumerate possible answers as many as possible. Since workers neither necessarily provide correct answers nor provide exactly the same answers even if the answers indicate the same idea, we propose a two-stage quality control method consisting of the answer clustering stage and the reliability estimation stage.
Shunsuke Kajimura, Yukino Baba, Hiroshi Kajino, Hisashi Kashima
HCOMP2
2014 Instance-Privacy Preserving Crowdsourcing
abstract
Crowdsourcing is a technique to outsource tasks to a number of workers. Although crowdsourcing has many advantages, it gives rise to the risk that sensitive information may be leaked, which has limited the spread of its popularity. Task instances (data workers receive to process tasks) often contain sensitive information, which can be extracted by workers. For example, in an audio transcription task, an audio file corresponds to an instance, and the content of the audio (e.g., the abstract of a meeting) can be sensitive information. In this paper, we propose a quantitative analysis framework for the instance privacy problem. The proposed framework supplies us performance measures of instance privacy preserving protocols. As a case study, we apply the proposed framework to an instance clipping protocol and analyze the properties of the protocol. The protocol preserves privacy by clipping instances to limit the amount of information workers obtain. The results show that the protocol can balance task performance and instance privacy preservation. They also show that the proposed measure is consistent with standard measures, which validates the proposed measure.
Hiroshi Kajino, Yukino Baba, Hisashi Kashima
HCOMP2
2014 Crowdordering
Toshiko Matsui, Yukino Baba, Toshihiro Kamishima, Hisashi Kashima
PAKDD (2)2
2014 Leveraging non-expert crowdsourcing workers for improper task detection in crowdsourcing marketplaces
Yukino Baba, Hisashi Kashima, Kei Kinoshita, Goushi Yamaguchi, Yosuke Akiyoshi
Expert Syst. Appl.1
2013 Leveraging Crowdsourcing to Detect Improper Tasks in Crowdsourcing Marketplaces
abstract
Controlling the quality of tasks is a major challenge in crowdsourcing marketplaces. Most of the existing crowdsourcing services prohibit requesters from posting illegal or objectionable tasks. Operators in the marketplaces have to monitor the tasks continuously to find such improper tasks; however, it is too expensive to manually investigate each task. In this paper, we present the reports of our trial study on automatic detection of improper tasks to support the monitoring of activities by marketplace operators. We perform experiments using real task data from a commercial crowdsourcing marketplace and show that the classifier trained by the operator judgments achieves high accuracy in detecting improper tasks. In addition, to reduce the annotation costs of the operator and improve the classification accuracy, we consider the use of crowdsourcing for task annotation. We hire a group of crowdsourcing (non-expert) workers to monitor posted tasks, and incorporate their judgments into the training data of the classifier. By applying quality control techniques to handle the variability in worker reliability, our results show that the use of non-expert judgments by crowdsourcing workers in combination with expert judgments improves the accuracy of detecting improper crowdsourcing tasks.
Yukino Baba, Hisashi Kashima, Kei Kinoshita, Goushi Yamaguchi, Yosuke Akiyoshi
IAAI1
2013 Accurate Integration of Crowdsourced Labels Using Workers' Self-reported Confidence Scores
Satoshi Oyama, Yukino Baba, Yuko Sakurai, Hisashi Kashima
IJCAI2
2013 Statistical quality estimation for general crowdsourcing tasks
abstract
One of the biggest challenges for requesters and platform providers of crowdsourcing is quality control, which is to expect high-quality results from crowd workers who are neither necessarily very capable nor motivated. A common approach to tackle this problem is to introduce redundancy, that is, to request multiple workers to work on the same tasks. For simple multiple-choice tasks, several statistical methods to aggregate the multiple answers have been proposed. However, these methods cannot always be applied to more general tasks with unstructured response formats such as article writing, program coding, and logo designing, which occupy the majority on most crowdsourcing marketplaces. In this paper, we propose an unsupervised statistical quality estimation method for such general crowdsourcing tasks. Our method is based on the two-stage procedure; multiple workers are first requested to work on the same tasks in the creation stage, and then another set of workers review and grade each artifact in the review stage. We model the ability of each author and the bias of each reviewer, and propose a two-stage probabilistic generative model using the graded response model in the item response theory. Experiments using several general crowdsourcing tasks show that our method outperforms popular vote aggregation methods, which implies that our method can deliver high quality results with lower costs.
Yukino Baba, Hisashi Kashima
KDD1
2010 Extraction of Places Related to Flickr Tags
abstract
Geographic information systems use databases to map keywords to places. These databases are currently most often created by using a top-down approach based on the geographic definitions. However, there is a problem with this approach in that these databases only contain location definitions such as addresses and place names, which does not allow for searches using keywords other than these words. Additionally, they do not give any information on the popularity, e.g., which is more popular among the places indexed by the same keyword. A bottom-up approach, based on the actual usage of words, can address these problems. We propose a method to aggregate tagging data and extract places related to a tag using the pair of a tag and a geo-tagged photo. We target the co-occurrence of a tag and the geolocation and represent the places related to a tag as a probability distribution over the longitudes and latitudes. We applied our method to data on the photo sharing service Flickr and experimentally confirmed that our method made it possible to highly-accurately extract places related to tags. Our direct bottom-up approach enables the extraction of place information that is not obtained by using traditional top-down approaches.
Yukino Baba, Fuyuki Ishikawa, Shinichi Honiden
ECAI1