Masaki Matsubara

dblp:30/8788 · DBLP profile ↗
← Back
27ranked-venue papers in the field
1as first author
6since 2021 · last 2024
0000-0003-1950-683XORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 19 (1 first)Information Retrieval & Web Search · 4Data Mining & Knowledge Discovery · 2Database Systems & Data Management · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2024 Automatic sleep stage classification for sleep apnea patients using an in-home sleep electroencephalography device
abstract
With the rising awareness of the critical role sleep plays in both health and social well-being, the demand for sleep studies is rapidly increasing.Automatic sleep stage classification is a fundamental part of sleep measurement, and machine learning models have been developed to assist in this process. These models achieve accuracy comparable to that of technicians when using data from healthy individuals. However, sleep patterns in individuals with sleep disorders, such as sleep apnea syndrome (SAS), one of the most common sleep disorders, differ from those of healthy individuals. As a result, existing models trained on healthy individuals’ data do not achieve sufficient accuracy when applied to SAS patients. This is a barrier to clinical application.A recent study using in-home EEG devices showed that technicians can accurately classify sleep stages in SAS cases by considering surrounding epochs. Based on this, we developed a model dedicated to SAS patients that incorporates the temporal context of relevant epochs.We found that this context-aware model significantly improved classification accuracy compared to models that only focused on the target epoch. In the training process using data from 76 severe SAS cases, the model based solely on single-epoch data achieved an accuracy of 71.5%, while the model considering the surrounding epochs achieved an accuracy of 73.7%. The classification accuracy improved across all stages except N3.This approach appears to capture the frequent sleep stage transitions characteristic of SAS.
Saki Tsumoto, Jaehoon Seol, Kazumasa Horie, Fusae Kawana, Morie Tominaga, Shigeru Chiba, Hideaki Kondo, Hiroyuki Yoshimine, Masaki Matsubara, Atsuyuki Morishima, Masashi Yanagisawa, Hiroyuki Kitagawa
IEEE Big Data9
2024 An Adaptive Feature Selection Method for Learning-to-Enumerate Problem
Satoshi Horikawa, Chiyonosuke Nemoto, Keishi Tajima, Masaki Matsubara, Atsuyuki Morishima
ECIR (3)4
2022 Image Geolocation by Non-Expert Crowd Workers with an Expert Strategy
abstract
Identifying the location where a photo was taken is an important operation in many applications such as the disaster response if it is not associated with the location information. This process usually is done by experts who are familiar with the geolocation. However it is not guaranteed that we can find such expert workers. In this paper, we explore an approach to improve the quality of image geolocation with an workflow implement experts’ strategy for the image geolocations. The result of preliminary experiment suggested that our approach is effective in improving the accuracy of non-expert geolocation.
Seungun Kim, Masaki Matsubara, Atsuyuki Morishima
IEEE Big Data2
2022 Multi-Armed Bandit Approach to Qualification Task Assignment across Multi Crowdsourcing Platforms
abstract
Many existing optimization approaches deal with task assignments on one single crowdsourcing platform. This paper addresses the difficulties of the optimal platform selection for qualification tasks on one single platform. We proposed a novel approach about assigning qualification tasks to workers iteratively on multiple platforms to maximize the total number of collected qualified workers on a limited budget. We applied Multi-Armed Bandit (MAB) algorithms to create strategies for the platform selections to achieve this goal. The conducted experiments revealed that (1) the optimal platform is not always trivial, and (2) the strategies created by MAB algorithms can achieve high-quality assignments under different settings which also satisfied different requesters’ needs.
Yunyi Xiao, Yu Yamashita, Hiroyoshi Ito, Masaki Matsubara, Atsuyuki Morishima
IEEE Big Data4
2021 A Skill-based Worksharing Approach for Microtask Assignment
abstract
When selecting workers in microtask crowdsourcing platforms, requesters select qualified workers by looking at the evaluation results for the tasks in the past or by conducting qualifying tests for the tasks. As a result, they choose workers whose skill levels are above some threshold. This sometimes limits the number of workers who perform the tasks, which has a negative effect for both of requesters and workers. In this paper, we explore an approach to increasing the work opportunities for many workers, by finding task assignment based on the estimated skill level of workers and the difficulty level of tasks. We show the result of a preliminary experiment to discuss the potential and limitation of this approach.
Kanta Negishi, Hiroyoshi Ito, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData3
2021 Does Multi-Hop Crowdsourcing Work? A Case Study on Collecting COVID-19 Local Information
abstract
The coronavirus disease 2019 (COVID-19) pandemic has spread across the globe from the beginning of 2020 and people worldwide have been receiving news about the same from government offices, press conferences and various other media outlets. The COVID-19 Information Watcher Project started in 2020 to collect and organize reliable information sources worldwide. However, it is difficult to automatically identify reliable information sources in foreign countries for several reasons. First, what kind of information sources are reliable heavily depend on each county situation. In some countries people trust their government’s official information but in other countries they do not. Secondly, such reliable information sources often provide information in their local languages. Reliable information sources are not necessarily top-ranked by search engines. Crowdsourcing is a promising way to deal with such a case. However, crowd-sourcing platforms do not cover crowds in all countries. In this study, we report some results of our attempt to collect local information regarding COVID-19 from several countries through multi-hop crowdsourcing, in which we allow crowd workers on a crowdsourcing platform to use other platforms in other countries. We show two case studies, Russia and Afghanistan. Our results show that the multi-hop crowdsourcing is a promising way to collect COVID-19 information from different countries.
Ying Zhong 0006, Masaki Kobayashi, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData3
2020 Analysis of Hand-drawn Maps of Places in Natural Disaster Pictures
abstract
Understanding the current situation of natural disaster damages is a critical step for an effective natural disaster responses, and many pictures uploaded after natural disasters are valuable resources for this purpose. However, many pictures are not associated with location information and it is not easy to connect the image content to locations on a map. There are two reasons for this difficulty. First, pictures are different in terms of directions and heights. Second, the situation of damaged areas in the picture may differ from before the natural disaster. Therefore, there is a mismatch between map fragments and the pictures taken after a disaster. This paper explores the potential of human computation to solve this problem. For this study, we asked people to draw a birds-eye view map of the place in a picture and compared the map with the correct map fragments to see the characteristics of such drawn maps.
Seungun Kim, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData2
2020 Human-in-the-loop Approach towards Dual Process AI Decisions
abstract
How to develop AI systems that can explain how they made decisions is one of the important and hot topics today. Inspired by the dual-process theory in psychology, this paper proposes a human-in-the-loop approach to develop System-2 AI that makes an inference logically and outputs interpretable explanation. Our proposed method first asks crowd workers to raise understandable features of objects of multiple classes and collect training data from the Internet to generate classifiers for the features. Logical decision rules with the set of generated classifiers can explain why each object is of a particular class. In our preliminary experiment, we applied our method to an image classification of Asian national flags and examined the effectiveness and issues of our method. In our future studies, we plan to combine the System-2 AI with System-1 AI (e.g., neural networks) to efficiently output decisions.
Hikaru Uchida, Masaki Matsubara, Kei Wakabayashi, Atsuyuki Morishima
IEEE BigData2
2019 Incentive Design for Crowdsourced Development of Selective AI for Human and Machine Data Processing: A Case Study
abstract
The most typical approach today to data processing which does not have proven algorithms is to first request humans to provide labels to a small set of data and then develop artificial intelligences (AIs) with the data to perform all the remaining tasks. This development is sometimes crowdsourced through platforms such as Kaggle. The approach, however, is not always effective; if the AI does not meet the quality requirement, we may have to give up the development and all the data items have to be done manually. In order to avoid this all-or-nothing situation, “selective” AI programs that perform tasks which they are confident to do will be effective. This study addresses the problem of designing an incentive structure for crowdsourcing the development of such selective AI programs. This paper shows the results of our real-world experiment with a stair-step incentive structure and the behavior of a worker who developed the AI agent under the incentive. This paper also discusses the limitations of the proposed incentive design.
Masafumi Hayashi, Masaki Kobayashi, Masaki Matsubara, Toshiyuki Amagasa, Atsuyuki Morishima
IEEE BigData3
2019 A Microtask Approach to Identifying Incomprehension for Facilitating Peer Learning
abstract
Peer learning is a known effective method of education, which involves people teaching each other. Usually, peer learning requires dedicated facilitators, but it is not always possible to have them in some situations such as in classes with a large number of students and in crowdsourcing settings. This paper addresses the question whether there is a way through which people can teach each other without dedicated facilitators. A key issue is to allow people to identify their own incomprehension. We applied a microtask approach to the problem and verified that the developed workflow was effective. The result of our preliminary experiment suggests that the approach is effective for identifying students' incomprehension. We also show that the identified incomprehensible parts allow people to teach each other.
Hinako Izumi, Masaki Matsubara, Chiemi Watanabe, Atsuyuki Morishima
IEEE BigData2
2019 Active Learning Strategies for Hierarchical Labeling Microtasks
abstract
This paper reports the result of a preliminary experiment on active learning strategies for the hierarchical labeling microtasks. A typical example of hierarchical labeling microtask consists of a set of labeling tasks for partitions of a large image; starting from the whole image, the workers choose to give a label or divide it into smaller ones. This paper shows the result of an experiment to compare several strategies for active learning in the setting. The result suggests that the difference in the strategies affects the performance in the early stage.
Kousuke Uo, Masaki Kobayashi, Masaki Matsubara, Yukino Baba, Atsuyuki Morishima
IEEE BigData3
2019 Towards Quality Assessment of Crowdworker Output Based on Behavioral Data
abstract
In this paper, we show preliminary results on the quality assessment of crowdworker output based on the movements of the mouse and the eyes while the task is performed. We assume that the mouse and the eyes stop longer if the quality is lower due to the lack of knowledge, or confidence, etc. Because the mouse- and eye-stopping duration follows lognormal distribution, we estimate its parameters (mean and standard deviation) to evaluate the quality. Results of preliminary experiments with 10 participants show that the parameters of correct outputs are different from those of incorrect ones. As compared to the task duration, which is often used as a feature for assessment, we have found that the mouse-and the eyestopping duration is advantageous and complementary for the assessment.
Shigeaki Yuasa, Takumi Nakai, Takanori Maruichi, Manuel Landsmann, Koichi Kise, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData6
2019 A Privacy-Preserving Similarity Search Scheme over Encrypted Word Embeddings
abstract
Recent evolution in cloud computing platforms have attracted the largest amount of data than ever before. Today, even the most sensitive data are being outsourced, thus, protection is essential to ensure that privacy is not traded for the convenience provided by cloud platforms. Traditional symmetric encryption schemes provide good protection; however, they ruin the merits of cloud computing. Attempts have been made to obtain a scheme where both functionality and protection can be achieved. However, features provided in existing searchable encryption schemes tend to be left behind the latest findings in the information retrieval (IR) area.
Daisuke Aritomo, Chiemi Watanabe, Masaki Matsubara, Atsuyuki Morishima
iiWAS3
2018 A Task Assignment Method Considering Inclusiveness and Activity Degree
abstract
Task assignment is one of the important issues in crowdsourcing. Most existing schemes consider both of task-centric measures (e.g., required abilities) and worker-centric measures (e.g., worker preference) at local (each task) assignment level, but consider only task-centric measures at the global (the whole workflow) assignment level, such as productivity and throughput. This paper proposes to introduce Inclusiveness and Activity Degree as the worker-centric measures at the global level. It is not trivial whether we can find assignment that are good in terms of all of the two task-centric measures (productivity and throughput) and the two worker-centric measures (Inclusiveness and Activity Degree) at the global level. This paper explains five assignments that are expected to increase some measures and shows Activity Degree Conscious Assignment can produce high Inclusiveness and Activity Degree and not significantly reduce the productivity and throughput assignment through a simulation.
Hirotaka Hashimoto, Masaki Matsubara, Yuhki Shiraishi, Daisuke Wakatsuki, Jianwei Zhang 0002, Atsuyuki Morishima
IEEE BigData2
2018 A Learning Effect by Presenting Machine Prediction as a Reference Answer in Self-correction
abstract
Can people learn from machines behavior in microtask based crowdsourcing? Can we train the machines as our mentor even without domain expertise? In this paper, we investigate how the task results improve concerning quality during and after presenting machine prediction as a reference answer in self-correction. Four reference types were examined in the experiment; Correct, Random, Machine prediction trained by correct answers, and that trained by human answers. Learning effects were observed only in presenting machine prediction, although those accuracy rates were far from correct (100%). Moreover, there were no learning effects in "Correct" and "Random". This suggests the following hypothesis: Since machine learners make some "models" for the problem, it is easier for humans to interpret the outputs of machine learners than the results without via them; it is more difficult to interpret not only random answers but also the correct answers in a case where the perfect interpretation of the problem is difficult. Furthermore, some workers answered with higher accuracy rate than machines in the post-test. Therefore, this strategy can be expected to be useful for bootstrapping solutions in the situation where unknown problems occur without expertise or at a low cost.
Masaki Matsubara, Masaki Kobayashi, Atsuyuki Morishima
IEEE BigData1
2018 Worker Classification based on Answer Pattern for Finding Typical Mistake Patterns
abstract
One of the problems in crowdsourcing is the development of appropriate instructions for workers. To improve task instructions, we must find typical mistake patterns. However, manually identifying these patterns is a cumbersome task. This study shows that a relatively simple approach classifying workers in terms of their understanding of task instructions is promising for addressing this issue. The verification results by domain experts suggest that the output is useful for improving task instructions.
Tomoya Mikami, Masaki Matsubara, Takashi Harada, Atsuyuki Morishima
IEEE BigData2
2018 A Cache-based Approach to Dynamic Switching between Different Dataflows in Crowdsourcing
abstract
At times, a composite dataflow needs rerunning in crowdsourcing for various reasons, even when the dataflow may be half complete. Rerunning the dataflow requires more time and incurs monetary costs for the additional work that would need to be completed by crowd workers. This time and cost may be reduced by reusing complete or intermediate results in the previous run. However, at times, such results cannot be used as is (e.g., when the dataflow has been changed), and some additional tasks need to be completed in the old dataflow in order to make them reusable in the new dataflow. The benefit of reusing these results in the previous run may or may not be worth the cost of these additional tasks. This paper gives a general framework for formulating this problem, and proposed a method to estimate the additional costs. The simulation result shows that it is worth devising optimization techniques to identify feasible (namely, cost-effective) plans.
Yusuke Suzuki, Masaki Matsubara, Keishi Tajima, Toshiyuki Amagasa, Atsuyuki Morishima
IEEE BigData2
2018 Finding Evidences by Crowdsourcing
abstract
Crowdsourcing is a promising tool involving multiple people in completing tasks that are difficult to complete by an individual, a small team or a computer. Ensuring the quality of the results is also one of the primary problems in crowdsourcing. One of the major approaches to improve the data quality to aggregate answers from more than one workers. This study explores a different approach - we ask workers to prove facts. We devise a general framework for collecting and ranking evidence-based proofs. The experiments results show that the proposed framework works and how diverse the collected proofs are. Our results clearly indicate that the crowd-based approach to prove facts is promising.
Nadeesha Wijerathna, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData2
2018 Identification of Important Images for Understanding Web Pages
abstract
People have become increasingly dependent on the Web. However, it is difficult for visually impaired individuals to understand the content, especially if the web page contains images. The images can be more accessible by adding alt text, but it is known that alt text added through the current automation techniques are not necessarily helpful. Crowdsourcing is a promising approach for it, but adding alt text to all the images on the Web requires a tremendous amount of effort. In addition, too many alt texts of images also increase the difficulty of reading. Therefore, it is crucial to select important images for understanding web pages. This paper presents the results of our preliminary experiments to identify important images for understanding web pages. We adopted a crowdsourcing approach with two microtask designs. The results of our study demonstrated that (1) there are several types of images that are difficult to automatically assess the importance; thus, human-in-the-loop can be a promising approach to identify important ones in the types of images; and (2) there are microtask designs that can result in the similar results but different to each other in terms of other criteria. The results suggested that the identification of important images for understanding web pages is an interesting problem to address.
Ying Zhong 0006, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData2
2018 CrowdSheet: An Easy-To-Use One-Stop Tool for Writing and Executing Complex Crowdsourcing
Rikuya Suzuki, Tetsuo Sakaguchi, Masaki Matsubara, Hiroyuki Kitagawa, Atsuyuki Morishima
CAiSE3
2018 Efficient Pipeline Processing of Crowdsourcing Workflows
abstract
This paper addresses the pipeline processing of sequential workflows in crowdsourcing. Sequential workflows consisting of several subtasks are ubiquitous in crowdsourcing. Our approach is to control the budget distribution to subtasks in order to balance the execution speed of the subtasks and to improve throughput of overall sequential workflows. As we cannot control the price for earlier steps retrospectively in the stepwise batch execution, we explore pipeline processing schemes. Our experimental results show that our pipeline processing scheme with price control achieves significantly higher throughput of sequential workflows.
Ken Mizusawa, Keishi Tajima, Masaki Matsubara, Toshiyuki Amagasa, Atsuyuki Morishima
CIKM3
2018 An Empirical Study on Short- and Long-Term Effects of Self-Correction in Crowdsourced Microtasks
abstract
Self-correction for crowdsourced tasks is a two-stage setting that allows a crowd worker to review the task results of other workers; the worker is then given a chance to update his/her results according to the review.Self-correction was proposed as an approach complementary to statistical algorithms in which workers independently perform the same task. It can provide higher-quality results with few additional costs. However, thus far, the effects have only been demonstrated in simulations, and empirical evaluations are needed. In addition, as self-correction gives feedback to workers, an interesting question arises: whether perceptual learning is observed in self-correction tasks. This paper reports our experimental results on self-corrections with a real-world crowdsourcing service.The empirical results show the following: (1) Self-correction is effective for making workers reconsider their judgments. (2) Self-correction is more effective if workers are shown task results produced by higher-quality workers during the second stage. (3) Perceptual learning effect is observed in some cases. Self-correction can give feedback that shows workers how to provide high-quality answers in future tasks.The findings imply that we can construct a positive loop to improve the quality of workers effectively.We also analyze in which cases perceptual learning can be observed with self-correction in crowdsourced microtasks.
Masaki Kobayashi, Hiromi Morita, Masaki Matsubara, Nobuyuki Shimizu, Atsuyuki Morishima
HCOMP3
2018 Skill-and-Stress-Aware Assignment of Crowd-Worker Groups to Task Streams
abstract
Worker-task assignments represent one of the critical issues in crowdsourcing, as they affect the quality of task results. This study addresses the problem of forming worker groups assigned to the same task in a task stream that requires more than one worker. We introduce a worker-group queue model that covers practical and common scenarios for task-stream crowdsourcing, and compare three strategies in terms of the skill balance among worker groups, the quality of the final outputs, the number of worker re-assignments of workers, and psychological stress felt by workers. We found that one of the compared strategies that employs multiple worker queues yields good results based on these measures.
Katsumi Kumai, Masaki Matsubara, Yuhki Shiraishi, Daisuke Wakatsuki, Jianwei Zhang 0002, Takeaki Shionome, Hiroyuki Kitagawa, Atsuyuki Morishima
HCOMP2
2018 CrowdSheet: Instant Implementation and Out-of-Hand Execution of Complex Crowdsourcing
abstract
We demonstrate CrowdSheet, a spreadsheet interface for implementing complex crowdsourcing. Despite its appeal, adoption of the spreadsheet paradigm is associated with two nontrivial problems: (1) how to design the interface, which must be a natural extension of existing spreadsheets, while guaranteeing a reasonable expressive power, and (2) how to incorporate techniques for improving data quality without sacrificing the spreadsheet's easy-to-use feature. In this demo, we show three things. First, CrowdSheet allows non IT experts to easily implement crowdsourcing applications with complex workflows. Second, CrowdSheet adopts only new two spreadsheet functions to implement a fairly wide range of real-world applications. Third, its modular architecture gives CrowdSheet a declarative feature that lets users choose alternative plans for improving data quality, while keeping the CrowdSheet description simple.
Rikuya Suzuki, Tetsuo Sakaguchi, Masaki Matsubara, Hiroyuki Kitagawa, Atsuyuki Morishima
ICDE3
2017 Crowd-based best-effort number estimation
abstract
Determining numbers is a fundamental operation in many applications. Although the main objective is to determine the exact number, a complete enumeration is not always possible. Therefore, different methods have been employed depending on the situation, such as Fermi estimation and complete enumeration. This paper introduces a method for crowd-based best-effort number estimation, which seamlessly integrates Fermi estimation and complete enumeration. The principle underlying of the method is defining a table called the Fermi-estimation table, which represents both task decomposition and Fermi estimation. This paper also shows our preliminary results and discusses some of the future prospects.
Yuzuki Furuhashi, Masaki Matsubara, Atsuyuki Morishima
IEEE BigData2
2017 A crowd-in-the-loop approach for generating conference programs with microtasks
abstract
Creating session programs for large-scale academic conferences is a cumbersome task for program committees. Our goal is to establish a crowd-in-the-loop method for creating session programs. In the proposed method, crowd workers are both paper authors and PC members. In contrast to existing approaches, our method is unique in that we use microtasks as much as possible to minimize interactions between workers. This paper presents an overview of our approach, as well as the results of a preliminary experiment on the bidding tasks for sessions by authors. We found that authors do not necessarily bid for appropriate sessions, meaning we need a mechanism to properly modify the bidding results.
Masaki Matsubara, Keishi Tajima, Atsuyuki Morishima
IEEE BigData2
2011 Development of a method for automatic basso continuo playing
Masahiro Niitsuma, Masaki Matsubara, Masaki Oono, Hiroaki Saito 0001
Inf. Process. Manag.2