VLDB 2026 Research / reviewers in the wild / expert
Jean Y. Song
dblp:216/0041
· DBLP profile ↗
23ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-4379-3971ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 19 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LingoQ: Bridging the Gap between EFL Learning and Work through AI-Generated Work-Related QuizzesabstractNon-native English speakers performing English-related tasks at work struggle to sustain EFL learning, despite their motivation. Often, study materials are disconnected from their work context. Our formative study revealed that reviewing work-related English becomes burdensome with current systems, especially after work. Although workers rely on LLM-based assistants to address their immediate needs, these interactions may not directly contribute to their English skills. We present LingoQ, an AI-mediated system that allows workers to practice English using quizzes generated from their LLM queries during work. LingoQ leverages these on-the-fly queries using AI to generate personalized quizzes that workers can review and practice on their smartphones. We conducted a three-week deployment study with 28 EFL workers to evaluate LingoQ. Participants valued the quality-assured, work-situated quizzes and constantly engaging with the app during the study. This active engagement improved self-efficacy and led to learning gains for beginners and, potentially, for intermediate learners. Drawing on these results, we discuss design implications for leveraging workers’ growing reliance on LLMs to foster proficiency and engagement while respecting work boundaries and ethics. Yeonsun Yang, Sang Won Lee 0002, Jean Y. Song, Sangdoo Yun, Young-Ho Kim |
CHI | 3 |
| 2026 | Caught by Surprise, Caught by Culture: Bridging Facial Expression's Recognition and Interpretation of Surprise Across CulturesabstractFacial expressions are powerful signals of human emotion, shaping both human–human and human–computer interaction. As interactive technologies, from adaptive interfaces to emotion-aware agents, become more pervasive, systems are increasingly expected to recognize and respond to users’ emotions naturally. But what if a system misreads your face? Such misinterpretation is particularly likely when cultural differences in emotion perception are overlooked. This problem may be compounded by the fact that most facial emotion recognition (FER) models are trained on datasets that reflect the norms of a particular cultural group that assume universality, limiting their reliability in multicultural contexts. Surprise, in particular, is an emotion whose valence can be either positive or negative depending on context, making it a critical case for investigating cultural bias in FER. To address this, we examined how cultural background shapes the recognition and valence interpretation of surprise facial expressions among South Korean (N=36) and American (N=34) participants. Participants labeled 200 facial expressions (surprise and fear), rated their perceived valence, and described personal experiences of surprise. Results show that South Korean-labeled surprise expressions exhibited stronger negative Action Unit (AU) activation and lower valence ratings, whereas American-labeled ones showed more balanced or positive facial cues. Qualitative accounts further revealed that South Koreans framed surprise as tense or socially cautious, while Americans viewed it as open and situationally flexible. These findings bridge recognition and interpretation in cross-cultural emotion research and highlight the need for culturally adaptive FER systems that can interpret ambiguous emotions like surprise more inclusively. Ha Eun Jang, Yeonsun Yang, Jean Y. Song |
IUI | 3 |
| 2026 | ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational InteractionsabstractFrom purchasing a gift to deciding on a hobby, unfamiliar decisions—decisions without domain knowledge and experience—are frequent and significant. The complexity and uncertainty of such decisions demand unique approaches to information seeking, understanding, and decision-making. Our formative study highlights that in the current workflow, users want to start by discovering broad and relevant domain information evenly and simultaneously, quickly address emerging inquiries, and gain personalized standards to assess information found. We present ChoiceMates, an interactive multi-agent system designed to address these needs by enabling users to engage with a dynamic set of LLM agents each presenting a unique experience in the domain. Unlike existing multi-agent systems that automate tasks with agents, the user orchestrates agents to assist their decision-making process in each turn, through chatting with all agents, with a tagged subset of agents, or calling in new agents into the space. By comparing ChoiceMates with a web search condition and a multi-agent framework (n=12), we show that ChoiceMates enables a more confident, satisfactory decision-making with better situation understanding than web search, and higher decision quality than a commercial multi-agent framework. We further illustrate how participants utilized ChoiceMates to make unfamiliar decisions, providing insights into designing a more controllable and collaborative multi-agent system. Jeongeon Park, Bryan Min, Kihoon Son, Jean Y. Song, Xiaojuan Ma, Juho Kim 0001 |
IUI | 4 |
| 2024 | Exploring Intervention Techniques to Alleviate Negative Emotions during Video Content Moderation Tasks as a Worker-centered Task DesignabstractVideos are dynamic and multi-modal compared to other types of content, making automatic filtering difficult, which is why content moderators play a crucial role. However, video content moderators are exposed to more profound emotional labor because videos contain rich visual information, sometimes including even harmful content, such as violent or terrifying scenes. In this work, we explore the effect of six intervention techniques on alleviating negative emotions during video content moderation tasks. We conducted one online crowdsourcing experiment and two controlled user studies to find out that (i) interleaving with positive videos or (ii) cartoonization could significantly reduce negative emotions in the moderators. Participants reported that the advantages of these approaches are in helping reduce negative emotions at the time of moderation while existing approaches focus on post-task activities (e.g., relaxation, talking with others, or getting a hobby). We discuss the applicability of our findings to broader tasks, including improvement in intervention techniques. Dokyun Lee, Sangeun Seo, Chanwoo Park, Sunjun Kim, Buru Chang, Jean Y. Song |
Conference on Designing Interactive Systems | 6 |
| 2024 | FLUID-IoT : Flexible and Fine-Grained Access Control in Shared IoT Environments via Multi-user UI DistributionabstractThe rapid growth of the Internet of Things (IoT) in shared spaces has led to an increasing demand for sharing IoT devices among multiple users. Yet, existing IoT platforms often fall short by offering an all-or-nothing approach to access control, not only posing security risks but also inhibiting the growth of the shared IoT ecosystem. This paper introduces FLUID-IoT, a framework that enables flexible and granular multi-user access control, even down to the User Interface (UI) component level. Leveraging a multi-user UI distribution technique, FLUID-IoT transforms existing IoT apps into centralized hubs that selectively distribute UI components to users based on their permission levels. Our performance evaluation, encompassing coverage, latency, and memory consumption, affirm that FLUID-IoT can be seamlessly integrated with existing IoT platforms and offers adequate performance for daily IoT scenarios. An in-lab user study further supports that the framework is intuitive and user-friendly, requiring minimal training for efficient utilization. Sunjae Lee, Minwoo Jeong, Daye Song, Junyoung Choi 0002, Seoyun Son, Jean Y. Song, Insik Shin |
CHI | 6 |
| 2024 | Find the Bot!: Gamifying Facial Emotion Recognition for Both Human Training and Machine Learning Data CollectionabstractFacial emotion recognition (FER) constitutes an essential social skill for both humans and machines to interact with others. To this end, computer interfaces serve as valuable tools for training individuals to improve FER abilities, while also serving as tools for gathering labels to train FER machine learning datasets. However, existing tools have limitations on the scope and methods of training non-clinical populations and also on collecting labels for machines. In this study, we introduce Find the Bot!, an integrated game that effectively engages the general population to support not only human FER learning on spontaneous expressions but also the collection of reliable judgment-based labels. We incorporated design guidelines from gamification, education, and crowdsourcing literature to engage and motivate players. Our evaluation (N=59) shows that the game encourages players to learn emotional social norms on perceived facial expressions with a high agreement rate, facilitating effective FER learning and reliable label collection all while enjoying gameplay. Yeonsun Yang, Ahyeon Shin, Huidam Woo, John Joon Young Chung, Jean Y. Song |
CHI | 6 |
| 2024 | SERENUS: Alleviating Low-Battery Anxiety Through Real-time, Accurate, and User-Friendly Energy Consumption Prediction of Mobile ApplicationsabstractLow-battery anxiety has emerged as a result of growing dependence on mobile devices, where the anxiety arises when the battery level runs low. While battery life can be extended through power-efficient hardware and software optimization techniques, low-battery anxiety will still remain a phenomenon as long as mobile devices rely on batteries. In this paper, we investigate how an accurate real-time energy consumption prediction at the application-level can improve the user experience in low-battery situations. We present Serenus, a mobile system framework specifically tailored to predict the energy consumption of each mobile application and present the prediction in a user-friendly manner. We conducted user studies using Serenus to verify that highly accurate energy consumption predictions can effectively alleviate low-battery anxiety by assisting users in planning their application usage based on the remaining battery life. We summarize requirements to mitigate users’ anxiety, guiding the design of future mobile system frameworks. Sera Lee, Dae R. Jeong, Junyoung Choi 0002, Jaeheon Kwak, Seoyun Son, Jean Y. Song, Insik Shin |
UIST | 6 |
| 2023 | It is Okay to be Distracted: How Real-time Transcriptions Facilitate Online Meeting with DistractionabstractOnline meetings are indispensable in collaborative remote work environments, but they are vulnerable to distractions due to their distributed and location-agnostic nature. While distraction often leads to a decrease in online meeting quality due to loss of engagement and context, natural multitasking has positive tradeoff effects, such as increased productivity within a given time unit. In this study, we investigate the impact of real-time transcriptions (i.e., full-transcripts, summaries, and keywords) as a solution to help facilitate online meetings during distracting moments while still preserving multitasking behaviors. Through two rounds of controlled user studies, we qualitatively and quantitatively show that people can better catch up with the meeting flow and feel less interfered with when using real-time transcriptions. The benefits of real-time transcriptions were more pronounced after distracting activities. Furthermore, we reveal additional impacts of real-time transcriptions (e.g., supporting recalling contents) and suggest design implications for future online meeting platforms where these could be adaptively provided to users with different purposes. Seoyun Son, Junyoung Choi 0002, Sunjae Lee, Jean Y. Song, Insik Shin |
CHI | 4 |
| 2023 | ModSandbox: Facilitating Online Community Moderation Through Error Prediction and Improvement of Automated RulesabstractDespite the common use of rule-based tools for online content moderation, human moderators still spend a lot of time monitoring them to ensure they work as intended. Based on surveys and interviews with Reddit moderators who use AutoModerator, we identified the main challenges in reducing false positives and false negatives of automated rules: not being able to estimate the actual effect of a rule in advance and having difficulty figuring out how the rules should be updated. To address these issues, we built ModSandbox, a novel virtual sandbox system that detects possible false positives and false negatives of a rule and visualizes which part of the rule is causing issues. We conducted a comparative, between-subject study with online content moderators to evaluate the effect of ModSandbox in improving automated rules. Results show that ModSandbox can support quickly finding possible false positives and false negatives of automated rules and guide moderators to improve them to reduce future errors. Jean Y. Song, Jisoo Lee, Juho Kim 0001 |
CHI | 1 |
| 2023 | Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsabstractSupervised learning of image classifiers distills human knowledge into a parametric model fθthrough pairs of images and corresponding labels $\left\{ {\left( {{X_i},{Y_i}} \right)} \right\}_{i = 1}^N$. We argue that this simple and widely used representation of human knowledge neglects rich auxiliary information from the annotation procedure, such as the time-series of mouse traces and clicks left after image selection. Our insight is that such annotation byproducts Z provide approximate human attention that weakly guides the model to focus on the foreground cues, reducing spurious correlations and discouraging shortcut learning. To verify this, we create ImageNet-AB and COCO-AB. They are ImageNet and COCO training sets enriched with sample-wise annotation byproducts, collected by replicating the respective original annotation tasks. We refer to the new paradigm of training models with annotation byproducts as learning using annotation byproducts (LUAB). We show that a simple multitask loss for regressing Z together with Y already improves the generalisability and robustness of the learned models. Compared to the original supervised learning, LUAB does not require extra annotation costs. ImageNet-AB and COCO-AB are at github.com/naverai/NeglectedFreeLunch. Dongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung, Minsuk Chang, Sangdoo Yun, Jean Y. Song, Seong Joon Oh |
ICCV | 7 |
| 2022 | Promptiverse: Scalable Generation of Scaffolding Prompts Through Human-AI Hybrid Knowledge Graph AnnotationabstractOnline learners are hugely diverse with varying prior knowledge, but most instructional videos online are created to be one-size-fits-all. Thus, learners may struggle to understand the content by only watching the videos. Providing scaffolding prompts can help learners overcome these struggles through questions and hints that relate different concepts in the videos and elicit meaningful learning. However, serving diverse learners would require a spectrum of scaffolding prompts, which incurs high authoring effort. In this work, we introduce Promptiverse, an approach for generating diverse, multi-turn scaffolding prompts at scale, powered by numerous traversal paths over knowledge graphs. To facilitate the construction of the knowledge graphs, we propose a hybrid human-AI annotation tool, Grannotate. In our study (N=24), participants produced 40 times more on-par quality prompts with higher diversity, through Promptiverse and Grannotate, compared to hand-designed prompts. Promptiverse presents a model for creating diverse and adaptive learning experiences online. Yoonjoo Lee, John Joon Young Chung, Tae Soo Kim 0002, Jean Y. Song, Juho Kim 0001 |
CHI | 4 |
| 2022 | A-mash: providing single-app illusion for multi-app use through user-centric UI mashupabstractMobile apps offer a variety of features that greatly enhance user experience. However, users still often find it difficult to use mobile apps in the way they want. For example, it is not easy to use multiple apps simultaneously on a small screen of a smartphone. In this paper, we present A-Mash, a mobile platform that aims to simplify the way of interacting with multiple apps concurrently to the level of using a single app only. A key feature of A-Mash is that users can mash up the UIs of different existing mobile apps on a single screen according to their preferences. To this end, A-Mash 1) extracts UIs from unmodified existing apps (dynamic UI extraction) and 2) embeds extracted UIs from different apps into a single wrapper app (cross-process UI embedding), while 3) making all these processes hidden from the users (transparent execution environment). To the best of our knowledge, A-Mash is the first work to enable UIs of different unmodified legacy apps to seamlessly integrate and synchronize on a single screen, providing an illusion as if they were developed as a single app. A-Mash offers great potential for a number of useful usage scenarios. For instance, a user can mashup UIs of different IoT administration apps to create an all-in-one IoT device controller or one can mashup today's headlines from different news and magazine apps to craft one's own news headline collection. In addition, A-Mash can be extended to an AR space, in which users can map UI elements of different mobile apps to physical objects inside their AR scenes. Our evaluation of the A-Mash prototype implemented in Android OS demonstrates that A-Mash successfully supports the mashup of various existing mobile apps with little or no performance bottleneck. We also conducted in-depth user studies to assess the effectiveness of the A-Mash in real-world use cases. Sunjae Lee, Hoyoung Kim, Sijung Kim, Hyosu Kim, Jean Y. Song, Steven Y. Ko, Sangeun Oh, Insik Shin |
MobiCom | 6 |
| 2022 | XDesign: Integrating Interface Design into Explainable AI EducationabstractWe introduce XDesign, a web-based interactive platform that guides learners through a multi-stage design process for creating user-centered explanations of AI models. Results from a course deployment show that students were able to identify concrete user needs in interacting with explanations, highlight user tasks to support the needs, and design a user interface that aids the tasks. Hyungyu Shin, Nabila Sindi, Yoonjoo Lee, Jaeryoung Ka, Jean Y. Song, Juho Kim 0001 |
SIGCSE (2) | 5 |
| 2021 | Personalizing Ambience and Illusionary Presence: How People Use "Study with me" Videos to Create Effective Studying Environmentsabstract“Study with me” videos contain footage of people studying for hours, in which social components like conversations or informational content like instructions are absent. Recently, they became increasingly popular on video-sharing platforms. This paper provides the first broad look into what “study with me” videos are and how people use them. We analyzed 30 “study with me” videos and conducted 12 interviews with their viewers to understand their motivation and viewing practices. We identified a three-factor model that explains the mechanism for shaping a satisfactory studying experience in general. One of the factors, a well-suited ambience, was difficult to achieve because of two common challenges: external conditions that prevent studying in study-friendly places and extra cost needed to create a personally desired ambience. We found that the viewers used “study with me” videos to create a personalized ambience at a lower cost, to find controllable peer pressure, and to get emotional support. These findings suggest that the viewers self-regulate their learning through watching “study with me” videos to improve efficiency even when studying alone at home. Yoonjoo Lee, John Joon Young Chung, Jean Y. Song, Minsuk Chang, Juho Kim 0001 |
CHI | 3 |
| 2021 | Crowdsourcing More Effective Initializations for Single-Target Trackers Through Automatic Re-queryingabstractIn single-target video object tracking, an initial bounding box is drawn around a target object and propagated through a video. When this bounding box is provided by a careful human expert, it is expected to yield strong overall tracking performance that can be mimicked at scale by novice crowd workers with the help of advanced quality control methods. However, we show through an investigation of 900 crowdsourced initializations that such quality control strategies are inadequate for this task in two major ways: first, the high level of redundancy in these methods (e.g., averaging multiple responses to reduce error) is unnecessary, as 23% of crowdsourced initializations perform just as well as the gold-standard initialization. Second, even nearly perfect initializations can lead to degraded long-term performance due to the complexity of object tracking. Considering these findings, we evaluate novel approaches for automatically selecting bounding boxes to re-query, and introduce Smart Replacement, an efficient method that decides whether to use the crowdsourced replacement initialization. Stephan J. Lemmer, Jean Y. Song, Jason J. Corso |
CHI | 2 |
| 2021 | Human-in-the-loop Pose Estimation via Shared AutonomyabstractReliable, efficient shared autonomy requires balancing human operation and robot automation on complex tasks, such as dexterous manipulation. Adding to the difficulty of shared autonomy is a robot’s limited ability to perceive the 6 degree-of-freedom pose of objects, which is essential to perform manipulations those objects afforded. Inspired by Monte Carlo Localization, we propose a generative human-in-the-loop approach to estimating object pose. We characterize the performance of our mixed-initiative 3D registration approach using 2D pointing devices via a user study. Seeking an analog for Fitts’s Law for 3D registration, we introduce a new evaluation framework that takes the entire registration process into account instead of only the outcome. When combined with estimates of registration confidence, we posit that mixed-initiative registration will reduce the human workload while maintaining or even improving final pose estimation accuracy. Zhefan Ye, Jean Y. Song, Zhiqiang Sui, Stephen Hart, Jorge Vilchis, Walter S. Lasecki, Odest Chadwicke Jenkins |
IUI | 2 |
| 2021 | FLUID-XP: flexible user interface distribution for cross-platform experienceabstractBeing able to use a single app across multiple devices can bring novel experiences to the users in various domains including entertainment and productivity. For instance, a user of a video editing app would be able to use a smart pad as a canvas and a smartphone as a remote toolbox so that the toolbox does not occlude the canvas during editing. However, existing approaches do not properly support the single-app multi-device execution due to several limitations, including high development cost, device heterogeneity, and high performance requirement. In this paper, we introduce FLUID-XP, a novel cross-platform multi-device system that enables UIs of a single app to be executed across heterogeneous platforms, while overcoming the limitations of previous approaches. FLUID-XP provides flexible, efficient, and seamless interactions by addressing three main challenges: i) how to transparently enable a single-display app to use multiple displays, ii) how to distribute UIs across heterogeneous devices with minimal network traffic, and iii) how to optimize the UI distribution process when multiple UIs have different distribution requirements. Our experiments with a working prototype of FLUID-XP on Android confirm that FLUID-XP successfully supports a variety of unmodified real-world apps across heterogeneous platforms (Android, iOS, and Linux). We also conduct a lab study with 25 participants to demonstrate the effectiveness of FLUID-XP with real users. Sunjae Lee, Hayeon Lee, Hoyoung Kim, Jeong Woon Choi, Yuseung Lee, Seono Lee, Ahyeon Kim, Jean Y. Song, Sangeun Oh, Steven Y. Ko, Insik Shin |
MobiCom | 9 |
| 2020 | Improving Crowd-Supported GUI Testing with Structural GuidanceabstractCrowd testing is an emerging practice in Graphical User Interface (GUI) testing, where developers recruit a large number of crowd testers to test GUI features. It is often easier and faster than a dedicated quality assurance team, and its output is more realistic than that of automated testing. However, crowds of testers working in parallel tend to focus on a small set of commonly-used User Interface (UI) navigation paths, which can lead to low test coverage and redundant effort. In this paper, we introduce two techniques to increase crowd testers' coverage: interactive event-flow graphs and GUI-level guidance. The interactive event-flow graphs track and aggregate every tester's interactions into a single directed graph that visualizes the cases that have already been explored. Crowd testers can interact with the graphs to find new navigation paths and increase the coverage of the created tests. We also use the graphs to augment the GUI (GUI-level guidance) to help testers avoid only exploring common paths. Our evaluation with 30 crowd testers on 11 different test pages shows that the techniques can help testers avoid redundant effort while also increasing untrained testers' coverage by 55%. These techniques can help us develop more robust software that works in more mission-critical settings not only by performing more thorough testing with the same effort that has been put in before but also by integrating them into different parts of the development pipeline to make more reliable software in the early development stage. Yan Chen 0033, Maulishree Pandey, Jean Y. Song, Walter S. Lasecki, Steve Oney |
CHI | 3 |
| 2020 | C-Reference: Improving 2D to 3D Object Pose Estimation Accuracy via Crowdsourced Joint Object EstimationabstractConverting widely-available 2D images and videos, captured using an RGB camera, to 3D can help accelerate the training of machine learning systems in spatial reasoning domains ranging from in-home assistive robots to augmented reality to autonomous vehicles. However, automating this task is challenging because it requires not only accurately estimating object location and orientation, but also requires knowing currently unknown camera properties (e.g., focal length). A scalable way to combat this problem is to leverage people's spatial understanding of scenes by crowdsourcing visual annotations of 3D object properties. Unfortunately, getting people to directly estimate 3D properties reliably is difficult due to the limitations of image resolution, human motor accuracy, and people's 3D perception (i.e., humans do not "see" depth like a laser range finder). In this paper, we propose a crowd-machine hybrid approach that jointly uses crowds' approximate measurements of multiple in-scene objects to estimate the 3D state of a single target object. Our approach can generate accurate estimates of the target object by combining heterogeneous knowledge from multiple contributors regarding various different objects that share a spatial relationship with the target object. We evaluate our joint object estimation approach with 363 crowd workers and show that our method can reduce errors in the target object's 3D location estimation by over 40%, while requiring only $35$% as much human time. Our work introduces a novel way to enable groups of people with different perspectives and knowledge to achieve more accurate collective performance on challenging visual annotation tasks. Jean Y. Song, John Joon Young Chung, David F. Fouhey, Walter S. Lasecki |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2020 | FourEyes: Leveraging Tool Diversity as a Means to Improve Aggregate Accuracy in CrowdsourcingabstractCrowdsourcing is a common means of collecting image segmentation training data for use in a variety of computer vision applications. However, designing accurate crowd-powered image segmentation systems is challenging, because defining object boundaries in an image requires significant fine motor skills and hand-eye coordination, which makes these tasks error-prone. Typically, special segmentation tools are created and then answers from multiple workers are aggregated to generate more accurate results. However, individual tool designs can bias how and where people make mistakes, resulting in shared errors that remain even after aggregation. In this article, we introduce a novel crowdsourcing approach that leverages tool diversity as a means of improving aggregate crowd performance. Our idea is that given a diverse set of tools, answer aggregation done across tools can help improve the collective performance by offsetting systematic biases induced by the individual tools themselves. To demonstrate the effectiveness of the proposed approach, we design four different tools and present FourEyes, a crowd-powered image segmentation system that uses aggregation across different tools. We then conduct a series of studies that evaluate different aggregation conditions and show that using multiple tools can significantly improve aggregate accuracy. Furthermore, we investigate the idea of applying post-processing for multi-tool aggregation in terms of correction mechanism. We introduce a novel region-based method for synthesizing more accurate bounds for image segmentation tasks through averaging surrounding annotations. In addition, we explore the effect of adjusting the threshold parameter of an EM-based aggregation method. Our results suggest that not only the individual tool’s design, but also the correction mechanism, can affect the performance of multi-tool aggregation. This article extends a work presented at ACM IUI 2018 [46] by providing a novel region-based error-correction method and additional in-depth evaluation of the proposed approach. Jean Y. Song, Raymond Fok, Juho Kim 0001, Walter S. Lasecki |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2019 | Popup: reconstructing 3D video using particle filtering to aggregate crowd responsesabstractCollecting a sufficient amount of 3D training data for autonomous vehicles to handle rare, but critical, traffic events (e.g., collisions) may take decades of deployment. Abundant video data of such events from municipal traffic cameras and video sharing sites (e.g., YouTube) could provide a potential alternative, but generating realistic training data in the form of 3D video reconstructions is a challenging task beyond the current capabilities of computer vision. Crowdsourcing the annotation of necessary information could bridge this gap, but the level of accuracy required to obtain usable reconstructions makes this task nearly impossible for non-experts. In this paper, we propose a novel hybrid intelligence method that combines annotations from workers viewing different instances (video frames) of the same target (3D object), and uses particle filtering to aggregate responses. Our approach can leveraging temporal dependencies between video frames, enabling higher quality through more aggressive filtering. The proposed method results in a 33% reduction in the relative error of position estimation compared to a state-of-the-art baseline. Moreover, our method enables skipping (self-filtering) challenging annotations, reducing the total annotation time for hard-to-annotate frames by 16%. Our approach provides a generalizable means of aggregating more accurate crowd responses in settings where annotation is especially challenging or error-prone. Jean Y. Song, Stephan J. Lemmer, Michael Xieyang Liu, Shiyan Yan, Juho Kim 0001, Jason J. Corso, Walter S. Lasecki |
IUI | 1 |
| 2019 | Efficient Elicitation Approaches to Estimate Collective Crowd AnswersabstractWhen crowdsourcing the creation of machine learning datasets, statistical distributions that capture diverse answers can represent ambiguous data better than a single best answer. Unfortunately, collecting distributions is expensive because a large number of responses need to be collected to form a stable distribution. Despite this, the efficient collection of answer distributions-that is, ways to use less human effort to collect estimates of the eventual distribution that would be formed by a large group of responses-is an under-studied topic. In this paper, we demonstrate that this type of estimation is possible and characterize different elicitation approaches to guide the development of future systems. We investigate eight elicitation approaches along two dimensions: annotation granularity and estimation perspective. Annotation granularity is varied by annotating i) a single "best" label, ii) all relevant labels, iii) a ranking of all relevant labels, or iv) real-valued weights for all relevant labels. Estimation perspective is varied by prompting workers to either respond with their own answer or an estimate of the answer(s) that they expect other workers would provide. Our study collected ordinal annotations on the emotional valence of facial images from 1,960 crowd workers and found that, surprisingly, the most fine-grained elicitation methods were not the most accurate, despite workers spending more time to provide answers. Instead, the most efficient approach was to ask workers to choose all relevant classes that others would have selected. This resulted in a 21.4% reduction in the human time required to reach the same performance as the baseline (i.e., selecting a single answer with their own perspective). By analyzing cases in which finer-grained annotations degraded performance, we contribute to a better understanding of the trade-offs between answer elicitation approaches. Our work makes it more tractable to use answer distributions in large-scale tasks such as ML training, and aims to spark future work on techniques that can efficiently estimate answer distributions. John Joon Young Chung, Jean Y. Song, Sindhu Kutty, Sungsoo Ray Hong, Juho Kim 0001, Walter S. Lasecki |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2018 | Two Tools are Better Than One: Tool Diversity as a Means of Improving Aggregate Crowd PerformanceabstractCrowdsourcing is a common means of collecting image segmentation training data for use in a variety of computer vision applications. However, designing accurate crowd-powered image segmentation systems is challenging because defining object boundaries in an image requires significant fine motor skills and hand-eye coordination, which makes these tasks error-prone. Typically, special segmentation tools are created and then answers from multiple workers are aggregated to generate more accurate results. However, individual tool designs can bias how and where people make mistakes, resulting in shared errors that remain even after aggregation. In this paper, we introduce a novel crowdsourcing workflow that leverages multiple tools for the same task to increase output accuracy by reducing systematic error biases introduced by the tools themselves. When a task can no longer be broken down into more-tractable subtasks (the conventional approach taken by microtask crowdsourcing), our multi-tool approach can be used to further improve accuracy by assigning different tools to different workers. We present a series of studies that evaluate our multi-tool approach and show that it can significantly improve aggregate accuracy in semantic image segmentation. Jean Y. Song, Raymond Fok, Alan Lundgard, Juho Kim 0001, Walter S. Lasecki |
IUI | 1 |