Tong Steven Sun

dblp:323/4248 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7298-0333ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 De-Decay: Defusing Computer Vision Model Degradation through Scalable and Actionable Human-Data Alignment
abstract
Computer Vision (CV) models can become outdated after deployment as real-world data evolves, requiring intensive attention from AI engineers to address degraded performance through tasks like data relabeling to update models with new human perceptions. Interactive human-in-the-loop systems have considerable potential to enhance model-steering practices. However, such workflows reveal two challenges: (1) scalability, where labor demands increase with data size, and (2) actionability, where human insights do not readily transform into model revisions. Based on our formative study (S1) on the current challenges faced by CV professionals, we developed De-Decay, an end-to-end Human-Data Alignment system offering scalable label-less assessment and actionable insight transformation . This enables engineers to investigate degradation and auto-retrain models with AI support, such as image clustering and regeneration. Our summative study (S2) showed that De-Decay helped engineers effectively identify and address CV degradation. We discuss how future research can enhance scalability and actionability in AI evaluation systems for aligning AI behaviors with human mental models.
Tong Steven Sun, Huining Feng, Jinwei Ye, Sangdoo Yun, Young-Ho Kim, Sungsoo Ray Hong
ACM Trans. Interact. Intell. Syst.1
2024 ShadowMagic: Designing Human-AI Collaborative Support for Comic Professionals' Shadowing
abstract
Shadowing allows artists to convey realistic volume and emotion of characters in comic colorization. While AI technologies have the potential to improve professionals’ shadowing experience, current practice is manual and time-consuming. To understand how we can improve their shadowing experience, we conducted interviews with 5 professionals. We found that professionals’ level of engagement can vary depending on semantics, such as characters’ faces or hair. We also found they spent time on shadow “landscaping”—deciding where to put big shadow regions to make a realistic volumetric presentation—while the final results can dramatically vary depending on their “staging” and “attention guiding” needs. We found they would accept AI suggestions for less engaging semantic parts or landscaping, while they would need to have the capability to adjust details. Based on our observations, we built ShadowMagic that (1) generates AI-driven shadows based on typically used light directions, (2) enables a user to selectively choose the results depending on the semantics, and (3) allows users to finish shadow areas by themselves for further perfection. Through a summative evaluation with 5 professionals, we found that they were significantly more satisfied with our AI-driven results than a baseline. We also found ShadowMagic’s “step by step” workflow helps participants more easily adopt AI-driven results. We conclude by providing implications.
Amrita Ganguly, Chuan Yan, John Joon Young Chung, Tong Steven Sun, Yoon Kiheon, Yotam I. Gingold, Sungsoo Ray Hong
UIST4
2024 3DPFIX: Improving Remote Novices' 3D Printing Troubleshooting through Human-AI Collaboration Design
abstract
The widespread consumer-grade 3D printers and learning resources online enable novices to self-train in remote settings. While troubleshooting plays an essential part of 3D printing, the process remains challenging for many remote novices even with the help of well-developed online sources, such as online troubleshooting archives and online community help. We conducted a formative study with 76 active 3D printing users to learn how remote novices leverage online resources in troubleshooting and their challenges. We found that remote novices cannot fully utilize online resources. For example, the online archives statically provide general information, making it hard to search and relate their unique cases with existing descriptions. Online communities can potentially ease their struggles by providing more targeted suggestions, but a helper who can provide custom help is rather scarce, making it hard to obtain timely assistance. We propose 3DPFIX, an interactive 3D troubleshooting system powered by the pipeline to facilitate Human-AI Collaboration, designed to improve novices' 3D printing experiences and thus help them easily accumulate their domain knowledge. We built 3DPFIX that supports automated diagnosis and solution-seeking. 3DPFIX was built upon shared dialogues about failure cases from Q&A discourses accumulated in online communities. We leverage social annotations (i.e., comments) to build an annotated failure image dataset for AI classifiers and extract a solution pool. Our summative study revealed that using 3DPFIX helped participants spend significantly less effort in diagnosing failures and finding a more accurate solution than relying on their common practice. We also found that 3DPFIX users learn about 3D printing domain-specific knowledge. We discuss the implications of leveraging community-driven data in developing future Human-AI Collaboration designs.
Nahyun Kwon, Tong Steven Sun, Liang Zhao 0002, Xu Wang 0016, Jeeeun Kim, Sungsoo Ray Hong
Proc. ACM Hum. Comput. Interact.2
2023 Designing a Direct Feedback Loop between Humans and Convolutional Neural Networks through Local Explanations
abstract
The local explanation provides heatmaps on images to explain how Convolutional Neural Networks (CNNs) derive their output. Due to its visual straightforwardness, the method has been one of the most popular explainable AI (XAI) methods for diagnosing CNNs. Through our formative study (S1), however, we captured ML engineers' ambivalent perspective about the local explanation as a valuable and indispensable envision in building CNNs versus the process that exhausts them due to the heuristic nature of detecting vulnerability. Moreover, steering the CNNs based on the vulnerability learned from the diagnosis seemed highly challenging. To mitigate the gap, we designed DeepFuse, the first interactive design that realizes the direct feedback loop between a user and CNNs in diagnosing and revising CNN's vulnerability using local explanations. DeepFuse helps CNN engineers to systemically search "unreasonable" local explanations and annotate the new boundaries for those identified as unreasonable in a labor-efficient manner. Next, it steers the model based on the given annotation such that the model doesn't introduce similar mistakes. We conducted a two-day study (S2) with 12 experienced CNN engineers. Using DeepFuse, participants made a more accurate and "reasonable" model than the current state-of-the-art. Also, participants found the way DeepFuse guides case-based reasoning can practically improve their current practice. We provide implications for design that explain how future HCI-driven design can move our practice forward to make XAI-driven insights more actionable.
Tong Steven Sun, Shubham Khaladkar, Sijia Liu 0001, Liang Zhao 0002, Young-Ho Kim, Sungsoo Ray Hong
Proc. ACM Hum. Comput. Interact.1
2022 RES: A Robust Framework for Guiding Visual Explanation
abstract
Despite the fast progress of explanation techniques in modern Deep Neural Networks (DNNs) where the main focus is handling "how to generate the explanations", advanced research questions that examine the quality of the explanation itself (e.g., "whether the explanations are accurate") and improve the explanation quality (e.g., "how to adjust the model to generate more accurate explanations when explanations are inaccurate") are still relatively under-explored. To guide the model toward better explanations, techniques in explanation supervision - which add supervision signals on the model explanation - have started to show promising effects on improving both the generalizability as and intrinsic interpretability of Deep Neural Networks. However, the research on supervising explanations, especially in vision-based applications represented through saliency maps, is in its early stage due to several inherent challenges: 1) inaccuracy of the human explanation annotation boundary, 2) incompleteness of the human explanation annotation region, and 3) inconsistency of the data distribution between human annotation and model explanation maps. To address the challenges, we propose a generic RES framework for guiding visual explanation by developing a novel objective that handles inaccurate boundary, incomplete region, and inconsistent distribution of human annotations, with a theoretical justification on model generalizability. Extensive experiments on two real-world image datasets demonstrate the effectiveness of the proposed framework on enhancing both the reasonability of the explanation and the performance of the backbone DNNs model.
Tong Steven Sun, Guangji Bai, Siyi Gu, Sungsoo Ray Hong, Liang Zhao 0002
KDD2
2022 Aligning Eyes between Humans and Deep Neural Network through Interactive Attention Alignment
abstract
While Deep Neural Networks (DNNs) are deriving the major innovations through their powerful automation, we are also witnessing the peril behind automation as a form of bias, such as automated racism, gender bias, and adversarial bias. As the societal impact of DNNs grows, finding an effective way to steer DNNs to align their behavior with the human mental model has become indispensable in realizing fair and accountable models. While establishing the way to adjust DNNs to "think like humans'' is in pressing need, there have been few approaches aiming to capture how "humans would think'' when DNNs introduce biased reasoning in seeing a new instance. We propose Interactive Attention Alignment (IAA), a framework that uses the methods for visualizing model attention, such as saliency maps, as an interactive medium that humans can leverage to unveil the cases of DNN's biased reasoning and directly adjust the attention. To realize more effective human-steerable DNNs than state-of-the-art, IAA introduces two novel devices. First, IAA uses Reasonability Matrix to systematically identify and adjust the cases of biased attention. Second, IAA applies GRADIA, a computational pipeline designed for effectively applying the adjusted attention to jointly maximize attention quality and prediction accuracy. We evaluated Reasonability Matrix in Study 1 and GRADIA in Study 2 in the gender classification problem. In Study 1, we found applying Reasonability Matrix in bias detection can significantly improve the perceived quality of model attention from human eyes than not applying Reasonability Matrix. In Study 2, we found using GRADIA significantly improves (1) the human-assessed perceived quality of model attention and (2) model performance in scenarios where the training samples are limited. Based on our observation in the two studies, we present implications for future design in the problem space of social computing and interactive data annotation toward achieving a human-centered steerable AI.
Tong Steven Sun, Liang Zhao 0002, Sungsoo Ray Hong
Proc. ACM Hum. Comput. Interact.2
2021 GNES: Learning to Explain Graph Neural Networks
abstract
In recent years, graph neural networks (GNNs) and the research on their explainability are experiencing rapid developments and achieving significant progress. Many methods are proposed to explain the predictions of GNNs, focusing on “how to generate explanations” However, research questions like “whether the GNN explanations are inaccurate”, “what if the explanations are inaccurate”, and “how to adjust the model to generate more accurate explanations” have not been well explored. To address the above questions, this paper proposes a GNN Explanation Supervision (GNES)1framework to adaptively learn how to explain GNNs more correctly. Specifically, our framework jointly optimizes both model prediction and model explanation by enforcing both whole graph regularization and weak supervision on model explanations. For the graph regularization, we propose a unified explanation formulation for both node-level and edge-level explanations by enforcing the consistency between them. The node- and edge-level explanation techniques we propose are also generic and rigorously demonstrated to cover several existing major explainers as special cases. Extensive experiments on five real-world datasets across two application domains demonstrate the effectiveness of the proposed model on improving the reasonability of the explanation while still keep or even improve the backbone GNNs model performance.1Code available at: https://github.com/YuyangGao/GNES.
Tong Steven Sun, Rishab Bhatt, Dazhou Yu, Sungsoo Ray Hong, Liang Zhao 0002
ICDM2