Seongmin Lee 0007

dblp:317/5565 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-1950-5004ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation
abstract
The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However, its complexity makes it difficult for non-experts to learn. Existing resources often lack interactivity, rely on static descriptions of simplified architectures, or fail to reflect models’ behavior with real data. To address this gap, we introduce Transformer Explainer, an interactive visualization tool for non-experts to learn Transformers. The tool integrates an overview illustrating the Transformer’s data flow with on-demand explanations that gradually reveal mathematical details. Smooth transitions across abstraction levels highlight the interplay between high-level structures and low-level operations. Running a live GPT-2 instance directly in the browser, Transformer Explainer empowers learners to experiment with custom input and hyperparameters without setup, observing next-token predictions in real time. A 90-participant user study showed that our tool offered significant advantages in improving user understanding and engagement. Transformer Explainer has attracted over 490,000 users.
Aeree Cho, Grace C. Kim, Alexander Karpekov, Seongmin Lee 0007, Alec Helbling, Benjamin Hoover, Zijie J. Wang, Minsuk Kahng, Polo Chau
CHI4
2025 LLM Attributor: Interactive Visual Attribution for LLM Generation
abstract
While large language models (LLMs) have shown remarkable capability to generate convincing text across diverse domains, concerns around its potential risks have highlighted the importance of understanding the rationale behind text generation. We present LLM ATTRIBUTOR, a Python library that provides interactive visualizations for training data attribution of an LLM’s text generation. Our library offers a new way to quickly attribute an LLM’s text generation to training data points to inspect model behaviors, enhance its trustworthiness, and compare model-generated text with user-provided text. Thanks to LLM ATTRIBUTOR’s broad support for computational notebooks, users can easily integrate it into their workflow to interactively visualize attributions of their models.
Seongmin Lee 0007, Zijie J. Wang, Aishwarya Chakravarthy, Alec Helbling, Shengyun Peng, Mansi Phute, Polo Chau, Minsuk Kahng
AAAI1
2025 TRANSFORMER EXPLAINER: Interactive Learning of Text-Generative Models
abstract
Transformers have revolutionized machine learning, yet their inner workings remain opaque to many. We present TRANSFORMER EXPLAINER, an interactive visualization tool designed for non-experts to learn about Transformers through the GPT-2 model. Our tool helps users understand complex Transformer concepts by integrating a model overview and smooth transitions across abstraction levels of math operations and model structures. It runs a live GPT-2 model locally in the user’s browser, empowering users to experiment with their own input and observe in real-time how the internal components and parameters of the Transformer work together to predict the next tokens. 125,000 users have used our open-source tool at https://poloclub.github.io/ transformer-explainer/.
Aeree Cho, Grace C. Kim, Alexander Karpekov, Alec Helbling, Zijie J. Wang, Seongmin Lee 0007, Benjamin Hoover, Polo Chau
AAAI6
2025 Probing LLM Hallucination from Within: Perturbation-Driven Method via Internal Knowledge
Seongmin Lee 0007, Hsiang Hsu, Chun-Fu Chen 0001, Polo Chau
IEEE Big Data1
2025 Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
abstract
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical.Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connections with safety are often overlooked in prior surveys.We present the first survey that bridges this gap, introducing a unified framework that connects safety-focused interpretation methods, the safety enhancements they inform, and the tools that operationalize them.Our novel taxonomy, organized by LLM workflow stages, summarizes nearly 70 works at their intersections.We conclude with open challenges and future directions.This timely survey helps researchers and practitioners navigate key advancements for safer, more interpretable LLMs.
Seongmin Lee 0007, Aeree Cho, Grace C. Kim, Shengyun Peng, Mansi Phute, Polo Chau
EMNLP1
2025 Shape it Up! Restoring LLM Safety during Finetuning
abstract
Finetuning large language models (LLMs) enables user-specific customization but introduces important safety risks: even a few harmful examples can compromise safety alignment. A common mitigation strategy is to update the model more strongly on examples deemed safe, while downweighting or excluding those flagged as unsafe. However, because safety context can shift within a single example, updating the model equally on both harmful and harmless parts of a response is suboptimal — an atomic treatment we term static safety shaping. In contrast, we propose dynamic safety shaping (DSS), a dynamic shaping framework that uses fine-grained safety signals to reinforce learning from safe segments of a response while suppressing unsafe content. To enable such fine-grained control during finetuning, we introduce a key insight: guardrail models, traditionally used for filtering, can be repurposed to evaluate partial responses, tracking how safety risk evolves throughout the response, segment by segment. This leads to the Safety Trajectory Assessment of Response (STAR), a token-level signal that enables shaping to operate dynamically over the training sequence. Building on this, we present ★DSS, a DSS method guided by STAR scores that robustly mitigates finetuning risks and delivers substantial safety improvements across diverse threats, datasets, and model families, all without compromising capability on intended tasks. We encourage future safety research to build on dynamic shaping principles for stronger mitigation against evolving finetuning risks. Our code is publicly available at https://github.com/poloclub/star-dss
Shengyun Peng, Jianfeng Chi, Seongmin Lee 0007, Polo Chau
NeurIPS4
2024 Effective Guidance for Model Attention with Simple Yes-no Annotations
abstract
Modern deep learning models often make predictions by focusing on irrelevant areas, leading to biased performance and limited generalization. Existing methods aimed at rectifying model attention require explicit labels for irrelevant areas or complex pixel-wise ground truth attention maps. We present Crayon (Correcting Reasoning with Annotations of Yes Or No), offering effective, scalable, and practical solutions to rectify model attention using simple yes-no annotations. Crayon empowers classical and modern model interpretation techniques to identify and guide model reasoning: Crayon-Attention directs classic interpretations based on saliency maps to focus on relevant image regions, while Crayon-Pruning removes irrelevant neurons identified by modern concept-based methods to mitigate their influence. Through extensive experiments with both quantitative and human evaluation, we showcase Crayon’s effectiveness, scalability, and practicality in refining model attention. Crayon achieves state-of-the-art performance, outperforming 12 methods across 3 benchmark datasets, surpassing approaches that require more complex annotations.
Seongmin Lee 0007, Ali Payani, Polo Chau
IEEE Big Data1
2024 Interactive Visual Learning for Stable Diffusion
Seongmin Lee 0007, Benjamin Hoover, Hendrik Strobelt, Zijie J. Wang, Shengyun Peng, Austin P. Wright, Haekyu Park, Haoyang Yang, Polo Chau
IJCAI1
2024 Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
abstract
Diffusion-based generative models’ impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion’s complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/.
Seongmin Lee 0007, Benjamin Hoover, Hendrik Strobelt, Zijie J. Wang, Shengyun Peng, Austin P. Wright, Haekyu Park, Haoyang Yang, Polo Chau
IEEE VIS1
2024 VG: Automatic Grading of D3 Visualizations
abstract
Manually grading D3 data visualizations is a challenging endeavor, and is especially difficult for large classes with hundreds of students. Grading an interactive visualization requires a combination of interactive, quantitative, and qualitative evaluation that are conventionally done manually and are difficult to scale up as the visualization complexity, data size, and number of students increase. We present VISGRADER, a first-of-its kind automatic grading method for D3 visualizations that scalably and precisely evaluates the data bindings, visual encodings, interactions, and design specifications used in a visualization. Our method enhances students' learning experience, enabling them to submit their code frequently and receive rapid feedback to better inform iteration and improvement to their code and visualization design. We have successfully deployed our method and auto-graded D3 submissions from more than 4000 students in a visualization course at Georgia Tech, and received positive feedback for expanding its adoption.
Matthew Hull, Vivian Pednekar, Hannah Murray, Nimisha Roy, Emmanuel Tung, Susanta Routray, Connor Guerin, Zijie J. Wang, Seongmin Lee 0007, Max Mahdi Roozbahani, Polo Chau
IEEE Trans. Vis. Comput. Graph.10
2023 Concept Evolution in Deep Learning Training: A Unified Interpretation Framework and Discoveries
abstract
We present ConceptEvo, a unified interpretation framework for deep neural networks (DNNs) that reveals the inception and evolution of learned concepts during training. Our work addresses a critical gap in DNN interpretation research, as existing methods primarily focus on post-training interpretation. ConceptEvo introduces two novel technical contributions: (1) an algorithm that generates a unified semantic space, enabling side-by-side comparison of different models during training, and (2) an algorithm that discovers and quantifies important concept evolutions for class predictions. Through a large-scale human evaluation and quantitative experiments, we demonstrate that ConceptEvo successfully identifies concept evolutions across different models, which are not only comprehensible to humans but also crucial for class predictions. ConceptEvo is applicable to both modern DNN architectures, such as ConvNeXt, and classic DNNs, such as VGGs and InceptionV3.
Haekyu Park, Seongmin Lee 0007, Benjamin Hoover, Austin P. Wright, Omar Shaikh, Rahul Duggal, Nilaksh Das, Judy Hoffman, Polo Chau
CIKM2
2022 VIsCUIT: Visual Auditor for Bias in CNN Image Classifier
abstract
CNN image classifiers are widely used, thanks to their efficiency and accuracy. However, they can suffer from biases that impede their practical applications. Most existing bias investigation techniques are either inapplicable to general image classification tasks or require significant user efforts in perusing all data subgroups to manually specify which data attributes to inspect. We present VIsCUIT, an interactive visualization system that reveals how and why a CNN classifier is biased. VIsCUIT visually summarizes the subgroups on which the classifier underperforms and helps users discover and characterize the cause of the underperformances by revealing image concepts responsible for activating neurons that contribute to misclassifications. VIsCUIT runs in modern browsers and is opensource, allowing people to easily access and extend the tool to other model architectures and datasets. VIsCUIT is available at the following public demo link: https://poloclub.github.io/VisCUIT. A video demo is available at https://youtu.be/eNDbSyM4R_4.
Seongmin Lee 0007, Judy Hoffman, Zijie J. Wang, Polo Chau
CVPR1