Sungeun An

dblp:150/1886 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Finding the Conversation: A Method for Scoring Documents for Natural Conversation Content
Robert J. Moore, Sungeun An, Jay Pankaj Gala, Divyesh Jadav
CHI2
2024 Adversarially Exploring Vulnerabilities in LLMs to Evaluate Social Biases
abstract
Generative AI has caused a paradigm shift in the area of Artificial Intelligence (AI) and as such has inspired much new research, especially on Large Language Models (LLMs). LLMs are transforming how people interact with computers in service-oriented fields in both the consumer (for example: retail, travel, education, healthcare) and enterprise (customer care, field service, sales, marketing, etc.) spaces. One barrier to widespread adoption is the current unpredictability of LLM behavior: users must trust that LLM-based services and systems are accurate, fair, and unbiased. Model responses that exhibit biases related to race, social status, and other sensitive topics can have serious consequences, ranging from lack of trust in the model to adverse social implications for consumers, all the way to damage to the reputations of the corporations that provide them. This study explores how to uncover biases related to social stigmas in LLM output, by using an adversarial prompt-based approach. Discovering model vulnerabilities of this type is a nontrivial task due to the large search space, making it resource-intensive. We present an evaluation framework for probing and analyzing the behaviors of multiple LLMs systematically. We use a curated set of adversarial prompts with a focus on uncovering biased responses to prompts associated with social attributes.
Yuya Jeremy Ong, Jay Pankaj Gala, Sungeun An, Robert J. Moore, Divyesh Jadav
IEEE Big Data3
2024 Data-Prep-Kit: getting your data ready for LLM application development
abstract
Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible open-source data preparation toolkit called Data Prep Kit (DPK). DPK is architected and designed to enable users to scale their data preparation to their needs. With DPK they can prepare data on a local machine or effortlessly scale to run on a cluster with thousands of CPU Cores. DPK comes with a highly scalable, yet extensible set of modules that transform natural language and code data. If the user needs additional transforms, they can be easily developed using extensive DPK support for transform creation. These modules can be used independently or pipelined to perform a series of operations. In this paper, we describe DPK architecture and show its performance from a small scale to a very large number of CPUs. The modules from DPK have been used for the preparation of Granite Models [1] [2]. We believe DPK is a valuable contribution to the AI community to easily prepare data to enhance the performance of their LLM models or to fine-tune models with Retrieval-Augmented Generation (RAG).
Boris Lublinsky, Alexy Roytman, Shivdeep Singh, Constantin Adam, Abdulhamid Adebayo, Sungeun An, Yuan Chi Chang, Xuan-Hong Dang, Nirmit Desai, Michele Dolfi, Hajar Emami-Gohari, Revital Eres, Takuya Goto, Dhiraj Joshi, Yan Koyfman, Mohammad Nassar, Hima Patel, Paramesvaran Selvam, Syed Yousaf Shah, Saptha Surendran, Daiki Tsuzuku, Petros Zerfos, Shahrokh Daijavad
IEEE Big Data7
2024 Understanding is a Two-Way Street: User-Initiated Repair on Agent Responses and Hearing in Conversational Interfaces
abstract
Although methods for repairing prior turns in natural conversation are critical for enabling mutual understanding, or successful communication, these methods are seldom built into conversational user interfaces systematically. Chatbots and voice assistants tend to ask users to paraphrase what they said if it was not understood, but users cannot do the same if they encounter trouble in understanding what the agent said. Understanding is a one-way street in most (intent-based) conversation-like interfaces. An exception to this is Moore and Arar (2019), who demonstrate nine types of user-initiated repair on agent responses that are common in natural conversation and who have shown that users will employ these repair features correctly in text-based interfaces if taught. In this small-scale study, we test these user-initiated repairs (in second position) in a voice-based interface. With understanding-oriented repairs, we found that participants employed them much the same way in text and voice. In addition, we examine some hearing- and speaking-oriented repairs that emerged from the use of our novel multi-modal interface. We found that participants used them to manage troubles specific to the voice modality. Analysis of user logs and transcripts suggests that user-initiated repair features are valuable components of conversational interfaces.
Robert J. Moore, Sungeun An, Olivia H. Marrese
Proc. ACM Hum. Comput. Interact.2
2023 A Scalable Architecture for Conducting A/B Experiments in Educational Settings
abstract
A/B experiments are commonly used in research to compare the effects of changing one or more variables in two different experimental groups-a control group and a treatment group. While the benefits of using A/B experiments are widely known and accepted in education, there is less agreement on an approach to creating software infrastructure systems to assist in rapidly conducting such experiments in the field. To assist in alleviating this gap, we are creating a software infrastructure for A/B experiments that allows researchers to conduct experiments and automatically analyze their results for an education-focused ecology-based conceptual modeling platform.
Andrew Hornback, Stephen Buckley, John Kos, Scott Bunin, Sungeun An, David A. Joyner, Ashok K. Goel 0001
L@S5
2023 The IBM natural conversation framework: a new paradigm for conversational UX design
abstract
User interfaces that take human conversation as their interaction metaphor work fundamentally differently than those that employ spatial metaphors, such as a desktop or a page. While the fundamenta...
Robert J. Moore, Sungeun An
Hum. Comput. Interact.2
2022 Effects of Guidance on Learning About Ill-defined Problems
Sungeun An, Emily Weigel, Ashok K. Goel 0001
ITS1
2021 Cognitive Strategies for Parameter Estimation in Model Exploration
Sungeun An, Spencer Rugaber, Emily Weigel, Ashok K. Goel 0001
CogSci1
2021 Recognizing Novice Learner's Modeling Behaviors
Sungeun An, William Broniec, Spencer Rugaber, Emily Weigel, Jennifer Hammock, Ashok K. Goel 0001
ITS1
2020 Scientific Modeling Using Large Scale Knowledge
Sungeun An, Robert Bates, Jennifer Hammock, Spencer Rugaber, Emily Weigel, Ashok K. Goel 0001
AIED (2)1
2019 Learning by doing: Supporting experimentation in inquiry-based modeling
Sungeun An, Robert Bates, Jennifer Hammock, Spencer Rugaber, Emily Weigel, Ashok K. Goel 0001
CogSci1
2018 VERA: Popularizing Science Through AI
Sungeun An, Robert Bates, Jennifer Hammock, Spencer Rugaber, Ashok K. Goel 0001
AIED (2)1