EDBT 2026 Demo / reviewers in the wild / expert
Joon Sung Park 0001
dblp:208/8975-1
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-5036-4409ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 14 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Finetuning LLMs for Human Behavior Prediction in Social Science ExperimentsabstractLarge language models (LLMs) offer a powerful opportunity to simulate the results of social science experiments.In this work, we demonstrate that finetuning LLMs directly on individual-level responses from past experiments meaningfully improves the accuracy of such simulations across diverse social science domains.We construct SOCSCI210 via an automatic pipeline, a dataset comprising 2.9 million responses from 400,491 participants in 210 open-source social science experiments.Through finetuning, we achieve multiple levels of generalization.In completely unseen studies, our strongest model, SOCRATES-QWEN-14B, produces predictions that are 26% more aligned with distributions of human responses to diverse outcome questions under varying conditions relative to its base model (Qwen2.5-14B),outperforming GPT-4o by 13%.By finetuning on a subset of conditions in a study, generalization to new unseen conditions is particularly robust, improving by 71%.Since SOCSCI210 contains rich demographic information, we reduce demographic parity difference, a measure of bias, by 10.6% through finetuning.Because social sciences routinely generate rich, topicspecific datasets, our findings indicate that finetuning on such data could enable more accurate simulations for experimental hypothesis screening.We release our data, models and finetuning code at stanfordhci.github.io/socrates. Akaash Kolluri, Shengguang Wu, Joon Sung Park 0001, Michael S. Bernstein |
EMNLP | 3 |
| 2025 | Creating General User Models from Computer Use
Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park 0001, Diyi Yang, Michael S. Bernstein |
UIST | 5 |
| 2023 | Holistic Evaluation of Text-to-Image ModelsabstractThe stunning qualitative improvement of text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on image-text alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/latest and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase Michihiro Yasunaga, Chenlin Meng, Yifan Mai 0001, Joon Sung Park 0001, Agrim Gupta, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei 0001, Jiajun Wu 0001, Stefano Ermon, Percy Liang |
NeurIPS | 5 |
| 2023 | Generative Agents: Interactive Simulacra of Human BehaviorabstractBelievable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents: computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent’s experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty-five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors. For example, starting with only a single user-specified notion that one agent wants to throw a Valentine’s Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture—observation, planning, and reflection—each contribute critically to the believability of agent behavior. By fusing large language models with computational interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior. Joon Sung Park 0001, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein |
UIST | 1 |
| 2023 | Working With AI to Persuade: Examining a Large Language Model's Ability to Generate Pro-Vaccination MessagesabstractArtificial Intelligence (AI) is a transformative force in communication and messaging strategy, with potential to disrupt traditional approaches. Large language models (LLMs), a form of AI, are capable of generating high-quality, humanlike text. We investigate the persuasive quality of AI-generated messages to understand how AI could impact public health messaging. Specifically, through a series of studies designed to characterize and evaluate generative AI in developing public health messages, we analyze COVID-19 pro-vaccination messages generated by GPT-3, a state-of-the-art instantiation of a large language model. Study 1 is a systematic evaluation of GPT-3's ability to generate pro-vaccination messages. Study 2 then observed peoples' perceptions of curated GPT-3-generated messages compared to human-authored messages released by the CDC (Centers for Disease Control and Prevention), finding that GPT-3 messages were perceived as more effective, stronger arguments, and evoked more positive attitudes than CDC messages. Finally, Study 3 assessed the role of source labels on perceived quality, finding that while participants preferred AI-generated messages, they expressed dispreference for messages that were labeled as AI-generated. The results suggest that, with human supervision, AI can be used to create effective public health messages, but that individuals prefer their public health messages to come from human institutions rather than AI sources. We propose best practices for assessing generative outputs of large language models in future social science research and ways health professionals can use AI systems to augment public health messaging. Elise Karinshak, Sunny Xun Liu, Joon Sung Park 0001, Jeffrey T. Hancock |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsabstractWhose labels should a machine learning (ML) algorithm learn to emulate? For ML tasks ranging from online comment toxicity to misinformation detection to medical diagnosis, different groups in society may have irreconcilable disagreements about ground truth labels. Supervised ML today resolves these label disagreements implicitly using majority vote, which overrides minority groups’ labels. We introduce jury learning, a supervised ML approach that resolves these disagreements explicitly through the metaphor of a jury: defining which people or groups, in what proportion, determine the classifier’s prediction. For example, a jury learning model for online toxicity might centrally feature women and Black jurors, who are commonly targets of online harassment. To enable jury learning, we contribute a deep learning architecture that models every annotator in a dataset, samples from annotators’ models to populate the jury, then runs inference to classify. Our architecture enables juries that dynamically adapt their composition, explore counterfactuals, and visualize dissent. A field evaluation finds that practitioners construct diverse juries that alter 14% of classification outcomes. Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park 0001, Kayur Patel, Jeffrey T. Hancock, Tatsunori B. Hashimoto, Michael S. Bernstein |
CHI | 3 |
| 2022 | Social Simulacra: Creating Populated Prototypes for Social Computing SystemsabstractSocial computing prototypes probe the social behaviors that may arise in an envisioned system design. This prototyping practice is currently limited to recruiting small groups of people. Unfortunately, many challenges do not arise until a system is populated at a larger scale. Can a designer understand how a social system might behave when populated, and make adjustments to the design before the system falls prey to such challenges? We introduce social simulacra, a prototyping technique that generates a breadth of realistic social interactions that may emerge when a social computing system is populated. Social simulacra take as input the designer’s description of a community’s design—goal, rules, and member personas—and produce as output an instance of that design with simulated behavior, including posts, replies, and anti-social behaviors. We demonstrate that social simulacra shift the behaviors that they generate appropriately in response to design changes, and that they enable exploration of “what if?” scenarios where community members or moderators intervene. To power social simulacra, we contribute techniques for prompting a large language model to generate thousands of distinct community members and their social interactions with each other; these techniques are enabled by the observation that large language models’ training data already includes a wide variety of positive and negative behavior on social media platforms. In evaluations, we show that participants are often unable to distinguish social simulacra from actual community behavior and that social computing designers successfully refine their social computing designs when using social simulacra. Joon Sung Park 0001, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein |
UIST | 1 |
| 2022 | Power Dynamics and Value Conflicts in Designing and Maintaining Socio-Technical Algorithmic ProcessesabstractHow do power dynamics and value conflicts affect our ability to design and maintain socio-technical algorithmic processes? In this paper, we study the SIGCHI student volunteer (SV) selection process that uses a weighted semi-randomized algorithm to recruit a desired pool of volunteers. Our interviews with the community members showed that the process is complex and socio-technical; the algorithm's outputs are interpreted and adjusted by the conference organizers to reflect the community values while ensuring the selection of effective volunteers to help with organizing the conference. This provides a stage in which the power dynamics and value conflicts among the stakeholders play salient roles in determining how the process was perceived and envisioned. For instance, non-organizers of the conference found the algorithm used in the selection process to be a power-balancer that places a check on the organizers who oversee the process. However, even with a participatory process to elicit the algorithm's weights, the power dynamics and value conflicts between the participants made it difficult to reach a consensus on what the SV selection process should consider and prioritize. Our findings highlight the importance of value transparency -- the type of transparency that focuses on explaining why a decision was made rather than how it was made -- as a mechanism for resolving such conflicts. Based on our findings, we lay out design recommendations that can guide communities to better design and maintain algorithmic socio-technical processes over time in the face of power dynamics and value conflicts. Joon Sung Park 0001, Karrie Karahalios, Niloufar Salehi, Motahhare Eslami |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | Measuring the Prevalence of Anti-Social Behavior in Online CommunitiesabstractWith increasing attention to online anti-social behaviors such as personal attacks and bigotry, it is critical to have an accurate accounting of how widespread anti-social behaviors are. In this paper, we empirically measure the prevalence of anti-social behavior in one of the world's most popular online community platforms. We operationalize this goal as measuring the proportion of unmoderated comments in the 97 most popular communities on Reddit that violate eight widely accepted platform norms. To achieve this goal, we contribute a human-AI pipeline for identifying these violations and a bootstrap sampling method to quantify measurement uncertainty. We find that 6.25% (95% Confidence Interval [5.36%, 7.13%]) of all comments in 2016, and 4.28% (95% CI [2.50%, 6.26%]) in 2020, are violations of these norms. Most anti-social behaviors remain unmoderated: moderators only removed one in twenty violating comments in 2016, and one in ten violating comments in 2020. Personal attacks were the most prevalent category of norm violation; pornography and bigotry were the most likely to be moderated, while politically inflammatory comments and misogyny/vulgarity were the least likely to be moderated. This paper offers a method and set of empirical results for tracking these phenomena as both the social practices (e.g., moderation) and technical practices (e.g., design) evolve. Joon Sung Park 0001, Joseph Seering, Michael S. Bernstein |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Understanding the Representation and Representativeness of Age in AI Data SetsabstractA diverse representation of different demographic groups in AI training data sets is important in ensuring that the models will work for a large range of users. To this end, recent efforts in AI fairness and inclusion have advocated for creating AI data sets that are well-balanced across race, gender, socioeconomic status, and disability status. In this paper, we contribute to this line of work by focusing on the representation of age by asking whether older adults are represented proportionally to the population at large in AI data sets. We examine publicly-available information about 92 face data sets to understand how they codify age as a case study to investigate how the subjects' ages are recorded and whether older generations are represented. We find that older adults are very under-represented; five data sets in the study that explicitly documented the closed age intervals of their subjects included older adults (defined as older than 65 years), while only one included oldest-old adults (defined as older than 85 years). Additionally, we find that only 24 of the data sets include any age-related information in their documentation or metadata, and that there is no consistent method followed across these data sets to collect and record the subjects' ages. We recognize the unique difficulties in creating representative data sets in terms of age, but raise it as an important dimension that researchers and engineers interested in inclusive AI should consider. Joon Sung Park 0001, Michael S. Bernstein, Robin Brewer, Ece Kamar, Meredith Ringel Morris |
AIES | 1 |
| 2021 | Mixed Abilities and Varied Experiences: a group autoethnography of a virtual summer internshipabstractThe COVID-19 pandemic forced many people to convert their daily work lives to a “virtual” format where everyone connected remotely from their home. In this new, virtual environment, accessibility barriers changed, in some respects for the better (e.g., more flexibility) and in other aspects, for the worse (e.g., problems including American Sign Language interpreters over video calls). Microsoft Research held its first cohort of all virtual interns in 2020. We the authors, full time and intern members and affiliates of the Ability Team, a research team focused on accessibility, reflect on our virtual work experiences as a team consisting of members with a variety of abilities, positions, and seniority during the summer intern season. Through our autoethnographic method, we provide a nuanced view into the experiences of a mixed-ability, virtual team, and how the virtual setting affected the team’s accessibility. We then reflect on these experiences, noting the successful strategies we used to promote access and the areas in which we could have further improved access. Finally, we present guidelines for future virtual mixed-ability teams looking to improve access. Kelly Mack, Maitraye Das, Dhruv Jain, Danielle Bragg, John C. Tang, Andrew Begel, Erin Beneteau, Josh Urban Davis, Abraham Glasser, Joon Sung Park 0001, Venkatesh Potluri |
ASSETS | 10 |
| 2021 | "We Just Use What They Give Us": Understanding Passenger User Perspectives in Smart HomesabstractWith a plethora of off-the-shelf smart home devices available commercially, people are increasingly taking a do-it-yourself approach to configuring their smart homes. While this allows for customization, users responsible for smart home configuration often end up with more control over the devices than other household members. This separates those who introduce new functionality to the smart home (pilot users) from those who do not (passenger users). To investigate the prevalence and impact of pilot-passenger user relationships, we conducted a Mechanical Turk survey and a series of one-hour interviews. Our results suggest that pilot-passenger relationships are common in multi-user households and shape how people form habits around devices. We find from interview data that smart homes reflect the values of their pilot users, making it harder for passenger users to incorporate their devices into daily life. We conclude the paper with design recommendations to improve passenger and pilot user experience. Vinay Koshy, Joon Sung Park 0001, Ti-Chung Cheng, Karrie Karahalios |
CHI | 2 |
| 2020 | Random, Messy, Funny, Raw: Finstas as Intimate Reconfigurations of Social MediaabstractAmong many young people, the creation of a finsta-a portmanteau of "fake" and "Instagram" which describes secondary Instagram accounts-provides an outlet to share emotional, low-quality, or indecorous content with their close friends. To study why people create and maintain finstas, we conducted a qualitative study through interviews with finsta users and content analysis of video bloggers exposing their finsta on YouTube. We found that one way that young people deal with mounting social pressures is by reconfiguring online platforms and changing their purposes, norms, expectations, and currencies. Carving out smaller spaces accessible only to close friends allows users the opportunity for a more unguarded, vulnerable, and unserious performance. Drawing on feminist theory, we term this process intimate reconfiguration. Through this reconfiguration finsta users repurpose an existing and widely-used social platform to create opportunities for more meaningful and reciprocal forms of social support. Sijia Xiao, Danaé Metaxa, Joon Sung Park 0001, Karrie Karahalios, Niloufar Salehi |
CHI | 3 |
| 2020 | Emotional Amplification During Live-Streaming: Evidence from Comments During and After News EventsabstractLive streaming services allow people to concurrently consume and comment on media events with other people in real time. Durkheim's theory of "collective effervescence" suggests that face-to-face encounters in ritual events conjure emotional arousal, so people often feel happier and more excited while watching events like the Super Bowl with family and friends through the television than if they were alone. Does a stronger emotional intensity also occur in live streaming? Using a large-scale dataset of comments posted to news and media events on YouTube, we address this question by examining emotional intensity in live comments versus those produced retrospectively. Results reveal that live comments are overall more emotionally intense than retrospective comments across all temporal periods and all event types examined. Findings support the emotional amplification hypothesis and provide preliminary evidence for shared attention theory in explaining the amplification effect. These findings have important implications for live streaming platforms to optimize resources for content moderation and to improve psychological well-being for content moderators, and more broadly as society grapples with using technology to stay connected during social distancing required by the COVID-19 pandemic. Mufan Luo, Tiffany W. Hsu, Joon Sung Park 0001, Jeffrey T. Hancock |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2019 | Search Media and Elections: A Longitudinal Investigation of Political Search ResultsabstractConcern about algorithmically-curated content and its impact on democracy is reaching a fever pitch worldwide. But relative to the role of social media in electoral processes, the role of search results has received less public attention. We develop a theoretical conceptualization of search results as a form of media-search media-and analyze search media in the context of political partisanship in the six months leading up to the 2018 U.S. midterm elections. Our empirical analyses use a total of over 4 million URLs, scraped daily from Google search queries for all candidates running for federal office in the United States in 2018. In our first set of analyses we characterize the nature of search media from the data collected in terms of the types of URLs present and the stability of search results over time. In our second, we annotate URLs' top-level domains with existing measures of political partisanship, examining trends by incumbency, election outcome, and other election characteristics. Among other findings, we note that partisanship trends in search media are largely similar for content about candidates from the two major political parties, whereas there are substantial differences in search media for incumbent versus challenger candidates. This work suggests that longitudinal, systematic audits of search media can reflect real-world political trends. We conclude with implications for web search designers and consumers of political content online. Danaé Metaxa, Joon Sung Park 0001, James A. Landay, Jeffrey T. Hancock |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2019 | A Slow Algorithm Improves Users' Assessments of the Algorithm's AccuracyabstractWith computational algorithms making an increasing number of deeply consequential, and often problematic judgments on our behalf, there is a growing interest in slowing down technology to encourage users to reflect on judgments made by algorithms. Prior work in slow technology has established slowness as an agent of reflection and serendipity; however, it has been unclear whether this waiting time actually helps users gain useful insight or any other benefits as they make judgments using an algorithm. To this end, we conducted a series of online and in-person between-subject user studies in which we isolate the impact of an algorithm's speed on how users incorporate the algorithm's advice when making judgments in the context of simple visual recognition tasks. We find that our participants followed good quality algorithms more and bad quality algorithms somewhat less if the response time of the algorithm is slower. Furthermore, qualitative analysis of the in-person study interviews reveals that the waiting was not time wasted, but was often used to reflect on the task and the estimation process of themselves and the algorithm, and to compare and reevaluate the two processes. Based on these findings, we outline design implications of future algorithmic systems. Joon Sung Park 0001, Rick Barber, Alex Kirlik, Karrie Karahalios |
Proc. ACM Hum. Comput. Interact. | 1 |