Eunice Jun

dblp:168/3483 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-4050-4284ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 PriorWeaver: Prior Elicitation via Iterative Dataset Construction
abstract
In Bayesian analysis, prior elicitation, or the process of facilitating the expression of one’s beliefs to inform statistical modeling, is an essential yet challenging step. Analysts often have beliefs about real-world variables and their relationships. However, existing tools require analysts to translate these beliefs and express them indirectly as probability distributions over model parameters. We present PriorWeaver, an interactive visualization system that facilitates prior elicitation through iterative dataset construction and refinement. Analysts visually express their assumptions about individual variables and their relationships. Under the hood, these assumptions create a dataset used to derive statistical priors. Prior predictive checks then help analysts compare the priors to their assumptions. In a lab study with 17 participants new to Bayesian analysis, we compare PriorWeaver to a baseline incorporating existing techniques. Compared to the baseline, PriorWeaver gave participants greater control, clarity, and confidence, leading to priors that were better aligned with their expectations.
Yuwei Xiao, Shuai Ma 0005, Antti Oulasvirta, Eunice Jun
CHI4
2026 Causality and Semantic Separation
abstract
The design of scientific experiments deserves its own variation of formal verification to catch cases where scientists made important mistakes, such as forgetting to take confounding variables into account. One of the most fundamental underpinnings of science is causality , or what it means for interventions in the world to cause other outcomes, as formalized by computer scientists like Judea Pearl. However, these ideas had not previously been made rigorous to the standards of the programming-languages community, where one expects a (syntactic) program analysis to be proved sound with respect to a natural semantics. In the domain of causality, as the relevant “program analysis,” we focus on d -separation, a classic condition on graphs that can be used to decide when the design of an experiment controls for sufficiently many confounding variables, even though the reason that this condition works is often unintuitive. Our central result (mechanized in Rocq) is that d -separation exactly coincides with a novel semantic definition inspired by noninterference from the theory of security. This characterization provides a structural semantic foundation for d -separation and helps explain why the graph-theoretic condition is correct, independently of probabilistic assumptions. For each given automated test on the quality of an experiment design, our theorem justifies an associated method for falsifying the world-modeling hypothesis behind the experiment.
Anna Zhang, Qinglan Luo, London Bielicke, Eunice Jun, Adam Chlipala
Proc. ACM Program. Lang.4
2025 Dreamcrafter: Immersive Editing of 3D Radiance Fields Through Flexible, Generative Inputs and Outputs
Cyrus Vachha, Yixiao Kang, Zach Dive, Ashwat Chidambaram, Anik Gupta, Eunice Jun, Björn Hartmann
CHI6
2025 Beyond Code Generation: LLM-supported Exploration of the Program Design Space
abstract
In this work, we explore explicit Large Language Model (LLM)powered support for the iterative design of computer programs.Program design, like other design activity, is characterized by navigating a space of alternative problem formulations and associated solutions in an iterative fashion.LLMs are potentially powerful tools in helping this exploration; however, by default, code-generation LLMs deliver code that represents a particular point solution.This obscures the larger space of possible alternatives, many of which might be preferable to the LLM's default interpretation and its generated code.We contribute an IDE that supports program design through generating and showing new ways to frame problems alongside alternative solutions, tracking design decisions, and identifying implicit decisions made by either the programmer or the LLM.In a user study, we find that with our IDE, users combine and parallelize design phases to explore a broader design space-but also struggle to keep up with LLM-originated changes to code and other information overload.These findings suggest a core challenge for future IDEs that support program design through higher-level instructions given to LLM-based agents: carefully managing attention and deciding what information agents should surface to program designers and when.
J. D. Zamfirescu-Pereira, Eunice Jun, Michael Terry, Qian Yang 0004, Björn Hartmann
CHI2
2025 Flowco: Mixed-Initiative Authoring of Reliable End-to-End Data Analyses via Dataflow Graphs and LLMs
Stephen N. Freund, Brooke Simon, Emery D. Berger, Eunice Jun
UIST4
2024 rTisane: Externalizing conceptual models for data analysis prompts reconsideration of domain assumptions and facilitates statistical modeling
abstract
Statistical models should accurately reflect analysts’ domain knowledge about variables and their relationships. While recent tools let analysts express these assumptions and use them to produce a resulting statistical model, it remains unclear what analysts want to express and how externalization impacts statistical model quality. This paper addresses these gaps. We first conduct an exploratory study of analysts using a domain-specific language (DSL) to express conceptual models. We observe a preference for detailing how variables relate and a desire to allow, and then later resolve, ambiguity in their conceptual models. We leverage these findings to develop rTisane, a DSL for expressing conceptual models augmented with an interactive disambiguation process. In a controlled evaluation, we find that analysts reconsidered their assumptions, self-reported externalizing their assumptions accurately, and maintained analysis intent with rTisane. Additionally, rTisane enabled some analysts to author statistical models they were unable to specify manually. For others, rTisane resulted in models that better fit the data or enabled iterative improvement.
Eunice Jun, Edward Misback, Jeffrey Heer, René Just
CHI1
2023 Understanding and Supporting Debugging Workflows in Multiverse Analysis
abstract
Multiverse analysis—a paradigm for statistical analysis that considers all combinations of reasonable analysis choices in parallel—promises to improve transparency and reproducibility. Although recent tools help analysts specify multiverse analyses, they remain difficult to use in practice. In this work, we identify debugging as a key barrier due to the latency from running analyses to detecting bugs and the scale of metadata processing needed to diagnose a bug. To address these challenges, we prototype a command-line interface tool, Multiverse Debugger, which helps diagnose bugs in the multiverse and propagate fixes. In a qualitative lab study (n=13), we use Multiverse Debugger as a probe to develop a model of debugging workflows and identify specific challenges, including difficulty in understanding the multiverse’s composition. We conclude with design implications for future multiverse analysis authoring systems.
Ken Gu, Eunice Jun, Tim Althoff
CHI2
2023 Odyssey: An Interactive Workbench for Expert-Driven Floating-Point Expression Rewriting
abstract
In recent years, researchers have proposed a number of automated tools to identify and improve floating-point rounding error in mathematical expressions. However, users struggle to effectively apply these tools. In this paper, we work with novices, experts, and tool developers to investigate user needs during the expression rewriting process. We find that users follow an iterative design process. They want to compare expressions on multiple input ranges, integrate and guide various rewriting tools, and understand where errors come from. We organize this investigation’s results into a three-stage workflow and implement that workflow in a new, extensible workbench dubbed Odyssey. Odyssey enables users to: (1) diagnose problems in an expression, (2) generate solutions automatically or by hand, and (3) tune their results. Odyssey tracks a working set of expressions and turns a state-of-the-art automated tool “inside out,” giving the user access to internal heuristics, algorithms, and functionality. In a user study, Odyssey enabled five expert numerical analysts to solve challenging rewriting problems where state-of-the-art automated tools fail. In particular, the experts unanimously praised Odyssey’s novel support for interactive range modification and local error visualization.
Edward Misback, Caleb C. Chan, Brett Saiki, Eunice Jun, Zachary Tatlock, Pavel Panchekha
UIST4
2022 Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data Relationships
abstract
Proper statistical modeling incorporates domain theory about how concepts relate and details of how data were measured. However, data analysts currently lack tool support for recording and reasoning about domain assumptions, data collection, and modeling choices in an integrated manner, leading to mistakes that can compromise scientific validity. For instance, generalized linear mixed-effects models (GLMMs) help answer complex research questions, but omitting random effects impairs the generalizability of results. To address this need, we present Tisane, a mixed-initiative system for authoring generalized linear models with and without mixed-effects. Tisane introduces a study design specification language for expressing and asking questions about relationships between variables. Tisane contributes an interactive compilation process that represents relationships in a graph, infers candidate statistical models, and asks follow-up questions to disambiguate user queries to construct a valid model. In case studies with three researchers, we find that Tisane helps them focus on their goals and assumptions while avoiding past mistakes.
Eunice Jun, Audrey Seo, Jeffrey Heer, René Just
CHI1
2022 Hypothesis Formalization: Empirical Findings, Software Limitations, and Design Implications
abstract
Data analysis requires translating higher level questions and hypotheses into computable statistical models. We present a mixed-methods study aimed at identifying the steps, considerations, and challenges involved in operationalizing hypotheses into statistical models, a process we refer to as hypothesis formalization . In a formative content analysis of 50 research papers, we find that researchers highlight decomposing a hypothesis into sub-hypotheses, selecting proxy variables, and formulating statistical models based on data collection design as key steps. In a lab study, we find that analysts fixated on implementation and shaped their analyses to fit familiar approaches, even if sub-optimal. In an analysis of software tools, we find that tools provide inconsistent, low-level abstractions that may limit the statistical models analysts use to formalize hypotheses. Based on these observations, we characterize hypothesis formalization as a dual-search process balancing conceptual and statistical considerations constrained by data and computation and discuss implications for future tools.
Eunice Jun, Melissa Birchfield, Nicole de Moura, Jeffrey Heer, René Just
ACM Trans. Comput. Hum. Interact.1
2021 Longitudinal Observational Evidence of the Impact of Emotion Regulation Strategies on Affective Expression
abstract
The ability to regulate our emotions plays an important role in our psychological and physical health. Regulating emotions influences how and when emotions are expressed. We performed a large scale, longitudinal observational study to investigate the effect of emotion regulation ability on expressed affect. We found that expression of negative affect increased throughout the day. For people who suppress emotion this increase is slower that for those who do not. For those with stronger cognitive reappraisal abilities, though not significant, there was a trend for higher positive affect and negative affect increased significantly less steeply, suggesting that they might experience more positive and less negative affect. These results reflect some of the first results based on large scale, continuous tracking of behavioral expression of emotion longitudinally. Our results demonstrate the need to carefully consider the time of day and emotion regulation ability, in addition to gender and age, when attempting to automatically infer affective states for facial behavior.
Daniel McDuff, Eunice Jun, Kael Rowan, Mary Czerwinski
IEEE Trans. Affect. Comput.2
2019 Tea: A High-level Language and Runtime System for Automating Statistical Analysis
abstract
Though statistical analyses are centered on research questions and hypotheses, current statistical analysis tools are not. Users must first translate their hypotheses into specific statistical tests and then perform API calls with functions and parameters. To do so accurately requires that users have statistical expertise. To lower this barrier to valid, replicable statistical analysis, we introduce Tea, a high-level declarative language and runtime system. In Tea, users express their study design, any parametric assumptions, and their hypotheses. Tea compiles these high-level specifications into a constraint satisfaction problem that determines the set of valid statistical tests and then executes them to test the hypothesis. We evaluate Tea using a suite of statistical analyses drawn from popular tutorials. We show that Tea generally matches the choices of experts while automatically switching to non-parametric tests when parametric assumptions are not met. We simulate the effect of mistakes made by non-expert users and show that Tea automatically avoids both false negatives and false positives that could be produced by the application of incorrect statistical tests.
Eunice Jun, Maureen Daum, Jared Roesch, Sarah E. Chasins, Emery D. Berger, René Just, Katharina Reinecke
UIST1
2019 Latent Space Cartography: Visual Analysis of Vector Space Embeddings
abstract
Abstract Latent spaces—reduced‐dimensionality vector space embeddings of data, fit via machine learning—have been shown to capture interesting semantic properties and support data analysis and synthesis within a domain. Interpretation of latent spaces is challenging because prior knowledge, sometimes subtle and implicit, is essential to the process. We contribute methods for “latent space cartography”, the process of mapping and comparing meaningful semantic dimensions within latent spaces. We first perform a literature survey of relevant machine learning, natural language processing, and scientific research to distill common tasks and propose a workflow process. Next, we present an integrated visual analysis system for supporting this workflow, enabling users to discover, define, and verify meaningful relationships among data points, encoded within latent space dimensions. Three case studies demonstrate how users of our system can compare latent space variants in image generation, challenge existing findings on cancer transcriptomes, and assess a word embedding benchmark.
Yang Liu 0136, Eunice Jun, Qisheng Li, Jeffrey Heer
Comput. Graph. Forum2
2019 Circadian Rhythms and Physiological Synchrony: Evidence of the Impact of Diversity on Small Group Creativity
abstract
Circadian rhythms determine daily sleep cycles, mood, and cognition. Depending on an individual's circadian preference, or chronotype (i.e.,"early birds" and "night owls"), the rhythms shift earlier or later in the day. Early birds experience circadian arousal peaks earlier in the morning than night owls. Prior work has shown that individuals are more effective at analytic tasks during their peak arousal times but are more creative during their off-peak times. We investigate if these findings hold true for small groups. We find that time of day and a group's majority chronotype impact performance on analytic and creative tasks. Physiological synchrony among group members positively predicts group satisfaction. Specifically, homogeneous groups perform worse on all tasks regardless of time of day, but they achieve greater physiological synchrony and feel more satisfied as a group. Based on these findings, we present and advocate for a temporal dimension of group diversity.
Eunice Jun, Daniel McDuff, Mary Czerwinski
Proc. ACM Hum. Comput. Interact.1
2018 The potential for scientific outreach and learning in mechanical turk experiments
abstract
The global reach of online experiments and their wide adoption in fields ranging from political science to computer science poses an underexplored opportunity for learning at scale: the possibility of participants learning about the research to which they contribute data. We conducted three experiments on Amazon's Mechanical Turk to evaluate whether participants of paid online experiments are interested in learning about research, what information they find most interesting, and whether providing them with such information actually leads to learning gains. Our findings show that 40% of our participants on Mechanical Turk actively sought out post-experiment learning opportunities despite having already received their financial compensation. Participants expressed high interest in a range of research topics, including previous research and experimental design. Finally, we find that participants comprehend and accurately recall facts from post-experiment learning opportunities. Our findings suggest that Mechanical Turk can be a valuable platform for learning at scale and scientific outreach.
Eunice Jun, Morelle Arian, Katharina Reinecke
L@S1
2018 Digestif: Promoting Science Communication in Online Experiments
abstract
Online experiments allow researchers to collect data from large, demographically diverse global populations. Unlike in-lab studies, however, online experiments often fail to inform participants about the research to which they contribute. This paper is the first to investigate barriers that prevent researchers from providing such science communication in online experiments. We found that the main obstacles preventing researchers from including such information are assumptions about participant disinterest, limited time, concerns about losing anonymity, and concerns about experimental bias. Researchers also noted the dearth of tools to help them close the information loop with their study participants. Based on these findings, we formulated design requirements and implemented Digestif, a new web-based tool that supports researchers in providing their participants with science communication pages. Our evaluation shows that Digestif's scaffolding, examples, and nudges to focus on participants make researchers more aware of their participants' curiosity about research and more likely to disclose pertinent research information.
Eunice Jun, Blue A. Jo, Nigini Oliveira, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.1
2017 Citizen Science Opportunities in Volunteer-Based Online Experiments
abstract
Online experimentation with volunteers could be described as a form of citizen science in which participants take part in behavioral studies without financial compensation. However, while citizen science projects aim to improve scientific understanding, volunteer-based online experiment platforms currently provide minimal possibilities for research involvement and learning. The goal of this paper is to uncover opportunities for expanding participant involvement and learning in the research process. Analyzing comments from 8,288 volunteers who took part in four online experiments on LabintheWild, we identified six themes that reveal needs and opportunities for closer interaction between researchers and participants. Our findings demonstrate opportunities for research involvement, such as engaging participants in refining experiment implementations, and learning opportunities, such as providing participants with possibilities to learn about research aims. We translate these findings into ideas for the design of future volunteer-based online experiment platforms that are more mutually beneficial to citizen scientists and researchers.
Nigini Oliveira, Eunice Jun, Katharina Reinecke
CHI2
2017 Types of Motivation Affect Study Selection, Attention, and Dropouts in Online Experiments
abstract
Understanding whether and how motivation affects participation in online experiments is critical because who contributes and how they contribute can affect the validity of findings. Analyzing data from 7,674 participants across three different studies on the volunteer-based online experiment platform LabintheWild, we identified five motivation types for participating: boredom, comparison, fun, science, and self-learning. We found that these motivation types affect study selection, attention, and dropouts. Participants who were highly motivated by boredom paid less attention and were more likely to dropout than those who were motivated by the possibility of contributing to science. We additionally show that motivation can impact study results and suggest how researchers can take participants' motivation into account when designing and analyzing data from volunteer-based online experiments.
Eunice Jun, Gary Hsieh, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.1
2015 Big Foot: Using the Size of a Virtual Foot to Scale Gap Width
abstract
Spatial perception research in the real world and in virtual environments suggests that the body (e.g., hands) plays a role in the perception of the scale of the world. However, little research has closely examined how varying the size of virtual body parts may influence judgments of action capabilities and spatial layout. Here, we questioned whether changing the size of virtual feet would affect judgments of stepping over and estimates of the width of a gap. Participants viewed their disembodied virtual feet as small or large and judged both their ability to step over a gap and the size of gaps shown in the virtual world. Foot size affected both affordance judgments and size estimates such that those with enlarged virtual feet estimated they could step over larger gaps and that the extent of the gap was smaller. Shrunken feet led to the perception of a reduced ability to step over a gap and smaller estimates of width. The results suggest that people use their visually perceived foot size to scale virtual spaces. Regardless of foot size, participants felt that they owned the feet rendered in the virtual world. Seeing disembodied, but motion-tracked, virtual feet affected spatial judgments, suggesting that the presentation of a single tracked body part is sufficient to produce similar effects on perception, as has been observed with the presence of fully co-located virtual self-avatars or other body parts in the past.
Eunice Jun, Jeanine K. Stefanucci, Sarah H. Creem-Regehr, Michael Geuss, William B. Thompson
ACM Trans. Appl. Percept.1