Ian Drosos

dblp:208/7388 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-3475-2609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
abstract
LLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Human Feedback), state-of-the-art LLMs like GPT4 and Llama3 are expected to have strong alignment with human preferences when prompted for a quality judgement, such as the coherence of a text. While this seems beneficial, it is not clear whether the assessments by an LLM-as-a-judge constitute only an evaluation based on the instructions in the prompts, or reflect its preference for high-quality data similar to its fine-tune data. To investigate how much influence prompting the LLMs-as-a-judge has on the alignment of AI judgements to human judgements, we analyze prompts with increasing levels of instructions about the target quality of an evaluation, for several LLMs-as-a-judge. Further, we compare to a prompt-free method using model perplexity as a quality measure instead. We aggregate a taxonomy of quality criteria commonly used across state-of-the-art evaluations with LLMs and provide this as a rigorous benchmark of models as judges. Overall, we show that the LLMs-as-a-judge benefit only little from highly detailed instructions in prompts and that perplexity can sometimes align better with human judgements than prompting, especially on textual quality.
Bhuvanashree Murugadoss, Christian Pölitz, Ian Drosos, Vu Le 0002, Nick McKenna, Carina Negreanu, Chris Parnin, Advait Sarkar
AAAI3
2025 The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
abstract
The rise of Generative AI (GenAI) in knowledge workflows raises questions about its impact on critical thinking skills and practices.We survey 319 knowledge workers to investigate 1) when and how they perceive the enaction of critical thinking when using GenAI, and 2) when and why GenAI affects their effort to do so.Participants shared 936 first-hand examples of using GenAI in work tasks.Quantitatively, when considering both task-and user-specific factors, a user's task-specific self-confidence and confidence in GenAI are predictive of whether critical thinking is enacted and the effort of doing so in GenAI-assisted tasks.Specifically, higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking.Qualitatively, GenAI shifts the nature of critical thinking toward information verification, response integration, and task stewardship.Our insights reveal new design challenges and opportunities for developing GenAI tools for knowledge work.
Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, Nicholas C. Wilson
CHI4
2024 Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
abstract
LLM-powered tools like ChatGPT Data Analysis, have the potential to help users tackle the challenging task of data analysis programming, which requires expertise in data processing, programming, and statistics. However, our formative study (n=15) uncovered serious challenges in verifying AI-generated results and steering the AI (i.e., guiding the AI system to produce the desired output). We developed two contrasting approaches to address these challenges. The first (Stepwise) decomposes the problem into step-by-step subgoals with pairs of editable assumptions and code until task completion, while the second (Phasewise) decomposes the entire problem into three editable, logical phases: structured input/output assumptions, execution plan, and code. A controlled, within-subjects experiment (n=18) compared these systems against a conversational baseline. Users reported significantly greater control with the Stepwise and Phasewise systems, and found intervention, correction, and verification easier, compared to the baseline. The results suggest design guidelines and trade-offs for AI-assisted data analysis tools.
Majeed Kazemitabaar, Jack Williams 0001, Ian Drosos, Tovi Grossman, Austin Z. Henley, Carina Negreanu, Advait Sarkar
UIST3
2023 FxD: a functional debugger for dysfunctional spreadsheets
abstract
Recent enhancements to the spreadsheet formula language and intelligent spreadsheet interfaces allow spreadsheet users to build more complex spreadsheets in systematic ways (e.g., via functional abstractions). However, users have been slow to adopt such features, partly due to the absence of corresponding improvements in tools such as editors and debuggers. In this paper, we present FxD, a novel spreadsheet debugging interface, which provides structured information needed for spreadsheet users to debug formulas in systematic ways through affordances such as the ability to step into the execution of dependencies and provide contextual information to users based on the current context. An in-vitro, within-subject (n=12) experiment revealed that, even though using FxD did not lead to faster debugging, participants reported qualitative improvements (e.g., feelings of efficiency and capability) when debugging with it. Further, participants were more satisfied with the amount of information provided by FxD and felt that it would enhance their existing debugging workflows. Our results have implications for the design of debuggers for spreadsheets and for functional programming languages in general.
Ian Drosos, Nicholas C. Wilson, Andrew D. Gordon 0001, Sruti Srinivasa Ragavan, Jack Williams 0001
VL/HCC1
2023 COLDECO: An End User Spreadsheet Inspection Tool for AI-Generated Code
abstract
Code-generating large language models (LLMs) are transforming programming. Their capability to generate multi-step solutions provides even non-programmers a mechanism to harness the power of coding. Non-programmers often use spreadsheets to manage tabular data, as they offer an intuitive understanding of data manipulation and formula out-comes. Considering that LLMs can generate complex, potentially incorrect code, our focus is on enabling user trust in the accuracy of LLM-generated code. We present ColDeco, the first end-user inspection tool for comprehending code produced by LLMs for tabular data tasks. ColDeco integrates two new features for inspection with a grid-based interface. First, users can decompose a generated solution into intermediate helper columns to understand how the problem is solved step by step. Second, users can interact with a filtered table of summary rows, which highlight interesting cases in the program. We evaluate our tool using a within-subjects user study (n=24) where participants are asked to verify the correctness of programs generated by an LLM. We found that while all features are independently useful, participants preferred them in combination. Users especially noted the usefulness of helper columns, but wanted more transparency in how summary rows are generated to assist with understanding and trusting them. Users also highlighted the application of ColDeco in collaborative settings for explaining and understanding existing formulas.
Kasra Ferdowsifard, Jack Williams 0001, Ian Drosos, Andrew D. Gordon 0001, Carina Negreanu, Nadia Polikarpova, Advait Sarkar, Benjamin G. Zorn
VL/HCC3
2022 The Design Space of Livestreaming Equipment Setups: Tradeoffs, Challenges, and Opportunities
abstract
Livestreaming has grown popular in recent years, with millions of people broadcasting themselves making digital art, playing games, programming, and doing other activities on sites like Twitch and YouTube. While many researchers have studied the actions of both streamers and their viewers, to our knowledge there has been no comprehensive analysis of the actual hardware and software equipment used in livestreaming. In this survey paper we present a holistic overview of modern livestreaming equipment in 2022 by analyzing 40 videos where streamers talk about various aspects of their setups. We categorized their equipment choices into a design space with ten dimensions: computer, software, stream control, encoding, cameras, lighting, video accessories, microphones, audio mixers, and audio accessories. We found that each streamer must make tradeoffs between lower- and higher-fidelity options within each dimension. Our design space analysis can inform ideas for future streaming support tools and, more broadly, tools for remote collaboration and learning via live video. As more of us work and learn online, we are in essence becoming amateur livestreamers, so understanding how professional streamers use their equipment to effectively engage their audiences might help us also engage better with our coworkers and classmates.
Ian Drosos, Philip J. Guo
Conference on Designing Interactive Systems1
2021 Streamers Teaching Programming, Art, and Gaming: Cognitive Apprenticeship, Serendipitous Teachable Moments, and Tacit Expert Knowledge
abstract
Livestreaming is now a popular way for programmers, artists, and gamers to teach their craft online. In this paper we propose the idea that streaming can enable cognitive apprenticeship, a form of teaching where an expert works on authentic tasks while thinking aloud to explain their creative process. To understand how streamers teach in this naturalistic way, we performed a content analysis of 20 stream videos across four popular categories: web development, data science, digital art, and gaming. We discovered four kinds of serendipitous teachable moments that are reminiscent of cognitive apprenticeship: 1) creators encountered unexpected errors that led to improvised problem solving, 2) they generated improvised examples on-the-fly, 3) they sometimes went on insightful tangents, 4) they paused to give high-level advice that was contextualized within the work they were currently performing. We also found missed opportunities for additional teachable moments due to creators not being able to express their tacit (unspoken) expert knowledge because of pattern irreducibility, context dependence, and routinization.
Ian Drosos, Philip J. Guo
VL/HCC1
2020 Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data Scientists
abstract
Data wrangling is a difficult and time-consuming activity in computational notebooks, and existing wrangling tools do not fit the exploratory workflow for data scientists in these environments. We propose a unified interaction model based on programming-by-example that generates readable code for a variety of useful data transformations, implemented as a Jupyter notebook extension called Wrex. User study results demonstrate that data scientists are significantly more effective and efficient at data wrangling with Wrex over manual programming. Qualitative participant feedback indicates that Wrex was useful and reduced barriers in having to recall or look up the usage of various data transform functions. The synthesized code allowed data scientists to verify the intended data transformation, increased their trust and confidence in Wrex, and fit seamlessly within their cell-based notebook workflows. This work suggests that presenting readable code to professional data scientists is an indispensable component of offering data wrangling tools in notebooks.
Ian Drosos, Titus Barik, Philip J. Guo, Robert DeLine, Sumit Gulwani
CHI1
2020 The Design Space of Computational Notebooks: An Analysis of 60 Systems in Academia and Industry
abstract
Computational notebooks such as Jupyter are now used by millions of data scientists, machine learning engineers, and computational researchers to do exploratory and end-user programming. In recent years, dozens of different notebook systems have been developed across academia and industry. However, we still lack an understanding of how their individual designs relate to one another and what their tradeoffs are. To provide a holistic view of this rapidly-emerging landscape, we performed, to our knowledge, the first comprehensive design analysis of dozens of notebook systems. We analyzed 60 notebooks (16 academic papers, 29 industry products, and 15 experimental/R&D projects) and formulated a design space that succinctly captures variations in system features. Our design space covers 10 dimensions that include diverse ways of importing data, editing code and prose, running code, and publishing notebook outputs. We conclude by suggesting ways for researchers to push future projects beyond the current bounds of this space.
Sam Lau, Ian Drosos, Julia M. Markel, Philip J. Guo
VL/HCC2
2018 Comparing developer-provided to user-provided tests for fault localization and automated program repair
abstract
To realistically evaluate a software testing or debugging technique, it must be run on defects and tests that are characteristic of those a developer would encounter in practice. For example, to determine the utility of a fault localization or automated program repair technique, it could be run on real defects from a bug tracking system, using real tests that are committed to the version control repository along with the fixes. Although such a methodology uses real tests, it may not use tests that are characteristic of the information a developer or tool would have in practice. The tests that a developer commits after fixing a defect may encode more information than was available to the developer when initially diagnosing the defect.
René Just, Chris Parnin, Ian Drosos, Michael D. Ernst
ISSTA3
2018 Aiding Collaborative Reuse of Computational Notebooks with Annotated Cell Folding
abstract
Computational notebooks aim to support collaborative data analysis by combining code, visualizations, and text in a single easily shared document. Yet, as notebooks evolve and grow they often become difficult to navigate or understand, discouraging sharing and reuse. We present the design and evaluation of a Jupyter Notebook extension providing facilities for annotated cell folding. Through a lab study and multi-week deployment we find cell folding aids notebook navigation and comprehension, not only by the original author, but also by collaborators viewing the notebook in a meeting or revising it on their own. However, in some cases cell folding encouraged collaborators to overlook folded sections or spend longer reviewing a notebook before editing it. These findings extend our understanding of code folding's trade-offs to a new medium and demonstrate its benefits for everyday collaboration. We conclude by discussing how dynamic reorganization can support sharing and reuse of computational notebooks.
Adam Rule, Ian Drosos, Aurélien Tabard, James D. Hollan
Proc. ACM Hum. Comput. Interact.2
2017 HappyFace: Identifying and predicting frustrating obstacles for learning programming at scale
abstract
Unnecessary obstacles limit learning in cognitively-complex domains such as computer programming. With a lack of appropriate feedback mechanisms, novice programmers can experience frustration and disengage from the learning experience. In large-scale educational settings, the struggles of learners are often invisible to the learning infrastructure and learners have limited ability to seek help. In this paper, we perform a large-scale collection of code snippets from an online learn-to-code platform, Python Tutor, and collect a frustration rating through a light-weight learner feedback mechanism. We then devise a technique that can automatically identify sources of frustration based on participants labeling their frustration levels. We found 3 factors that best predicted novice programmers' frustration state: syntax errors, using niche language features, and understanding code with high complexity. Additionally, we found evidence that we could predict sources of frustration. Based on these results, we believe an embedded feedback mechanism can lead to future intervention systems.
Ian Drosos, Philip J. Guo, Chris Parnin
VL/HCC1