EDBT 2026 Demo / reviewers in the wild / expert
Yu Huang 0015
dblp:39/6301-15
· DBLP profile ↗
41ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0003-2730-5077ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 27 · 3 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EyeMulator: Improving Code Language Models by Mimicking Human Visual AttentionabstractYifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Zhang 0013, Chen Huang 0006, Yueke Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang 0015 |
ACL (1) | 8 |
| 2026 | How Well Can 3D Accessibility Guidelines Support XR Development? An Interview Study with XR Practitioners in IndustryabstractWhile accessibility (a11y) guidelines exist for 3D games and virtual worlds, their applicability to extended reality (XR)’s unique interaction paradigms (e.g., spatial tracking, kinesthetic interactions) remains unexplored. XR practitioners need practical guidance to successfully implement a11y guidelines under real-world constraints. We present the first evaluation of existing 3D a11y guidelines applied to XR development through semi-structured interviews with 25 XR practitioners across diverse organization contexts. We assessed 20 commonly-agreed a11y guidelines from six major resources across visual, motor, cognitive, speech, and hearing domains, comparing practitioners’ development practices against guideline applicability to XR. Our investigation reveals that guidelines can be highly effective when designed as transformation catalysts rather than compliance checklists, but fundamental mismatches exist between existing 3D guidelines and XR requirements, creating both implementation barriers and design gaps. This work provides foundational insights towards developing a11y guidelines and support tools that address XR’s distinct characteristics. Daniel Killough, Tiger F. Ji, Kexin Zhang 0002, Yaxin Hu 0002, Yu Huang 0015, Ruofei Du, Yuhang Zhao 0001 |
CHI | 5 |
| 2026 | EyeLayer: Integrating Human Attention Patterns into LLM-Based Code SummarizationabstractCode summarization is the task of generating natural language descriptions of source code, which is critical for software comprehension and maintenance. While large language models (LLMs) have achieved remarkable progress on this task, an open question remains: can human expertise in code understanding further guide and enhance these models? We propose EyeLayer, a lightweight attention-augmentation module that incorporates human eye-gaze patterns, as a proxy of human expertise, into LLM-based code summarization. EyeLayer models human attention during code reading via a Multimodal Gaussian Mixture, redistributing token embeddings based on learned parameters \((\mu _i, \sigma _i^2)\) that capture where and how intensively developers focus. This design enables learning generalizable attention priors from eye-tracking data and incorporating them into LLMs seamlessly, without disturbing existing representations. We evaluate EyeLayer across diverse model families (i.e., LLaMA-3.2, Qwen3, and CodeBERT) covering different scales and architectures. EyeLayer consistently outperforms strong fine-tuning baselines across standard metrics, achieving gains of up to 13.17% on BLEU-4. These results demonstrate that human gaze patterns encode complementary attention signals that enhance the semantic focus of LLMs and transfer effectively across diverse models for code summarization. Yifan Zhang 0013, Kevin Leach, Yu Huang 0015 |
ICPC | 4 |
| 2026 | Context-aware code summary generation
Chia-Yi Su, Aakash Bansal, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
J. Syst. Softw. | 3 |
| 2026 | Contribution Patterns in Open Source Software for Social Good: Dynamics, Individuals, and Impact CSCW010abstractOpen Source Software for Social Good (OSS4SG), a specialized segment within the Open Source Software (OSS) domain, is gaining increasing recognition for its focus on addressing societal challenges and delivering positive social impact. Learning about how contributors engage with OSS4SG is crucial to its sustainability, as the long-term success of these projects relies heavily on active and ongoing contributor participation. However, no study has yet examined the dynamics of contributors within OSS4SG. To fill this gap, we analyzed over 2.2 million commits made by 5,860 contributors to both OSS4SG and general OSS projects on GitHub, identifying contribution patterns and factors influencing sustained contribution to OSS4SG. We found that although OSS4SG contributors tend to show lower overall contribution intensity and shorter active lifespans, their activity during engaged periods is relatively more regular compared to contributions to general OSS. In addition, contributors from developing regions (e.g., Africa) or women are more likely to start with and continue contributing to OSS4SG, despite their overall contribution levels being lower than those of others. Based on these insights, we propose targeted strategies to increase contributions to OSS4SG projects to maximize their social impact to benefit society and harness their potential to foster broader participation in open source, ultimately enhancing the sustainability of the whole community. Zihan Fang 0001, Yueke Zhang, Thomas Zimmermann 0001, Denae Ford, Yu Huang 0015 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2026 | ComCat: Expertise-Guided Context Generation to Enhance Code ComprehensionabstractSoftware maintenance constitutes a substantial portion of the total lifetime costs of software, with a significant portion attributed to code comprehension. Software comprehension is eased by documentation such as comments that summarize and explain code. We present ComCat , an approach to automate comment generation by augmenting Large Language Models (LLMs) with expertise-guided context to target the annotation of source code with comments that improve comprehension. Our approach enables the selection of the most relevant and informative comments for a given snippet or file containing source code. We develop the ComCat pipeline to comment C/C++ files by (1) automatically identifying suitable locations in which to place comments, (2) predicting the most helpful type of comment for each location, and (3) generating a comment based on the selected location and comment type. In a human subject evaluation, we demonstrate that ComCat -generated comments significantly improve developer code comprehension across three indicative software engineering tasks by up to 13% for 80% of participants. In addition, we demonstrate that ComCat -generated comments are at least as accurate and readable as human-generated comments and are preferred over standard ChatGPT-generated comments for up to 92% of snippets of code. Furthermore, we develop and release a dataset containing source code snippets, human-written comments, and human-annotated comment categories. ComCat leverages LLMs to offer a significant improvement in code comprehension across a variety of human software engineering tasks. Skyler Grandel, Scott Thomas Andersen, Yu Huang 0015, Kevin Leach |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2026 | Investigating the Feasibility of Conducting Webcam-Based Eye-Tracking Studies in Code ComprehensionabstractResearchers in Software Engineering (SE) often use onsite screen-mounted eye-tracking experiments to investigate programmers’ visual attention patterns in various programming activities. The pandemic and the difficulty of recruiting many participants, especially those with special expertise in SE, have hastened the shift towards conducting eye-tracking studies offsite, which use integrated webcams to track participants’ gaze in natural settings. This study compares the efficacy of a webcam-based eye tracker to a research-focused screen-mounted eye tracker in code comprehension tasks. We conducted onsite experiments with 49 participants, each using both types of eye trackers simultaneously to assess the webcam-based eye tracker’s capability to capture visual patterns at general, semantic, and token levels and detect individual differences. Additionally, we conducted offsite experiments with 10 participants to supplement the findings. Our findings indicate that while the webcam-based eye tracker effectively captures programmers’ semantic comprehension, but faces challenges in accurately identifying cognitive patterns at a more detailed token level in onsite settings. Furthermore, the elevated noise observed in real-world offsite conditions significantly limits the tracker’s reliability for drawing accurate conclusions. Participants also encountered challenges with calibration and task initiation, highlighting areas for improvement in conducting webcam-based eye-tracking studies offsite in the future.This study investigates the feasibility of webcam-based eye-tracking studies in SE, offers insights to enhance the accuracy of webcam-based eye-tracking in programming potentially, and provides guidelines for future webcam-based eye-tracking study designs. Zihan Fang 0001, Robert Wallace, Zachary Karas, Toby Jia-Jun Li, Collin McMillan, Yu Huang 0015 |
IEEE Trans. Software Eng. | 6 |
| 2025 | MalMixer: Few-Shot Malware Classification with Retrieval-Augmented Semi-Supervised LearningabstractRecent growth and proliferation of malware have tested practitioners’ ability to promptly classify new samples according to malware families. In contrast to labor-intensive reverse engineering efforts, machine learning approaches have demonstrated increased speed and accuracy. However, most existing deep-learning malware family classifiers must be calibrated using a large number of samples that are painstakingly manually analyzed before training. Furthermore, as novel malware samples arise that are beyond the scope of the training set, additional reverse engineering effort must be employed to update the training set. The sheer volume of new samples found in the wild creates substantial pressure on practitioners’ ability to reverse engineer enough malware to adequately train modern classifiers.In this paper, we present MALMIXER, a malware family classifier using semi-supervised learning that achieves high accuracy with sparse training data. We present a domain-knowledge-aware data augmentation technique for malware feature representations, enhancing few-shot performance of semi-supervised malware family classification. We show that MALMIXER achieves state-of-the-art performance in few-shot malware family classification settings. Our research confirms the feasibility and effectiveness of lightweight, domain-knowledge-aware data augmentation methods for malware features and shows the capabilities of similar semi-supervised classifiers in addressing malware classification issues. Yifan Zhang 0013, Yu Huang 0015, Kevin Leach |
EuroS&P | 3 |
| 2025 | Who's Pushing the Code? An Exploration of GitHub ImpersonationabstractGitHub is one of the largest open-source software (OSS) communities for software development and collaboration. Impersonation in the OSS communities refers to the malicious act of assuming another user's identity, often aiming to gain unauthorized access to code, manipulate project outcomes, or spread misinformation. With several recent real-world attacks resulting from impersonation, this issue is becoming more and more concerning within the OSS community. We present the first exploration of the impact of impersonation in GitHub. Specifically, we conduct structured interviews with 17 real-world OSS contributors about their perception of impersonation and corresponding mitigations. Our study reveals that, in general, GitHub users lack awareness of impersonation and underestimate the severity of its implications. After witnessing a demo of impersonation, they show significant concern for the OSS community. Meanwhile, we also demonstrate that the current best practices (i.e., commit signing) that might mitigate impersonation must be improved to encourage use and adoption. We also present and discuss participant perceptions of potential ways to mitigate GitHub impersonation. We collect a dataset comprising 12.5 million commits to investigate the current status of impersonation. Interestingly, we find out that currently impersonation cannot be easily detected. We observe that existing commit histories treat impersonation behavior identically to pull request events, resulting in a lack of detection methods for impersonation. Yueke Zhang, Anda Liang, Pamela J. Wisniewski, Fengwei Zhang, Kevin Leach, Yu Huang 0015 |
ICSE | 7 |
| 2025 | Optimizing Code Runtime Performance Through Context-Aware Retrieval-Augmented GenerationabstractOptimizing software performance through automated code refinement offers a promising avenue for enhancing execution speed and efficiency. Despite recent advancements in LLMs, a significant gap remains in their ability to perform indepth program analysis. This study introduces AutoPatch, an in-context learning approach designed to bridge this gap by enabling LLMs to automatically generate optimized code. Inspired by how programmers learn and apply knowledge to optimize software, AutoPatch incorporates three key components: (1) an analogy-driven framework to align LLM optimization with human cognitive processes, (2) a unified approach that integrates historical code examples and CFG analysis for context-aware learning, and (3) an automated pipeline for generating optimized code through in-context prompting. Experimental results demonstrate that AutoPatch achieves a$\mathbf{7. 3 \%}$improvement in execution efficiency over GPT-4o across common generated executable code, highlighting its potential to advance automated program runtime optimization. Manish Acharya, Yifan Zhang 0013, Kevin Leach, Yu Huang 0015 |
ICPC | 4 |
| 2025 | Programmers' Visual Attention on Function Call Graphs During Code SummarizationabstractThis paper studies programmer visual attention on code as it relates to underlying function call graphs during code summarization. Programmer visual attention refers to where people look when performing a software engineering task, and code summarization is the task of writing a natural language description about a section of source code. Prior work has studied programmers’ visual attention during code summarization, with the vast majority of research effort placed on details in single functional units of code. There have not been any techniques developed to understand code comprehension at the project level due to the difficulty of this task, despite the nature of most real-world methods as embedded within complex project context. This paper focuses on the visual attention paid to the call graph context in which a method sits. We analyze visual attention coverage of call graphs with graph-based metrics, such as the depth that programmers traverse or the amount of coverage they attain. We use these metrics, among other means, to reevaluate an existing dataset from a previous eye-tracking study of programmers (n = 10) that considered basic properties of programmer visual attention in a project context. We then created a new dataset (n = 12) using the same procedures specifically for this paper, resulting in a total of 88 hours of recorded visual behavior on source code. We used our proposed metrics to analyze how participants’ visual strategies correlated with their code summary quality, and confidence in their summaries. Interestingly, we found that higher coverage of the call graph was associated with decreases in both summary quality and participants’ confidence. Samantha McLoughlin, Zachary Karas, Robert Wallace, Aakash Bansal, Collin McMillan, Yu Huang 0015 |
ASE | 6 |
| 2025 | CodeACT-R: A Cognitive Simulation Framework for Human Attention in Code ReadingabstractReading code is a fundamental activity in both software engineering and computer science education. Understanding the cognitive processes involved in reading code is crucial for identifying effective cognitive strategies, which can inform teaching methods and tooling support for developers. However, collecting large human subject eye tracking datasets, especially for programming tasks, is often costly and time-consuming, limiting its scalability and applicability. To address this issue, we present CodeACT-R, the first cognitive simulation framework tailored for code reading, based on the well-established Adaptive Control of Thought—Rational (ACT-R) architecture from cognitive science. CodeACT-R simulates how humans read code and requires only a small, manageable amount of human data to initiate the simulator design, offering a cost-effective and scalable alternative to traditional data collection methods like eye tracking.Specifically, we first collected real human visual attention data from 48 programmers reading code using eye tracking. These data were then used to develop CodeACT-R, enabling the simulation of human-like code reading behaviors. Our evaluation demonstrates that CodeACT-R is capable of simulating visual attention patterns (i.e., scanpaths) that closely resemble real-world human attention patterns, also accounting for up to 87% of observed pattern variations. Yueke Zhang, Zihan Fang 0001, J. Gregory Trafton, Daniel Levin 0001, Kevin Leach, Yu Huang 0015 |
ASE | 6 |
| 2025 | Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code ModificationabstractThis paper presents a study of using large language models (LLMs) in modifying existing code. While LLMs for generating code have been widely studied, their role in code modification remains less understood. Although “prompting” serves as the primary interface for developers to communicate intents to LLMs, constructing effective prompts for code modification introduces challenges different from generation. Prior work suggests that natural language summaries may help scaffold this process, yet such approaches have been validated primarily in narrow domains like SQL rewriting. This study investigates two prompting strategies for LLM-assisted code modification: Direct Instruction Prompting, where developers describe changes explicitly in free-form language, and SummaryMediated Prompting, where changes are made by editing the generated summaries of the code. We conducted an exploratory study with 15 developers who completed modification tasks using both techniques across multiple scenarios. Our findings suggest that developers followed an iterative workflow: understanding the code, localizing the edit, and validating outputs through execution or semantic reasoning. Each prompting strategy presented tradeoffs: direct instruction prompting was more flexible and easier to specify, while summary-mediated prompting supported comprehension, prompt scaffolding, and control. Developers’ choice of strategy was shaped by task goals and context, including urgency, maintainability, learning intent, and code familiarity. These findings highlight the need for more usable prompt interactions, including adjustable summary granularity, reliable summarycode traceability, and consistency in generated summaries. Ningzhi Tang, Emory Smith, Yu Huang 0015, Collin McMillan, Toby Jia-Jun Li |
VL/HCC | 3 |
| 2025 | Programmer Visual Attention During Context-Aware Code SummarizationabstractProgrammer attention represents the visual focus of programmers on parts of the source code in pursuit of programming tasks. The focus of current research in modeling this programmer attention has been on using mouse cursors, keystrokes, or eye tracking equipment to map areas in a snippet of code. These approaches have traditionally only mapped attention for a single method. However, there is a knowledge gap in the literature because programming tasks such as source code summarization require programmers to use contextual knowledge that can only be found in other parts of the project, not only in a single method. To address this knowledge gap, we conducted an in-depth human study with 10 Java programmers, where each programmer generated summaries for 40 methods from five large Java projects over five one-hour sessions. We used eye tracking equipment to map the visual attention of programmers while they wrote the summaries. We also rate the quality of each summary. We found eye-gaze patterns and metrics that define common behaviors between programmer attention during context-aware code summarization. Specifically, we found that programmers need to read up to 35% fewer words (p$\boldsymbol{ \lt }$0.01) over the whole session, and revisit 13% fewer words (p$ \lt $0.03) as they summarize each method during a session, while maintaining the quality of summaries. We also found that the amount of source code a participant looks at correlates with a higher quality summary, but this trend follows a bell-shaped curve, such that after a threshold reading more source code leads to a significant decrease (p$\boldsymbol{ \lt }$0.01) in the quality of summaries. We also gathered insight into the type of methods in the project that provide the most contextual information for code summarization based on programmer attention. Specifically, we observed that programmers spent a majority of their time looking at methods inside the same class as the target method to be summarized. Surprisingly, we found that programmers spent significantly less time looking at methods in the call graph of the target method. We discuss how our empirical observations may aid future studies towards modeling programmer attention and improving context-aware automatic source code summarization. Robert Wallace, Aakash Bansal, Zachary Karas, Ningzhi Tang, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
IEEE Trans. Software Eng. | 5 |
| 2024 | Breaking the Flow: A Study of Interruptions During Software Engineering ActivitiesabstractIn software engineering, interruptions during tasks can have significant implications for productivity and well-being. While previous studies have investigated the effect of interruptions on productivity, to the best of our knowledge, no prior work has yet distinguished the effect of different types of interruptions on software engineering activities. Yimeng Ma, Yu Huang 0015, Kevin Leach |
ICSE | 2 |
| 2024 | ChatGPT Giving Relationship Advice - How Reliable Is It?abstractIn the evolving realm of natural language processing (NLP), generative AI models like ChatGPT are increasingly utilized across various applications. Among the possible purposes, many people are considering asking ChatGPT for relationship advice. However, the lack of in-depth examination of ChatGPT's response quality could be concerning when it is used for personal topics like mental health issues and intimate relationship problems. In these topics, a piece of misleading advice could cause harmful repercussions. In response to people's growing interest in using ChatGPT as a relationship advisor, our research evaluates ChatGPT's proficiency in discerning relationship advice. Specifically, we investigate its alignment with human judgements. We conducted our analysis with 13,138 Reddit posts about intimate relationship problems to examine the overall alignment. Furthermore, we investigate ChatGPT's consistency in judging intimate relationship advice by re-prompting identical queries. Our results indicate a significant disparity between ChatGPT and human judgments, with the model displaying inconsistency in its own decisions. Our findings emphasize the need for comprehensive insights into ChatGPT's mechanisms for intimacy problems and future improvements in its proficiency in helping people's relationship struggles. Haonan Hou, Kevin Leach, Yu Huang 0015 |
ICWSM | 3 |
| 2024 | Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code SummarizationabstractRecent language models have demonstrated proficiency in summarizing source code. However, as in many other domains of machine learning, language models of code lack sufficient explainability --- informally, we lack a formulaic or intuitive understanding of what and how models learn from code. Explainability of language models can be partially provided if, as the models learn to produce higher-quality code summaries, they also align in deeming the same code parts important as those identified by human programmers. In this paper, we report negative results from our investigation of explainability of language models in code summarization through the lens of human comprehension. We measure human focus on code using eye-tracking metrics such as fixation counts and duration in code summarization tasks. To approximate language model focus, we employ a state-of-the-art model-agnostic, black-box, perturbation-based approach, SHAP (SHapley Additive exPlanations), to identify which code tokens influence that generation of summaries. Using these settings, we find no statistically significant relationship between language models' focus and human programmers' attention. Furthermore, alignment between model and human foci in this setting does not seem to dictate the quality of the LLM-generated summaries. Our study highlights an inability to align human focus with SHAP-based model focus measures. This result calls for future investigation of multiple open questions for explainable language models for code summarization and software engineering tasks in general, including the training mechanisms of language models for code, whether there is an alignment between human and model attention on code, whether human attention can improve the development of language models, and what other model focus measures are appropriate for improving explainability. Yifan Zhang 0013, Zachary Karas, Collin McMillan, Kevin Leach, Yu Huang 0015 |
ICPC | 6 |
| 2024 | Developer Behaviors in Validating and Repairing LLM-Generated Code Using IDE and Eye TrackingabstractThe increasing use of large language model (LLM)-powered code generation tools, such as GitHub Copilot, is transforming software engineering practices. This paper investigates how developers validate and repair code generated by Copilot and examines the impact of code provenance awareness during these processes. We conducted a lab study with 28 participants tasked with validating and repairing Copilot-generated code in three software projects. Participants were randomly divided into two groups: one informed about the provenance of LLM-generated code and the other not. We collected data on IDE interactions, eye-tracking, cognitive workload assessments, and conducted semi-structured interviews. Our results indicate that, without explicit information, developers often fail to identify the LLM origin of the code. Developers exhibit LLM-specific behaviors such as frequent switching between code and comments, different attentional focus, and a tendency to delete and rewrite code. Being aware of the code’s provenance led to improved performance, increased search efforts, more frequent Copilot usage, and higher cognitive workload. These findings enhance our understanding of developer interactions with LLM-generated code and inform the design of tools for effective human-LLM collaboration in software development. Ningzhi Tang, Meng Chen 0020, Zheng Ning, Aakash Bansal, Yu Huang 0015, Collin McMillan, Toby Jia-Jun Li |
VL/HCC | 5 |
| 2024 | "Math is a pain!": Understanding challenges and needs of the Machine Learning community on Stack OverflowabstractStack Overflow (SO) is a widely recognized online question-and-answer platform for programming, which has also fostered a substantial community dedicated to machine learning (ML), providing a space for both novices and experts to exchange ideas and find solutions to ML-related problems. However, as a relative minority of this online programming platform, research has demonstrated lower engagement in the ML community, but it remains largely unexplored to understand what hinders the engagement and contribution from ML users' perspectives. This paper presents an empirical study based on 22 hours of semi-structured interviews and 131 survey responses with users on SO and reveals the key factors that may lead to the lower response rate and extended waiting time for ML questions on SO, which includes the unique quality requirement for posting ML questions, the discrepancy between time invested and benefits gained, the dispersed nature of the ML community across various platforms and the desired improvement for SO. Moreover, the qualitative study reveals a declining friendliness in SO's culture over time; the subsequent quantitative study corroborates that newcomers frequently encounter stress when posting and answering ML questions, even though this stress diminishes with increased experience. Additionally, we also explored the potential influence of generative AI tools (e.g., ChatGPT) on online question-and-answer platforms, specifically focusing on ML Q&A. We hope the results of this study can pave the way for enhancing the experience of ML users on online platforms, ultimately facilitating improved knowledge exchange and collaboration within the ML domain. Zihan Fang 0001, Yu Huang 0015 |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2024 | A Tale of Two Comprehensions? Analyzing Student Programmer Attention during Code SummarizationabstractCode summarization is the task of creating short, natural language descriptions of source code. It is an important part of code comprehension and a powerful method of documentation. Previous work has made progress in identifying where programmers focus in code as they write their own summaries (i.e., Writing). However, there is currently a gap in studying programmers’ attention as they read code with pre-written summaries (i.e., Reading). As a result, it is currently unknown how these two forms of code comprehension compare: Reading and Writing. Also, there is a limited understanding of programmer attention with respect to program semantics. We address these shortcomings with a human eye-tracking study ( n = 27) comparing Reading and Writing. We examined programmers’ attention with respect to fine-grained program semantics, including their attention sequences (i.e., scan paths). We find distinctions in programmer attention across the comprehension tasks, similarities in reading patterns between them, and differences mediated by demographic factors. This can help guide code comprehension in both computer science education and automated code summarization. Furthermore, we mapped programmers’ gaze data onto the Abstract Syntax Tree to explore another representation of human attention. We find that visual behavior on this structure is not always consistent with that on source code. Zachary Karas, Aakash Bansal, Yifan Zhang 0013, Toby Jia-Jun Li, Collin McMillan, Yu Huang 0015 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | A Controlled Experiment in Age and Gender Bias When Reading Technical Articles in Software EngineeringabstractOnline platforms and communities are a critical part of modern software engineering, yet are often affected by human biases. While previous studies investigated human biases and their potential harms against the efficiency and fairness of online communities, they have mainly focused on the open source andQ & Aplatforms, such asGitHubandStack Overflow, but overlooked the audience-focused online platforms for delivering programming and SE-related technical articles, where millions of software engineering practitioners share, seek for, and learn from high-quality software engineering articles (i.e.,technical articlesfor SE). Furthermore, most of the previous work has revealed gender and race bias, but we have little knowledge about the effect of age on software engineering practice. In this paper, we propose to investigate the effect of authors’ demographic information (gender and age) on the evaluation of technical articles on software engineering and potential behavioral differences among participants. We conducted a survey-based and controlled human study and collected responses from 540 participants to investigate developers’ evaluation of technical articles for software engineering. By controlling the gender and age of the author profiles of technical articles for SE, we found that raters tend to have more positive content depth evaluations for younger male authors when compared to older male authors and that male participants conduct technical article evaluations faster than female participants, consistent with prior study findings. Surprisingly, different from other software engineering evaluation activities (e.g., code review, pull request, etc.), we did not find a significant difference in the genders of authors on the evaluation outcome of technical articles in SE. Anda Liang, Emerson R. Murphy-Hill, Westley Weimer, Yu Huang 0015 |
IEEE Trans. Software Eng. | 4 |
| 2023 | Leveraging Evidence Theory to Improve Fault Localization: An Exploratory Study
Yueke Zhang, Kevin Leach, Yu Huang 0015 |
ESEM | 3 |
| 2023 | Modeling Programmer Attention as Scanpath PredictionabstractThis paper launches a new effort at modeling programmer attention by predicting eye movement scanpaths. Programmer attention refers to what information people intake when performing programming tasks. Models of programmer attention refer to machine prediction of what information is important to people. Models of programmer attention are important because they help researchers build better interfaces, assistive technologies, and more human-like AI. For many years, researchers in SE have built these models based on features such as mouse clicks, key logging, and IDE interactions. Yet the holy grail in this area is scanpath prediction - the prediction of the sequence of eye fixations a person would take over a visual stimulus. A person's eye movements are considered the most concrete evidence that a person is taking in a piece of information. Scanpath prediction is a notoriously difficult problem, but we believe that the emergence of lower-cost, higheraccuracy eye tracking equipment and better large language models of source code brings a solution within grasp. We present an eye tracking experiment with 27 programmers and a prototype scanpath predictor to present preliminary results and obtain early community feedback. Aakash Bansal, Chia-Yi Su, Zachary Karas, Yifan Zhang 0013, Yu Huang 0015, Toby Jia-Jun Li, Collin McMillan |
ASE | 5 |
| 2023 | A Four-Year Study of Student Contributions to OSS vs. OSS4SG with a Lightweight InterventionabstractModern software engineering practice and training increasingly rely on Open Source Software (OSS). The recent growth in demand for professional software engineers has led to increased contributions to, and usage of, OSS. However, there is limited understanding of the factors affecting how developers, and how new or student developers in particular, decide which OSS projects to contribute to, a process critical to OSS sustainability, access, adoption, and growth. To better understand OSS contributions from the developers of tomorrow, we conducted a four-year study with 1,361 students investigating the life cycle of their contributions (from project selection to pull request acceptance). During the study, we also delivered a lightweight intervention to promote the awareness of open source projects for social good (OSS4SG), OSS projects that have positive impacts in other domains. Using both quantitative and qualitative methods, we analyze student experience reports and the pull requests they submit. Compared to general OSS projects, we find significant differences in project selection (𝑝 < 0.0001, effect size = 0.84), student motivation (𝑝 < 0.01, effect size = 0.13), and increased pull-request acceptance rates for OSS4SG contributions. We also find that our intervention correlates with increased student contributions to OSS4SG (𝑝 < 0.0001, effect size = 0.38). Finally, we analyze correlations of factors such as gender or working with a partner. Our findings may help improve the experience for new developers participating in OSS4SG and the quality of their contributions. We also hope our work helps educators, project leaders, and contributors to build a mutually-beneficial framework for the future growth of OSS4SG. Zihan Fang 0001, Madeline Endres, Thomas Zimmermann 0001, Denae Ford, Westley Weimer, Kevin Leach, Yu Huang 0015 |
ESEC/SIGSOFT FSE | 7 |
| 2023 | Function Call Graph Context Encoding for Neural Source Code SummarizationabstractSource code summarization is the task of writing natural language descriptions of source code. The primary use of these descriptions is in documentation for programmers. Automatic generation of these descriptions is a high value research target due to the time cost to programmers of writing these descriptions themselves. In recent years, a confluence of software engineering and artificial intelligence research has made inroads into automatic source code summarization through applications of neural models of that source code. However, an Achilles’ heel to a vast majority of approaches is that they tend to rely solely on the context provided by the source code being summarized. But empirical studies in program comprehension are quite clear that the information needed to describe code much more often resides in the context in the form of Function Call Graph surrounding that code. In this paper, we present a technique for encoding this call graph context for neural models of code summarization. We implement our approach as a supplement to existing approaches, and show statistically significant improvement over existing approaches. In a human study with 20 programmers, we show that programmers perceive generated summaries to generally be as accurate, readable, and concise as human-written summaries. Aakash Bansal, Zachary Eberhart, Zachary Karas, Yu Huang 0015, Collin McMillan |
IEEE Trans. Software Eng. | 4 |
| 2023 | CirFix: Automated Hardware Repair and its Real-World ApplicationsabstractThis article presents CirFix, a framework for automatically repairing defects in hardware designs implemented in languages like Verilog. We propose a novel fault localization approach based on assignments to wires and registers, and a fitness function tailored to the hardware domain to bridge the gap between software-level automated program repair and hardware descriptions. We also present a benchmark suite of 32 defect scenarios corresponding to a variety of hardware projects. Overall, CirFix produces plausible repairs for 21/32 and correct repairs for 16/32 of the defect scenarios. Additionally, we evaluate CirFix's fault localization independently through a human study (n = 41), and find that the approach may be a beneficial debugging aid for complex multi-line hardware defects. Priscila Santiesteban, Yu Huang 0015, Westley Weimer, Hammad Ahmad |
IEEE Trans. Software Eng. | 2 |
| 2022 | CirFix: automatically repairing defects in hardware design codeabstractThis paper presents CirFix, a framework for automatically repairing defects in hardware designs implemented in languages like Verilog. We propose a novel fault localization approach based on assignments to wires and registers, and a fitness function tailored to the hardware domain to bridge the gap between software-level automated program repair and hardware descriptions. We also present a benchmark suite of 32 defect scenarios corresponding to a variety of hardware projects. Overall, CirFix produces plausible repairs for 21/32 and correct repairs for 16/32 of the defect scenarios. This repair rate is comparable to that of successful program repair approaches for software, indicating CirFix is effective at bringing over the benefits of automated program repair to the hardware domain for the first time. Hammad Ahmad, Yu Huang 0015, Westley Weimer |
ASPLOS | 2 |
| 2021 | Leaving My Fingerprints: Motivations and Challenges of Contributing to OSS for Social GoodabstractWhen inspiring software developers to contribute to open source software, the act is often referenced as an opportunity to build tools to support the developer community. However, that is not the only charge that propels contributions-growing interest in open source has also been attributed to software developers deciding to use their technical skills to benefit a common societal good. To understand how developers identify these projects, their motivations for contributing, and challenges they face, we conducted 21 semi-structured interviews with OSS for Social Good (OSS4SG) contributors. From our interview analysis, we identified themes of contribution styles that we wanted to understand at scale by deploying a survey to over 5765 OSS and Open Source Software for Social Good contributors. From our quantitative analysis of 517 responses, we find that the majority of contributors demonstrate a distinction between OSS4SG and OSS. Likewise, contributors described definitions based on what societal issue the project was to mitigate and who the outcomes of the project were going to benefit. In addition, we find that OSS4SG contributors focus less on benefiting themselves by padding their resume with new technology skills and are more interested in leaving their mark on society at statistically significant levels. We also find that OSS4SG contributors evaluate the owners of the project significantly more than OSS contributors. These findings inform implications to help contributors identify high societal impact projects, help project maintainers reduce barriers to entry, and help organizations understand why contributors are drawn to these projects to sustain active participation. Yu Huang 0015, Denae Ford, Thomas Zimmermann 0001 |
ICSE | 1 |
| 2021 | Connecting the dots: rethinking the relationship between code and prose writing with functional connectivityabstractMedical imaging studies of software engineering have risen in popularity and may reveal the neural underpinnings of coding activities. To date, however, all studies in computer science venues have treated brain regions independently and in isolation. Since most complex neural activity involves coordination among multiple regions, previous analyses may overlook neural behavior. Zachary Karas, Andrew Jahn, Westley Weimer, Yu Huang 0015 |
ESEC/SIGSOFT FSE | 4 |
| 2021 | Toward an Objective Measure of Developers' Cognitive ActivitiesabstractUnderstanding how developers carry out different computer science activities with objective measures can help to improve productivity and guide the use and development of supporting tools in software engineering. In this article, we present two controlled experiments involving 112 students to explore multiple computing activities (code comprehension, code review, and data structure manipulations) using three different objective measures including neuroimaging (functional near-infrared spectroscopy (fNIRS) and functional magnetic resonance imaging (fMRI)) and eye tracking. By examining code review and prose review using fMRI, we find that the neural representations of programming languages vs. natural languages are distinct. We can classify which task a participant is undertaking based solely on brain activity, and those task distinctions are modulated by expertise. We leverage insights from the psychological notion of spatial ability to decode the neural representations of several fundamental data structures and their manipulations using fMRI, fNIRS, and eye tracking. We examine list, array, tree, and mental rotation tasks and find that data structure and spatial operations use the same focal regions of the brain but to different degrees: they are related but distinct neural tasks. We demonstrate best practices and describe the implication and tradeoffs between fMRI, fNIRS, eye tracking, and self-reporting for software engineering research. Zohreh Sharafi, Yu Huang 0015, Kevin Leach, Westley Weimer |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2020 | Trustworthiness Perceptions in Code Review: An Eye-tracking StudyabstractBackground: Automated program repair and other bug-fixing approaches are gaining attention in the software engineering community. Automation shows promise in reducing bug fixing costs. However, many developers express reluctance about accepting machine-generated patches into their codebases. Ian Bertram, Jack Hong, Yu Huang 0015, Westley Weimer, Zohreh Sharafi |
ESEM | 3 |
| 2020 | Neurological divide: an fMRI study of prose and code writingabstractSoftware engineering involves writing new code or editing existing code. Recent efforts have investigated the neural processes associated with reading and comprehending code --- however, we lack a thorough understanding of the human cognitive processes underlying code writing. While prose reading and writing have been studied thoroughly, that same scrutiny has not been applied to code writing. In this paper, we leverage functional brain imaging to investigate neural representations of code writing in comparison to prose writing. We present the first human study in which participants wrote code and prose while undergoing a functional magnetic resonance imaging (fMRI) brain scan, making use of a full-sized fMRI-safe QWERTY keyboard. Ryan Krueger, Yu Huang 0015, Tyler Santander, Westley Weimer, Kevin Leach |
ICSE | 2 |
| 2020 | A Human Study of Comprehension and Code SummarizationabstractSoftware developers spend a great deal of time reading and understanding code that is poorly-documented, written by other developers, or developed using differing styles. During the past decade, researchers have investigated techniques for automatically documenting code to improve comprehensibility. In particular, recent advances in deep learning have led to sophisticated summary generation techniques that convert functions or methods to simple English strings that succinctly describe that code's behavior. However, automatic summarization techniques are assessed using internal metrics such as BLEU scores, which measure natural language properties in translational models, or ROUGE scores, which measure overlap with human-written text. Unfortunately, these metrics do not necessarily capture how machine-generated code summaries actually affect human comprehension or developer productivity. Sean Stapleton, Yashmeet Gambhir, Alexander LeClair, Zachary Eberhart, Westley Weimer, Kevin Leach, Yu Huang 0015 |
ICPC | 7 |
| 2020 | Biases and differences in code review using medical imaging and eye-tracking: genders, humans, and machinesabstractCode review is a critical step in modern software quality assurance, yet it is vulnerable to human biases. Previous studies have clarified the extent of the problem, particularly regarding biases against the authors of code,but no consensus understanding has emerged. Advances in medical imaging are increasingly applied to software engineering, supporting grounded neurobiological explorations of computing activities, including the review, reading, and writing of source code. In this paper, we present the results of a controlled experiment using both medical imaging and also eye tracking to investigate the neurological correlates of biases and differences between genders of humans and machines (e.g., automated program repair tools) in code review. We find that men and women conduct code reviews differently, in ways that are measurable and supported by behavioral, eye-tracking and medical imaging data. We also find biases in how humans review code as a function of its apparent author, when controlling for code quality. In addition to advancing our fundamental understanding of how cognitive biases relate to the code review process, the results may inform subsequent training and tool design to reduce bias. Yu Huang 0015, Kevin Leach, Zohreh Sharafi, Nicholas McKay, Tyler Santander, Westley Weimer |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Distilling neural representations of data structure manipulation using fMRI and fNIRSabstractData structures permeate many aspects of software engineering, but their associated human cognitive processes are not thoroughly understood. We leverage medical imaging and insights from the psychological notion of spatial ability to decode the neural representations of several fundamental data structures and their manipulations. In a human study involving 76 participants, we examine list, array, tree, and mental rotation tasks using both functional near-infrared spectroscopy (fNIRS) and functional magnetic resonance imaging (fMRI). We find a nuanced relationship: data structure and spatial operations use the same focal regions of the brain but to different degrees. They are related but distinct neural tasks. In addition, more difficult computer science problems induce higher cognitive load than do problems of pure spatial reasoning. Finally, while fNIRS is less expensive and more permissive, there are some computing-relevant brain regions that only fMRI can reach. Yu Huang 0015, Ryan Krueger, Tyler Santander, Xiao-Su Hu, Kevin Leach, Westley Weimer |
ICSE | 1 |
| 2017 | Daehr: A Discriminant Analysis Framework for Electronic Health Record Data and an Application to Early Detection of Mental Health DisordersabstractElectronic health records (EHR) provide a rich source of temporal data that present a unique opportunity to characterize disease patterns and risk of imminent disease. While many data-mining tools have been adopted for EHR-based disease early detection, linear discriminant analysis (LDA) is one of the most commonly used statistical methods. However, it is difficult to train an accurate LDA model for early disease diagnosis when too few patients are known to have the target disease. Furthermore, EHR data are heterogeneous with significant noise. In such cases, the covariance matrices used in LDA are usually singular and estimated with a large variance. This article presents Daehr , an extension of the LDA framework using electronic health record data to address these issues. Beyond existing LDA analyzers, we propose Daehr to (1) eliminate the data noise caused by the manual encoding of EHR data and (2) lower the variance of parameter (covariance matrices) estimation for LDA models when only a few patients’ EHR are available for training. To achieve these two goals, we designed an iterative algorithm to improve the covariance matrix estimation with embedded data-noise/parameter-variance reduction for LDA. We evaluated Daehr extensively using the College Health Surveillance Network, a large, real-world EHR dataset. Specifically, our experiments compared the performance of LDA to three baselines (i.e., LDA and its derivatives) in identifying college students at high risk for mental health disorders from 23 U.S. universities. Experimental results demonstrate Daehr significantly outperforms the three baselines by achieving 1.4%--19.4% higher accuracy and a 7.5%--43.5% higher F1-score. Haoyi Xiong, Jinghe Zhang, Yu Huang 0015, Kevin Leach, Laura E. Barnes |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | Assessing social anxiety using gps trajectories and point-of-interest dataabstractMental health problems are highly prevalent and appear to be increasing in frequency and severity among the college student population. The upsurge in mobile and wearable wireless technologies capable of intense, longitudinal tracking of individuals, provide valuable opportunities to examine temporal patterns and dynamic interactions of key variables in mental health research. In this paper, we present a feasibility study leveraging non-invasive mobile sensing technology to passively assess college students' social anxiety, one of the most common disorders in the college student population. We have first developed a smartphone application to continuously track GPS locations of college students, then we built an analytic infrastructure to collect the GPS trajectories and finally we analyzed student behaviors (e.g. studying or staying at home) using Point-Of-Interest (POI). The whole framework supports intense, longitudinal, dynamic tracking of college students to evaluate how their anxiety and behaviors change in the college campus environment. The collected data provides critical information about how students' social anxiety levels and their mobility patterns are correlated. Our primary analysis based on 18 college students demonstrated that social anxiety level is significantly correlated with places students' visited and location transitions. Yu Huang 0015, Haoyi Xiong, Kevin Leach, Philip Chow, Karl C. Fua, Bethany A. Teachman, Laura E. Barnes |
UbiComp | 1 |
| 2016 | Sensus: a cross-platform, general-purpose system for mobile crowdsensing in human-subject studiesabstractThe burden of entry into mobile crowdsensing (MCS) is prohibitively high for human-subject researchers who lack a technical orientation. As a result, the benefits of MCS remain beyond the reach of research communities (e.g., psychologists) whose expertise in the study of human behavior might advance applications and understanding of MCS systems. This paper presents Sensus, a new MCS system for human-subject studies that bridges the gap between human-subject researchers and MCS methods. Sensus alleviates technical burdens with on-device, GUI-based design of sensing plans, simple and efficient distribution of sensing plans to study participants, and uniform participant experience across iOS and Android devices. Sensing plans support many hardware and software sensors, automatic deployment of sensor-triggered surveys, and double-blind assignment of participants within randomized controlled trials. Sensus offers these features to study designers without requiring knowledge of markup and programming languages. We demonstrate the feasibility of using Sensus within two human-subject studies, one in psychology and one in engineering. Feedback from non-technical users indicates that Sensus is an effective and low-burden system for MCS-based data collection and analysis. Haoyi Xiong, Yu Huang 0015, Laura E. Barnes, Matthew S. Gerber |
UbiComp | 2 |
| 2015 | M-SEQ: Early detection of anxiety and depression via temporal orders of diagnoses in electronic health dataabstractAccording to a 2014 Spring American College Health Association Survey, almost 50% of college students reported feeling things were hopeless and that it was difficult to function within the last 12 months. More than 80% reported feeling overwhelmed and exhausted by their responsibilities. This critical subpopulation of Americans is facing significant levels of mental health disorders, challenging colleges to provide accessible and high quality behavioral health care. However, psychiatric disorders are frequently unrecognized in primary care settings, posing physical, emotional, economic, and social burdens to patients and others. Towards the goal of earlier identification and treatment of mental health disorders, this paper proposes M-SEQ, an early detection framework for anxiety/depression using electronic health data from primary care visit sequences. Specifically, compared to existing methods that predict a future disease state using frequency of diagnoses in a patient's medical history, we hypothesize that future disease might also be correlated with the temporal orders of diagnoses. Thus, M-SEQ first discovers a set of diagnosis codes that are discriminative of anxiety/depression, and then extracts each diagnosis pair from each patient's health record to represent the temporal orders of diagnoses. Further, it incorporates the extracted temporal order information with the existing representation to predict whether a patient is at risk of anxiety/depression. We evaluate M-SEQ using the electronic health record (EHR) data of 213,112 college students from 10 schools participating in the College Health Surveillance Network (CHSN) from January 1, 2011 through December 31, 2014. The experimental results shows that our framework can detect a future diagnosis of anxiety and depression based on the primary care visit data up to 3 months in advance, with approximately 1%-4.5% higher accuracy, compared to baseline methods using frequency of diagnoses. Jinghe Zhang, Haoyi Xiong, Yu Huang 0015, Kevin Leach, Laura E. Barnes |
IEEE BigData | 3 |
| 2015 | Using island-style bi-directional intra-CLB routing in low-power FPGAsabstractIncreased clustering in Field Programmable Gate Arrays (FPGAs) has shifted a larger fraction of the overall routing load into the configurable logic blocks (CLBs), reducing usage of the costly global interconnect. However, increases in CLB size introduce additional overheads inside CLBs, which can limit the savings gained by minimizing the global interconnect use, motivating more efficient intra-CLB routing. This paper explores different topologies for the intra-CLB connectivity and identifies how the optimal local-CLB interconnect changes for different FPGA architecture and circuit parameters. This work compares area, delay, and energy for two intra-CLB topologies: multiplexer-based routing and island-style bi-directional routing, similar to the global FPGA interconnect, but used inside the CLB (which we call a mini-FPGA). The mini-FPGA style of local CLB interconnect prove to be favorable for minimum-energy operation, as they can reduce transistor count by as much as 62%, and consume as much as 77.9% less energy. Multipexer-based CLBs have performance benefits by reducing delays by almost 3×. Multiplexer-based CLBs can consume less energy at nominal voltages, but only if additional measures are taken to limit power consumption in the multiplexers. A 130-nm CMOS test chip confirms that simulation results track measured data for mini-FPGA CLBs. Oluseyi A. Ayorinde, He Qi, Yu Huang 0015, Benton H. Calhoun |
FPL | 3 |
| 2015 | Optimizing energy efficient low-swing interconnect for sub-threshold FPGAsabstractFPGA interconnect traditionally dominates energy and delay, and designs such as low-swing interconnect have been proven to reduce the interconnect burden for low energy FPGAs. This paper presents an optimized low-swing interconnect for FPGAs operating in the sub-threshold region. We also address signal degradation along lengthy interconnect paths and examine strategies for inserting low-switching-threshold repeaters. A 130nm test chip implementing low-swing interconnect meshes with different circuit parameters is measured. The results show that optimization of the low-swing interconnect provides up to 60.2% lower energy-delay-product (EDP) than a straightforward, un-optimized low-swing design at VDD= 0.4V. Furthermore, the simulation results show that the optimized low-swing interconnect is 97.7% faster and 42.7% lower energy than a traditional uni-directional interconnect at VDD= 0.4V. He Qi, Oluseyi A. Ayorinde, Yu Huang 0015, Benton H. Calhoun |
FPL | 3 |