VLDB 2026 Research / reviewers in the wild / expert
Bonita Sharif
dblp:08/3315
· DBLP profile ↗
60ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0002-5178-7160ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 44 · 11 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 23 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analyzing Reading Behavior across Source Code and Stack Overflow for Method and Class Summarization
Alaa Ismail, Zachary Kozak, Bonita Sharif |
ETRA | 3 |
| 2026 | Do Developers Read Type Information? An Eye-Tracking Study on TypeScriptabstractStatically-annotated types have been shown to aid developers in a number of programming tasks, and this benefit holds true even when static type checking is not used. It is hypothesized that this is because developers use type annotations as in-code documentation. In this study, we aim to provide evidence that developers use type annotations as in-code documentation. Understanding this hypothesized use will help to understand how, and in what contexts, developers use type information; additionally, it may help to design better development tools and inform educational decisions. To provide this evidence, we conduct an eye tracking study with 26 undergraduate students to determine if they read type annotations during code comprehension and bug localization in the TypeScript language. We found that developers do not look directly at lines containing type annotations or type declarations more often when they are present, in either code summarization or bug localization tasks. The results have implications for tool builders to improve the availability of type information, the development community to build good standards for use of type annotations, and education to enforce deliberate teaching of reading patterns. Samuel W. Flint, Robert Dyer 0001, Bonita Sharif |
ICPC | 3 |
| 2026 | An exploratory eye tracking study on how developers classify and debug Python code in different paradigms
Samuel W. Flint, Jigyasa Chauhan, Niloofar Mansoor, Bonita Sharif, Robert Dyer 0001 |
Empir. Softw. Eng. | 4 |
| 2025 | Extending Support for Analyzing Eye Tracking Studies on Python Source Code in iTrace
Joshua Behler, Zachary Kozak, Kang-Il Park, Bonita Sharif, Jonathan I. Maletic |
ETRA | 4 |
| 2025 | How Developers Make Decisions When Choosing Issues and Reviewing Code: An Eye Tracking GitHub Study
Igor Scaliante Wiese, Jasmine Boyer, Ethan Rasgorshek, Gustavo Pinto 0001, Marco Aurélio Gerosa, Igor Steinmacher, Bonita Sharif |
ETRA | 7 |
| 2025 | Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research ScientistsabstractScientific software, defined as computer programs, scripts, or code used in scientific research, data analysis, modeling, or simulation, has become central to modern research. However, there is limited research on the readability and understandability of scientific code, both of which are vital for effective collaboration and reproducibility in scientific research. This study surveys 57 research scientists from various disciplines to explore their programming backgrounds, practices, and the challenges they face regarding code readability. Our findings reveal that most participants learn programming through self-study or on-the-job training, with$\mathbf{5 7. 9 \%}$lacking formal instruction in writing readable code. Scientists mainly use Python and$R$, relying on comments and documentation for readability. While most consider code readability essential for scientific reproducibility, they often face issues with inadequate documentation and poor naming conventions, with challenges including cryptic names and inconsistent conventions. Our findings also show low adoption of code quality tools and a trend towards utilizing large language models to improve code quality. These findings offer practical insights into enhancing coding practices and supporting sustainable development in scientific software. Alyssia Chen, Carol Wong, Bonita Sharif, Anthony Peruma |
ICPC | 3 |
| 2025 | Method Names in Jupyter Notebooks: An Exploratory StudyabstractMethod names play an important role in communicating the purpose and behavior of their functionality. Research has shown that high-quality names significantly improve code comprehension and the overall maintainability of software. However, these studies primarily focus on naming practices in traditional software development. There is limited research on naming patterns in Jupyter Notebooks, a popular environment for scientific computing and data analysis. In this exploratory study, we analyze the naming practices found in 691 methods across 384 Jupyter Notebooks, focusing on three key aspects: naming style conventions, grammatical composition, and the use of abbreviations and acronyms. Our findings reveal distinct characteristics of notebook method names, including a preference for conciseness and deviations from traditional naming patterns. We identified 68 unique grammatical patterns, with only 55.57 % of methods beginning with a verb. Further analysis revealed that half of the methods with return statements do not start with a verb. We also found that 30.39 % of method names contain abbreviations or acronyms, representing mathematical or statistical terms and image processing concepts, among others. We envision our findings contributing to developing specialized tools and techniques for evaluating and recommending high-quality names in scientific code and creating educational resources tailored to the notebook development community. Carol Wong, Gunnar Larsen, Rocky Huang, Bonita Sharif, Anthony Peruma |
ICPC | 4 |
| 2025 | Automated Fixation Error Correction to Support Eye Tracking Studies on Source CodeabstractA significant challenge in eye-tracking studies is detecting and fixing errors in data collection that happen for various reasons (drift, calibration issues, etc.). Many errors cannot be fully mitigated and require manual correction (which is intensively time-consuming) or automated correction. The work presented in this paper focuses on error correction, primarily on eye-tracking data on source code written in programming languages such as C++, Java, and C#. Many automated correction solutions are general-purpose, computationally inefficient, and use little information about the stimulus. To bridge this gap, we introduce srcGaze , a heuristic algorithm explicitly developed for correcting fixation gaze events in eye-tracking data from studies using source code as a stimulus. A golden dataset is manually constructed and verified to establish the heuristics. Results show a ≈40% improvement compared to no fixation correction. The approach has a multi-linear complexity and can correct over 44K fixations in approximately 6 seconds. Drew T. Guarnera, Joshua Behler, Bonita Sharif, Jonathan I. Maletic |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Integrating Lean Processes and Engineering Discipline into Work Culture Over 20 Years: An Experience ReportabstractThe paper presents an experience report for Don't Panic Labs (DPL), a software and consulting company, outlining the cultural shift in building software using lean processes and engineering principles over 20 years. The shift was intentional due to failures of prior projects. Challenges are identified along with opportunities to address them. An evolution of the company culture is explained via themes that were incorporated chronologically over time, which resulted in reducing errors in judgment, leading to reliable and predictable successes. Two factors, namely, the percentage of rework and co-creation sessions for efficient iteration of requirements, stand out as significantly impacting project success. Lessons learned with discussion and impact to the profession and education are presented. Doug Durham, Bonita Sharif |
ICSME | 2 |
| 2024 | Extending iTrace-Visualize to Support Token-based Heatmaps and Region of Interest Scarf Plots for Source CodeabstractThe iTrace Infrastructure is a suite of community eye-tracking tools that enables researchers to conduct eye-tracking studies on software projects in real development environments. The infrastructure consists of tools providing support for data gathering, post processing, and visualization. iTrace-Visualize provides researchers with a way to visualize gathered and post-processed eye-movement data. iTrace involves the analysis of more than just eye-movement data, and includes information gathered from the development environment and the source code. This work describes additions to iTrace-Visualize that provide researchers with visualizations of the gathered source code data. Specifically, a tokenized heatmap of the source code is presented, which shows the source code tokens that are viewed the most. Additionally, a region of interest scarf plot that details the timeline of what parts of the code a participant views is added as a new feature. A usage example comparing student and industry developers is presented to demonstrate the use of these tools. Demo Video: https://youtu.be/0iZcCC8CK94 Joshua Behler, Giovanni Villalobos, Julia Pangonis, Bonita Sharif, Jonathan I. Maletic |
VISSOFT | 4 |
| 2024 | Exploring How Developers Layout UML Class DiagramsabstractThe paper presents a video-based exploratory study that seeks to understand how developers modify UML class diagram layouts for better readability and comprehension of the system. Two diagram layouts showing a model subset from a Java open-source system are presented to six participants experienced in reading UML class diagrams. They are tasked to change the layout to make it easier for them to read and comprehend. The video is reviewed for major modifications to the layouts. Behaviors observed are presented. The eventual goal is to use this information to construct heuristics for automated layout algorithms based on semantics and architectural importance. Bonita Sharif, Nathaniel Liess, Jonathan I. Maletic |
VISSOFT | 1 |
| 2024 | Examining the Effects of Layout and Working Memory on UML Class Diagram Defect IdentificationabstractA controlled experiment investigating the effect layout has on how students find defects in UML class diagrams with respect to requirements is presented. Two layout schemes from prior literature namely, multi-cluster and orthogonal layouts, are compared with respect to two open source systems, Doxygen and Qt. The experiment is conducted with 89 students from two universities in a classroom lab setting. Each participant is placed in one of two groups where each group are given 2 defect detection tasks (with five sub-parts) with each task using one of the two layouts in each subject system. The only difference between groups is that the layouts were flipped between the two tasks. Feedback is collected after each task. A mental rotation and object memory task is conducted at the end of the two tasks to correlate their spatial and working memory skills to the task performance. Results indicate that the multi-cluster layout performed better in terms of accuracy of finding defects, but not significantly. There is also not much difference in time to find them. Furthermore, it is found that the object memory skills are sometimes correlated with the performance of the defect detection tasks. These results can be used to help improve the teaching of UML class diagram defect detection skills by incorporating clustered layouts and object memory tasks. In addition, they can help identify people who are best suited for finding critical defects in design. Bonita Sharif, Kang-Il Park, Michael P. DeJournett, Isaac Baysinger, Mohammed Aly, Jonathan I. Maletic |
VISSOFT | 1 |
| 2024 | An eye tracking study assessing source code readability rules for program comprehension
Kang-Il Park, Jack Johnson, Cole S. Peterson, Nishitha Yedla, Isaac Baysinger, Jairo Aponte, Bonita Sharif |
Empir. Softw. Eng. | 7 |
| 2023 | An eye tracking study assessing the impact of background styling in code editors on novice programmers' code understandingabstractBackground and Context: The designers of programming editors aimed at learners have long experimented with different styles of code presentation. The idea of syntax highlighting – coloring specific words – is very old. More recently, some editors (including text-, frame- and block-based editors) have added forms of scope highlighting – colored rectangles to represent programming scope – but there have been few studies to investigate whether this is beneficial for novices when reading and understanding program code. Kang-Il Park, Pierre Weill-Tessier, Neil Brown 0001, Bonita Sharif, Nikolaj Jensen, Michael Kölling |
ICER (1) | 4 |
| 2023 | iTrace-Visualize: Visualizing Eye-Tracking Data for Software Engineering StudiesabstractiTrace is community infrastructure that allows software engineering researchers to conduct eye-tracking studies on large realistic code bases. The iTrace infrastructure consists of a set of tools that assist with gathering, processing, and evaluating eye-tracking data on large software projects within an Integrated Development Environment (IDE). A typical eye-tracking study results in millions of raw gazes that are overwhelming to view and sort through. To help researchers view and comprehend this data, iTrace-Visualize is presented. This tool integrates information produced by the iTrace infrastructure into a dynamic video recording of the eye-tracking session. Eye fixations and the scan path between fixations are overlayed on the video. Additionally, the line being examined can be highlighted in the video. iTrace-Visualize lets a researcher replay eye fixations via a video overlay immediately after a study. This serves as quick validation of what was done during the study and can also provide quick insights into what the participants looked at. To illustrate iTrace-Visualize's capabilities, a small preliminary study is performed. Demo Video-https://youtu.be/c1hUFDmBM50 Joshua Behler, Gino Chiudioni, Alex Ely, Julia Pangonis, Bonita Sharif, Jonathan I. Maletic |
VISSOFT | 5 |
| 2023 | Studying Developer Eye Movements to Measure Cognitive Workload and Visual Effort for Expertise AssessmentabstractEye movement data provides valuable insights that help test hypotheses about a software developer's comprehension process. The pupillary response is successfully used to assess mental processing effort and attentional focus. Relatively little is known about the impact of expertise level in cognitive effort during programming tasks. This paper presents a quantitative analysis that compares the eye movements of 207 experts and novices collected while solving program comprehension tasks. The goal is to examine changes of developers' eye movement metrics in accordance with their expertise. The results indicate significant increase in pupil size with the novice group compared to the experts, explaining higher cognitive effort for novices. Novices also tend to have a significant number of fixations and higher gaze time compared to experts when they comprehend code. Moreover, a correlation study found that programming experience is still a powerful indicator when explaining expertise in this eye-tracking dataset among other expertise variables. Salwa D. Aljehane, Bonita Sharif, Jonathan I. Maletic |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Towards Modeling Human Attention from Eye Movements for Neural Source Code SummarizationabstractNeural source code summarization is the task of generating natural language descriptions of source code behavior using neural networks. A fundamental component of most neural models is an attention mechanism. The attention mechanism learns to connect features in source code to specific words to use when generating natural language descriptions. Humans also pay attention to some features in code more than others. This human attention reflects experience and high-level cognition well beyond the capability of any current neural model. In this paper, we use data from published eye-tracking experiments to create a model of this human attention. The model predicts which words in source code are the most important for code summarization. Next, we augment a baseline neural code summarization approach using our model of human attention. We observe an improvement in prediction performance of the augmented approach in line with other bio-inspired neural models. Aakash Bansal, Bonita Sharif, Collin McMillan |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | An Empirical Assessment on Merging and Repositioning of Static Analysis AlarmsabstractStatic analysis tools generate a large number of alarms that require manual inspection. In prior work, repositioning of alarms is proposed to (1) merge multiple similar alarms together and replace them by a fewer alarms, and (2) report alarms as close as possible to the causes for their generation. The premise is that the proposed merging and repositioning of alarms will reduce the manual inspection effort. To evaluate the premise, this paper presents an empirical study with 249 developers on the proposed merging and repositioning of static alarms. The study is conducted using static analysis alarms generated on$C$programs, where the alarms are representative of the merging vs. non-merging and repositioning vs. non-repositioning situations in real-life code. Developers were asked to manually inspect and determine whether assertions added corresponding to alarms in$C$code hold. Additionally, two spatial cognitive tests are also done to determine relationship in performance. The empirical evaluation results indicate that, in contrast to expectations, there was no evidence that merging and repositioning of alarms reduces manual inspection effort or improves the inspection accuracy (at times a negative impact was found). Results on cognitive abilities correlated with comprehension and alarm inspection accuracy. Niloofar Mansoor, Tukaram Muske, Alexander Serebrenik, Bonita Sharif |
SCAM | 4 |
| 2022 | Humans in Empirical Software Engineering Studies: An Experience ReportabstractThe use of human validation in software engineering methods, tools, and processes is crucial to understanding how these artifacts actually impact the people using them. In this paper, we report our experiences on two methods of data collection we have used in software engineering empirical studies, namely online questionnaire-based data collection and in-person eye tracking data collection using eye tracking equipment. The design and instrumentation challenges we faced are discussed with possible ways to mitigate them. We conclude with some guidelines and our vision for the future in human-centric studies in software engineering. Bonita Sharif, Niloofar Mansoor |
SANER | 1 |
| 2022 | Deja Vu: semantics-aware recording and replay of high-speed eye tracking and interaction data to support cognitive studies of software engineering tasks - methodology and analyses
Vlas Zyrianov, Cole S. Peterson, Drew T. Guarnera, Joshua Behler, Praxis Weston, Bonita Sharif, Jonathan I. Maletic |
Empir. Softw. Eng. | 6 |
| 2021 | From Novice to Expert: Analysis of Token Level Effects in a Longitudinal Eye Tracking StudyabstractProgram comprehension is a vital skill in software development. This work investigates program comprehension by examining the eye movement of novice programmers as they gain programming experience over the duration of a Java course. Their eye movement behavior is compared to the eye movement of expert programmers. Eye movement studies of natural text show that word frequency and length influence eye movement duration and act as indicators of reading skill. The study uses an existing longitudinal eye tracking dataset with 20 novice and experienced readers of source code. The work investigates the acquisition of the effects of token frequency and token length in source code reading as an indication of program reading skill. The results show evidence of the frequency and length effects in reading source code and the acquisition of these effects by novices. These results are then leveraged in a machine learning model demonstrating how eye movement can be used to estimate programming proficiency and classify novices from experts with 72% accuracy. Naser Al Madi, Cole S. Peterson, Bonita Sharif, Jonathan I. Maletic |
ICPC | 3 |
| 2020 | Automated Recording and Semantics-Aware Replaying of High-Speed Eye Tracking and Interaction Data to Support Cognitive Studies of Software Engineering TasksabstractThe paper introduces a fundamental technological problem with collecting high-speed eye tracking data while studying software engineering tasks in an integrated development environment. The use of eye trackers is quickly becoming an important means to study software developers and how they comprehend source code and locate bugs. High quality eye trackers can record upwards of 120 to 300 gaze points per second. However, it is not possible to map each of these points to a line and column position in a source code file (in the presence of scrolling and file switching) in real time at data rates over 60 gaze points per second without data loss. Unfortunately, higher data rates are more desirable as they allow for finer granularity and more accurate study analyses. To alleviate this technological problem, a novel method for eye tracking data collection is presented. Instead of performing gaze analysis in real time, all telemetry (keystrokes, mouse movements, and eye tracker output) data during a study is recorded as it happens. Sessions are then replayed at a much slower speed allowing for ample time to map gaze point positions to the appropriate file, line, and column to perform additional analysis. A description of the method and corresponding tool, Déjà Vu, is presented. An evaluation of the method and tool is conducted using three different eye trackers running at four different speeds (60Hz, 120Hz, 150Hz, and 300 Hz). This timing evaluation is performed in Visual Studio and Eclipse IDEs. Results show that Déjà Vu can playback 100% of the data recordings, correctly mapping the gaze to corresponding elements, making it a well-founded and suitable post processing step for future eye tracking studies in software engineering. Vlas Zyrianov, Drew T. Guarnera, Cole S. Peterson, Bonita Sharif, Jonathan I. Maletic |
ICSME | 4 |
| 2020 | A randomized controlled trial on the effects of embedded computer language switchingabstractPolyglot programming, the use of multiple programming languages during the development process, is common practice in modern software development. This study investigates this practice through a randomized controlled trial conducted under the context of database programming. Participants in the study were given coding tasks written in Java and one of three SQL-like embedded languages. One was plain SQL in strings, one was in Java only, and the third was a hybrid embedded language that was closer to the host language. We recorded 109 valid data points. Results showed significant differences in how developers of different experience levels code using polyglot techniques. Notably, less experienced programmers wrote correct programs faster in the hybrid condition (frequent, but less severe, switches), while more experienced developers that already knew both languages performed better in traditional SQL (less frequent but more complete switches). The results indicate that the productivity impact of polyglot programming is complex and experience level dependent. Phillip Merlin Uesbeck, Cole S. Peterson, Bonita Sharif, Andreas Stefik |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Studying Developer Reading Behavior on Stack Overflow during API Summarization TasksabstractStack Overflow is commonly used by software developers to help solve problems they face while working on software tasks such as fixing bugs or building new features. Recent research has explored how the content of Stack Overflow posts affects attraction and how the reputation of users attracts more visitors. However, there is very little evidence on the effect that visual attractors and content quantity have on directing gaze toward parts of a post, and which parts hold the attention of a user longer. Moreover, little is known about how these attractors help developers (students and professionals) answer comprehension questions. This paper presents an eye tracking study on thirty developers constrained to reading only Stack Overflow posts while summarizing four open source methods or classes. Results indicate that on average paragraphs and code snippets were fixated upon most often and longest. When ranking pages by number of appearance of code blocks and paragraphs, we found that while the presence of more code blocks did not affect number of fixations, the presence of increasing numbers of plain text paragraphs significantly drove down the fixations on comments. SO posts that were looked at only by students had longer fixation times on code elements within the first ten fixations. We found that 16 developer summaries contained 5 or more meaningful terms from SO posts they viewed. We discuss how our observations of reading behavior could benefit how users structure their posts. Jonathan Saddler, Cole S. Peterson, Sanjana Sama, Shruthi Nagaraj, Olga Baysal, Latifa Guerrouj, Bonita Sharif |
SANER | 7 |
| 2020 | A practical guide on conducting eye tracking studies in software engineering
Zohreh Sharafi, Bonita Sharif, Yann-Gaël Guéhéneuc, Andrew Begel, Roman Bednarik, Martha E. Crosby |
Empir. Softw. Eng. | 2 |
| 2020 | EMIP: The eye movements in programming datasetabstractA large dataset that contains the eye movements of N=216 programmers of different experience levels captured during two code comprehension tasks is presented. Data are grouped in terms of programming expertise (from none to high) and other demographic descriptors. Data were collected through an international collaborative effort that involved eleven research teams across eight countries on four continents. The same eye tracking apparatus and software was used for the data collection. The Eye Movements in Programming (EMIP) dataset is freely available for download. The varied metadata in the EMIP dataset provides fertile ground for the analysis of gaze behavior and may be used to make novel insights about code comprehension. Roman Bednarik, Teresa Busjahn, Agostino Gibaldi, Alireza Ahadi, Mária Bieliková, Martha E. Crosby, Kai Essig, Fabian Fagerholm, Ahmad Jbara, Raymond Lister, Pavel A. Orlov, James H. Paterson, Bonita Sharif, Teemu Sirkiä, Jan Stelovsky, Jozef Tvarozek, Hana Vrzakova, Ian van der Linde |
Sci. Comput. Program. | 13 |
| 2019 | Using developer eye movements to externalize the mental model used in code summarization tasksabstractEye movements of developers are used to speculate the mental cognition model (i.e., bottom-up or top-down) applied during program comprehension tasks. The cognition models examine how programmers understand source code by describing the temporary information structures in the programmer's short term memory. The two types of models that we are interested in are top-down and bottom-up. The top-down model is normally applied as-needed (i.e., the domain of the system is familiar). The bottom-up model is typically applied when a developer is not familiar with the domain or the source code. An eye-tracking study of 18 developers reading and summarizing Java methods is used as our dataset for analyzing the mental cognition model. The developers provide a written summary for methods assigned to them. In total, 63 methods are used from five different systems. The results indicate that on average, experts and novices read the methods more closely (using the bottom-up mental model) than bouncing around (using top-down). However, on average novices spend longer gaze time performing bottom-up (66s.) compared to experts (43s.) Nahla J. Abid, Jonathan I. Maletic, Bonita Sharif |
ETRA | 3 |
| 2019 | Visually analyzing eye movements on natural language texts and source code snippetsabstractIn this paper, we analyze eye movement data of 26 participants using a quantitative and qualitative approach to investigate how people read natural language text in comparison to source code. In particular, we use the radial transition graph visualization to explore strategies of participants during these reading tasks and extract common patterns amongst participants. We illustrate via examples how visualization can play a role at uncovering behavior of people while reading natural language text versus source code. Our results show that the linear reading order of natural text is only partially applicable to source code reading. We found patterns representing a linear order and also patterns that represent reading of the source code in execution order. Participants also focus more on those areas that are important to comprehend core functionality and we found that they skip unimportant constructs such as brackets. Tanja Blascheck, Bonita Sharif |
ETRA | 2 |
| 2019 | Factors influencing dwell time during source code reading: a large-scale replication experimentabstractThe paper partially replicates and extends a previous study by Busjahn et al. [4] on the factors influencing dwell time during source code reading, where source code element type and frequency of gaze visits are studied as factors. Unlike the previous study, this study focuses on analyzing eye movement data in large open source Java projects. Five experts and thirteen novices participated in the study where the main task is to summarize methods. The results examine semantic line-level information that developers view during summarization. We find no correlation between the line length and the total duration of time spent looking on the line even though it exists between a token's length and the total fixation time on the token reported in prior work. The first fixations inside a method are more likely to be on a method's signature, a variable declaration, or an assignment compared to the other fixations inside a method. In addition, it is found that smaller methods tend to have shorter overall fixation duration for the entire method, but have significantly longer duration per line in the method. The analysis provides insights into how source code's unique characteristics can help in building more robust methods for analyzing eye movements in source code and overall in building theories to support program comprehension on realistic tasks. Cole S. Peterson, Nahla J. Abid, Corey A. Bryant, Jonathan I. Maletic, Bonita Sharif |
ETRA | 5 |
| 2019 | Developer reading behavior while summarizing Java methods: size and context mattersabstractAn eye-tracking study of 18 developers reading and summarizing Java methods is presented. The developers provide a written summary for methods assigned to them. In total, 63 methods are used from five different systems. Previous studies on this topic use only short methods presented in isolation usually as images. In contrast, this work presents the study in the Eclipse IDE allowing access to all the source code in the system. The developer can navigate via scrolling and switching files while writing the summary. New eye-tracking infrastructure allows for this improvement in the study environment. Data collected includes eye gazes on source code, written summaries, and time to complete each summary. Unlike prior work that concluded developers focus on the signature the most, these results indicate that they tend to focus on the method body more than the signature. Moreover, both experts and novices tend to revisit control flow terms rather than reading them for a long period. They also spend a significant amount of gaze time and have higher gaze visits when they read call terms. Experts tend to revisit the body of the method significantly more frequently than its signature as the size of the method increases. Moreover, experts tend to write their summaries from source code lines that they read the most. Nahla J. Abid, Bonita Sharif, Natalia Dragan, Hend Alrasheed, Jonathan I. Maletic |
ICSE | 2 |
| 2019 | An Empirical Study Assessing Source Code Readability in ComprehensionabstractSoftware developers spend a significant amount of time reading source code. If code is not written with readability in mind, it impacts the time required to maintain it. In order to alleviate the time taken to read and understand code, it is important to consider how readable the code is. The general consensus is that source code should be written to minimize the time it takes for others to read and understand it. In this paper, we conduct a controlled experiment to assess two code readability rules: nesting and looping. We test 32 Java methods in four categories: ones that follow/do not follow the readability rule and that are correct/incorrect. The study was conducted online with 275 participants. The results indicate that minimizing nesting decreases the time a developer spends reading and understanding source code, increases confidence about the developer's understanding of the code, and also suggests that it improves their ability to find bugs. The results also show that avoiding the do-while statement had no significant impact on level of understanding, time spent reading and understanding, confidence in understanding, or ease of finding bugs. It was also found that the better knowledge of English a participant had, the more their readability and comprehension confidence ratings were affected by the minimize nesting rule. We discuss the implications of these findings for code readability and comprehension. Sergio Lubo, Nishitha Yedla, Jairo Aponte, Bonita Sharif |
ICSME | 5 |
| 2019 | Exploring Eye Tracking Data on Source Code via Dual Space AnalysisabstractEye tracking is a frequently used technique to collect data capturing users' strategies and behaviors in processing information. Understanding how programmers navigate through a large number of classes and methods to find bugs is important to educators and practitioners in software engineering. However, the eye tracking data collected on realistic codebases is massive compared to traditional eye tracking data on one static page. The same content may appear in different areas on the screen with users scrolling in an Integrated Development Environment (IDE). Hierarchically structured content and fluid method position compose the two major challenges for visualization. We present a dual-space analysis approach to explore eye tracking data by leveraging existing software visualizations and a new graph embedding visualization. We use the graph embedding technique to quantify the distance between two arbitrary methods, which offers a more accurate visualization of distance with respect to the inherent relations, compared with the direct software structure and the call graph. The visualization offers both naturalness and readability showing time-varying eye movement data in both the content space and the embedded space, and provides new discoveries in developers' eye tracking behaviors. Jianxin Sun 0001, Cole S. Peterson, Bonita Sharif, Hongfeng Yu 0001 |
VISSOFT | 4 |
| 2018 | iTrace: eye tracking infrastructure for development environmentsabstractThe paper presents iTrace, an eye tracking infrastructure, that enables eye tracking in development environments such as Visual Studio and Eclipse. Software developers work with software that is comprised of numerous source code files. This requires frequent switching between project artifacts during program understanding or debugging activities. Additionally, the amount of content contained within each artifact can be quite large and require scrolling or navigation of the content. Current approaches to eye tracking are meant for fixed stimuli and struggle to capture context during these activities. iTrace overcomes these limitations allowing developers to work in realistic settings during an eye tracking study. The iTrace architecture is presented along with several use cases of where it can be used by researchers. A short video demonstration is available at https://youtu.be/AmrLWgw4OEs Drew T. Guarnera, Corey A. Bryant, Ashwin Mishra, Jonathan I. Maletic, Bonita Sharif |
ETRA | 5 |
| 2017 | iTraceVis: Visualizing Eye Movement Data Within EclipseabstractThe paper presents iTraceVis, an eye tracking visualization component for iTrace, a gaze-aware Eclipse plugin. The visualization component is designed to work with data generated from iTrace after an eye tracking session. iTrace provides us with an automatic mapping of raw eye gaze on corresponding source code elements according to their hierarchy in the abstract syntax graph even in the presence of scrolling and context switching between files. This feature provides us with the ability to visualize eye tracking data in large source code files and is not just restricted to visualize only a method at a time i.e., something that fits on the screen. Due to the enormous size and richness of data collected from the eye tracker, visualizations help both the researcher and the developer to comprehend what transpired during an eye tracking session. iTraceVis currently supports four visualization views - heat map, gaze skyline, static gaze map, and dynamic gaze map. In order to determine the usefulness of the visualizations, we conduct a pilot user study with 10 senior students. We present an existing eye-tracking developer session to the study participants and ask them a series of questions while they interacted with iTraceVis and its various views in Eclipse. The results indicate that the visualizations do indeed help users understand the visualized session represented by the data, and also provide insight into where the visualizations can be improved as part of future work. Benjamin Clark, Bonita Sharif |
VISSOFT | 2 |
| 2017 | Eye movements in software traceability link recovery
Bonita Sharif, John Meinken, Timothy Shaffer, Huzefa H. Kagdi |
Empir. Softw. Eng. | 1 |
| 2017 | Eye gaze and interaction contexts for change tasks - Observations and potential
Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001 |
J. Syst. Softw. | 4 |
| 2016 | Towards automating fixation correction for source codeabstractDuring eye-tracking studies there is a possibility for the actual fixation to shift a little when recorded. The cause of this shift could be due to various reasons such as the accuracy of the calibration or drift. Researchers usually correct fixations manually. Manual corrections are error prone especially if done on large samples for extended periods. There is also no guarantee that two corrections done by different people on the same data set will be consistent with each other. In order to solve this problem, we introduce an attempt at automatically correcting fixations that uses a variable offset for groups of fixations. Our focus is on source code, which is read differently than natural language requiring an algorithm that adapts to these differences. We introduce a Hill Climbing algorithm that shifts fixations to a best-fit location based on a scoring function. In order to evaluate the algorithm's effectiveness, we compare the automatically corrected fixations against a set of manually corrected ones, giving us an accuracy of 89%. These findings are discussed with additional ways to improve the algorithm. Christopher Palmer, Bonita Sharif |
ETRA | 2 |
| 2016 | iTrace: Overcoming the Limitations of Short Code Examples in Eye Tracking ExperimentsabstractSummary form only given. Eye trackers are being used by software engineering researchers to study how developers work. In this technical briefing, we give an overview of eye tracking and how it can help researchers to conduct their own studies. Eye tracking studies are done on a single screen of text and there is no support for scrolling or switching between files. This scenario is impractical to study developers as they actually work on large software artifacts. To overcome this an Eclipse plugin, iTrace, is introduced that monitors developers eye movements even in the presence of scrolling and file switching within an IDE. In addition, it automatically maps the eye gaze to source code elements. Existing work using iTrace is presented followed by a scenario of how to setup and run an eye tracking study. Data filtering, data cleaning, and data analysis are also discussed. Bonita Sharif, Jonathan I. Maletic |
ICSME | 1 |
| 2016 | Analyzing developer sentiment in commit logsabstractThe paper presents an analysis of developer commit logs for GitHub projects. In particular, developer sentiment in commits is analyzed across 28,466 projects within a seven year time frame. We use the Boa infrastructure's online query system to generate commit logs as well as files that were changed during the commit. We analyze the commits in three categories: large, medium, and small based on the number of commits using a sentiment analysis tool. In addition, we also group the data based on the day of week the commit was made and map the sentiment to the file change history to determine if there was any correlation. Although a majority of the sentiment was neutral, the negative sentiment was about 10% more than the positive sentiment overall. Tuesdays seem to have the most negative sentiment overall. In addition, we do find a strong correlation between the number of files changed and the sentiment expressed by the commits the files were part of. Future work and implications of these results are discussed. Vinayak Sinha, Alina Lazar, Bonita Sharif |
MSR | 3 |
| 2016 | Studying developer gaze to empower software engineering research and practiceabstractA new research paradigm is proposed that leverages developer eye gaze to improve the state of the art in software engineering research and practice. The vision of this new paradigm for use on software engineering tasks such as code summarization, code recommendations, prediction, and continuous traceability is described. Based on this new paradigm, it is foreseen that new benchmarks will emerge based on developer gaze. The research borrows from cognitive psychology, artificial intelligence, information retrieval, and data mining. It is hypothesized that new algorithms will be discovered that work with eye gaze data to help improve current IDEs, thus improving developer productivity. Conducting empirical studies using an eye tracker will lead to inventing, evaluating, and applying innovative methods and tools that use eye gaze to support the developer. The implications and challenges of this paradigm for future software engineering research is discussed. Bonita Sharif, Benjamin Clark, Jonathan I. Maletic |
SIGSOFT FSE | 1 |
| 2016 | An empirical study on the effect of 3D visualization for project tasks and resources
Khaled Jaber, Bonita Sharif, Chang Liu 0028 |
J. Syst. Softw. | 2 |
| 2015 | Eye-Tracking Metrics in Software EngineeringabstractEye-tracking studies are getting more prevalent in software engineering. Researchers often use different metrics when publishing their results in eye-tracking studies. Even when the same metrics are used, they are given different names, causing difficulties in comparing studies. To encourage replications and facilitate advancing the state of the art, it is important that the metrics used by researchers be clearly and consistently defined in the literature. There is therefore a need for a survey of eye-tracking metrics to support the (future) goal of standardizing eye-tracking metrics. This paper seeks to bring awareness to the use of different metrics along with practical suggestions on using them. It compares and contrasts various eye-tracking metrics used in software engineering. It also provides definitions for common metrics and discusses some metrics that the software engineering community might borrow from other fields. Zohreh Sharafi, Timothy Shaffer, Bonita Sharif, Yann-Gaël Guéhéneuc |
APSEC | 3 |
| 2015 | Eye movements in code reading: relaxing the linear orderabstractCode reading is an important skill in programming. Inspired by the linearity that people exhibit while natural language text reading, we designed local and global gaze-based measures to characterize linearity (left-to-right and top-to-bottom) in reading source code. Unlike natural language text, source code is executable and requires a specific reading approach. To validate these measures, we compared the eye movements of novice and expert programmers who were asked to read and comprehend short snippets of natural language text and Java programs. Our results show that novices read source code less linearly than natural language text. Moreover, experts read code less linearly than novices. These findings indicate that there are specific differences between reading natural language and source code, and suggest that non-linear reading skills increase with expertise. We discuss the implications for practitioners and educators. Teresa Busjahn, Roman Bednarik, Andrew Begel, Martha E. Crosby, James H. Paterson, Carsten Schulte 0001, Bonita Sharif, Sascha Tamm |
ICPC | 7 |
| 2015 | Tracing software developers' eyes and interactions for change tasksabstractWhat are software developers doing during a change task? While an answer to this question opens countless opportunities to support developers in their work, only little is known about developers' detailed navigation behavior for realistic change tasks. Most empirical studies on developers performing change tasks are limited to very small code snippets or are limited by the granularity or the detail of the data collected for the study. In our research, we try to overcome these limitations by combining user interaction monitoring with very fine granular eye-tracking data that is automatically linked to the underlying source code entities in the IDE. In a study with 12 professional and 10 student developers working on three change tasks from an open source system, we used our approach to investigate the detailed navigation of developers for realistic change tasks. The results of our study show, amongst others, that the eye tracking data does indeed capture different aspects than user interaction data and that developers focus on only small parts of methods that are often related by data flow. We discuss our findings and their implications for better developer tool support. Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2015 | iTrace: enabling eye tracking on software artifacts within the IDE to support software engineering tasksabstractThe paper presents iTrace, an Eclipse plugin that implicitly records developers' eye movements while they work on change tasks. iTrace is the first eye tracking environment that makes it possible for researchers to conduct eye tracking studies on large software systems. An overview of the design and architecture is presented along with features and usage scenarios. iTrace is designed to support a variety of eye trackers. The design is flexible enough to record eye movements on various types of software artifacts (Java code, text/html/xml documents, diagrams), as well as IDE user interface elements. The plugin has been successfully used for software traceability tasks and program comprehension tasks. iTrace is also applicable to other tasks such as code summarization and code recommendations based on developer eye movements. A short video demonstration is available at https://youtu.be/3OUnLCX4dXo. Timothy Shaffer, Jenna DiVincenzo, Braden Walters, Sebastian C. Müller, Michael Falcone, Bonita Sharif |
ESEC/SIGSOFT FSE | 6 |
| 2014 | An eye-tracking study assessing the comprehension of c++ and Python source codeabstractA study to assess the effect of programming language on student comprehension of source code is presented, comparing the languages of C++ and Python in two task categories: overview and find bug tasks. Eye gazes are tracked while thirty-eight students complete tasks and answer questions. Results indicate no significant difference in accuracy or time, however there is a significant difference reported on the rate at which students look at buggy lines of code. These results start to provide some direction as to the effect programming language might have in introductory programming classes. Rachel Turner, Michael Falcone, Bonita Sharif, Alina Lazar |
ETRA | 3 |
| 2014 | Eye tracking in computing educationabstractThe methodology of eye tracking has been gradually making its way into various fields of science, assisted by the diminishing cost of the associated technology. In an international collaboration to open up the prospect of eye movement research for programming educators, we present a case study on program comprehension and preliminary analyses together with some useful tools. Teresa Busjahn, Carsten Schulte 0001, Bonita Sharif, Simon, Andrew Begel, Michael Hansen, Roman Bednarik, Pavel A. Orlov, Petri Ihantola, Galina Shchekotova, Maria Antropova |
ICER | 3 |
| 2014 | Capturing software traceability links from developers' eye gazesabstractThe paper presents a novel approach for recovering software traceability links from developers' eye gazes. An eye tracker is used to capture eye gazes while developers perform software maintenance tasks within the Eclipse IDE. An algorithm is presented that establishes a set of traceability links from the eye-gaze data of several developer sessions. A preliminary study assesses the feasibility and validity of the approach. The links generated by the approach were validated by another set of developers. Results indicate that our algorithm achieves strong recall when developers accurately perform bug-localization tasks. Braden Walters, Timothy Shaffer, Bonita Sharif, Huzefa H. Kagdi |
ICPC | 3 |
| 2014 | Improving the accuracy of duplicate bug report detection using textual similarity measuresabstractThe paper describes an improved method for automatic duplicate bug report detection based on new textual similarity features and binary classification. Using a set of new textual features, inspired from recent text similarity research, we train several binary classification models. A case study was conducted on three open source systems: Eclipse, Open Office, and Mozilla to determine the effectiveness of the improved method. A comparison is also made with current state-of-the-art approaches highlighting similarities and differences. Results indicate that the accuracy of the proposed method is better than previously reported research with respect to all three systems. Alina Lazar, Sarah Ritchey, Bonita Sharif |
MSR | 3 |
| 2014 | Generating duplicate bug datasetsabstractAutomatic identification of duplicate bug reports is an important research problem in the mining software repositories field. This paper presents a collection of bug datasets collected, cleaned and preprocessed for the duplicate bug report identification problem. The datasets were extracted from open-source systems that use Bugzilla as their bug tracking component and contain all the bugs ever submitted. The systems used are Eclipse, Open Office, NetBeans and Mozilla. For each dataset, we store both the initial data and the cleaned data in separate collections in a mongoDB document-oriented database. For each dataset, in addition to the bug data collections downloaded from bug repositories, the database includes a set of all pairs of duplicate bugs together with randomly selected pairs of non-duplicate bugs. Such a dataset is useful as input for classification models and forms a good base to support replications and comparisons by other researchers. We used a subset of this data to predict duplicate bug reports but the same data set may also be used to predict bug priorities and severity. Alina Lazar, Sarah Ritchey, Bonita Sharif |
MSR | 3 |
| 2013 | OnionUML: An Eclipse plug-in for visualizing UML class diagrams in onion graph notationabstractThis paper presents OnionUML, an Eclipse plug-in that reduces the number of visible classes in a UML class diagram while preserving structure and semantics of the UML elements. Compaction of class elements is done using onion graph notation. The goal is that developers will be able to view and understand subsystems of a large software system while being able to visualize how that subsystem fits into the whole system. Michael Falcone, Bonita Sharif |
ICPC | 2 |
| 2013 | An empirical study assessing the effect of seeit 3D on comprehensionabstractA study to assess the effect of SeeIT 3D, a software visualization tool is presented. Six different tasks in three different task categories are assessed in the context of a large open-source system. Ninety-seven subjects were recruited from three different universities to participate in the study. Two methods of data collection: traditional questionnaires and an eye-tracker were used. The main goal was to determine the impact and added benefit of SeeIT 3D while performing typical software tasks within the Eclipse IDE. Results indicate that SeeIT 3D performs significantly better in one task category namely overview tasks but takes significantly longer when completing bug fixing tasks. Scores obtained by the subjects in the SeeIT 3D group are 13% better and 45% faster for overview tasks. Bonita Sharif, Grace Jetty, Jairo Aponte, Esteban Parra |
VISSOFT | 1 |
| 2013 | The impact of identifier style on effort and comprehension
Dave W. Binkley, Marcia Davis, Dawn J. Lawrie, Jonathan I. Maletic, Christopher Morrell, Bonita Sharif |
Empir. Softw. Eng. | 6 |
| 2012 | An eye-tracking study on the role of scan time in finding source code defectsabstractAn eye-tracking study is presented that investigates how individuals find defects in source code. This work partially replicates a previous eye-tracking study by Uwano et al. [2006]. In the Uwano study, eye movements are used to characterize the performance of individuals in reviewing source code. Their analysis showed that subjects who did not spend enough time initially scanning the code tend to take more time finding defects. The study here follows a similar setup with added eye-tracking measures and analyses on effectiveness and efficiency of finding defects with respect to eye gaze. The subject pool is larger and is comprised of a varied skill level. Results indicate that scanning significantly correlates with defect detection time as well as visual effort on relevant defect lines. Results of the study are compared and contrasted to the Uwano study. Bonita Sharif, Michael Falcone, Jonathan I. Maletic |
ETRA | 1 |
| 2011 | Empirical assessment of UML class diagram layouts based on architectural importanceabstractThe paper presents a family of experiments that investigate the effectiveness of different layout techniques for class diagrams in the Unified Modeling Language (UML). Three different layout schemes are examined based on architectural importance of class stereotypes. The premise is that layout techniques for UML class diagrams significantly impact comprehension. Both traditional questionnaire-based studies as well as eye-tracking studies are done to quantitatively measure the performance of subjects solving specific software maintenance tasks. The main contribution is the detailed empirical validation of a set of layout techniques with respect to a variety of software maintenance tasks. Results indicate that layout plays a significant role in the comprehension of UML class diagrams. In particular, there is a significant improvement in accuracy, time, and visual effort for one particular layout scheme, namely multi-cluster. The end goal is to determine effective ways to adjust the layout of existing UML class diagrams to support program comprehension during maintenance. Bonita Sharif |
ICSM | 1 |
| 2010 | The Effects of Layout on Detecting the Role of Design PatternsabstractA controlled experiment investigating the effect layout has on how students identify design pattern roles in UML class diagrams is presented. Two layout schemes, multi-cluster and orthogonal, are compared with respect to three open source systems and four design patterns. Seventeen students were asked a series of eight design pattern role detection (comprehension) questions for each layout, followed by eight preference rating questions. Results indicate a significant improvement in role detection accuracy with the multi-cluster layout for the strategy pattern and a significant improvement in detection time with the multi-cluster layout for all four patterns. Preference ratings significantly favored the multi-cluster layout for pattern role detection ease. These results can be used to help improve the teaching of design patterns. Bonita Sharif, Jonathan I. Maletic |
CSEE&T | 1 |
| 2010 | An eye tracking study on the effects of layout in understanding the role of design patternsabstractThe effect of layout in the comprehension of design pattern roles in UML class diagrams is assessed. This work replicates and extends a previous study using questionnaires but uses an eye tracker to gather additional data. The purpose of the replication is to gather more insight into the eye gaze behavior not evident from questionnaire-based methods. Similarities and differences between the studies are presented. Four design patterns are examined in two layout schemes in the context of three open source systems. Fifteen participants answered a series of eight design pattern role detection questions. Results show a significant improvement in role detection accuracy and visual effort with a certain layout for the Strategy and Observer patterns and a significant improvement in role detection time for all four patterns. Eye gaze data indicates classes participating in a design pattern act like visual beacons when they are in close physical proximity and follow the canonical layout, even though they violate some general graph aesthetics. Bonita Sharif, Jonathan I. Maletic |
ICSM | 1 |
| 2010 | An Eye Tracking Study on camelCase and under_score Identifier StylesabstractAn empirical study to determine if identifier-naming conventions (i.e., camelCase and under_score) affect code comprehension is presented. An eye tracker is used to capture quantitative data from human subjects during an experiment. The intent of this study is to replicate a previous study published at ICPC 2009 (Binkley et al.) that used a timed response test method to acquire data. The use of eye-tracking equipment gives additional insight and overcomes some limitations of traditional data gathering techniques. Similarities and differences between the two studies are discussed. One main difference is that subjects were trained mainly in the underscore style and were all programmers. While results indicate no difference in accuracy between the two styles, subjects recognize identifiers in the underscore style more quickly. Bonita Sharif, Jonathan I. Maletic |
ICPC | 1 |
| 2009 | An empirical study on the comprehension of stereotyped UML class diagram layoutsabstractAn empirical study is presented that investigates how stereotype based layouts impact the comprehension of UML class diagrams. This work replicates a previous study using eye-tracking equipment but uses online questionnaires instead. Subjects were given two types of tasks: one addressing UML syntax and the other addressing software design. Three different layout strategies are compared. Along with general aesthetics, the layouts are primarily organized by class stereotypes of control, boundary, and entity. A confidence value for each question was collected from the subjects to help validate the categorization of subjects. Results of the study are compared and contrasted to the eye-tracking study done with the same tasks and layouts. Results show a significant improvement in performance in both types of tasks with the multi-cluster stereotyped layouts. Bonita Sharif, Jonathan I. Maletic |
ICPC | 1 |
| 2007 | Mining Software Repositories for Traceability LinksabstractAn approach to recover/discover traceability links between software artifacts via the examination of a software system's version history is presented. A heuristic-based approach that uses sequential-pattern mining is applied to the commits in software repositories for uncovering highly frequent co-changing sets of artifacts (e.g., source code and documentation). If different types of files are committed together with high frequency then there is a high probability that they have a traceability link between them. The approach is evaluated on a number of versions of the open source system KDE. As a validation step, the discovered links are used to predict similar changes in the newer versions of the same system. The results show highly precision predictions of certain types of traceability links. Huzefa H. Kagdi, Jonathan I. Maletic, Bonita Sharif |
ICPC | 3 |