Kostadin Damevski

dblp:d/KostadinDamevski · DBLP profile ↗
← Back
51ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-7799-2026ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 37 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Systems, architecture and hardware · 6 · 3 first-authorHuman-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Psycholinguistic analyses in software engineering text: A systematic mapping study
abstract
Context: A deeper understanding of human factors in software engineering (SE) is essential for improving team collaboration, decision-making, and productivity. Communication channels like code reviews and chats provide insights into developers’ psychological and emotional states. While large language models excel at text analysis, they often lack transparency and precision. Psycholinguistic tools like Linguistic Inquiry and Word Count (LIWC) offer clearer, interpretable insights into cognitive and emotional processes exhibited in text. Despite its wide use in SE research, no comprehensive mapping study of LIWC’s use has been conducted. Objective: We examine the importance of psycholinguistic tools, particularly LIWC, and provide a thorough analysis of its current and potential future applications in SE research. Methods: We conducted a systematic mapping study of six prominent databases, identifying 43 SE-related papers using LIWC. Our analysis focuses on five research questions: RQ1. How was LIWC employed in SE studies, and for what purposes?, RQ2. What datasets were analyzed using LIWC?, RQ3: What Behavioral Software Engineering (BSE) concepts were studied using LIWC? RQ4: How often has LIWC been evaluated in SE research?, RQ5: What concerns were raised about adopting LIWC in SE? Results: Our findings reveal a wide range of applications, including analyzing team communication to detect developer emotions and personality, developing ML models to predict deleted Stack Overflow posts, and more recently comparing AI-generated and human-written text. LIWC has been primarily used with data from project management platforms (e.g., GitHub) and Q&A forums (e.g., Stack Overflow). Key BSE concepts include Communication , Organizational Climate , and Positive Psychology . 26 of 43 papers did not formally evaluate LIWC. Concerns were raised about some limitations, including difficulty handling SE-specific vocabulary. Conclusion: We highlight the potential of psycholinguistic tools and their limitations, and present new use cases for advancing research on human factors in SE (e.g., bias in human-LLM conversations).
Amirali Sajadi, Kostadin Damevski, Preetha Chatterjee
Inf. Softw. Technol.2
2026 A weekly project newsletter to engage episodic participants in OSS projects
Christian Novalski, Ghalian Fayyadh, Christopher Chavez, Kostadin Damevski
J. Syst. Softw.4
2025 Detecting Natural Emotions in Virtual Reality Through Facial Movement Analysis
abstract
Facial expression recognition (FER) plays a key role in enabling emotionally responsive systems in virtual reality (VR). However, most existing FER datasets in VR are constructed from acted emotions, which may not accurately reflect how emotions are naturally expressed. Prior work suggests that natural emotions differ from acted ones in terms of intensity, subtlety, and temporal dynamics. This mismatch raises concerns about the real-world performance of FER models trained solely on acted data.
Terens Tare, Rahat Rizvi Rahman, Hee Yun Choi, Joonghyo Lim, Go Eun Lee, Seungmoo Lee, Chungyean Cho, Kostadin Damevski
VRST8
2025 Do LLMs consider security? an empirical study on responses to programming questions
abstract
Abstract The widespread adoption of conversational LLMs for software development has raised new security concerns regarding the safety of LLM-generated content. Our motivational study outlines ChatGPT’s potential in volunteering context-specific information to the developers, promoting safe coding practices. Motivated by this finding, we conduct a study to evaluate the degree of security awareness exhibited by three prominent LLMs: Claude 3, GPT-4, and Llama 3. We prompt these LLMs with Stack Overflow questions that contain vulnerable code to evaluate whether they merely provide answers to the questions or if they also warn users about the insecure code, thereby demonstrating a degree of security awareness. Further, we assess whether LLM responses provide information about the causes, exploits, and the potential fixes of the vulnerability, to help raise users’ awareness. Our findings show that all three models struggle to accurately detect and warn users about vulnerabilities, achieving a detection rate of only 12.6% to 40% across our datasets. We also observe that the LLMs tend to identify certain types of vulnerabilities related to sensitive information exposure and improper input neutralization much more frequently than other types, such as those involving external control of file names or paths. Furthermore, when LLMs do issue security warnings, they often provide more information on the causes, exploits, and fixes of vulnerabilities compared to Stack Overflow responses. Finally, we provide an in-depth discussion on the implications of our findings, and demonstrated a CLI-based prompting tool that can be used to produce more secure LLM responses.
Amirali Sajadi, Binh Le, Kostadin Damevski, Preetha Chatterjee
Empir. Softw. Eng.4
2025 Towards higher quality software vulnerability data using LLM-based patch filtering
Charlie Dil, Hui Chen 0001, Kostadin Damevski
J. Syst. Softw.3
2025 Improving Data Curation of Software Vulnerability Patches through Uncertainty Quantification
abstract
The changesets (or patches) that fix open source software vulnerabilities form critical datasets for various machine learning security-enhancing applications, such as automated vulnerability patching and silent fix detection. These patch datasets are derived from extensive collections of historical vulnerability fixes, maintained in databases like the Common Vulnerabilities and Exposures list and the National Vulnerability Database. However, since these databases focus on rapid notification to the security community, they contain significant inaccuracies and omissions that have a negative impact on downstream software security quality assurance tasks.In this paper, we propose an approach employing Uncertainty Quantification (UQ) to curate datasets of publicly-available software vulnerability patches. Our methodology leverages machine learning models that incorporate UQ to differentiate between patches based on their potential utility. We begin by evaluating a number of popular UQ techniques, including Vanilla, Monte Carlo Dropout, and Model Ensemble, as well as homoscedastic and heteroscedastic models of noise. Our findings indicate that Model Ensemble and heteroscedastic models are the best choices for vulnerability patch datasets. Based on these UQ modeling choices, we propose a heuristic that uses UQ to filter out lower quality instances and select instances with high utility value from the vulnerability dataset. Using our approach, we observe an improvement in predictive performance and a significant reduction of model training time (i.e., energy consumption) for a state-of-the-art vulnerability prediction model.
Hui Chen 0001, Yunhua Zhao, Kostadin Damevski
IEEE Trans. Software Eng.3
2024 Utilizing Real-World Software Vulnerabilities to Enhance Secure Programming Education
abstract
This research paper describes a study of using real-world vulnerabilities to motivate computer science students to-wards learning secure programming. Given the rise in cybersecurity incidents due to programming errors, there is a pressing need to improve programmers' secure programming skills. Despite educators' numerous efforts towards this goal, communicating the importance of this training to students remains a challenge. Grounding on the theory of intrinsic motivation, we propose that exposing students to authentic, relatable vulnerabilities can significantly enhance their learning orientation towards secure programming. Our approach involves selecting vulnerabilities from the National Vulnerability Database that are both relatable to students and understandable without extensive external context. These vulnerabilities are transformed into comprehensive course modules, each featuring a demonstrative video, source code snippets of the vulnerability and its patch, and associated developer communications about the vulnerability. We assess the impact of one of our course modules on students' learning disposition through a study conducted in two universities in an identical setting. The study results indicate that students appreciate seeing real-world vulnerabilities in detail, especially the video we recorded reproducing the vulnerability, and that they gain in self-efficacy after completing the module.
Denise Daniels, Joon-Suk Lee, Hui Chen 0001, Kostadin Damevski
FIE4
2024 Uncovering the Causes of Emotions in Software Developer Communication Using Zero-shot LLMs
abstract
Understanding and identifying the causes behind developers' emotions (e.g., Frustration caused by 'delays in merging pull requests') can be crucial towards finding solutions to problems and fostering collaboration in open-source communities. Effectively identifying such information in the high volume of communications across the different project channels, such as chats, emails, and issue comments, requires automated recognition of emotions and their causes. To enable this automation, large-scale software engineering-specific datasets that can be used to train accurate machine learning models are required. However, such datasets are expensive to create with the variety and informal nature of software projects' communication channels.
Mia Mohammad Imran, Preetha Chatterjee, Kostadin Damevski
ICSE3
2024 Shedding Light on Software Engineering-specific Metaphors and Idioms
abstract
Use of figurative language, such as metaphors and idioms, is common in our daily-life communications, and it can also be found in Software Engineering (SE) channels, such as comments on GitHub. Automatically interpreting figurative language is a challenging task, even with modern Large Language Models (LLMs), as it often involves subtle nuances. This is particularly true in the SE domain, where figurative language is frequently used to convey technical concepts, often bearing developer affect (e.g., 'spaghetti code). Surprisingly, there is a lack of studies on how figurative language in SE communications impacts the performance of automatic tools that focus on understanding developer communications, e.g., bug prioritization, incivility detection. Furthermore, it is an open question to what extent state-of-the-art LLMs interpret figurative expressions in domain-specific communication such as software engineering. To address this gap, we study the prevalence and impact of figurative language in SE communication channels. This study contributes to understanding the role of figurative language in SE, the potential of LLMs in interpreting them, and its impact on automated SE communication analysis. Our results demonstrate the effectiveness of fine-tuning LLMs with figurative language in SE and its potential impact on automated tasks that involve affect. We found that, among three state-of-the-art LLMs, the best improved fine-tuned versions have an average improvement of 6.66% on a GitHub emotion classification dataset, 7.07% on a GitHub incivility classification dataset, and 3.71% on a Bugzilla bug report prioritization dataset.
Mia Mohammad Imran, Preetha Chatterjee, Kostadin Damevski
ICSE3
2024 Customizing ChatGPT to Help Computer Science Principles Students Learn Through Conversation
abstract
This paper explores leveraging conversational agents, specifically ChatGPT, to enhance the introduction of computing, focused on the Advanced Placement Computer Science Principles (CSP) course in secondary schools. Despite the potential benefits for diverse student audiences, little research has investigated their effectiveness and engagement in this context. We examine the customization of ChatGPT for secondary school CSP students, assessing its impact on exploratory searches for learning CSP concepts. Results from 20 high school students in grades 10-12 (ages 15-18) in a CSP course indicate that students preferred a customized ChatGPT, with its terminology more suitable to secondary school level, examples more understandable, and better connections to personal experiences compared to standard ChatGPT.
Matthew Frazier, Kostadin Damevski, Lori L. Pollock
ITiCSE (1)2
2024 Incivility in Open Source Projects: A Comprehensive Annotated Dataset of Locked GitHub Issue Threads
abstract
In the dynamic landscape of open source software (OSS) development, understanding and addressing incivility within issue discussions is crucial for fostering healthy and productive collaborations. This paper presents a curated dataset of 404 locked GitHub issue discussion threads and 5961 individual comments, collected from 213 OSS projects. We annotated the comments with various categories of incivility using Tone Bearing Discussion Features (TBDFs), and, for each issue thread, we annotated the triggers, targets, and consequences of incivility. We observed that Bitter frustration, Impatience, and Mocking are the most prevalent TBDFs exhibited in our dataset. The most common triggers, targets, and consequences of incivility include Failed use of tool/code or error messages, People, and Discontinued further discussion, respectively. This dataset can serve as a valuable resource for analyzing incivility in OSS and improving automated tools to detect and mitigate such behavior.
Ramtin Ehsani, Mia Mohammad Imran, Robert Zita, Kostadin Damevski, Preetha Chatterjee
MSR4
2023 Towards Understanding Emotions in Informal Developer Interactions: A Gitter Chat Study
abstract
Emotions play a significant role in teamwork and collaborative activities like software development. While researchers have analyzed developer emotions in various software artifacts (e.g., issues, pull requests), few studies have focused on understanding the broad spectrum of emotions expressed in chats. As one of the most widely used means of communication, chats contain valuable information in the form of informal conversations, such as negative perspectives about adopting a tool. In this paper, we present a dataset of developer chat messages manually annotated with a wide range of emotion labels (and sub-labels), and analyze the type of information present in those messages. We also investigate the unique signals of emotions specific to chats and distinguish them from other forms of software communication. Our findings suggest that chats have fewer expressions of Approval and Fear but more expressions of Curiosity compared to GitHub comments. We also notice that Confusion is frequently observed when discussing programming-related information such as unexpected software behavior. Overall, our study highlights the potential of mining emotions in developer chats for supporting software maintenance and evolution tools.
Amirali Sajadi, Kostadin Damevski, Preetha Chatterjee
ESEC/SIGSOFT FSE2
2022 Fast Changeset-based Bug Localization with BERT
abstract
Automatically localizing software bugs to the changesets that induced them has the potential to improve software developer efficiency and to positively affect software quality. To facilitate this automation, a bug report has to be effectively matched with source code changes, even when a significant lexical gap exists between natural language used to describe the bug and identifier naming practices used by developers. To bridge this gap, we need techniques that are able to capture software engineering-specific and project-specific semantics in order to detect relatedness between the two types of documents that goes beyond exact term matching. Popular transformer-based deep learning architectures, such as BERT, excel at leveraging contextual information, hence appear to be a suitable candidate for the task. However, BERT-like models are computationally expensive, which precludes them from being used in an environment where response time is important.
Agnieszka Ciborowska, Kostadin Damevski
ICSE2
2022 Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
abstract
Emotions (e.g., Joy, Anger) are prevalent in daily software engineering (SE) activities, and are known to be significant indicators of work productivity (e.g., bug fixing efficiency). Recent studies have shown that directly applying general purpose emotion classification tools to SE corpora is not effective. Even within the SE domain, tool performance degrades significantly when trained on one communication channel and evaluated on another (e.g, StackOverflow vs. GitHub comments). Retraining a tool with channel-specific data takes significant effort since manually annotating a large dataset of ground truth data is expensive.
Mia Mohammad Imran, Yashasvi Jain, Preetha Chatterjee, Kostadin Damevski
ASE4
2022 Grouping related stack overflow comments for software developer recommendation
Viral Sheth, Kostadin Damevski
Autom. Softw. Eng.2
2022 Using clarification questions to improve software developers' Web search
Mia Mohammad Imran, Kostadin Damevski
Inf. Softw. Technol.2
2021 Automatic Extraction of Opinion-based Q&A from Online Developer Chats
abstract
Virtual conversational assistants designed specifically for software engineers could have a huge impact on the time it takes for software engineers to get help. Research efforts are focusing on virtual assistants that support specific software development tasks such as bug repair and pair programming. In this paper, we study the use of online chat platforms as a resource towards collecting developer opinions that could potentially help in building opinion Q&A systems, as a specialized instance of virtual assistants and chatbots for software engineers. Opinion Q&A has a stronger presence in chats than in other developer communications, thus mining them can provide a valuable resource for developers in quickly getting insight about a specific development topic (e.g., What is the best Java library for parsing JSON?). We address the problem of opinion Q&A extraction by developing automatic identification of opinion-asking questions and extraction of participants' answers from public online developer chats. We evaluate our automatic approaches on chats spanning six programming communities and two platforms. Our results show that a heuristic approach to opinion-asking questions works well (.87 precision), and a deep learning approach customized to the software domain outperforms heuristics-based, machine-learning-based and deep learning for answer extraction in community question answering.
Preetha Chatterjee, Kostadin Damevski, Lori L. Pollock
ICSE2
2021 Automatically Selecting Follow-up Questions for Deficient Bug Reports
abstract
The availability of quality information in bug reports that are created daily by software users is key to rapidly fixing software faults. Improving incomplete or deficient bug reports, which are numerous in many popular and actively developed open source software projects, can make software maintenance more effective and improve software quality. In this paper, we propose a system that addresses the problem of bug report incompleteness by automatically posing follow-up questions, intended to elicit answers that add value and provide missing information to a bug report. Our system is based on selecting follow-up questions from a large corpus of already posted follow-up questions on GitHub. To estimate the best follow-up question for a specific deficient bug report we combine two metrics based on: 1) the compatibility of a follow-up question to a specific bug report; and 2) the utility the expected answer to the follow-up question would provide to the deficient bug report. Evaluation of our system, based on a manually annotated held-out data set, indicates improved performance over a set of simple and ablation baselines. A survey of software developers confirms the held-out set evaluation result that about half of the selected follow-up questions are considered valid. The survey also indicates that the valid follow-up questions are useful and can provide new information to a bug report most of the time, and are specific to a bug report some of the time.
Mia Mohammad Imran, Agnieszka Ciborowska, Kostadin Damevski
MSR3
2021 Automatically Identifying the Quality of Developer Chats for Post Hoc Use
abstract
Software engineers are crowdsourcing answers to their everyday challenges on Q&A forums (e.g., Stack Overflow) and more recently in public chat communities such as Slack, IRC, and Gitter. Many software-related chat conversations contain valuable expert knowledge that is useful for both mining to improve programming support tools and for readers who did not participate in the original chat conversations. However, most chat platforms and communities do not contain built-in quality indicators (e.g., accepted answers, vote counts). Therefore, it is difficult to identify conversations that contain useful information for mining or reading, i.e., conversations of post hoc quality. In this article, we investigate automatically detecting developer conversations of post hoc quality from public chat channels. We first describe an analysis of 400 developer conversations that indicate potential characteristics of post hoc quality, followed by a machine learning-based approach for automatically identifying conversations of post hoc quality. Our evaluation of 2,000 annotated Slack conversations in four programming communities (python, clojure, elm, and racket) indicates that our approach can achieve precision of 0.82, recall of 0.90, F-measure of 0.86, and MCC of 0.57. To our knowledge, this is the first automated technique for detecting developer conversations of post hoc quality.
Preetha Chatterjee, Kostadin Damevski, Nicholas A. Kraft, Lori L. Pollock
ACM Trans. Softw. Eng. Methodol.2
2020 Software-related Slack Chats with Disentangled Conversations
abstract
More than ever, developers are participating in public chat communities to ask and answer software development questions. With over ten million daily active users, Slack is one of the most popular chat platforms, hosting many active channels focused on software development technologies, e.g., python, react. Prior studies have shown that public Slack chat transcripts contain valuable information, which could provide support for improving automatic software maintenance tools or help researchers understand developer struggles or concerns.
Preetha Chatterjee, Kostadin Damevski, Nicholas A. Kraft, Lori L. Pollock
MSR2
2020 Automatically identifying valid API versions for software development tutorials on the Web
abstract
Abstract Online tutorials are a valuable source of community‐created information used by numerous developers to learn new APIs and techniques. Once written, tutorials are rarely actively curated and can become dated over time. Tutorials often reference APIs that change rapidly, and deprecated classes, methods, and fields can render tutorials inapplicable to newer releases of the API. Newer tutorials may not be compatible with older APIs that are still in use. In this paper, we first empirically study the tutorial versioning problem, confirming its presence in popular tutorials on the Web. We subsequently propose a technique, based on similar techniques in the literature, for automatically detecting the applicable API version ranges of tutorials, given access to the official API documentation they reference. The proposed technique identifies each API mention in a tutorial and maps the mention to the corresponding API element in the official documentation. The version of the tutorial is determined by combining the version ranges of all of the constituent API mentions. Our technique's precision varies from 61% to 89% and recall varies from 42% to 84% based on different levels of granularity of API mentions and different problem constraints. We observe API methods are the most challenging to accurately disambiguate due to method overloading. As the API mentions in tutorials are often redundant, and each mention of a specific API element commonly occurs several times in a tutorial, the distance of the predicted version range from the true version range is low: 3.61 on average for the tutorials in our sample.
Manziba Akanda Nishi, Kostadin Damevski
J. Softw. Evol. Process.2
2020 Changeset-Based Topic Modeling of Software Repositories
abstract
The standard approach to applying text retrieval models to code repositories is to train models on documents representing program elements. However, code changes lead to model obsolescence and to the need to retrain the model from the latest snapshot. To address this, we previously introduced an approach that trains a model on documents representing changesets from a repository and demonstrated its feasibility for feature location. In this paper, we expand our work by investigating: a second task (developer identification), the effects of including different changeset parts in the model, the repository characteristics that affect the accuracy of our approach, and the effects of the time invariance assumption on evaluation results. Our results demonstrate that our approach is as accurate as the standard approach for projects with most changes localized to a subset of the code, but less accurate when changes are highly distributed throughout the code. Moreover, our results demonstrate that context and messages are key to the accuracy of changeset-based models and that the time invariance assumption has a statistically significant effect on evaluation results, providing overly-optimistic results. Our findings indicate that our approach is a suitable alternative to the standard approach, providing comparable accuracy while eliminating retraining costs.
Christopher S. Corley, Kostadin Damevski, Nicholas A. Kraft
IEEE Trans. Software Eng.2
2019 Using Automated Prompts for Student Reflection on Computer Security Concepts
abstract
Reflection is known to be an effective means to improve students' learning. In this paper, we aim to foster meaningful reflection via prompts in computer science courses with a significant practical, software development component. To this end we develop an instructional strategy and system that automatically delivers prompts to students based on their commits in a source code repository. The system allows for prompts that instigate reflection in students to be timely with respect to students' work, and delivered automatically, thus easily scaling up the strategy.
Hui Chen 0001, Agnieszka Ciborowska, Kostadin Damevski
ITiCSE3
2019 Exploratory study of slack Q&A chats as a mining source for software engineering tools
abstract
Modern software development communities are increasingly social. Popular chat platforms such as Slack host public chat communities that focus on specific development topics such as Python or Ruby-on-Rails. Conversations in these public chats often follow a Q&A format, with someone seeking information and others providing answers in chat form. In this paper, we describe an exploratory study into the potential use-fulness and challenges of mining developer Q&A conversations for supporting software maintenance and evolution tools. We designed the study to investigate the availability of information that has been successfully mined from other developer communications, particularly Stack Overflow. We also analyze characteristics of chat conversations that might inhibit accurate automated analysis. Our results indicate the prevalence of useful information, including API mentions and code snippets with descriptions, and several hurdles that need to be overcome to automate mining that information.
Preetha Chatterjee, Kostadin Damevski, Lori L. Pollock, Vinay Augustine, Nicholas A. Kraft
MSR2
2019 Characterizing duplicate code snippets between stack overflow and tutorials
abstract
Developers are usually unaware of the quality and lineage of information available on popular Web resources, leading to potential maintenance problems and license violations when reusing code snippets from these resources. In this paper, we study the duplication of code snippets between two popular sources of software development information: the Stack Overflow Q a significant number (31%) of answers that contained a duplicate code block were chosen as the accepted answer. Qualitative analysis reveals that developers commonly use Stack Overflow to ask clarifying questions about code they reused from tutorials, and copy code snippets from tutorials to provide answers to questions.
Manziba Akanda Nishi, Agnieszka Ciborowska, Kostadin Damevski
MSR3
2019 Modeling hierarchical usage context for software exceptions based on interaction data
Hui Chen 0001, Kostadin Damevski, David C. Shepherd, Nicholas A. Kraft
Autom. Softw. Eng.2
2019 Modeling stack overflow tags and topics as a hierarchy of concepts
Hui Chen 0001, John Coogle, Kostadin Damevski
J. Syst. Softw.3
2018 Predicting future developer behavior in the IDE using topic models
abstract
Interaction data, gathered from developers' daily clicks and key presses in the IDE, has found use in both empirical studies and in recommendation systems for software engineering. We observe that this data has several characteristics, common across IDEs:
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock
ICSE1
2018 Detecting and characterizing developer behavior following opportunistic reuse of code snippets from the web
abstract
Modern software development is social and relies on many online resources and tools. In this paper, we study opportunistic code reuse from the Web, e.g., when developers copy code snippets from popular Q&A sites and paste them into their projects. Our focus is the behavior of developers following opportunistic code reuse, which reveals the success or failure of the action. We study developer behavior via a large, representative dataset of micro-interactions in the IDE. Our analysis of developer behavior exhibited in this dataset confirms laboratory study observations that code reuse from the Web is followed by heavy editing, in some cases by a rapid undo, and rarely by the execution of tests.
Agnieszka Ciborowska, Nicholas A. Kraft, Kostadin Damevski
MSR3
2018 Scalable code clone detection and search based on adaptive prefix filtering
Manziba Akanda Nishi, Kostadin Damevski
J. Syst. Softw.2
2018 Predicting Future Developer Behavior in the IDE Using Topic Models
abstract
While early software command recommender systems drew negative user reaction, recent studies show that users of unusually complex applications will accept and utilize command recommendations. Given this new interest, more than a decade after first attempts, both the recommendation generation (backend) and the user experience (frontend) should be revisited. In this work, we focus on recommendation generation. One shortcoming of existing command recommenders is that algorithms focus primarily on mirroring the short-term past,-i.e., assuming that a developer who is currently debugging will continue to debug endlessly. We propose an approach to improve on the state of the art by modeling future task context to make better recommendations to developers. That is, the approach can predict that a developer who is currently debugging may continue to debug OR may edit their program. To predict future development commands, we applied Temporal Latent Dirichlet Allocation, a topic model used primarily for natural language, to software development interaction data (i.e., command streams). We evaluated this approach on two large interaction datasets for two different IDEs, Microsoft Visual Studio and ABB Robot Studio. Our evaluation shows that this is a promising approach for both predicting future IDE commands and producing empirically-interpretable observations.
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Nicholas A. Kraft, Lori L. Pollock
IEEE Trans. Software Eng.1
2017 Behavior Metrics for Prioritizing Investigations of Exceptions
abstract
Many software development teams collect product defect reports, which can either be manually submitted or automatically created from product logs. Periodically, the teams use the collected defect reports to prioritize which defect to address next. We present a set of behavior-based metrics that can be used in this process. These metrics are based on the insight that development teams can estimate user inconvenience from user and application behavior in interaction logs. To estimate user inconvenience, the behavior metrics capture important user and application behavior after exceptions (the defects of interest in our case). We validated these metrics through a survey of how developers would incorporate the behavior metrics into their prioritization decisions. We found that developers change their priority of investigating an exception about 31% of the time after including the behavior metrics in the priority decision. These findings provide evidence that behavior metrics provide a promising advance towards prioritizing application exceptions.
Zack Coker, Kostadin Damevski, Claire Le Goues, Nicholas A. Kraft, David C. Shepherd, Lori L. Pollock
ICSME2
2017 What information about code snippets is available in different software-related documents? An exploratory study
abstract
A large corpora of software-related documents is available on the Web, and these documents offer the unique opportunity to learn from what developers are saying or asking about the code snippets that they are discussing. For example, the natural language in a bug report provides information about what is not functioning properly in a particular code snippet. Previous research has mined information about code snippets from bug reports, emails, and Q&A forums. This paper describes an exploratory study into the kinds of information that is embedded in different software-related documents. The goal of the study is to gain insight into the potential value and difficulty of mining the natural language text associated with the code snippets found in a variety of software-related documents, including blog posts, API documentation, code reviews, and public chats.
Preetha Chatterjee, Manziba Akanda Nishi, Kostadin Damevski, Vinay Augustine, Lori L. Pollock, Nicholas A. Kraft
SANER3
2017 Reconstructing and evolving software architectures using a coordinated clustering framework
Sheikh Motahar Naim, Kostadin Damevski, Mahmud Shahriar Hossain
Autom. Softw. Eng.2
2017 Mining Sequences of Developer Interactions in Visual Studio for Usage Smells
abstract
In this paper, we present a semi-automatic approach for mining a large-scale dataset of IDE interactions to extract usage smells, i.e., inefficient IDE usage patterns exhibited by developers in the field. The approach outlined in this paper first mines frequent IDE usage patterns, filtered via a set of thresholds and by the authors, that are subsequently supported (or disputed) using a developer survey, in order to form usage smells. In contrast with conventional mining of IDE usage data, our approach identifies time-ordered sequences of developer actions that are exhibited by many developers in the field. This pattern mining workflow is resilient to the ample noise present in IDE datasets due to the mix of actions and events that these datasets typically contain. We identify usage patterns and smells that contribute to the understanding of the usability of Visual Studio for debugging, code search, and active file navigation, and, more broadly, to the understanding of developer behavior during these software development activities. Among our findings is the discovery that developers are reluctant to use conditional breakpoints when debugging, due to perceived IDE performance problems as well as due to the lack of error checking in specifying the conditional.
Kostadin Damevski, David C. Shepherd, Johannes Schneider 0002, Lori L. Pollock
IEEE Trans. Software Eng.1
2016 Interactive exploration of developer interaction traces using a hidden markov model
abstract
Using IDE usage data to analyze the behavior of software developers in the field, during the course of their daily work, can lend support to (or dispute) laboratory studies of developers. This paper describes a technique that leverages Hidden Markov Models (HMMs) as a means of mining high-level developer behavior from low-level IDE interaction traces of many developers in the field. HMMs use dual stochastic processes to model higher-level hidden behavior using observable input sequences of events. We propose an interactive approach of mining interpretable HMMs, based on guiding a human expert in building a high quality HMM in an iterative, one state at a time, manner. The final result is a model that is both representative of the field data and captures the field phenomena of interest. We apply our HMM construction approach to study debugging behavior, using a large IDE interaction dataset collected from nearly 200 developers at ABB, Inc. Our results highlight the different modes and constituent actions in debugging, exhibited by the developers in our dataset.
Kostadin Damevski, Hui Chen 0001, David C. Shepherd, Lori L. Pollock
MSR1
2016 A field study of how developers locate features in source code
Kostadin Damevski, David C. Shepherd, Lori L. Pollock
Empir. Softw. Eng.1
2015 How and When to Transfer Software Engineering Research via Extensions
abstract
It is often reported that there is a large gap between software engineering research and practice, with little transfer from research to practice. While this is true in general, one transfer technique is increasingly breaking down this barrier: extensions to integrated development environments (IDEs). With the proliferation of app stores for IDEs and increasing transfer effort from researchers several research-based extensions have seen significant adoption. In this talk we'll discuss our experience transferring code search research, which currently is in the top 5% of Visual Studio extensions with over 13,000 downloads, as well as other research techniques transferred via extensions such as NCrunch, FindBugs, Code Recommenders, Mylyn, and Instasearch. We'll use the lessons learned from our transfer experience to provide case study evidence as to best practices for successful transfer, supplementing it with the quantitative evidence offered by app store and usage data across the broader set of extensions. The goal of this 30 minute talk is to provide researchers with a realistic view on which research techniques can be transferred to practice as well as concrete steps to execute such a transfer.
David C. Shepherd, Kostadin Damevski, Lori L. Pollock
ICSE (2)2
2015 Exploring the use of deep learning for feature location
abstract
Deep learning models can infer complex patterns present in natural language text. Relative to n-gram models, deep learning models can capture more complex statistical patterns based on smaller training corpora. In this paper we explore the use of a particular deep learning model, document vectors (DVs), for feature location. DVs seem well suited to use with source code, because they both capture the influence of context on each term in a corpus and map terms into a continuous semantic space that encodes semantic relationships such as synonymy. We present preliminary results that show that a feature location technique (FLT) based on DVs can outperform an analogous FLT based on latent Dirichlet allocation (LDA) and then suggest several directions for future work on the use of deep learning models to improve developer effectiveness in feature location.
Christopher S. Corley, Kostadin Damevski, Nicholas A. Kraft
ICSME2
2015 Scaling up evaluation of code search tools through developer usage metrics
abstract
Code search is a fundamental part of program understanding and software maintenance and thus researchers have developed many techniques to improve its performance, such as corpora preprocessing and query reformulation. Unfortunately, to date, evaluations of code search techniques have largely been in lab settings, while scaling and transitioning to effective practical use demands more empirical feedback from the field. This paper addresses that need by studying metrics based on automatically-gathered anonymous field data from code searches to infer user satisfaction. We describe techniques for addressing important concerns, such as how privacy is retained and how the overhead on the interactive system is minimized. We perform controlled user and field studies which identify metrics that correlate with user satisfaction, enabling the future evaluation of search tools through anonymous usage data. In comparing our metrics to similar metrics used in Internet search we observe differences in the relationship of some of the metrics to user satisfaction. As we further explore the data, we also present a predictive multi-metric model that achieves accuracy of over 70% in determining query satisfaction.
Kostadin Damevski, David C. Shepherd, Lori L. Pollock
SANER1
2014 A teaching model for development of sensor-driven mobile applications
abstract
This paper concerns teaching computer science undergraduate students to develop sophisticated sensor-driven mobile applications, which students find interesting and motivating. Computer science students commonly adopt a trial-and-error application development process. However, indeterminacy inherent in sensor data makes the trial-and-error approach difficult, which frustrates students and impairs learning. In addition, the complexity of modern mobile devices' development environment and numerous APIs can further undo the motivating effect that these types of applications bring. To address these challenges, we propose a teaching model for sensor-driven mobile application development. The model features an application development process and a set of supporting tools and programs. The model provides a structured way for students to deal with the indeterminacy of sensor data and the complex development environments and results in a positive and supportive learning experience for the students. A case study of applying the model in an upper-level computer science elective course has shown it to be effective.
Hui Chen 0001, Kostadin Damevski
ITiCSE2
2014 How developers use multi-recommendation system in local code search
abstract
Developers often start programming tasks by searching for relevant code in their local codebase. Previous research suggests that 88% of manually-composed queries retrieve no relevant results. Many searches fail because existing search tools depend solely on string matching with a manually-composed query, which cannot find semantically-related code. To solve this problem, researchers proposed query recommendation techniques to help developers compose queries without the extensive knowledge of the codebase under search. However, few of these techniques are empirically evaluated by the usage data from real-world developers. To fill this gap, we studied several query recommendation techniques by extending Sando and conducting a longitudinal field study. Our study shows that over 30% of all queries were adopted from recommendation; and recommended queries retrieved results 7% more often than manual queries.
Xi Ge, David C. Shepherd, Kostadin Damevski, Emerson R. Murphy-Hill
VL/HCC3
2013 Teaching cyber-physical systems to computer scientists via modeling and verification
abstract
The greater versatility and increasingly smaller sizes of computing, sensing, and networking devices have resulted in a new computing paradigm called Cyber-Physical Systems (CPSs), which integrates computation and sensing into physical processes producing a wealth of exciting applications in many domains of life, such as transportation, medicine, and agriculture. In order to equip students with the essential knowledge and skills to be successful in the future, this paradigm requires an expansion in the scope of computer science curricula to enable students to understand and overcome the complexity inherent in CPSs. In this paper, we describe our experience with teaching CPS via a set of course modules that rely heavily on modeling and verification. By using the popular Android platform, we aim to engage students to successfully build CPS applications while enhancing their understanding of intellectually challenging concepts.
Kostadin Damevski, Badreldin Altayeb, Hui Chen 0001, David Walter
SIGCSE1
2012 Sando: an extensible local code search framework
abstract
Developers heavily rely on Local Code Search (LCS)---the execution of a text-based search on a single code base---to find starting points in software maintenance tasks. While LCS approaches commonly used by developers are based on lexical matching and often result in failed searches or irrelevant results, developers have not yet migrated to the various research approaches that have made significant advancements in LCS. We hypothesize that two of the major reasons for this lack of migration are as follows. First, developers do not know which approach is the best, due to a lack of comparative field studies and the discrepancies in the underlying LCS process that these research approaches address. Second, developers lack access to a stable implementation of most of the research approaches. To address these issues, we studied a number of LCS approaches, distilled the general component structure underlying these approaches and, based on this structure, developed a LCS tool and framework, called Sando. Currently used by developers at ABB, Inc. and elsewhere, Sando also supports the flexible extension of its components to rapidly disseminate research advancements, and allows for user-based evaluation of competing approaches.
David C. Shepherd, Kostadin Damevski, Bartosz Ropski, Thomas Fritz 0001
SIGSOFT FSE2
2011 Offline enforcement of contracts for high-performance computing
abstract
Abstract Design by contract is a well‐known software design methodology which enhances the quality of software through assertions expressed at the interface level. The overhead averse nature of High‐Performance Computing (HPC) applications often precludes the use of design by contract due to its potential overhead, especially when handling the large data sizes common in HPC. Our approach is to reduce the overhead of design by contract, by postponing (or offloading) across time and space the enforcement of contracts. We argue that the semantic implications of this approach are not significant, while leading to a large potential overhead reduction. A reduced overhead strategy to contract implementation may be necessary for a wider acceptance of this useful software engineering primitive. We apply our approach to contracts developed for software components based on the CCA (Common Component Architecture) model, which targets HPC applications. Copyright © 2011 John Wiley & Sons, Ltd.
Kostadin Damevski
Concurr. Comput. Pract. Exp.1
2009 Application-aware management of parallel simulation collections
abstract
This paper presents a system deployed on parallel clusters to manage a collection of parallel simulations that make up a computational study. It explores how such a system can extend traditional parallel job scheduling and resource allocation techniques to incorporate knowledge specific to the study.
Siu Yau, Vijay Karamcheti, Denis Zorin, Kostadin Damevski, Steven G. Parker
PPoPP4
2008 Result reuse in design space exploration: A study in system support for interactive parallel computing
abstract
This paper presents a system supporting reuse of simulation results in multi-experiment computational studies involving independent simulations and explores the benefits of such reuse. Using a SCIRun-based defibrillator device simulation code (DefibSim) and the SimX system for computational studies, this paper demonstrates how aggressive reuse between and within computational studies can enable interactive rates for such studies on a moderate-sized 128-node processor cluster; a brute-force approach to the problem would require two thousand nodes or more on a massively parallel machine for similar performance. Key to realizing these performance improvements is exploiting optimization opportunities that present themselves at the level of the overall workflow of the study as opposed to focusing on individual simulations. Such global optimization approaches are likely to become increasingly important with the shift towards interactive and universal parallel computing.
Shi-Man Yau, Kostadin Damevski, Vijay Karamcheti, Steven G. Parker, Denis Zorin
IPDPS2
2007 CCALoop: scalable design of a distributed component framework
abstract
No abstract available.
Kostadin Damevski, Ashwin Deepak Swaminathan, Steven G. Parker
HPDC1
2006 Data redistribution and remote method invocation for coupled components
Felipe Bertrand, Randall Bramley, David E. Bernholdt, James Arthur Kohl, Alan Sussman, Jay Walter Larson, Kostadin Damevski
J. Parallel Distributed Comput.7
2004 Imprecise Exceptions in Distributed Parallel Components
Kostadin Damevski, Steven G. Parker
Euro-Par1
2004 SCIRun2: A CCA Framework for High Performance Computing
abstract
We present an overview of the SCIRun2 parallel component framework. SCIRun2 is based on the common component architecture (CCA) as stated by R. Armstrong et al. (1999) and the SCI Institutes' SCIRun by C. Johnson and S. Parker (1999). SCIRun2 supports distributed computing through distributed objects. Parallel components are managed transparently over an M/spl times/N method invocation and data redistribution subsystem. A meta component model based on CCA is used to accommodate multiple component models such as CCA, CORBA and Dataflow. A group of monitoring components built on top of the TAU toolkit as stated in Advanced Computing Laboratory (1999) evaluate the performance of the other components.
Kostadin Damevski, Venkatanand Venkatachalapathy, Steven G. Parker
HIPS2