Akalanka Galappaththi

dblp:173/5868 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-6756-6610ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 An Empirical Study of API Misuses of Data-Centric Libraries
abstract
Developers rely on third-party library Application Programming Interfaces (APIs) when developing software. However, libraries typically come with assumptions and API usage constraints, whose violation results in API misuse. API misuses may result in crashes or incorrect behavior. Even though API misuse is a well-studied area, a recent study of API misuse of deep learning libraries showed that the nature of these misuses and their symptoms are different from misuses of traditional libraries, and as a result highlighted potential shortcomings of current misuse detection tools. We speculate that these observations may not be limited to deep learning API misuses but may stem from the data-centric nature of these APIs. Data-centric libraries often deal with diverse data structures, intricate processing workflows, and a multitude of parameters, which can make them inherently more challenging to use correctly. Therefore, understanding the potential misuses of these libraries is important to avoid unexpected application behavior. To this end, this paper contributes an empirical study of API misuses of five data-centric libraries that cover areas such as data processing, numerical computation, machine learning, and visualization. We identify misuses of these libraries by analyzing data from both Stack Overflow and GitHub. Our results show that many of the characteristics of API misuses observed for deep learning libraries extend to misuses of the data-centric library APIs we study. We also find that developers tend to misuse APIs from data-centric libraries, regardless of whether the API directive appears in the documentation. Overall, our work exposes the challenges of API misuse in data-centric libraries, rather than only focusing on deep learning libraries. Our collected misuses and their characterization lay groundwork for future research to help reduce misuses of these libraries.
Akalanka Galappaththi, Sarah Nadi, Christoph Treude
ESEM1
2022 Does This Apply to Me? An Empirical Study of Technical Context in Stack Overflow
abstract
Stack Overflow has become an essential technical resource for developers. However, given the vast amount of knowledge available on Stack Overflow, finding the right information that is relevant for a given task is still challenging, especially when a developer is looking for a solution that applies to their specific requirements or technology stack. Clearly marking answers with their technical context, i.e., the information that characterizes the technologies and assumptions needed for this answer, is potentially one way to improve navigation. However, there is no information about how often such context is mentioned, and what kind of information it might offer. In this paper, we conduct an empirical study to understand the occurrence of technical context in Stack Overflow answers and comments, using tags as a proxy for technical context. We specifically focus on additional context, where answers/comments mention information that is not already discussed in the question. Our results show that nearly half of our studied threads contain at least one additional context. We find that almost 50% of the additional context are either a library/framework, a programming language, a tool/application, an API, or a database. Overall, our findings show the promise of using additional context as navigational cues.
Akalanka Galappaththi, Sarah Nadi, Christoph Treude
MSR1
2021 Automatically Annotating Sentences for Task-specific Bug Report Summarization
abstract
There is a need to summarize bug reports as they can become long due to many comments from conversations between developers and various DevOps tools. Although automated approaches to bug report summarization have been developed, we believe they aim at the wrong target - getting as close as possible to a gold-standard summary. Instead, researchers should create automated bug report annotation approaches that allow project members to create summaries based on their task-specific information needs. We present such an approach.
Akalanka Galappaththi, John Anvik, Rafat Bin Islam
ASE1
2019 Feature Evaluation for Automatic Bug Report Summarization (S)
abstract
Bug reports can be lengthy due to long descriptions and long conversation threads.Automatic summarization of the text in a bug report can reduce the time spent by software project members on understanding the content of a bug report.Our work further examines Rastkar et al.'s use of a logistic regression model to determine which sentences from the text of a bug report should be extracted for creating a summary.Using their publicly available bug report corpus, which contains manually annotated bug reports, we examined two aspects regarding the features used by the model.First, we examined how much of a reduction occurs in the precision and recall if some of the more complex features are not used.Second, we examined how the use of different feature combinations affects the precision and recall of the models.We found that the absence of some of the complex features resulted in a modest decrease in precision and recall, and confirmed that some features, such as sentence length, were the most significant features for bug report summarization.
Akalanka Galappaththi, John Anvik
SEKE1