VLDB 2026 Research / reviewers in the wild / expert
Themistoklis G. Diamantopoulos
dblp:134/7482 · also Themistoklis Diamantopoulos
· DBLP profile ↗
27ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-0520-7225ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 26 · 9 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Online Malware Detection Using Process Resource Utilization Metrics
Themistoklis G. Diamantopoulos, Dimosthenis Natsos, Andreas L. Symeonidis |
SANER | 1 |
| 2025 | Towards Effective Issue Assignment using Online Machine LearningabstractEfficient issue assignment in software development relates to faster resolution time, resources optimization, and reduced development effort. To this end, numerous systems have been developed to automate issue assignment, including AI and machine learning approaches. Most of them, however, often solely focus on a posteriori analyses of textual features (e.g. issue titles, descriptions), disregarding the temporal characteristics of software development. Thus, they fail to adapt as projects and teams evolve, such cases of team evolution, or project phase shifts (e.g. from development to maintenance). To incorporate such cases in the issue assignment process, we propose an Online Machine Learning methodology that adapts to the evolving characteristics of software projects. Our system processes issues as a data stream, dynamically learning from new data and adjusting in real time to changes in team composition and project requirements. We incorporate metadata such as issue descriptions, components and labels and leverage adaptive drift detection mechanisms to identify when model re-evaluation is necessary. Upon assessing our methodology on a set of software projects, we conclude that it can be effective on issue assignment, while meeting the evolving needs of software teams. Athanasios Michailoudis, Themistoklis G. Diamantopoulos, Antonios Favvas, Andreas L. Symeonidis |
EASE | 2 |
| 2025 | Towards an Interpretable Analysis for Estimating the Resolution Time of Software IssuesabstractLately, software development has become a predominantly online process, as more teams host and monitor their projects remotely. Sophisticated approaches employ issue tracking systems like Jira, predicting the time required to resolve issues and effectively assigning and prioritizing project tasks. Several methods have been developed to address this challenge, widely known as bug-fix time prediction, yet they exhibit significant limitations. Most consider only textual issue data and/or use techniques that overlook the semantics and metadata of issues (e.g., priority or assignee expertise). Many also fail to distinguish actual development effort from administrative delays, including assignment and review phases, leading to estimates that do not reflect the true effort needed. In this work, we build an issue monitoring system that extracts the actual effort required to fix issues on a per-project basis. Our approach employs topic modeling to capture issue semantics and leverages metadata (components, labels, priority, issue type, assignees) for interpretable resolution time analysis. Final predictions are generated by an aggregated model, enabling contributors to make informed decisions. Evaluation across multiple projects shows the system can effectively estimate resolution time and provide valuable insights. Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Davide Tosi, Martina Tropeano, Andreas L. Symeonidis |
EASE | 2 |
| 2025 | Accelerating Educational Assessment in Software Engineering through Human-AI CollaborationabstractThe continuously improving capabilities of Artificial Intelligence (AI) systems are rapidly establishing them as the de facto choice for complex tasks and sophisticated reasoning. Evaluation tasks, for example, are particularly challenging, since they demand expert judgment, contextual analysis, and nuanced decision-making. Educational assessment represents a critical instance of such a complex cognitive task, where scalability challenges force institutions to choose between efficiency and assessment quality, as enrollment outpaces faculty capacity. While large language models seem promising for assisting in this context, they often lack domain-specific knowledge and pedagogical context, which are essential for effective assessment. This paper presents a systematic methodology for human-AI collaboration that addresses these limitations and achieves scalable efficiency, while preserving instructor autonomy. We validate our approach in the Software Engineering education domain, an especially demanding testbed, which requires assessment across multiple technical artifacts that combine objective correctness with subjective design quality. The system is tested against 30 software engineering projects, across different software engineering artifact types. Our evaluation demonstrates significant efficiency improvement against the current (manual) assessment approach, indicating that the systematic provision of domain knowledge can enable AI assistance in complex educational evaluation tasks. Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
ICTAI | 2 |
| 2025 | Extracting Fix Patterns for Static Analysis Violations Based on Collective Developer KnowledgeabstractABSTRACT Introduction Much of the effort spent on software development is allocated on detecting and fixing bugs or, more generally, violations that lead to erroneous or inefficient code. Although static analysis tools aspire to automate bug detection, their usage is typically limited to code style rules and typical violations, while they only provide generic instructions for bug fixing. As a result, contemporary approaches focus on extracting bug‐fix patterns from source code revisions (commits). Methodology In this work, we build a methodology that extracts fix patterns for violations detected by the PMD static analysis tool. In contrast to current approaches, which employ specific data sources and focus on particular violations, our system extracts source code edits from multiple GitHub repositories and maps them into 36 different types of violations. By employing a detailed syntax tree representation and a tree edit distance technique, we build a similarity scheme for source code edits, which is used to group them into clusters. We employ two clustering algorithms, K‐medoids and DBSCAN, and further optimize them using a purity metric to produce clusters that correspond to specific fixes. Results Our evaluation indicates that DBSCAN extracts more cohesive clusters (purity more than 0.9), which effectively target specific PMD rules, while K‐medoids extracts generic clusters (purity around 0.7) that pinpoint common edits. Conclusion Finally, upon analyzing the diversity of the sources from which fixes are derived (commits and repositories), we conclude that they are generic enough to reflect how similar issues are dealt with by the developer community. Michael Karatzas, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
Softw. Pract. Exp. | 2 |
| 2024 | Write me this Code: An Analysis of ChatGPT Quality for Producing Source CodeabstractDevelopers nowadays are increasingly turning to large language models (LLMs) like ChatGPT to assist them with coding tasks, inspired by the promise of efficiency and the advanced capabilities they offer. However, this raises important questions about the ease of integration and the safety of incorporating these tools into the development process. To investigate these questions, this paper examines a set of ChatGPT conversations. Upon annotating the conversations according to the intent of the developer, we focus on two critical aspects: firstly, the ease with which developers can produce suitable source code using ChatGPT, and, secondly, the quality aspects of the generated source code, determined by the compliance to standards and best practices. We research both the quality of the generated code itself and its impact on the project of the developer. Our results indicate that ChatGPT can be a useful tool for software development when used with discretion. Konstantinos Moratis, Themistoklis G. Diamantopoulos, Dimitrios-Nikitas Nastos, Andreas L. Symeonidis |
MSR | 2 |
| 2023 | Towards Readability-Aware Recommendations of Source Code Snippets
Athanasios Michailoudis, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
ICSOFT | 2 |
| 2023 | Towards Interpretable Monitoring and Assignment of Jira Issues
Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
ICSOFT | 2 |
| 2023 | Semantically-enriched Jira Issue Tracking DataabstractCurrent state of practice dictates that software developers host their projects online and employ project management systems to monitor the development of product features, keep track of bugs, and prioritize task assignments. The data stored in these systems, if their semantics are extracted effectively, can be used to answer several interesting questions, such as finding who is the most suitable developer for a task, what the priority of a task should be, or even what is the actual workload of the software team. To support researchers and practitioners that work towards these directions, we have built a system that crawls data from the Jira management system, performs topic modeling on the data to extract useful semantics and stores them in a practical database schema. We have used our system to retrieve and analyze 656 projects of the Apache Software Foundation, comprising data from more than a million Jira issues. Themistoklis G. Diamantopoulos, Dimitrios-Nikitas Nastos, Andreas L. Symeonidis |
MSR | 1 |
| 2023 | Automated issue assignment using topic modelling on Jira issue tracking dataabstractAbstract As more and more software teams use online issue tracking systems to collaborate on software projects, the accurate assignment of new issues to the most suitable contributors may have significant impact on the success of the project. As a result, several research efforts have been directed towards automating this process to save considerable time and effort. However, most approaches focus mainly on software bugs and employ models that do not sufficiently take into account the semantics and the non‐textual metadata of issues and/or produce models that may require manual tuning. A methodology that extracts both textual and non‐textual features from different types of issues is designed, providing a Jira dataset that involves not only bugs but also new features, issues related to documentation, patches, etc. Moreover, the semantics of issue text are effectively captured by employing a topic modelling technique that is optimised using the assignment result. Finally, this methodology aggregates probabilities from a set of individual models to provide the final assignment. Upon evaluating this approach in an automated issue assignment setting using a dataset of Jira issues, the authors conclude that it can be effective for automated issue assignment. Themistoklis G. Diamantopoulos, Nikolaos Saoulidis, Andreas L. Symeonidis |
IET Softw. | 1 |
| 2022 | Semantic Code Search in Software Repositories using Neural Machine TranslationabstractAbstract Nowadays, software development is accelerated through the reuse of code snippets found online in question-answering platforms and software repositories. In order to be efficient, this process requires forming an appropriate query and identifying the most suitable code snippet, which can sometimes be challenging and particularly time-consuming. Over the last years, several code recommendation systems have been developed to offer a solution to this problem. Nevertheless, most of them recommend API calls or sequences instead of reusable code snippets. Furthermore, they do not employ architectures advanced enough to exploit the semantics of natural language and code in order to form the optimal query from the question posed. To overcome these issues, we propose CodeTransformer, a code recommendation system that provides useful, reusable code snippets extracted from open-source GitHub repositories. By employing a neural network architecture that comprises advanced attention mechanisms, our system effectively understands and models natural language queries and code snippets in a joint vector space. Upon evaluating CodeTransformer quantitatively against a similar system and qualitatively using a dataset from Stack Overflow, we conclude that our approach can recommend useful and reusable snippets to developers. Evangelos Papathomas, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
FASE | 2 |
| 2021 | Software Task Importance Prediction based on Project Management Data
Themistoklis G. Diamantopoulos, Christiana Galegalidou, Andreas L. Symeonidis |
ICSOFT | 1 |
| 2020 | Extracting Semantics from Question-Answering Services for Snippet ReuseabstractNowadays, software developers typically search online for reusable solutions to common programming problems. However, forming the question appropriately, and locating and integrating the best solution back to the code can be tricky and time consuming. As a result, several mining systems have been proposed to aid developers in the task of locating reusable snippets and integrating them into their source code. Most of these systems, however, do not model the semantics of the snippets in the context of source code provided. In this work, we propose a snippet mining system, named StackSearch, that extracts semantic information from Stack Overlow posts and recommends useful and in-context snippets to the developer. Using a hybrid language model that combines Tf-Idf and fastText, our system effectively understands the meaning of the given query and retrieves semantically similar posts. Moreover, the results are accompanied with useful metadata using a named entity recognition technique. Upon evaluating our system in a set of common programming queries, in a dataset based on post links, and against a similar tool, we argue that our approach can be useful for recommending ready-to-use snippets to the developer. Themistoklis G. Diamantopoulos, Nikolaos Oikonomou, Andreas L. Symeonidis |
FASE | 1 |
| 2020 | Employing Contribution and Quality Metrics for Quantifying the Software Development ProcessabstractThe full integration of online repositories in contemporary software development promotes remote work and collaboration. Apart from the apparent benefits, online repositories offer a deluge of data that can be utilized to monitor and improve the software development process. Towards this direction, we have designed and implemented a platform that analyzes data from GitHub in order to compute a series of metrics that quantify the contributions of project collaborators, both from a development as well as an operations (communication) perspective. We analyze contributions throughout the projects' lifecycle and track the number of coding violations, this way aspiring to identify cases of software development that need closer monitoring and (possibly) further actions to be taken. In this context, we have analyzed the 3000 most popular GitHub Java projects and provide the data to the community. Themistoklis G. Diamantopoulos, Michail Papamichail, Thomas Karanikiotis, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis |
MSR | 1 |
| 2020 | Towards Analyzing Contributions from Software Repositories to Optimize Issue AssignmentabstractMost software teams nowadays host their projects online and monitor software development in the form of issues/tasks. This process entails communicating through comments and reporting progress through commits and closing issues. In this context, assigning new issues, tasks or bugs to the most suitable contributor largely improves efficiency. Thus, several automated issue assignment approaches have been proposed, which however have major limitations. Most systems focus only on assigning bugs using textual data, are limited to projects explicitly using bug tracking systems, and may require manually tuning parameters per project. In this work, we build an automated issue assignment system for GitHub, taking into account the commits and issues of the repository under analysis. Our system aggregates feature probabilities using a neural network that adapts to each project, thus not requiring manual parameter tuning. Upon evaluating our methodology, we conclude that it can be efficient for automated issue assignment. Vasileios Matsoukas, Themistoklis G. Diamantopoulos, Michail Papamichail, Andreas L. Symeonidis |
QRS | 2 |
| 2019 | npm Packages as Ingredients: A Recipe-based Approach
Kyriakos C. Chatzidimitriou, Michail Papamichail, Themistoklis G. Diamantopoulos, Napoleon-Christos I. Oikonomou, Andreas L. Symeonidis |
ICSOFT | 3 |
| 2019 | Towards Extracting the Role and Behavior of Contributors in Open-source Projects
Michail Papamichail, Themistoklis G. Diamantopoulos, Vasileios Matsoukas, Christos L. Athanasiadis, Andreas L. Symeonidis |
ICSOFT | 2 |
| 2019 | Towards mining answer edits to extract evolution patterns in stack overflowabstractThe current state of practice dictates that in order to solve a problem encountered when building software, developers ask for help in online platforms, such as Stack Overflow. In this context of collaboration, answers to question posts often undergo several edits to provide the best solution to the problem stated. In this work, we explore the potential of mining Stack Overflow answer edits to extract common patterns when answering a post. In particular, we design a similarity scheme that takes into account the text and code of answer edits and cluster edits according to their semantics. Upon applying our methodology, we provide frequent edit patterns and indicate how they could be used to answer future research questions. Assessing our approach indicates that it can be effective for identifying commonly applied edits, thus illustrating the transformation path from the initial answer to the optimal solution. Themistoklis G. Diamantopoulos, Maria-Ioanna Sifaki, Andreas L. Symeonidis |
MSR | 1 |
| 2019 | A Mechanism for Automatically Summarizing Software Functionality from Source CodeabstractWhen developers search online to find software components to reuse, they usually first need to understand the container projects/libraries, and subsequently identify the required functionality. Several approaches identify and summarize the offerings of projects from their source code, however they often require that the developer has knowledge of the underlying topic modeling techniques; they do not provide a mechanism for tuning the number of topics, and they offer no control over the top terms for each topic. In this work, we use a vectorizer to extract information from variable/method names and comments, and apply Latent Dirichlet Allocation to cluster the source code files of a project into different semantic topics. The number of topics is optimized based on their purity with respect to project packages, while topic categories are constructed to provide further intuition and Stack Exchange tags are used to express the topics in more abstract terms. Christos Psarras, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
QRS | 2 |
| 2019 | Measuring the reusability of software components using static analysis metrics and reuse rate information
Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
J. Syst. Softw. | 2 |
| 2018 | Summarizing Software API Usage Examples Using Clustering TechniquesabstractAs developers often use third-party libraries to facilitate software development, the lack of proper API documentation for these libraries undermines their reuse potential. And although several approaches extract usage examples for libraries, they are usually tied to specific language implementations, while their produced examples are often redundant and are not presented as concise and readable snippets. In this work, we propose a novel approach that extracts API call sequences from client source code and clusters them to produce a diverse set of source code snippets that effectively covers the target API. We further construct a summarization algorithm to present concise and readable snippets to the users. Upon evaluating our system on software libraries, we indicate that it achieves high coverage in API methods, while the produced snippets are of high quality and closely match handwritten examples. Nikolaos Katirtzis, Themistoklis G. Diamantopoulos, Charles Sutton |
FASE | 2 |
| 2018 | npm-miner: an infrastructure for measuring the quality of the npm registryabstractAs the popularity of the JavaScript language is constantly increasing, one of the most important challenges today is to assess the quality of JavaScript packages. Developers often employ tools for code linting and for the extraction of static analysis metrics in order to assess and/or improve their code. In this context, we have developed npn-miner, a platform that crawls the npm registry and analyzes the packages using static analysis tools in order to extract detailed quality metrics as well as high-level quality attributes, such as maintainability and security. Our infrastructure includes an index that is accessible through a web interface, while we have also constructed a dataset with the results of a detailed analysis for 2000 popular npm packages. Kyriakos C. Chatzidimitriou, Michail Papamichail, Themistoklis G. Diamantopoulos, Michail Tsapanos, Andreas L. Symeonidis |
MSR | 3 |
| 2017 | Towards Modeling the User-perceived Quality of Source Code using Static Analysis Metrics
Valasia Dimaridou, Alexandros-Charalampos Kyprianidis, Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
ICSOFT | 4 |
| 2017 | From requirements to source code: a Model-Driven Engineering approach for RESTful web services
Christoforos Zolotas, Themistoklis G. Diamantopoulos, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis |
Autom. Softw. Eng. | 2 |
| 2016 | QualBoa: reusability-aware recommendations of source code componentsabstractContemporary software development processes involve finding reusable software components from online repositories and integrating them to the source code, both to reduce development time and to ensure that the final software project is of high quality. Although several systems have been designed to automate this procedure by recommending components that cover the desired functionality, the reusability of these components is usually not assessed by these systems. In this work, we present QualBoa, a recommendation system for source code components that covers both the functional and the quality aspects of software component reuse. Upon retrieving components, QualBoa provides a ranking that involves not only functional matching to the query, but also a reusability score based on configurable thresholds of source code metrics. The evaluation of QualBoa indicates that it can be effective for recommending reusable source code. Themistoklis G. Diamantopoulos, Klearchos Thomopoulos, Andreas L. Symeonidis |
MSR | 1 |
| 2016 | User-Perceived Source Code Quality Estimation Based on Static Analysis MetricsabstractThe popularity of open source software repositories and the highly adopted paradigm of software reuse have led to the development of several tools that aspire to assess the quality of source code. However, most software quality estimation tools, even the ones using adaptable models, depend on fixed metric thresholds for defining the ground truth. In this work we argue that the popularity of software components, as perceived by developers, can be considered as an indicator of software quality. We present a generic methodology that relates quality with source code metrics and estimates the quality of software components residing in popular GitHub repositories. Our methodology employs two models: a one-class classifier, used to rule out low quality code, and a neural network, that computes a quality score for each software component. Preliminary evaluation indicates that our approach can be effective for identifying high quality software components in the context of reuse. Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
QRS | 2 |
| 2015 | Employing Source Code Information to Improve Question-Answering in Stack OverflowabstractNowadays, software development has been greatly influenced by question-answering communities, such as Stack Overflow. A new problem-solving paradigm has emerged, as developers post problems they encounter that are then answered by the community. In this paper, we propose a methodology that allows searching for solutions in Stack Overflow, using the main elements of a question post, including not only its title, tags, and body, but also its source code snippets. We describe a similarity scheme for these elements and demonstrate how structural information can be extracted from source code snippets and compared to further improve the retrieval of questions. The results of our evaluation indicate that our methodology is effective on recommending similar question posts allowing community members to search without fully forming a question. Themistoklis G. Diamantopoulos, Andreas L. Symeonidis |
MSR | 1 |