Andreas L. Symeonidis

dblp:01/2872 · DBLP profile ↗
← Back
72ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0003-0235-6046ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 34 · 17 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 15 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FAIR Data Management from Collection to Exploitation: The RAISE Suite Project
Evdokimos I. Konstantinidis, Gorka Epelde, Dimosthenis Natsos, Despoina Petsani, Anastasia Valtopoulou, Elli Papadopoulou, Mikel Hernandez, Dimitris Bamidis, Panagiotis G. Sarigiannidis, Andreas L. Symeonidis, Alexandros Chatzigeorgiou, Panagiotis Bamidis
DATA (2)10
2026 Towards Online Malware Detection Using Process Resource Utilization Metrics
Themistoklis G. Diamantopoulos, Dimosthenis Natsos, Andreas L. Symeonidis
SANER3
2026 QuAVA: A privacy-aware architecture for conversational desktop Content Retrieval systems
abstract
Question Answering (QA) and Content Retrieval (CR) systems have experienced a boost in performance in recent years leveraging state-of-the-art Transformer models to process user expressions and retrieve and extract information requested. Despite the constant language understanding improvements, very little effort has been put into the design of such systems for personal desktop use, where data are kept locally and are not sent to cloud services and decisions and outputs are transparent and explainable to the user. To that end, we present QuAVA, a conversational desktop content retrieval assistant, designed on four pillars: privacy and security, explainability, low-resource requirements, and multi-source data fusion. QuAVA is a data and privacy-preserving assistant that enables users to access their private data such as files, emails, and message exchanges, conversationally and transparently. The proposed architecture automatically extracts and preprocesses content from various sources and organizes it in a 3-layered hierarchical structure, namely a topic, a subtopic, and a content layer by employing ML algorithms for clustering and labeling. This way, users can navigate and access information via a set of conversation rules embedded in the assistant. We conduct a qualitative comparison analysis of the QuAVA architecture with other well-established QA and CR architectures against the four pillars defined, as well as privacy tests, and conclude that QuAVA is the only -to our knowledge- virtual assistant that successfully satisfies them. • Design and development of a desktop Content Retrieval assistant. • Private and secure Content Retrieval system. • Personal desktop assistant for Content Retrieval for everyday use. • Hierarchical structure for explainable and transparent Virtual Assistants.
Nikolaos Malamas, Andreas L. Symeonidis, John B. Theocharis
Comput. Speech Lang.2
2025 Towards Effective Issue Assignment using Online Machine Learning
abstract
Efficient issue assignment in software development relates to faster resolution time, resources optimization, and reduced development effort. To this end, numerous systems have been developed to automate issue assignment, including AI and machine learning approaches. Most of them, however, often solely focus on a posteriori analyses of textual features (e.g. issue titles, descriptions), disregarding the temporal characteristics of software development. Thus, they fail to adapt as projects and teams evolve, such cases of team evolution, or project phase shifts (e.g. from development to maintenance). To incorporate such cases in the issue assignment process, we propose an Online Machine Learning methodology that adapts to the evolving characteristics of software projects. Our system processes issues as a data stream, dynamically learning from new data and adjusting in real time to changes in team composition and project requirements. We incorporate metadata such as issue descriptions, components and labels and leverage adaptive drift detection mechanisms to identify when model re-evaluation is necessary. Upon assessing our methodology on a set of software projects, we conclude that it can be effective on issue assignment, while meeting the evolving needs of software teams.
Athanasios Michailoudis, Themistoklis G. Diamantopoulos, Antonios Favvas, Andreas L. Symeonidis
EASE4
2025 Towards an Interpretable Analysis for Estimating the Resolution Time of Software Issues
abstract
Lately, software development has become a predominantly online process, as more teams host and monitor their projects remotely. Sophisticated approaches employ issue tracking systems like Jira, predicting the time required to resolve issues and effectively assigning and prioritizing project tasks. Several methods have been developed to address this challenge, widely known as bug-fix time prediction, yet they exhibit significant limitations. Most consider only textual issue data and/or use techniques that overlook the semantics and metadata of issues (e.g., priority or assignee expertise). Many also fail to distinguish actual development effort from administrative delays, including assignment and review phases, leading to estimates that do not reflect the true effort needed. In this work, we build an issue monitoring system that extracts the actual effort required to fix issues on a per-project basis. Our approach employs topic modeling to capture issue semantics and leverages metadata (components, labels, priority, issue type, assignees) for interpretable resolution time analysis. Final predictions are generated by an aggregated model, enabling contributors to make informed decisions. Evaluation across multiple projects shows the system can effectively estimate resolution time and provide valuable insights.
Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Davide Tosi, Martina Tropeano, Andreas L. Symeonidis
EASE5
2025 Towards a Defense-in-Depth Approach for Securing Collaborative Cloud Infrastructures
Dimosthenis Natsos, Andreas L. Symeonidis
SEAA (3)2
2025 Accelerating Educational Assessment in Software Engineering through Human-AI Collaboration
abstract
The continuously improving capabilities of Artificial Intelligence (AI) systems are rapidly establishing them as the de facto choice for complex tasks and sophisticated reasoning. Evaluation tasks, for example, are particularly challenging, since they demand expert judgment, contextual analysis, and nuanced decision-making. Educational assessment represents a critical instance of such a complex cognitive task, where scalability challenges force institutions to choose between efficiency and assessment quality, as enrollment outpaces faculty capacity. While large language models seem promising for assisting in this context, they often lack domain-specific knowledge and pedagogical context, which are essential for effective assessment. This paper presents a systematic methodology for human-AI collaboration that addresses these limitations and achieves scalable efficiency, while preserving instructor autonomy. We validate our approach in the Software Engineering education domain, an especially demanding testbed, which requires assessment across multiple technical artifacts that combine objective correctness with subjective design quality. The system is tested against 30 software engineering projects, across different software engineering artifact types. Our evaluation demonstrates significant efficiency improvement against the current (manual) assessment approach, indicating that the systematic provision of domain knowledge can enable AI assistance in complex educational evaluation tasks.
Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
ICTAI3
2025 RQ2DSL: An AI Agent for Automated Domain-Specific Language Grammar Generation from Requirements
abstract
Domain-Specific Language (DSL) design is a foundational yet labor-intensive task in software engineering (SE). However, recent advances in Large Language Models (LLMs) offer the potential to automate key aspects of DSL development, including grammar generation, paving the way for wider DSL adoption. This paper presents RQ2DSL, an AI agent for automated grammar generation that transforms requirements (Functional and Non-Functional) into syntactically correct DSL grammars using LLMs. The agent's methodology comprises an optimally defined LLM context that contains a TextX knowledge base, reference grammars, syntactic examples and erroraware guidance with explicit validation rules, supported by an iterative refinement process that leverages syntax validation feedback. Comprehensive experiments were conducted against DSLs of varying complexity levels (simple, medium, high, ultrahigh), with scalable requirement sets ranging from 4 to 30 specifications, demonstrating that RQ2DSL maintains robust performance across simple$(100 \%)$, medium$(92 \%)$and high (92%) complexity levels. However, a critical complexity threshold emerges at ultra-high complex DSLs, where RQ2DSL experiences a notable performance degradation to 75%, while reducedcontext approaches lead to a dramatic drop in success rate.
Theodoros Tsampouris, Emmanouil G. Tsardoulias, Konstantinos Panayiotou, Andreas L. Symeonidis
ICTAI4
2025 Knee-cartilage segmentation from MR images using Multi-view Hypergraph Convolutional Neural Networks
abstract
Abstract Leveraging the increased capacities of hypergraphs to model complex data structures, we propose in this article the Multi-view Hyper-Graph Convolutional Network (MVHGCN) to yield automated knee-joint cartilage segmentations from MRIs. The main properties of our approach are presented as follows: 1) Node features are obtained from multi-view (MV) acquisitions, corresponding to different feature extractors or image modalities. 2) Node embeddings are generated using a distributive MV convolution scheme which combines the various view-specific convolutions. These results are aggregated via an attention-based fusion module to automatically learn the weights of the different views. 3) Our model integrates both local and global level learning, simultaneously. Local hypergraph convolutions explore the relationships across the spatially aligned node libraries, while global hypergraph convolutions search for global affinities between nodes located at different positions within the image. 4) We propose two different blending schemes to combine local and global convolutions, namely, the cross-talk (CT) and the collaborative (COL) blending units, respectively. Using these units as building blocks, we construct the MVHGCN model, a deep network with enhanced feature representation and learning capabilities. The suggested segmentation method is evaluated on the publicly available Osteoarthritis Initiative (OAI) cohort. Specifically, we have designed a thorough experimental setup, including parameter sensitivity analysis and comparative results against a series of existing traditional methods, deep CNN models, and graph convolutional networks. The results show that MVHGCN outperforms the competing methods, achieving an overall cartilage segmentation score of $$\mathcal {DSC} = 95.81\%$$ and $$\mathcal {DSC} = 96.33\%$$ , for the CT and the COL blending, respectively.
Christos G. Chadoulos, John B. Theocharis, Andreas L. Symeonidis, Serafeim P. Moustakidis
Appl. Intell.3
2025 A Data-Driven Methodology for Quality Aware Code Fixing
abstract
In today’s rapidly changing software development landscape, ensuring code quality is essential to reliability, maintainability, and security among other aspects. Identifying code quality issues can be tackled; however, implementing code quality improvements can be a complex and time‐consuming task. To address this problem, we present a novel methodology designed to assist developers by suggesting alternative code snippets that not only match the functionality of the original code but also improve its quality based on predefined metrics. Our system is based on a language‐agnostic approach that allows the analysis of code snippets written in different programming languages. It employs advanced techniques to assess functional similarity and evaluates syntactic similarity, suggesting alternatives that minimize the need for extensive modification. The evaluation of our system on multiple axes demonstrates the effectiveness of our approach in providing usable code alternatives that are both functionally equivalent and syntactically similar to the original snippets, while significantly improving quality metrics. We argue that our methodology and tool can be valuable for the software engineering community, bridging the gap between the identification of code quality problems and the implementation of practical solutions that improve software quality.
Thomas Karanikiotis, Andreas L. Symeonidis
IET Softw.2
2025 Extracting Fix Patterns for Static Analysis Violations Based on Collective Developer Knowledge
abstract
ABSTRACT Introduction Much of the effort spent on software development is allocated on detecting and fixing bugs or, more generally, violations that lead to erroneous or inefficient code. Although static analysis tools aspire to automate bug detection, their usage is typically limited to code style rules and typical violations, while they only provide generic instructions for bug fixing. As a result, contemporary approaches focus on extracting bug‐fix patterns from source code revisions (commits). Methodology In this work, we build a methodology that extracts fix patterns for violations detected by the PMD static analysis tool. In contrast to current approaches, which employ specific data sources and focus on particular violations, our system extracts source code edits from multiple GitHub repositories and maps them into 36 different types of violations. By employing a detailed syntax tree representation and a tree edit distance technique, we build a similarity scheme for source code edits, which is used to group them into clusters. We employ two clustering algorithms, K‐medoids and DBSCAN, and further optimize them using a purity metric to produce clusters that correspond to specific fixes. Results Our evaluation indicates that DBSCAN extracts more cohesive clusters (purity more than 0.9), which effectively target specific PMD rules, while K‐medoids extracts generic clusters (purity around 0.7) that pinpoint common edits. Conclusion Finally, upon analyzing the diversity of the sources from which fixes are derived (commits and repositories), we conclude that they are generic enough to reflect how similar issues are dealt with by the developer community.
Michael Karatzas, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
Softw. Pract. Exp.3
2024 Write me this Code: An Analysis of ChatGPT Quality for Producing Source Code
abstract
Developers nowadays are increasingly turning to large language models (LLMs) like ChatGPT to assist them with coding tasks, inspired by the promise of efficiency and the advanced capabilities they offer. However, this raises important questions about the ease of integration and the safety of incorporating these tools into the development process. To investigate these questions, this paper examines a set of ChatGPT conversations. Upon annotating the conversations according to the intent of the developer, we focus on two critical aspects: firstly, the ease with which developers can produce suitable source code using ChatGPT, and, secondly, the quality aspects of the generated source code, determined by the compliance to standards and best practices. We research both the quality of the generated code itself and its impact on the project of the developer. Our results indicate that ChatGPT can be a useful tool for software development when used with discretion.
Konstantinos Moratis, Themistoklis G. Diamantopoulos, Dimitrios-Nikitas Nastos, Andreas L. Symeonidis
MSR4
2024 Forward-Oriented Programming: A meta-DSL for fast development of component libraries
Emmanouil Krasanakis, Andreas L. Symeonidis
Inf. Softw. Technol.2
2024 Adversarial robustness improvement for deep neural networks
abstract
Abstract Deep neural networks (DNNs) are key components for the implementation of autonomy in systems that operate in highly complex and unpredictable environments (self-driving cars, smart traffic systems, smart manufacturing, etc.). It is well known that DNNs are vulnerable to adversarial examples, i.e. minimal and usually imperceptible perturbations, applied to their inputs, leading to false predictions. This threat poses critical challenges, especially when DNNs are deployed in safety or security-critical systems, and renders as urgent the need for defences that can improve the trustworthiness of DNN functions. Adversarial training has proven effective in improving the robustness of DNNs against a wide range of adversarial perturbations. However, a general framework for adversarial defences is needed that will extend beyond a single-dimensional assessment of robustness improvement; it is essential to consider simultaneously several distance metrics and adversarial attack strategies. Using such an approach we report the results from extensive experimentation on adversarial defence methods that could improve DNNs resilience to adversarial threats. We wrap up by introducing a general adversarial training methodology, which, according to our experimental results, opens prospects for an holistic defence against a range of diverse types of adversarial perturbations.
Charis Eleftheriadis, Andreas L. Symeonidis, Panagiotis Katsaros
Mach. Vis. Appl.2
2024 SmAuto: A domain-specific-language for application development in smart environments
Konstantinos Panayiotou, Constantine Doumanidis, Emmanouil G. Tsardoulias, Andreas L. Symeonidis
Pervasive Mob. Comput.4
2023 Towards Readability-Aware Recommendations of Source Code Snippets
Athanasios Michailoudis, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
ICSOFT3
2023 Towards Interpretable Monitoring and Assignment of Jira Issues
Dimitrios-Nikitas Nastos, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
ICSOFT3
2023 Semantically-enriched Jira Issue Tracking Data
abstract
Current state of practice dictates that software developers host their projects online and employ project management systems to monitor the development of product features, keep track of bugs, and prioritize task assignments. The data stored in these systems, if their semantics are extracted effectively, can be used to answer several interesting questions, such as finding who is the most suitable developer for a task, what the priority of a task should be, or even what is the actual workload of the software team. To support researchers and practitioners that work towards these directions, we have built a system that crawls data from the Jira management system, performs topic modeling on the data to extract useful semantics and stores them in a practical database schema. We have used our system to retrieve and analyze 656 projects of the Apache Software Foundation, comprising data from more than a million Jira issues.
Themistoklis G. Diamantopoulos, Dimitrios-Nikitas Nastos, Andreas L. Symeonidis
MSR3
2023 Automated issue assignment using topic modelling on Jira issue tracking data
abstract
Abstract As more and more software teams use online issue tracking systems to collaborate on software projects, the accurate assignment of new issues to the most suitable contributors may have significant impact on the success of the project. As a result, several research efforts have been directed towards automating this process to save considerable time and effort. However, most approaches focus mainly on software bugs and employ models that do not sufficiently take into account the semantics and the non‐textual metadata of issues and/or produce models that may require manual tuning. A methodology that extracts both textual and non‐textual features from different types of issues is designed, providing a Jira dataset that involves not only bugs but also new features, issues related to documentation, patches, etc. Moreover, the semantics of issue text are effectively captured by employing a topic modelling technique that is optimised using the assignment result. Finally, this methodology aggregates probabilities from a set of individual models to provide the final assignment. Upon evaluating this approach in an automated issue assignment setting using a dataset of Jira issues, the authors conclude that it can be effective for automated issue assignment.
Themistoklis G. Diamantopoulos, Nikolaos Saoulidis, Andreas L. Symeonidis
IET Softw.3
2023 Calista: A deep learning-based system for understanding and evaluating website aesthetics
Alexandros Delitzas, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis
Int. J. Hum. Comput. Stud.3
2022 Semantic Code Search in Software Repositories using Neural Machine Translation
abstract
Abstract Nowadays, software development is accelerated through the reuse of code snippets found online in question-answering platforms and software repositories. In order to be efficient, this process requires forming an appropriate query and identifying the most suitable code snippet, which can sometimes be challenging and particularly time-consuming. Over the last years, several code recommendation systems have been developed to offer a solution to this problem. Nevertheless, most of them recommend API calls or sequences instead of reusable code snippets. Furthermore, they do not employ architectures advanced enough to exploit the semantics of natural language and code in order to form the optimal query from the question posed. To overcome these issues, we propose CodeTransformer, a code recommendation system that provides useful, reusable code snippets extracted from open-source GitHub repositories. By employing a neural network architecture that comprises advanced attention mechanisms, our system effectively understands and models natural language queries and code snippets in a joint vector space. Upon evaluating CodeTransformer quantitatively against a similar system and qualitatively using a dataset from Stack Overflow, we conclude that our approach can recommend useful and reusable snippets to developers.
Evangelos Papathomas, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
FASE3
2022 A Heuristic Approach towards Continuous Implicit Authentication
abstract
Smartphones nowadays handle large amounts of sensitive user information, since users exchange undisclosed information on an everyday basis. This generates the need for more effective authentication mechanisms, deviating from the traditional ones. In this direction, many research approaches are targeted towards continuous implicit authentication, on the basis of modelling the constant interaction of the user with the device. These approaches yield promising results, however certain improvements can be made by exploiting the sequential order of the predictions and the known performance metrics. In this work, we propose a heuristics algorithm, which, given a series of predictions from any continuous implicit authentication model, can ex-ploit the sequential order in order to fix any false predictions and improve the accuracy of the smartphone security system. Preliminary evaluation on several axes indicates that our approach can effectively improve any CIA model and achieve significantly better results.
Georgios Kalantzis 0002, Gerasimos Papakostas, Thomas Karanikiotis, Michail Papamichail, Andreas L. Symeonidis
IJCB5
2022 A Mechanism for Automatically Extracting Reusable and Maintainable Code Idioms from Software Repositories
Argyrios Papoudakis, Thomas Karanikiotis, Andreas L. Symeonidis
ICSOFT3
2022 A Methodology for Enabling NLP Capabilities on Edge and Low-Resource Devices
Andreas Goulas, Nikolaos Malamas, Andreas L. Symeonidis
NLDB3
2022 A mechanism for personalized Automatic Speech Recognition for less frequently spoken languages: the Greek case
Panagiotis Antoniadis, Emmanouil G. Tsardoulias, Andreas L. Symeonidis
Multim. Tools Appl.3
2021 Software Task Importance Prediction based on Project Management Data
Themistoklis G. Diamantopoulos, Christiana Galegalidou, Andreas L. Symeonidis
ICSOFT3
2021 Towards Automatically Generating a Personalized Code Formatting Mechanism
Thomas Karanikiotis, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis
ICSOFT3
2021 Embedding Rasa in edge Devices: Capabilities and Limitations
abstract
Over the past few years, there has been a boost in the use of commercial virtual assistants. Obviously, these proprietary tools are well-performing, however the functionality they offer is limited, users are ”vendor-locked”, while possible user privacy issues rise. In this paper we argue that low-cost, open hardware solutions may also perform well, given the proper setup. Specifically, we perform an initial assessment of a low-cost virtual agent employing the Rasa framework integrated into a Raspberry Pi 4. We set up three different architectures, discuss their capabilities and limitations and evaluate the dialogue system against three axes: assistant comprehension, task success and assistant usability. Our experiments show that our low-cost virtual assistant performs in a satisfactory manner, even when a small-sized training dataset is used.
Nikolaos Malamas, Andreas L. Symeonidis
KES2
2021 Optimizing Sales Forecasting in e-Commerce with ARIMA and LSTM Models
Konstantinos N. Vavliakis, Andreas Siailis, Andreas L. Symeonidis
WEBIST3
2021 Defining behaviorizeable relations to enable inference in semi-automatic program synthesis
Emmanouil Krasanakis, Andreas L. Symeonidis
J. Log. Algebraic Methods Program.2
2020 Extracting Semantics from Question-Answering Services for Snippet Reuse
abstract
Nowadays, software developers typically search online for reusable solutions to common programming problems. However, forming the question appropriately, and locating and integrating the best solution back to the code can be tricky and time consuming. As a result, several mining systems have been proposed to aid developers in the task of locating reusable snippets and integrating them into their source code. Most of these systems, however, do not model the semantics of the snippets in the context of source code provided. In this work, we propose a snippet mining system, named StackSearch, that extracts semantic information from Stack Overlow posts and recommends useful and in-context snippets to the developer. Using a hybrid language model that combines Tf-Idf and fastText, our system effectively understands the meaning of the given query and retrieves semantically similar posts. Moreover, the results are accompanied with useful metadata using a named entity recognition technique. Upon evaluating our system in a set of common programming queries, in a dataset based on post links, and against a similar tool, we argue that our approach can be useful for recommending ready-to-use snippets to the developer.
Themistoklis G. Diamantopoulos, Nikolaos Oikonomou, Andreas L. Symeonidis
FASE3
2020 Audio-based Near-Duplicate Video Retrieval with Audio Similarity Learning
abstract
In this work, we address the problem of audio-based near-duplicate video retrieval. We propose the Audio Similarity Learning (AuSiL) approach that effectively captures temporal patterns of audio similarity between video pairs. For the robust similarity calculation between two videos, we first extract representative audio-based video descriptors by leveraging transfer learning based on a Convolutional Neural Network (CNN) trained on a large scale dataset of audio events, and then we calculate the similarity matrix derived from the pairwise similarity of these descriptors. The similarity matrix is subsequently fed to a CNN network that captures the temporal structures existing within its content. We train our network following a triplet generation process and optimizing the triplet loss function. To evaluate the effectiveness of the proposed approach, we have manually annotated two publicly available video datasets based on the audio duplicity between their videos. The proposed approach achieves very competitive results compared to three state-of-the-art methods. Also, unlike the competing methods, it is very robust to the retrieval of audio duplicates generated with speed transformations.
Pavlos Avgoustinakis, Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Andreas L. Symeonidis, Ioannis Kompatsiaris
ICPR4
2020 A Data-driven Methodology towards Interpreting Readability against Software Properties
Thomas Karanikiotis, Michail Papamichail, Ioannis Gonidelis, Dimitra Karatza, Andreas L. Symeonidis
ICSOFT5
2020 Employing Contribution and Quality Metrics for Quantifying the Software Development Process
abstract
The full integration of online repositories in contemporary software development promotes remote work and collaboration. Apart from the apparent benefits, online repositories offer a deluge of data that can be utilized to monitor and improve the software development process. Towards this direction, we have designed and implemented a platform that analyzes data from GitHub in order to compute a series of metrics that quantify the contributions of project collaborators, both from a development as well as an operations (communication) perspective. We analyze contributions throughout the projects' lifecycle and track the number of coding violations, this way aspiring to identify cases of software development that need closer monitoring and (possibly) further actions to be taken. In this context, we have analyzed the 3000 most popular GitHub Java projects and provide the data to the community.
Themistoklis G. Diamantopoulos, Michail Papamichail, Thomas Karanikiotis, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis
MSR5
2020 Continuous Implicit Authentication through Touch Traces Modelling
abstract
Nowadays, the continuously increasing use of smart-phones as the primary way of dealing with day-to-day tasks raises several concerns mainly focusing on privacy and security. In this context and given the known limitations and deficiencies of traditional authentication mechanisms, a lot of research efforts are targeted towards continuous implicit authentication on the basis of behavioral biometrics. In this work, we propose a methodology towards continuous implicit authentication that refrains from the limitations imposed by small-scale and/or controlled environment experiments by employing a real-world application used widely by a large number of individuals. Upon constructing our models using Support Vector Machines, we introduce a confidence-based methodology, in order to strengthen the effectiveness and the efficiency of our approach. The evaluation of our methodology on a set of diverse scenarios indicates that our approach achieves good results both in terms of efficiency and usability.
Thomas Karanikiotis, Michail Papamichail, Kyriakos C. Chatzidimitriou, Napoleon-Christos I. Oikonomou, Andreas L. Symeonidis, Sashi K. Saripalle
QRS5
2020 Towards Analyzing Contributions from Software Repositories to Optimize Issue Assignment
abstract
Most software teams nowadays host their projects online and monitor software development in the form of issues/tasks. This process entails communicating through comments and reporting progress through commits and closing issues. In this context, assigning new issues, tasks or bugs to the most suitable contributor largely improves efficiency. Thus, several automated issue assignment approaches have been proposed, which however have major limitations. Most systems focus only on assigning bugs using textual data, are limited to projects explicitly using bug tracking systems, and may require manually tuning parameters per project. In this work, we build an automated issue assignment system for GitHub, taking into account the commits and issues of the repository under analysis. Our system aggregates feature probabilities using a neural network that adapts to each project, thus not requiring manual parameter tuning. Upon evaluating our methodology, we conclude that it can be efficient for automated issue assignment.
Vasileios Matsoukas, Themistoklis G. Diamantopoulos, Michail Papamichail, Andreas L. Symeonidis
QRS4
2020 A generic methodology for early identification of non-maintainable source code components through analysis of software releases
Michail Papamichail, Andreas L. Symeonidis
Inf. Softw. Technol.2
2020 Boosted seed oversampling for local community ranking
abstract
Local community detection is an emerging topic in network analysis that aims to detect well-connected communities encompassing sets of priorly known seed nodes. In this work, we explore the similar problem of ranking network nodes based on their relevance to the communities characterized by seed nodes. However, seed nodes may not be central enough or sufficiently many to produce high quality ranks. To solve this problem, we introduce a methodology we call seed oversampling, which first runs a node ranking algorithm to discover more nodes that belong to the community and then reruns the same ranking algorithm for the new seed nodes. We formally discuss why this process improves the quality of calculated community ranks if the original set of seed nodes is small and introduce a boosting scheme that iteratively repeats seed oversampling to further improve rank quality when certain ranking algorithm properties are met. Finally, we demonstrate the effectiveness of our methods in improving community relevance ranks given only a few random seed nodes of real-world network communities. In our experiments, boosted and simple seed oversampling yielded better rank quality than the previous neighborhood inflation heuristic, which adds the neighborhoods of original seed nodes to seeds.
Emmanouil Krasanakis, Emmanouil Schinas, Symeon Papadopoulos, Ioannis Kompatsiaris, Andreas L. Symeonidis
Inf. Process. Manag.5
2019 npm Packages as Ingredients: A Recipe-based Approach
Kyriakos C. Chatzidimitriou, Michail Papamichail, Themistoklis G. Diamantopoulos, Napoleon-Christos I. Oikonomou, Andreas L. Symeonidis
ICSOFT5
2019 Towards Extracting the Role and Behavior of Contributors in Open-source Projects
Michail Papamichail, Themistoklis G. Diamantopoulos, Vasileios Matsoukas, Christos L. Athanasiadis, Andreas L. Symeonidis
ICSOFT5
2019 Towards mining answer edits to extract evolution patterns in stack overflow
abstract
The current state of practice dictates that in order to solve a problem encountered when building software, developers ask for help in online platforms, such as Stack Overflow. In this context of collaboration, answers to question posts often undergo several edits to provide the best solution to the problem stated. In this work, we explore the potential of mining Stack Overflow answer edits to extract common patterns when answering a post. In particular, we design a similarity scheme that takes into account the text and code of answer edits and cluster edits according to their semantics. Upon applying our methodology, we provide frequent edit patterns and indicate how they could be used to answer future research questions. Assessing our approach indicates that it can be effective for identifying commonly applied edits, thus illustrating the transformation path from the initial answer to the optimal solution.
Themistoklis G. Diamantopoulos, Maria-Ioanna Sifaki, Andreas L. Symeonidis
MSR3
2019 A Mechanism for Automatically Summarizing Software Functionality from Source Code
abstract
When developers search online to find software components to reuse, they usually first need to understand the container projects/libraries, and subsequently identify the required functionality. Several approaches identify and summarize the offerings of projects from their source code, however they often require that the developer has knowledge of the underlying topic modeling techniques; they do not provide a mechanism for tuning the number of topics, and they offer no control over the top terms for each topic. In this work, we use a vectorizer to extract information from variable/method names and comments, and apply Latent Dirichlet Allocation to cluster the source code files of a project into different semantic topics. The number of topics is optimized based on their purity with respect to project packages, while topic categories are constructed to provide further intuition and Stack Exchange tags are used to express the topics in more abstract terms.
Christos Psarras, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
QRS3
2019 Cenote: A Big Data Management and Analytics Infrastructure for the Web of Things
abstract
In the era of Big Data, Cloud Computing and Internet of Things, most of the existing, integrated solutions that attempt to solve their challenges are either proprietary, limit functionality to a predefined set of requirements, or hide the way data are stored and accessed. In this work we propose Cenote, an open source Big Data management and analytics infrastructure for the Web of Things that overcomes the above limitations. Cenote is built on component-based software engineering principles and provides an all-inclusive solution based on components that work well individually.
Kyriakos C. Chatzidimitriou, Michail Papamichail, Napoleon-Christos I. Oikonomou, Dimitrios Lampoudis, Andreas L. Symeonidis
WI5
2019 Measuring the reusability of software components using static analysis metrics and reuse rate information
Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
J. Syst. Softw.3
2019 Of daemons and men: reducing false positive rate in intrusion detection systems with file system footprint analysis
George Mamalakis, Christos Diou, Andreas L. Symeonidis, Leonidas Georgiadis
Neural Comput. Appl.3
2018 npm-miner: an infrastructure for measuring the quality of the npm registry
abstract
As the popularity of the JavaScript language is constantly increasing, one of the most important challenges today is to assess the quality of JavaScript packages. Developers often employ tools for code linting and for the extraction of static analysis metrics in order to assess and/or improve their code. In this context, we have developed npn-miner, a platform that crawls the npm registry and analyzes the packages using static analysis tools in order to extract detailed quality metrics as well as high-level quality attributes, such as maintainability and security. Our infrastructure includes an index that is accessible through a web interface, while we have also constructed a dataset with the results of a detailed analysis for 2000 popular npm packages.
Kyriakos C. Chatzidimitriou, Michail Papamichail, Themistoklis G. Diamantopoulos, Michail Tsapanos, Andreas L. Symeonidis
MSR5
2018 Recommendation Systems in a Conversational Web
Konstantinos N. Vavliakis, Maria Th. Kotouza, Andreas L. Symeonidis, Pericles A. Mitkas
WEBIST3
2017 Robotic Applications Towards an Interactive Alerting System for Medical Purposes
abstract
Social consumer robots are slowly but strongly invading our everyday lives as their prices are becoming lower and lower, constituting them affordable for a wide range of civilians. There has been a lot of research concerning the potential applications of social robots, some of which may implement companionship or proxying technology-related tasks and assisting in everyday household endeavors, among others. In the current work, the RAPP framework is being used towards easily creating robotic applications suitable for utilization as a socially interactive alerting system with the employment of the NAO robot. The developed application stores events in an on-line calendar, directly via the robot or indirectly via a web environment, and asynchronously informs an end-user of imminent events.
Konstantinos Panayiotou, Sofia Reppou, George Karagiannis, Emmanouil G. Tsardoulias, Aristeidis G. Thallas, Andreas L. Symeonidis
CBMS6
2017 Towards Modeling the User-perceived Quality of Source Code using Static Analysis Metrics
Valasia Dimaridou, Alexandros-Charalampos Kyprianidis, Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
ICSOFT5
2017 From requirements to source code: a Model-Driven Engineering approach for RESTful web services
Christoforos Zolotas, Themistoklis G. Diamantopoulos, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis
Autom. Softw. Eng.4
2017 QATCH - An adaptive framework for software product quality assessment
Miltiadis G. Siavvas, Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis
Expert Syst. Appl.3
2016 QualBoa: reusability-aware recommendations of source code components
abstract
Contemporary software development processes involve finding reusable software components from online repositories and integrating them to the source code, both to reduce development time and to ensure that the final software project is of high quality. Although several systems have been designed to automate this procedure by recommending components that cover the desired functionality, the reusability of these components is usually not assessed by these systems. In this work, we present QualBoa, a recommendation system for source code components that covers both the functional and the quality aspects of software component reuse. Upon retrieving components, QualBoa provides a ranking that involves not only functional matching to the query, but also a reusability score based on configurable thresholds of source code metrics. The evaluation of QualBoa indicates that it can be effective for recommending reusable source code.
Themistoklis G. Diamantopoulos, Klearchos Thomopoulos, Andreas L. Symeonidis
MSR3
2016 User-Perceived Source Code Quality Estimation Based on Static Analysis Metrics
abstract
The popularity of open source software repositories and the highly adopted paradigm of software reuse have led to the development of several tools that aspire to assess the quality of source code. However, most software quality estimation tools, even the ones using adaptable models, depend on fixed metric thresholds for defining the ground truth. In this work we argue that the popularity of software components, as perceived by developers, can be considered as an indicator of software quality. We present a generic methodology that relates quality with source code metrics and estimates the quality of software components residing in popular GitHub repositories. Our methodology employs two models: a one-class classifier, used to rule out low quality code, and a neural network, that computes a quality score for each software component. Preliminary evaluation indicates that our approach can be effective for identifying high quality software components in the context of reuse.
Michail Papamichail, Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
QRS3
2015 Employing Source Code Information to Improve Question-Answering in Stack Overflow
abstract
Nowadays, software development has been greatly influenced by question-answering communities, such as Stack Overflow. A new problem-solving paradigm has emerged, as developers post problems they encounter that are then answered by the community. In this paper, we propose a methodology that allows searching for solutions in Stack Overflow, using the main elements of a question post, including not only its title, tags, and body, but also its source code snippets. We describe a similarity scheme for these elements and demonstrate how structural information can be extracted from source code snippets and compared to further improve the retrieval of questions. The results of our evaluation indicate that our methodology is effective on recommending similar question posts allowing community members to search without fully forming a question.
Themistoklis G. Diamantopoulos, Andreas L. Symeonidis
MSR2
2015 Identifying valid search engine ranking factors in a Web 2.0 and Web 3.0 context for building efficient SEO mechanisms
Themistoklis Mavridis, Andreas L. Symeonidis
Eng. Appl. Artif. Intell.2
2014 Towards the Design of User Friendly Search Engines for Software Projects
Rafaila Grigoriou, Andreas L. Symeonidis
NLDB2
2014 Bottom-up modeling of small-scale energy consumers for effective Demand Response Applications
Antonios C. Chrysopoulos, Christos Diou, Andreas L. Symeonidis, Pericles A. Mitkas
Eng. Appl. Artif. Intell.3
2014 Semantic analysis of web documents for the generation of optimal content
Themistoklis Mavridis, Andreas L. Symeonidis
Eng. Appl. Artif. Intell.2
2013 Event identification in web social media through named entity recognition and topic modeling
Konstantinos N. Vavliakis, Andreas L. Symeonidis, Pericles A. Mitkas
Data Knowl. Eng.2
2012 Identifying Webpage Semantics for Search Engine Optimization
Themistoklis Mavridis, Andreas L. Symeonidis
WEBIST2
2011 An integrated framework for enhancing the semantic transformation, editing and querying of relational databases
Konstantinos N. Vavliakis, Andreas L. Symeonidis, Georgios T. Karagiannis, Pericles A. Mitkas
Expert Syst. Appl.2
2010 Towards Understanding How Personality, Motivation, and Events Trigger Web User Activity
abstract
Web 2.0 provided internet users with a dynamic medium, where information is updated continuously and anyone can participate. Though preliminary analysis exists, there is still little understanding on what exactly stimulates users to actively participate, create and share content in online communities. In this paper we present a methodology that aspires to identify and analyze those events that trigger web user activity, content creation and sharing in Web 2.0. Our approach is based on user personality and motivation, and on the occurrence of events with a personal or global impact. The proposed methodology was applied on data collected from Flickr and analysis was performed through the use of statistics and data mining techniques.
Konstantinos N. Vavliakis, Andreas L. Symeonidis, Pericles A. Mitkas
Web Intelligence2
2009 An integrated infrastructure for monitoring and evaluating agent-based systems
Christos Dimou, Andreas L. Symeonidis, Pericles A. Mitkas
Expert Syst. Appl.2
2008 BioCrawler: An intelligent crawler for the semantic web
Alexandros Batzios, Christos Dimou, Andreas L. Symeonidis, Pericles A. Mitkas
Expert Syst. Appl.3
2008 Agent Mertacor: A robust design for dealing with uncertainty and variation in SCM environments
Kyriakos C. Chatzidimitriou, Andreas L. Symeonidis, Ioannis Kontogounis, Pericles A. Mitkas
Expert Syst. Appl.2
2007 Eikonomia-An Integrated Semantically Aware Tool for Description and Retrieval of Byzantine Art Information
abstract
Semantic annotation and querying is currently applied on a number of versatile disciplines, providing the added-value of such an approach and, consequently the need for more elaborate - either case-specific or generic - tools. In this context, we have developed Eikonomia: an integrated semantically-aware tool for the description and retrieval of Byzantine artwork Information. Following the needs of the ORMYLIA Art Diagnosis Center for adding semantics to their legacy data, an ontology describing Byzantine artwork based on CIDOC-CRM, along with the interfaces for synchronization to and from the existing RDBMS have been implemented. This ontology has been linked to a reasoning tool, while a dynamic interface for the automated creation of semantic queries in SPARQL was developed. Finally, all the appropriate interfaces were instantiated, in order to allow easy ontology manipulation, query results projection and restrictions creation.
Konstantinos N. Vavliakis, Andreas L. Symeonidis, Georgios T. Karagiannis, Pericles A. Mitkas
ICTAI (2)2
2007 Data mining for agent reasoning: A synergy for training intelligent agents
Andreas L. Symeonidis, Kyriakos C. Chatzidimitriou, Ioannis N. Athanasiadis, Pericles A. Mitkas
Eng. Appl. Artif. Intell.1
2007 A retraining methodology for enhancing agent intelligence
Andreas L. Symeonidis, Ioannis N. Athanasiadis, Pericles A. Mitkas
Knowl. Based Syst.1
2006 GeneCity: A Multi Agent Simulation Environment for Hereditary Diseases
abstract
Simulating the psycho-societal aspects of a human community is an issue always intriguing and challenging, aspiring us to help better understand, macroscopically, the way(s) humans behave. The mathematical models that have extensively been used for the analytical study of the various related phenomena prove inefficient, since they cannot conceive the notion of population heterogeneity, a parameter highly critical when it comes to community interactions. Following the more successful paradigm of artificial societies, coupled with multi-agent systems and other Artificial Intelligence primitives, and extending previous epidemiological research work, we have developed GeneCity: an extended agent community, where agents live and interact under the veil of a hereditary epidemic. The members of the community, which can be either healthy, carriers of a trait, or patients, exhibit a number of human-like social (and medical) characteristics: wealth, acceptance and influence, fear and knowledge, phenotype and reproduction ability. GeneCity provides a highly-configurable interface for simulating social environments and the way they are affected with the appearance of a hereditary disease, either Autosome or X-linked. This paper presents an analytical overview of the work conducted and examines a testhypothesis based on the spreading of Thalassaemia major.
Demetrios Eliades, Andreas L. Symeonidis, Pericles A. Mitkas
AICCSA2
2006 A Multi-Agent Simulation Framework for Spiders Traversing the Semantic Web
abstract
Although search engines traditionally use spiders for traversing and indexing the Web, there has not yet been any methodological attempt to model, deploy and test learning spiders. The flourishing of the semantic Web provides understandable information that may improve the accuracy of search engines. In this paper, we introduce BioSpider, an agent-based simulation framework for developing and testing autonomous, intelligent, semantically-focused Web spiders. BioSpider assumes a direct analogy of the problem at hand with a multi-variate ecosystem, where each member is self-maintaining. The population of the ecosystem comprises cooperative spiders incorporating communication, mobility and learning skills, striving to improve efficiency. Genetic algorithms and classifier rules have been employed for spider adaptation and learning. A set of experiments has been performed in order to qualitatively test the efficacy and applicability of the proposed approach
Christos Dimou, Alexandros Batzios, Andreas L. Symeonidis, Pericles A. Mitkas
Web Intelligence3
2005 Biotope: an integrated framework for simulating distributed multiagent computational systems
abstract
The study of distributed computational systems issues, such as heterogeneity, concurrency, control, and coordination, has yielded a number of models and architectures, which aspire to provide satisfying solutions to each of the above problems. One of the most intriguing and complex classes of distributed systems are computational ecosystems, which add an "ecological" perspective to these issues and introduce the characteristic of self-organization. Extending previous research work on self-organizing communities, we have developed Biotope, which is an agent simulation framework, where each one of its members is dynamic and self-maintaining. The system provides a highly configurable interface for modeling various environments as well as the "living" or computational entities that reside in them, while it introduces a series of tools for monitoring system evolution. Classifier systems and genetic algorithms have been employed for agent learning, while the dispersal distance theory has been adopted for agent replication. The framework has been used for the development of a characteristic demonstrator, where Biotope agents are engaged in well-known vital activities-nutrition, communication, growth, death-directed toward their own self-replication, just like in natural environments. This paper presents an analytical overview of the work conducted and concludes with a methodology for simulating distributed multiagent computational systems.
Andreas L. Symeonidis, E. Valtos, S. Seroglou, Pericles A. Mitkas
IEEE Trans. Syst. Man Cybern. Part A1
2003 Intelligent policy recommendations on enterprise resource planning by the use of agent technology and data mining techniques
Andreas L. Symeonidis, Dionisis D. Kehagias, Pericles A. Mitkas
Expert Syst. Appl.1