VLDB 2026 Research / reviewers in the wild / expert
Panagiotis Louridas
dblp:63/5819 · also Panos Louridas
· DBLP profile ↗
31ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-3971-4612ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 since 2021Security and privacy · 4Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Sequential Recommendations with LLMsabstractThe sequential recommendation problem has attracted considerable research attention in the past few years, leading to the rise of numerous recommendation models. In this work, we explore how Large Language Models (LLMs), which are nowadays introducing disruptive effects in many AI-based applications, can be used to build or improve sequential recommendation approaches. Specifically, we design three orthogonal approaches and hybrids of those to leverage the power of LLMs in different ways. In addition, we investigate the potential of each approach by focusing on its technical aspects and determining an array of alternative choices for each one. We conduct extensive experiments on three datasets and explore a large variety of configurations, including different language models and baseline recommendation models, to obtain a comprehensive picture of the performance of each approach. Among other observations, we highlight that initializing state-of-the-art sequential recommendation models such as BERT4Rec or SASRec with embeddings obtained from an LLM can lead to substantial performance gains in terms of accuracy. Furthermore, we find that fine-tuning an LLM for recommendation tasks enables it to learn not only the tasks but also the concepts of a domain to some extent. We also show that fine-tuning OpenAI GPT leads to considerably better performance than fine-tuning Google PaLM 2. Overall, our extensive experiments indicate a huge potential value of leveraging LLMs in future recommendation approaches. We publicly share the code and data of our experiments to ensure reproducibility. 1 Artun Boz, Wouter Zorgdrager, Zoe Kotti, Jesse Harte, Panagiotis Louridas, Vassilios Karakoidas, Dietmar Jannach, Marios Fragkoulis |
Trans. Recomm. Syst. | 5 |
| 2025 | PyTrim: A Practical Tool for Reducing Python Dependency BloatabstractDependency bloat is a persistent challenge in Python projects, which increases maintenance costs and security risks. While numerous tools exist for detecting unused dependencies in Python, removing these dependencies across the source code and configuration files of a project requires manual effort and expertise. To tackle this challenge we introduce PYTRIM, an end-to-end system to automate this process. PYTRIM eliminates unused imports and package declarations across a variety of file types, including Python source and configuration files such as requirements.txt and setup.py. PYTRIM’s modular design makes it agnostic to the source of dependency bloat information, enabling integration with any detection tool. Beyond its contribution when it comes to automation, PYTRIM also incorporates a novel dynamic analysis component that improves dependency detection recall. Our evaluation of PYTRIM’s end-to-end effectiveness on a ground-truth dataset of 37 merged pull requests from prior work, shows that PYTRIM achieves 98.3% accuracy in replicating human-made changes. To show its practical impact, we run PYTRIM on 971 open-source packages, identifying and trimming bloated dependencies in 39 of them. For each case, we submit a corresponding pull request, 14 of which have already been accepted and merged. PYTRIM is available as an open-source project, encouraging community contributions and further development.Video demonstration: https://youtu.be/LqTEdOUbJRICode repository: https://github.com/TrimTeam/PyTrim Konstantinos Karakatsanis, Georgios Alexopoulos, Ioannis Karyotakis, Foivos Timotheos Proestakis, Evangelos Talos, Panagiotis Louridas, Dimitris Mitropoulos |
ASE | 6 |
| 2024 | Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek ProverbsabstractProverbs carry wisdom transferred orally from generation to generation.Based on the place they were recorded, this study introduces a publicly-available and machine-actionable dataset of more than one hundred thousand Greek proverb variants.By quantifying the spatial distribution of proverbs, we show that the most widespread proverbs come from the mainland while the least widespread proverbs come primarily from the islands.By focusing on the least dispersed proverbs, we present the most frequent tokens per location and undertake a benchmark in geographical attribution, using text classification and regression (text geocoding).Our results show that this is a challenging task for which specific locations can be attributed more successfully compared to others.The potential of our resource and benchmark is showcased by two novel applications.First, we extracted terms moving the regression prediction toward the four cardinal directions.Second, we leveraged conformal prediction to attribute 3,676 unregistered proverbs with statistically rigorous predictions of locations each of these proverbs was possibly registered in. John Pavlopoulos, Panagiotis Louridas, Panagiotis Filos |
EMNLP | 2 |
| 2024 | Broken Windows: Exploring the Applicability of a Controversial Theory on Code QualityabstractIs the quality of existing code correlated with the quality of subsequent changes? According to the (controversial) broken windows theory, which inspired this study, disorder sets descriptive norms and signals behavior that further increases it. From a large code corpus, we examine whether code history does indeed affect the evolution of code quality. We examine C code quality metrics and Java code smells in specific files, and see whether subsequent commits by developers continue on that path. We check whether developers tailor the quality of their commits based on the quality of the file they commit to. Our results show that history matters, that developers behave differently depending on some aspects of the code quality they encounter, and that programming style inconsistency is not necessarily related to structural qualities. These findings have implications for both software practice and research. Software practitioners can emphasize current quality practices as these influence the code that will be developed in the future. Researchers in the field may replicate and extend the study to improve our understanding of the theory and its practical implications on artifacts, processes, and people. Diomidis Spinellis, Panagiotis Louridas, Maria Kechagia, Tushar Sharma 0001 |
ICSME | 2 |
| 2023 | Leveraging Large Language Models for Sequential RecommendationabstractSequential recommendation problems have received increasing attention in research during the past few years, leading to the inception of a large variety of algorithmic approaches. In this work, we explore how large language models (LLMs), which are nowadays introducing disruptive effects in many AI-based applications, can be used to build or improve sequential recommendation approaches. Specifically, we devise and evaluate three approaches to leverage the power of LLMs in different ways. Our results from experiments on two datasets show that initializing the state-of-the-art sequential recommendation model BERT4Rec with embeddings obtained from an LLM improves NDCG by 15-20% compared to the vanilla BERT4Rec model. Furthermore, we find that a simple approach that leverages LLM embeddings for producing recommendations, can provide competitive performance by highlighting semantically related items. We publicly share the code and data of our experiments to ensure reproducibility.1 Jesse Harte, Wouter Zorgdrager, Panagiotis Louridas, Asterios Katsifodimos, Dietmar Jannach, Marios Fragkoulis |
RecSys | 3 |
| 2022 | A Greek Parliament Proceedings Dataset for Computational Linguistics and Political AnalysisabstractLarge, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically from 1989 up to 2020. It consists of more than 1 million speeches with extensive meta-data, extracted from 5,355 parliamentary sitting record files. We explain how it was constructed and the challenges that had to be overcome. The dataset can be used for both computational linguistics and political analysis---ideally, combining the two. We present such an application, showing (i) how the dataset can be used to study the change of word usage through time, (ii) between significant historical events and political parties, (iii) by evaluating and employing algorithms for detecting semantic shifts. Konstantina Dritsa, Aikaterini Thoma, Ioannis Pavlopoulos, Panagiotis Louridas |
NeurIPS | 4 |
| 2021 | PyCG: Practical Call Graph Generation in PythonabstractCall graphs play an important role in different contexts, such as profiling and vulnerability propagation analysis. Generating call graphs in an efficient manner can be a challenging task when it comes to high-level languages that are modular and incorporate dynamic features and higher-order functions. Despite the language's popularity, there have been very few tools aiming to generate call graphs for Python programs. Worse, these tools suffer from several effectiveness issues that limit their practicality in realistic programs. We propose a pragmatic, static approach for call graph generation in Python. We compute all assignment relations between program identifiers of functions, variables, classes, and modules through an inter-procedural analysis. Based on these assignment relations, we produce the resulting call graph by resolving all calls to potentially invoked functions. Notably, the underlying analysis is designed to be efficient and scalable, handling several Python features, such as modules, generators, function closures, and multiple inheritance. We have evaluated our prototype implementation, which we call PyCG, using two benchmarks: a micro-benchmark suite containing small Python programs and a set of macro-benchmarks with several popular real-world Python packages. Our results indicate that PyCG can efficiently handle thousands of lines of code in less than a second (0.38 seconds for 1k LoC on average). Further, it outperforms the state-of-the-art for Python in both precision and recall: PyCG achieves high rates of precision ~99.2% and adequate recall ~69.9%. Finally, we demonstrate how PyCG can aid dependency impact analysis by showcasing a potential enhancement to GitHub's "security advisory" notification service using a real-world example. Vitalis Salis, Thodoris Sotiropoulos, Panagiotis Louridas, Diomidis Spinellis, Dimitris Mitropoulos |
ICSE | 3 |
| 2021 | CrossVul: a cross-language vulnerability dataset with commit dataabstractExamining the characteristics of software vulnerabilities and the code that contains them can lead to the development of more secure software. We present a dataset (∼1.4 GB) containing vulnerable source code files together with the corresponding, patched versions. Contrary to other existing vulnerability datasets, ours includes vulnerable files written in more than 40 programming languages. Each file is associated to (1) a Common Vulnerability Exposures identifier (CVE ID) and (2) the repository it came from. Further, our dataset can be the basis for machine learning applications that identify defects, as we show in specific examples. We also present a supporting dataset that contains commit messages derived from Git commits that serve as security patches. This dataset can be used to train ML models that in turn, can be used to detect security patch commits as we highlight in a specific use case. Georgios Nikitopoulos, Konstantina Dritsa, Panagiotis Louridas, Dimitris Mitropoulos |
ESEC/SIGSOFT FSE | 3 |
| 2021 | Code smell detection by deep direct-learning and transfer-learning
Tushar Sharma 0001, Vasiliki Efstathiou, Panagiotis Louridas, Diomidis Spinellis |
J. Syst. Softw. | 3 |
| 2020 | A Non-interactive Shuffle Argument with Low Trust Assumptions
Antonis Aggelakis, Prastudy Fauzi, Georgios Korfiatis, Panagiotis Louridas, Foteinos Mergoupis-Anagnou, Janno Siim, Michal Zajac 0001 |
CT-RSA | 4 |
| 2020 | A Dataset of Enterprise-Driven Open Source SoftwareabstractWe present a dataset of open source software developed mainly by enterprises rather than volunteers. This can be used to address known generalizability concerns, and, also, to perform research on open source business software development. Based on the premise that an enterprise's employees are likely to contribute to a project developed by their organization using the email account provided by it, we mine domain names associated with enterprises from open data sources as well as through white- and blacklisting, and use them through three heuristics to identify 17 264 enterprise GitHub projects. We provide these as a dataset detailing their provenance and properties. A manual evaluation of a dataset sample shows an identification accuracy of 89%. Through an exploratory data analysis we found that projects are staffed by a plurality of enterprise insiders, who appear to be pulling more than their weight, and that in a small percentage of relatively large projects development happens exclusively through enterprise insiders. Diomidis Spinellis, Zoe Kotti, Konstantinos Kravvaritis, Georgios Theodorou, Panagiotis Louridas |
MSR | 5 |
| 2020 | Search Engine Similarity Analysis: A Combined Content and Rankings Approach
Konstantina Dritsa, Thodoris Sotiropoulos, Haris Skarpetis, Panagiotis Louridas |
WISE (2) | 4 |
| 2019 | Time present and time past: analyzing the evolution of JavaScript code in the wildabstractJavaScript is one of the web's key building blocks. It is used by the majority of web sites and it is supported by all modern browsers. We present the first large-scale study of client-side JavaScript code over time. Specifically, we have collected and analyzed a dataset containing daily snapshots of JavaScript code coming from Alexa's Top 10000 web sites (~7.5 GB per day) for nine consecutive months, to study different temporal aspects of web client code. We found that scripts change often; typically every few days, indicating a rapid pace in web applications development. We also found that the lifetime of web sites themselves, measured as the time between JavaScript changes, is also short, in the same time scale. We then performed a qualitative analysis to investigate the nature of the changes that take place. We found that apart from standard changes such as the introduction of new functions, many changes are related to online configuration management. In addition, we examined JavaScript code reuse over time and especially the widespread reliance on third-party libraries. Furthermore, we observed how quality issues evolve by employing established static analysis tools to identify potential software bugs, whose evolution we tracked over time. Our results show that quality issues seem to persist over time, while vulnerable libraries tend to decrease. Dimitris Mitropoulos, Panagiotis Louridas, Vitalis Salis, Diomidis Spinellis |
MSR | 2 |
| 2019 | Live interactive queries to a software application's memory profileabstractMemory operations are critical to an application's reliability and performance. To reason about their correctness and track opportunities for optimisations, sophisticated instrumentation frameworks, such as Valgrind and Pin, have been developed. Both provide only limited facilities for analysing the collected data. This work presents a Valgrind's extension for examining a software applications’ dynamic memory profile through live interactive analysis with SQL. The Pico COllections Query Library (pico ql) module maps Valgrind's data structures that contain the instrumented application's memory operations metadata to a relational interface. Queries are type‐safe and the module imposes only a trivial overhead when idle. The authors evaluate the proposed approach on ten applications and through a qualitative study. They find 900 kb of undefined bytes in bzip2 that account for 12% of its total memory use and a performance‐critical code execution path in the Unix commands sort and uniq . The referenced functions are part of glibc and have been independently modified to boost the library's performance. The qualitative study has users rate the usefulness, usability, effort, correctness, and expressiveness of pico ql queries compared to Python scripts. The findings indicate that querying with pico ql incurs lower user effort. Marios Fragkoulis, Diomidis Spinellis, Panagiotis Louridas |
IET Softw. | 3 |
| 2019 | Defending Against Web Application Attacks: Approaches, Challenges and ImplicationsabstractSome of the most dangerous web attacks, such as Cross-Site Scripting and sql injection, exploit vulnerabilities in web applications that may accept and process data of uncertain origin without proper validation or filtering, allowing the injection and execution of dynamic or domain-specific language code. These attacks have been constantly topping the lists of various security bulletin providers despite the numerous countermeasures that have been proposed over the past 15 years. In this paper, we provide an analysis on various defense mechanisms against web code injection attacks. We propose a model that highlights the key weaknesses enabling these attacks, and that provides a common perspective for studying the available defenses. We then categorize and analyze a set of 41 previously proposed defenses based on their accuracy, performance, deployment, security, and availability characteristics. Detection accuracy is of particular importance, as our findings show that many defense mechanisms have been tested in a poor manner. In addition, we observe that some mechanisms can be bypassed by attackers with knowledge of how the mechanisms work. Finally, we discuss the results of our analysis, with emphasis on factors that may hinder the widespread adoption of defenses in practice. Dimitris Mitropoulos, Panagiotis Louridas, Michalis Polychronakis, Angelos D. Keromytis |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2018 | Enabling Global Big Data Computations
Damianos Chatziantoniou, Panagiotis Louridas |
DOLAP | 2 |
| 2018 | What are your programming language's energy-delay implications?abstractMotivation: Even though many studies examine the energy efficiency of hardware and embedded systems, those that investigate the energy consumption of software applications are still limited, and mostly focused on mobile applications. As modern applications become even more complex and heterogeneous a need arises for methods that can accurately assess their energy consumption. Stefanos Georgiou, Maria Kechagia, Panagiotis Louridas, Diomidis Spinellis |
MSR | 3 |
| 2018 | The exception handling riddle: An empirical study on the Android API
Maria Kechagia, Marios Fragkoulis, Panagiotis Louridas, Diomidis Spinellis |
J. Syst. Softw. | 3 |
| 2016 | The evolution of C programming practices: a study of the Unix operating system 1973-2015abstractTracking long-term progress in engineering and applied science allows us to take stock of things we have achieved, appreciate the factors that led to them, and set realistic goals for where we want to go. We formulate seven hypotheses associated with the long term evolution of C programming in the Unix operating system, and examine them by extracting, aggregating, and synthesising metrics from 66 snapshots obtained from a synthetic software configuration management repository covering a period of four decades. We found that over the years developers of the Unix operating system appear to have evolved their coding style in tandem with advancements in hardware technology, promoted modularity to tame rising complexity, adopted valuable new language features, allowed compilers to allocate registers on their behalf, and reached broad agreement regarding code formatting. The progress we have observed appears to be slowing or even reversing prompting the need for new sources of innovation to be discovered and followed. Diomidis Spinellis, Panagiotis Louridas, Maria Kechagia |
ICSE | 2 |
| 2015 | An Exploratory Study on the Evolution of C Programming in the Unix Operating SystemabstractContext: Numerous factors drive long term progress in programming practices. Goal: We study the evolution of C programming in the Unix operating system. Method: We extract, aggregate, and synthesize metrics from 66 snapshots obtained from an artificial software configuration management repository tracking the evolution of the Unix operating system over four decades. Results: C language programming practices appear to evolve over long term periods; our study identified some continuous trends with highly significant coefficients of determination. Many trends point toward increasing code quality through adherence to numerous programming guidelines, while some others indicate adoption that has reached maturity. In the area of commenting progress appears to have stalled. Conclusions: Studying the long term evolution of programming practices identifies areas where progress has been achieved along an expected path, as well as cases where there is room for improvement. Diomidis Spinellis, Panagiotis Louridas, Maria Kechagia |
ESEM | 2 |
| 2015 | Generating the Blueprints of the Java EcosystemabstractExamining a large number of software artifacts can provide the research community with data regarding quality and design. We present a dataset obtained by statically analyzing 22730 jar files taken from the Maven central archive, which is the de-facto application library repository for the Java ecosystem. For our analysis we used three popular static analysis tools that calculate metrics regarding object-oriented design, program size, and package design. The dataset contains the metrics results that every tool reports for every selected jar of the ecosystem. Our dataset can be used to produce interesting research results, such as measure the domain-specific language usage. Vassilios Karakoidas, Dimitris Mitropoulos, Panagiotis Louridas, Georgios Gousios, Diomidis Spinellis |
MSR | 3 |
| 2015 | A type-safe embedding of SQL into Java using the extensible compiler framework J%
Vassilios Karakoidas, Dimitris Mitropoulos, Panagiotis Louridas, Diomidis Spinellis |
Comput. Lang. Syst. Struct. | 3 |
| 2014 | Relational access to Unix kernel data structuresabstractState of the art kernel diagnostic tools like DTrace and Systemtap provide a procedural interface for expressing analysis tasks. We argue that a relational interface to kernel data structures can offer complementary benefits for kernel diagnostics. Marios Fragkoulis, Diomidis Spinellis, Panagiotis Louridas, Angelos Bilas |
EuroSys | 3 |
| 2014 | The bug catalog of the maven ecosystemabstractExamining software ecosystems can provide the research community with data regarding artifacts, processes, and communities. We present a dataset obtained from the Maven central repository ecosystem (approximately 265GB of data) by statically analyzing the repository to detect potential software bugs. For our analysis we used FindBugs, a tool that examines Java bytecode to detect numerous types of bugs. The dataset contains the metrics results that FindBugs reports for every project version (a JAR) included in the ecosystem. For every version we also stored specific metadata such as the JAR's size, its dependencies and others. Our dataset can be used to produce interesting research results, as we show in specific examples. Dimitris Mitropoulos, Vassilios Karakoidas, Panagiotis Louridas, Georgios Gousios, Diomidis Spinellis |
MSR | 3 |
| 2011 | Countering Code Injection Attacks: A Unified ApproachabstractPurpose The purpose of this paper is to propose a generic approach that prevents a specific class of code injection attacks (CIAs) in a novel way. Design/methodology/approach To defend against CIAs this approach involves detecting attacks by using location‐specific signatures to validate code statements. The signatures are unique identifiers that represent specific characteristics of a statement's execution. The key property that differentiates the scheme presented in this paper is that these characteristics do not depend entirely on the code statement, but also take into account elements from its execution context. Findings The approach was applied successfully to defend against attacks targeting structured query language (SQL), XML Path Language and JavaScript with positive results. Originality/value Despite many countermeasures that have been proposed the number of CIAs has been increasing. Malicious users seem to find new ways to introduce compromised embedded executable code to applications by using a variety of languages and techniques. Hence, a generic approach that defends against such attacks would be a useful countermeasure. This approach can defend attacks that involve both domain‐specific languages (e.g. SQL) and general purpose languages (e.g. JavaScript) and can be used both against client‐side and server‐side attacks. Dimitris Mitropoulos, Vassilios Karakoidas, Panagiotis Louridas, Diomidis Spinellis |
Inf. Manag. Comput. Secur. | 3 |
| 2010 | A market-based approach to managing the risk of peer-to-peer transactions
Stephanos Androutsellis-Theotokis, Diomidis Spinellis, Panagiotis Louridas, Kostas Stroggylos |
Comput. Networks | 3 |
| 2008 | Static Program Analysis for Java Card Applets
Vasilios Almaliotis, Alexandros Loizidis, Panagiotis Katsaros, Panagiotis Louridas, Diomidis Spinellis |
CARDIS | 4 |
| 2008 | Grid-Enabled Instrument Representation and ReservationabstractThe integration of instruments with the grid promotes scientific and business collaboration by allowing shared access to rare and expensive instrumentation, regardless of its physical location. Instruments on the grid are virtualized resources similar to storage and computing devices, and can be part of composite services or workflows concurrently invoked by multiple users in such service-oriented environments. The requirement for real-time, interactive control of instruments poses a need for exclusive or limited access by concurrent clients. Unfortunately, instrument modeling and standardization of instrument reservation and management are still in a very early stage. According to a number of use case requirements, which we review in this work, we propose a standards-compliant model for the representation of instruments on the grid, and we describe a set of primitives for their reservation. The definition of these constructs can be used for the normative description of such reservation capabilities and the future specification of a standardized protocol. Constantinos Kotsokalis, Tiziana Ferrari, Panagiotis Louridas, Elisabetta Ronchieri, Panayiotis Tsanakas |
eScience | 3 |
| 2008 | Power laws in softwareabstractA single statistical framework, comprising power law distributions and scale-free networks, seems to fit a wide variety of phenomena. There is evidence that power laws appear in software at the class and function level. We show that distributions with long, fat tails in software are much more pervasive than previously established, appearing at various levels of abstraction, in diverse systems and languages. The implications of this phenomenon cover various aspects of software engineering research and practice. Panagiotis Louridas, Diomidis Spinellis, Vasileios Vlachos |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2007 | A framework for the static verification of api calls
Diomidis Spinellis, Panagiotis Louridas |
J. Syst. Softw. | 2 |
| 2000 | A generic model for reflective designabstractRapid technological change has had an impact on the nature of software. This has led to new exigencies and to demands for software engineering paradigms that pay particular atttention to meeting them. We advocate that such demands can be met, at least in large parts, through the adoption of software engineering processes that are founded on a reflective stance. To this end, we turn our attention to the field of Design Rationale. We analyze and characterize Design Rationale approaches and show that despite surface differences between different approaches, they all tend to be variants of a relatively small set of static and dynamic affinities. We use the synthesis of static and dynamic affinities to develop a generic model for reflective design. The model is nonprescriptive and affects minimally the design process. It is context-independent and is intended to be used as a facilitator in participative design, supporting group communication and deliberation. The potential utility of the model is demonstrated through two examples, one from the world of business design and the other from programming language design Panagiotis Louridas, Peri Loucopoulos |
ACM Trans. Softw. Eng. Methodol. | 1 |