VLDB 2026 Research / reviewers in the wild / expert
Stefano Cirillo
dblp:231/4634
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-0201-2753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCoPE: Cross-Platform Discovery and RAG-Driven Profiling of Public Personal Data
Stefano Cirillo, Giuseppe Polese, Giandomenico Solimando, Nicola Zannone |
DBSec | 1 |
| 2026 | Phishing Detection in Web Domains: new intelligent tool leveraging the effectiveness of emerging Generative modelsabstractThe rapid growth of online services has heightened concerns about user protection from cyber threats, particularly phishing, which poses significant risks to cyber-social security. To this end, we propose a novel tool for phishing detection called U-Proof. Our tool uses both state-of-the-art LLMs and traditional ML models to detect phishing websites. In particular, we evaluate the phishing detection capabilities of different LLMs and compare them with several ML models to analyze the impact of different model architectures on the identification of phishing websites. For a comprehensive experimental evaluation, we use a combination of public and custom datasets. These include active phishing websites from September 2024, as well as URLs from banks and postal services. Furthermore, the tool includes explanations to enhance user awareness of phishing tactics, supporting broader educational efforts to reduce risks. Carmine Ambrosino, Maurizio Atzori, Stefano Cirillo, Domenico Desiato, Simona Ettari, Giuseppe Polese, Giandomenico Solimando |
WSDM | 3 |
| 2026 | Towards structure-aware AI: modeling and analyzing directed balanced cliques in signed graphs
Abdallah Tubaishat, Zahid Halim, Stefano Cirillo, Fawaz Khaled Alarfaj, Imad Rida, Sajid Anwar 0001 |
Inf. Sci. | 4 |
| 2025 | CADHE: Privacy-Preserving Medical Image Analysis Through Homomorphic Encrypted Convolutional Networks
Stefano Cirillo, Vincenzo Deufemia, Luigi Di Biasi, Giuseppe Polese, Giandomenico Solimando, Genny Tortora |
IEEE Big Data | 1 |
| 2025 | An RFD-based approach for concept drift detection in Machine Learning Systems
Loredana Caruccio, Stefano Cirillo, Giuseppe Polese, Roberto Stanzione |
EDBT | 2 |
| 2025 | Identifying fake reviews for refund purposes: Evaluating the effectiveness of a transfer-learning model against emerging Large Language ModelsabstractRecently, dishonest sellers are using social platforms to advertise products that can be purchased for free through a refund mechanism, which is based on the writing of five-star fake reviews. The aim is to increase product visibility by influencing their ranking compared to similar products. This mechanism is leading to a significant distortion of e-commerce platforms, eroding trust among customers and sellers. In this paper, we address the problem of identifying fake reviews, aiming to provide an approach for mitigating fraudulent practices that compromise the integrity and transparency of e-commerce platforms. We propose a supervised model tailored for identifying fake reviews for refund purposes and compare its performance with some of the most recent generative models. Since, to the best of our knowledge, no datasets exist in the literature suitable for fake review identification in the process of Purchasing, Requesting reviews, and Refunding a product, we first proposed a new dataset of fake and genuine reviews from Amazon, collected with the help of a domain expert. Then, we defined five other new datasets containing reviews automatically generated by language models. To interact with these models, we designed new prompt approaches specifically tailored to our goal, which exploit the iterative refinement behind these models for improving classification results. Experimental results demonstrated the effectiveness of the supervised model in detecting both types of fake reviews, outperforming state-of-the-art models with improvements ranging from 0.23 to 0.70 in terms of accuracy, precision, and recall. • The requesting reviews and refunding phenomenon (PRP process) is analyzed. • PRP reviews are collected from web pages and then validated by a domain expert. • A BERT-based model is successfully applied for properly identifying PRP fake reviews. • The proposed model is compared with Large Language Models and state-of-the-art models. • New datasets of real and automatically generated PRR reviews have been proposed. Loredana Caruccio, Gaetano Cimino, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese, Giandomenico Solimando |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Exploring the ability of emerging large language models to detect cyberbullying in social posts through new prompt-based classification approaches
Stefano Cirillo, Domenico Desiato, Giuseppe Polese, Giandomenico Solimando, Vijayan Sugumaran, Shanmugam Sundaramurthy |
Inf. Process. Manag. | 1 |
| 2025 | Non-blocking functional dependency discovery from data streams
Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese |
Inf. Sci. | 2 |
| 2025 | A hybrid approach combining images and questionnaires for early detection and severity assessment of Autism Spectrum DisorderabstractIn this research, we propose a novel integrated system for the early diagnosis and cognitive enhancement of infants with Autism Spectrum Disorder (ASD). The system combines two core modules: the Behavioral Analytic Module and the Cognitive Skill Enhancement Module. The Behavioral Analytic Module includes a Questionnaire Analysis Sub-module, which utilizes Random Forest classifiers to analyze questionnaire data, and an Image Analysis Sub-module, which employs a fine-tuned VGG16 Convolutional Neural Network to process facial images. These sub-modules independently assess ASD likelihood and combine their outputs to generate a comprehensive diagnosis using a weighted averaging technique. The Cognitive Skill Enhancement Module integrates interactive games and web-based animations designed to improve cognitive abilities and daily living skills in toddlers with ASD. Additionally, a reward system is incorporated to reinforcement learning outcomes, adaptively calculating rewards based on the infants’ progress. The proposed system aims to provide a holistic approach to ASD diagnosis and intervention, offering an effective tool for early detection and tailored cognitive development. The system’s efficacy is demonstrated through comparative analysis, showing a 93% improvement in diagnostic accuracy and a 92% enhancement in cognitive skill development among toddlers with ASD. S. C. Rajkumar, Stefano Cirillo, Yuvasini D, Luisa Solimando |
Image Vis. Comput. | 2 |
| 2024 | RYAN: A tool for explaining and visually analyzing the evolution of Relaxed Functional DependenciesabstractThe importance of exploiting profiling metadata, such as Relaxed Functional Dependencies (RFDs), to support advanced data processing tasks, continues to grow also due to the availability of algorithms capable of automatically extracting them from data. Nevertheless, in order to use this type of metadata in real-life contexts, it is also necessary to ensure their correct interpretation of their meaningfulness and their possible evolution over time. To this end, in this paper, we present a new tool that allows visual analysis and explainability of how discovery results evolve according to changes in the data. More specifically, it provides a comprehensive overview of the impact that data changes, by possibly analyzing in-depth affected dependencies and understanding motivations underlying their evolution through a textual explanation. The effectiveness of the proposed tool has been evaluated by conducting a user study, which highlighted RYAN’s capability to yield an intuitive visualization of RFD discovery results and to provide a clear explanation of the reasons that led to the evolution of RFDs. Loredana Caruccio, Stefano Cirillo, Gianpaolo Iuliano, Giuseppe Polese, Roberto Stanzione |
IEEE Big Data | 2 |
| 2024 | Privacy-Preserving Artificial Intelligence on Edge Devices: A Homomorphic Encryption ApproachabstractRecent advancements in privacy-preserving artificial intelligence (AI) have paved the way for enhanced privacy in computational processes. A standing challenge, however, is the robust privacy preservation in AI algorithms, especially when integrated into edge devices and Internet-of-Thing (IoT) infrastructures. Most prevailing solutions have adopted traditional encryption methods which, though secure, often introduce significant overhead and potential dips in accuracy. In this study, we put forth an innovative approach, utilizing the CKKS encryption scheme, aiming to harmoniously balance computational efficiency with stringent data privacy. By harnessing the capabilities of Full Homomorphic Encryption (FHE) under the CKKS scheme, we ensure the preservation of privacy, successfully curbing the inherent noise traditionally linked with accuracy reductions in similar encryption-oriented solutions. Through comprehensive experiments, our approach showcased its potential as a strong contender for privacy preservation, demonstrating commendable performance across all tests, affirming that FHE is indeed viable for devices with constrained computational power and energy resources. Muhammad Jahanzeb Khan, Bo Fang 0002, Gaetano Cimino, Stefano Cirillo, Lei Yang 0001, Dongfang Zhao 0001 |
ICWS | 4 |
| 2024 | Can ChatGPT provide intelligent diagnoses? A comparative study between predictive models and ChatGPT to define a new medical diagnostic botabstractIntelligent diagnosis processes rely on Artificial Intelligence (AI) techniques to provide possible diagnoses by analyzing patient data and medical information. To make accurate and quick diagnoses, it is possible to use AI tools to efficiently analyze huge amounts of data and find patterns that a clinician might miss. In recent years, new large language models (LLMs), such as ChatGPT and Google BARD, have shown remarkable capabilities in several domains, including intelligent diagnostics. This research aims to compare the performances of ChatGPT and traditional machine learning models for making diagnoses of low- and medium- risk diseases only based on their symptoms. On the basis of our study, we defined four research questions: RQ1) What are the benefits and limitations of using ChatGPT in intelligent diagnosis? RQ2) How do traditional machine learning approaches compare to ChatGPT for intelligent diagnosis? RQ3) How does ChatGPT compare with other LLMs and domain-specific natural language processing models in the intelligent diagnosis tasks?, and RQ4) What are the implications of the predictive models and ChatGPT for healthcare, and how can they be used to support people?. To answer these RQs, we first evaluate the performances of different engines of ChatGPT, also introducing a new prompt engineering methodology specifically tailored for achieving accurate diagnostic outcomes. Moreover, we compare these results with those achieved by different predictive models trained for intelligent diagnosis tasks, i.e., Google BARD, and two domain-specific NLP models. Finally, we propose a new interactive bot available for users that relies on the best-performing models evaluated in the previous steps. The experiments have been conducted using two medical datasets for disease prediction consisting of more than 100 symptoms associated with several diagnoses. Loredana Caruccio, Stefano Cirillo, Giuseppe Polese, Giandomenico Solimando, Shanmugam Sundaramurthy, Genny Tortora |
Expert Syst. Appl. | 2 |
| 2024 | A deep learning approach to classify country and value of modern coins
Stefano Cirillo, Giandomenico Solimando, Luca Virgili |
Neural Comput. Appl. | 1 |
| 2024 | Decentralized and Incremental Discovery of Relaxed Functional Dependencies Using Bitwise SimilarityabstractOver the past decade, there have been numerous extensions to the definition of Functional Dependency (fd), culminating in the introduction of Relaxed Functional Dependency (rfd), offering more flexible constraints compared to traditionalfds. This increased flexibility makesrfds well-suited for exploring and profiling data in datasets with lower data quality. However, efficiently identifyingrfds within dynamic data sources presents a significant challenge, as it requires processing an entire dataset from scratch whenever modifications occur. To tackle this problem, incremental discovery algorithms have been defined, but they often suffer when the frequency and the size of batches of updates increase. This article presents a new algorithm, namelyD-IndiBits, relying on a new decentralized architecture to balance the workload that drives the incremental discovery process ofIndiBits, which is based on bitwise operators for computing attribute similarities. Experiments demonstrateD-IndiBits's effectiveness compared tofdandrfddiscovery algorithms on both static and dynamic real-world data. With batches of modifications of sizes 10 k and 100 k,D-IndiBitsis capable of updating the set ofrfds in a few seconds, whereas all other approaches often employ more than 3 hours. Bernardo Breve, Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | REQUIRED: A Tool to Relax Queries through Relaxed Functional Dependencies
Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese, Roberto Stanzione |
EDBT | 2 |
| 2023 | IndiBits: Incremental Discovery of Relaxed Functional Dependencies using Bitwise SimilarityabstractOne of the main challenges in data profiling is to efficiently extract metadata from dynamic information sources, by avoiding the processing of the whole dataset from scratch upon modifications. In this paper, we present IndiBits, an algorithm for discovering relaxed functional dependencies (RFDs for short), which represent data relationships relying on approximate matching paradigms. IndiBits is able to dynamically infer and update the RFDs holding on a dataset upon modification operations performed on it. It exploits a binary representation of data similarities, a new validation method, and specific search methods, to dynamically update the set of RFDs, based on previously holding RFDs and the type of modifications performed over data. Experimental results demonstrate the effectiveness of IndiBits on real-world datasets, even in comparison with FD and RFD discovery algorithms in both static and dynamic scenarios. Bernardo Breve, Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese |
ICDE | 3 |
| 2023 | Malicious Account Identification in Social Network PlatformsabstractToday, people of all ages are increasingly using Web platforms for social interaction. Consequently, many tasks are being transferred over social networks, like advertisements, political communications, and so on, yielding vast volumes of data disseminated over the network. However, this raises several concerns regarding the truthfulness of such data and the accounts generating them. Malicious users often manipulate data to gain profit. For example, malicious users often create fake accounts and fake followers to increase their popularity and attract more sponsors, followers, and so on, potentially producing several negative implications that impact the whole society. To deal with these issues, it is necessary to increase the capability to properly identify fake accounts and followers. By exploiting automatically extracted data correlations characterizing meaningful patterns of malicious accounts, in this article we propose a new feature engineering strategy to augment the social network account dataset with additional features, aiming to enhance the capability of existing machine learning strategies to discriminate fake accounts. Experimental results produced through several machine learning models on account datasets of both the Twitter and the Instagram platforms highlight the effectiveness of the proposed approach toward the automatic discrimination of fake accounts. The choice of Twitter is mainly due to its strict privacy laws, and because its the only social network platform making data of their accounts publicly available. Loredana Caruccio, Gaetano Cimino, Stefano Cirillo, Domenico Desiato, Giuseppe Polese, Genny Tortora |
ACM Trans. Internet Techn. | 3 |
| 2022 | Investigating the COVID-19 vaccine discussions on Twitter through a multilayer network-based approach
Gianluca Bonifazi, Bernardo Breve, Stefano Cirillo, Enrico Corradini, Luca Virgili |
Inf. Process. Manag. | 3 |
| 2022 | Enhancing spatial perception through sound: mapping human movements into MIDIabstractAbstract Gestural expressiveness plays a fundamental role in the interaction with people, environments, animals, things, and so on. Thus, several emerging application domains would exploit the interpretation of movements to support their critical designing processes. To this end, new forms to express the people’s perceptions could help their interpretation, like in the case of music. In this paper, we investigate the user’s perception associated with the interpretation of sounds by highlighting how sounds can be exploited for helping users in adapting to a specific environment. We present a novel algorithm for mapping human movements into MIDI music. The algorithm has been implemented in a system that integrates a module for real-time tracking of movements through a sample based synthesizer using different types of filters to modulate frequencies. The system has been evaluated through a user study, in which several users have participated in a room experience, yielding significant results about their perceptions with respect to the environment they were immersed. Bernardo Breve, Stefano Cirillo, Mariano Cuofano, Domenico Desiato |
Multim. Tools Appl. | 2 |
| 2021 | Efficient Discovery of Functional Dependencies from Incremental DatabasesabstractWith the advent of Big Data there is an increasing necessity to incrementally mine information from data originating from sensors and other dynamic sources. Thus, it is necessary to devise algorithms capable of mining useful information upon possible evolutions of databases. Among these, there are certainly data profiling info, such as functional dependencies (fd for short), which are particularly useful for data integration and for assessing the quality of data. The incremental scenario requires the definition of search strategies and validation methods able to analyze only the portion of the dataset affected by the last changes. In this paper, we propose a new validation method, which exploits regular expressions and compressed data structures to efficiently verify whether a candidate fd holds on an updated version of the dataset. Experimental results demonstrate the effectiveness of the proposed method on real-world datasets adapted for incremental scenarios, also compared with a baseline incremental fd discovery algorithm. Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia, Giuseppe Polese |
iiWAS | 2 |
| 2021 | An intelligent system for focused crawling from Big Data sources
Ida Bifulco, Stefano Cirillo, Christian Esposito 0001, Roberta Guadagni, Giuseppe Polese |
Expert Syst. Appl. | 2 |
| 2019 | CHRAVAT - Chronology Awareness Visual Analytic ToolabstractNowadays, the amount of information spread over networks is extremely large, and many sensible data are granted by legitimate owners aiming to exploit different networking services. In particular, the majority of people give their own consent for processing personal data without understanding how network providers will manage them, and if they will be shared among different network providers. In this paper, we propose a tool exploiting visualization techniques in order to make a user aware of how his/her personal data are exchanged and shared during daily web browsing activities. In particular, the proposed tool enables a user to interactively visualize the communication flows during the aforesaid browsing process, and to discover possibly hidden network providers involved in it. Moreover, the graphical interface also provides real-time summary graphs, which show the amount of information acquired from the network. Finally, we performed several users studies aiming to analyse how the tool can improve the user's perception on the privacy issues that s/he is exposed to. Results demonstrate the effectiveness of the proposed tool. Stefano Cirillo, Domenico Desiato, Bernardo Breve |
IV (1) | 1 |
| 2018 | Discovery Multiple Data Structures in Big Data through Global Optimization and Clustering MethodsabstractIn this paper, we propose an approach to Big Data visualization, based on clustering techniques, in order to find a structure of them and to facilitate their visualization. However, the main problem of clustering is that sometimes converge to a local minimum showing only one solution, so an optimization of the K-means algorithm has been proposed with the aim to escape from local minimum and to visualize different solutions of the same problem. In particular, we use the K-means algorithm with multiple random starting points, in order to find several solutions to the same problem. This algorithm considers the data of the Italian calls for tenders, extracted through a crawling technique, and optimized through the proposed approach to obtain multiple solutions. These are used to achieve a repository of products that can be easily displayed and inquired during the formulation of an offer from a bidder company willing to participate to a call for tenders. The case study results show the feasibility and validity of the proposed approach. Ida Bifulco, Stefano Cirillo |
IV | 2 |