VLDB 2026 Research / reviewers in the wild / expert
Daniele Apiletti
dblp:11/5703
· DBLP profile ↗
21ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-0538-9775ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Computer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enforcing domain constraints through Lagrangian primal-dual learningabstractAbstract While neural networks have demonstrated remarkable predictive capabilities in various scenarios, they typically struggle to learn to avoid regions of the output space that are considered off-limits due to known domain-specific constraints. This paper presents a novel primal-dual learning approach inspired by augmented Lagrangian methods to address a priori output constraints for neural network predictions. Our solution encodes domain constraints for the output space by using a static Implicit Neural Representation and penalising the violation of these constraints at the loss level. This choice allows full flexibility in incorporating the constraints. We conduct extensive evaluations on several 2D and 3D synthetic datasets with different constraint topologies and two real-world datasets to evaluate the effectiveness of the proposed method. Our approach consistently outperforms methods without constraints in all experiments, yielding higher accuracy and lower constraint violations. Finally, our method shows superior performance improvements in scenarios with limited data availability, opening up potential benefits for various applications. These include geo-localisation tasks, where accurate positioning is crucial, and physical problems with theoretically-imposed constraints. Simone Monaco, Flavio Giobergia, Alkis Koudounas, Daniele Apiletti |
Neural Comput. Appl. | 4 |
| 2025 | A comparative study of neural ordinary differential equations and neural operators for modeling temporal dynamicsabstractAbstract Capturing the dynamics of relational systems is a key challenge in the natural sciences, with applications ranging from simulating molecular interactions to analyzing particle mechanics. Machine learning approaches have made significant progress in this area by using graph neural networks to learn and visualize spatial interactions effectively. Neural ordinary differential equations (Neural ODEs) and neural operators (NO) represent two distinct paradigms. However, a clear comparative understanding of when to prefer one over the other is still lacking. To address this gap, we present the first systematic comparison between two representative architectures: EGNO (Equivariant Graph Neural Operator) and SEGNO (Second-order Equivariant Graph Neural Ordinary Differential Equation). Through a series of experiments, we investigate their strengths and limitations in various simulation scenarios in the multi-step trajectory prediction tasks. Specifically, we employ rollout strategies and different input/output configurations, including multiple and irregularly sampled time steps. Our findings highlight a key trade-off between precision and stability that is central to model selection. SEGNO demonstrates superior robustness and stability over long prediction horizons, making it well-suited for tasks requiring reliable long-term forecasting. Conversely, EGNO offers higher precision during early stages of the trajectory and better leverages diverse training configurations, thanks to its discretization-invariant design. In summary, Neural Operators (EGNO) are preferable when short-term accuracy and data efficiency are critical, while Neural ODEs (SEGNO) are advantageous for scenarios demanding stable long-term predictions. This work not only clarifies the practical advantages of each approach but also lays the groundwork for informed model selection and future hybrid strategies in dynamical system modeling. Matteo Celia, Simone Monaco, Daniele Apiletti |
Neural Comput. Appl. | 3 |
| 2025 | Unsupervised Concept Drift Detection From Deep Learning Representations in Real-TimeabstractConcept drift is the phenomenon in which the underlying data distributions and statistical properties of a target domain change over time, leading to a degradation in model performance. Consequently, production models require continuous drift detection monitoring. Most drift detection methods to date are supervised, relying on ground-truth labels. However, they are inapplicable in many real-world scenarios, as true labels are often unavailable. Although recent efforts have proposed unsupervised drift detectors, many lack the accuracy required for reliable detection or are too computationally intensive for real-time use in high-dimensional, large-scale production environments. Moreover, they often fail to characterize or explain drift effectively. To address these limitations, we proposeDRIFTLENS, an unsupervised framework for real-time concept drift detection and characterization. Designed for deep learning classifiers handling unstructured data,DRIFTLENSleverages distribution distances in deep learning representations to enable efficient and accurate detection. Additionally, it characterizes drift by analyzing and explaining its impact on each label. Our evaluation across classifiers and data-types demonstrates thatDRIFTLENS(i) outperforms previous methods in detecting drift in 15/17 use cases; (ii) runs at least 5 times faster; (iii) produces drift curves that align closely with actual drift (correlation$\geq 0.85$); (iv) effectively identifies representative drift samples as explanations. Salvatore Greco, Bartolomeo Vacchetti, Daniele Apiletti, Tania Cerquitelli |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | DriftLens: A Concept Drift Detection Tool
Salvatore Greco, Bartolomeo Vacchetti, Daniele Apiletti, Tania Cerquitelli |
EDBT | 3 |
| 2024 | Explaining deep convolutional models by measuring the influence of interpretable features in image classificationabstractAbstract The accuracy and flexibility of Deep Convolutional Neural Networks (DCNNs) have been highly validated over the past years. However, their intrinsic opaqueness is still affecting their reliability and limiting their application in critical production systems, where the black-box behavior is difficult to be accepted. This work proposes EBAnO, an innovative explanation framework able to analyze the decision-making process of DCNNs in image classification by providing prediction-local and class-based model-wise explanations through the unsupervised mining of knowledge contained in multiple convolutional layers. EBAnO provides detailed visual and numerical explanations thanks to two specific indexes that measure the features’ influence and their influence precision in the decision-making process. The framework has been experimentally evaluated, both quantitatively and qualitatively, by (i) analyzing its explanations with four state-of-the-art DCNN architectures, (ii) comparing its results with three state-of-the-art explanation strategies and (iii) assessing its effectiveness and easiness of understanding through human judgment, by means of an online survey. EBAnO has been released as open-source code and it is freely available online. Francesco Ventura, Salvatore Greco, Daniele Apiletti, Tania Cerquitelli |
Data Min. Knowl. Discov. | 3 |
| 2024 | Hermes, a low-latency transactional storage for binary data streams from remote devices
Gabriele Scaffidi Militone, Daniele Apiletti, Giovanni Malnati |
Data Knowl. Eng. | 2 |
| 2023 | Combining fault-tolerant persistence and low-latency streaming access to binary data for AI modelsabstractIn many AI-enabled scenarios, such as video surveillance systems, besides requiring the data to be stored safely, human operators and AI models must also be able to access audio and video streams continuously while media files are still being collected. However, system throughput and latency are often limited by the use of transactionality to guarantee data persistence. This paper presents a solution providing both high ingestion rates with transactional data persistence and low-latency access to the stream during collection in near real-time. This enables the AI algorithms to be immediately applied as soon as the data is received. The binary data sources fit well with the audio and video capture of surveillance or similar systems, but the proposed solution can be extended through well-defined general interfaces. The scalability of the proposed approach is based on the microservice architecture. Using Apache Kafka and MongoDB replica sets, preliminary results show that the proposed solution provides up to 6 times larger throughput and 4.5 times lower latency than current standard multi-document transactions. Gabriele Scaffidi Militone, Daniele Apiletti, Giovanni Malnati |
IEEE Big Data | 2 |
| 2022 | A Dataset for Burned Area Delineation and Severity Estimation from Satellite ImageryabstractThe ability to correctly identify areas damaged by forest wildfires is essential to plan and monitor the restoration process and estimate the environmental damages after such catastrophic events. The wide availability of satellite data, combined with the recent development of machine learning and deep learning methodologies applied to the computer vision field, makes it extremely interesting to apply the aforementioned techniques to the field of automatic burned area detection. One of the main issues in such a context is the limited amount of labeled data, especially in the context of semantic segmentation. In this paper, we introduce a publicly available dataset for the burned area detection problem for semantic segmentation. The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel. Data were collected from the Sentinel-2 L2A satellite mission and the target labels were generated from the Copernicus Emergency Management Service (EMS) annotations, with five different severity levels, ranging from undamaged to completely destroyed. Finally, we report the benchmark values obtained by applying a Convolutional Neural Network on the proposed dataset to address the burned area identification problem. Luca Colomba, Alessandro Farasin, Simone Monaco, Salvatore Greco, Paolo Garza, Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
CIKM | 6 |
| 2022 | Trusting deep learning natural-language models via local and global explanationsabstractAbstract Despite the high accuracy offered by state-of-the-art deep natural-language models (e.g., LSTM, BERT), their application in real-life settings is still widely limited, as they behave like a black-box to the end-user. Hence, explainability is rapidly becoming a fundamental requirement of future-generation data-driven systems based on deep-learning approaches. Several attempts to fulfill the existing gap between accuracy and interpretability have been made. However, robust and specialized eXplainable Artificial Intelligence solutions, tailored to deep natural-language models, are still missing. We propose a new framework, named T-EBAnO, which provides innovative prediction-local and class-based model-global explanation strategies tailored to deep learning natural-language models. Given a deep NLP model and the textual input data, T-EBAnO provides an objective, human-readable, domain-specific assessment of the reasons behind the automatic decision-making process. Specifically, the framework extracts sets of interpretable features mining the inner knowledge of the model. Then, it quantifies the influence of each feature during the prediction process by exploiting the normalized Perturbation Influence Relation index at the local level and the novel Global Absolute Influence and Global Relative Influence indexes at the global level. The effectiveness and the quality of the local and global explanations obtained with T-EBAnO are proved on an extensive set of experiments addressing different tasks, such as a sentiment-analysis task performed by a fine-tuned BERT model and a toxic-comment classification task performed by an LSTM model. The quality of the explanations proposed by T-EBAnO, and, specifically, the correlation between the influence index and human judgment, has been evaluated by humans in a survey with more than 4000 judgments. To prove the generality of T-EBAnO and its model/task-independent methodology, experiments with other models (ALBERT, ULMFit) on popular public datasets (Ag News and Cola) are also discussed in detail. Francesco Ventura, Salvatore Greco, Daniele Apiletti, Tania Cerquitelli |
Knowl. Inf. Syst. | 3 |
| 2021 | Cyst segmentation on kidney tubules by means of U-Net deep-learning modelsabstractAutosomal dominant polycystic kidney disease (ADPKD) is one of the most widespread genetic disorders affecting the kidney. Nevertheless, there is still no cure for ADPKD. Domain experts test the effectiveness of different treatments by investigating how they can reduce the number and dimension of cysts on kidney tissues. Image processing of the microscope acquisitions is then an expensive but necessary operation currently performed by operators to determine and compare cyst size and quantity. In this work, we propose a deep learning algorithm for fast and accurate cysts detection in sequential 2-D images. Experiments on 507 RGB immunofluorescence images of 8 kidney tubules show that the proposed U-Net-based deep-learning solution can automatically segment images with increasing performance at larger cyst dimensions (Pr > 0.8, Re > 0.75 for cysts larger than 32 µm2). Such a reliable method performing an accurate cyst segmentation can be a valid support for researchers in optimising the effort to find new effective treatments for ADPKD. Simone Monaco, Nicole Bussola, Sara Buttò, Diego Sona, Daniele Apiletti, Giuseppe Jurman, Elisa Viola, Marco Chierici, Christodoulos Xinaris, Vincenzo Viola |
IEEE BigData | 5 |
| 2021 | Enhancing manufacturing intelligence through an unsupervised data-driven methodology for cyclic industrial processes
Tania Cerquitelli, Francesco Ventura, Daniele Apiletti, Elena Baralis, Enrico Macii, Massimo Poncino |
Expert Syst. Appl. | 3 |
| 2020 | Improving Wildfire Severity Classification of Deep Learning U-Nets from Satellite ImagesabstractUncontrolled wildfires are dangerous events capable of harming people safety. To contrast their increasing impact in recent years, a key task is an accurate detection of the affected areas and their damage assessment from satellite images. Current state-of-the-art solutions address such problem through a double convolutional neural network able to automatically detect wildfires in satellite acquisitions and associate a damage index from a defined scale. However, such deep-learning model performance is strongly dependent on many factors. In this work, we specifically focus on a key parameter, i.e., the loss function, exploited in the underlying neural networks. Besides the state-of-the-art solutions based on the Dice-MSE, among the many loss functions proposed in literature, we focus on the Binary Cross-Entropy (BCE) and the Intersection over Union (IoU), as two representatives of the distribution-based and region-based categories, respectively. Experiments show that the BCE loss function coupled with a double-step U-Net architecture provides better results than current state-of-the-art solutions on a public labeled dataset of European wildfires. Simone Monaco, Andrea Pasini, Daniele Apiletti, Luca Colomba, Paolo Garza, Elena Baralis |
IEEE BigData | 3 |
| 2020 | DSLE: A Smart Platform for Designing Data Science CompetitionsabstractDuring the last years an increasing number of university-level and post-graduation courses on Data Science have been offered. Practices and assessments need specific learning environments where learners could play with data samples and run machine learning and data mining algorithms. To foster learner engagement many closed-and open-source platforms support the design of data science competitions. However, they show limitations on the ability to handle private data, customize the analytics and evaluation processes, and visualize learners' activities and outcomes. This paper presents Data Science Lab Environment (DSLE, in short), a new open-source platform to design and monitor data science competitions. DSLE offers a easily configurable interface to share training and test data, design group works or individual sessions, evaluate the competition runs according to customizable metrics, manage public and private leaderboards, monitor participants' activities and their progress over time. The paper describes also a real experience of usage of DSLE in the context of a 1st-year M.Sc. course, which has involved around 160 students. Giuseppe Attanasio, Flavio Giobergia, Andrea Pasini, Francesco Ventura, Elena Baralis, Luca Cagliero, Paolo Garza, Daniele Apiletti, Tania Cerquitelli, Silvia Chiusano |
COMPSAC | 8 |
| 2016 | SaFe-NeC: A scalable and flexible system for network data characterizationabstractNowadays, large volumes of data and measurements are being continuously generated by computer and telecommunication networks, but such volumes make it difficult to extract meaningful knowledge from them. This paper presents SaFe-NeC, an innovative methodology for analyzing network traffic by exploiting data mining techniques, i.e. clustering and classification algorithms, focusing on self-learning capabilities of state-of-the-art scalable approaches. Self-learning algorithms, coupled with self-assessment indicators and domain-driven semantics enriching data mining results, are able to build a model of the data with minimal user intervention and highlight possibly meaningful interpretations to domain experts. Furthermore, a self-evolving model evaluation phase is included to continuously track the quality degradation of the model itself, whose rebuilding is triggered as soon as quality indicators fall below a threshold of tolerance. The proposed methodology can exploit the computational advantages of distributed computing frameworks, as the current implementation runs on Apache Spark. Preliminary experimental results on a real traffic dataset show the full potential of the proposed methodology to characterize network traffic data. Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Luca Venturini |
NOMS | 1 |
| 2016 | SeLINA: A Self-Learning Insightful Network AnalyzerabstractUnderstanding the behavior of a network from a large scale traffic dataset is a challenging problem. Big data frameworks offer scalable algorithms to extract information from raw data, but often require a sophisticated fine-tuning and a detailed knowledge of machine learning algorithms. To streamline this process, we propose self-learning insightful network analyzer (SeLINA), a generic, self-tuning, simple tool to extract knowledge from network traffic measurements. SeLINA includes different data analytics techniques providing self-learning capabilities to state-of-the-art scalable approaches, jointly with parameter auto-selection to off-load the network expert from parameter tuning. We combine both unsupervised and supervised approaches to mine data with a scalable approach. SeLINA embeds mechanisms to check if the new data fits the model, to detect possible changes in the traffic, and to, possibly automatically, trigger model rebuilding. The result is a system that offers human-readable models of the data with minimal user intervention, supporting domain experts in extracting actionable knowledge and highlighting possibly meaningful interpretations. SeLINA's current implementation runs on Apache Spark. We tested it on large collections of real-world passive network measurements from a nationwide ISP, investigating YouTube, and P2P traffic. The experimental results confirmed the ability of SeLINA to provide insights and detect changes in the data that suggest further analyses. Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Danilo Giordano, Marco Mellia, Luca Venturini |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2013 | Desidoo, a Big-Data Application to Join the Online and Real-World Marketplaces
Daniele Apiletti, Fabio Forno |
ADBIS (2) | 1 |
| 2012 | MaskedPainter: Feature selection for microarray data analysisabstractSelecting a small number of discriminative genes from thousands is a fundamental task in microarray data analysis. An effective feature selection allows biologists to investigate only a subset of genes instead of the entire set, thus avoiding insignificant, noisy, and redundant features. This paper presents the MaskedPainter feature selection method for gene expression data. The proposed method measures the ability of each gene to classify samples belonging to different classes and ranks genes by computing an overlap score. A density based technique is exploited to smooth the effects of outliers in the overlap score computation. Analogously to other approaches, the number of selected genes can be set by the user. However, our algorithm may automatically detect the minimum set of genes that yields the best classification coverage of training set samples. The effectiveness of our approach has been demonstrated through an empirical study on public microarray datasets with different characteristics. Experimental results show that the proposed approach yields a higher classification accuracy with respect to widely used feature selection techniques. Daniele Apiletti, Elena Baralis, Giulia Bruno, Alessandro Fiori |
Intell. Data Anal. | 1 |
| 2011 | Energy-saving models for wireless sensor networks
Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
Knowl. Inf. Syst. | 1 |
| 2009 | Characterizing network traffic by means of the NetMine framework
Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Vincenzo D'Elia |
Comput. Networks | 1 |
| 2009 | Real-Time Analysis of Physiological Data to Support Medical ApplicationsabstractThis paper presents a flexible framework that performs real-time analysis of physiological data to monitor people's health conditions in any context (e.g., during daily activities, in hospital environments). Given historical physiological data, different behavioral models tailored to specific conditions (e.g., a particular disease, a specific patient) are automatically learnt. A suitable model for the currently monitored patient is exploited in the real-time stream classification phase. The framework has been designed to perform both instantaneous evaluation and stream analysis over a sliding time window. To allow ubiquitous monitoring, real-time analysis could also be executed on mobile devices. As a case study, the framework has been validated in the intensive care scenario. Experimental validation, performed on 64 patients affected by different critical illnesses, demonstrates the effectiveness and the flexibility of the proposed framework in detecting different severity levels of monitored people's clinical situations. Daniele Apiletti, Elena Baralis, Giulia Bruno, Tania Cerquitelli |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2007 | SAPhyRA: Stream Analysis for Physiological Risk AssessmentabstractAdvances in technology allow the continuous physiological monitoring of people using noninvasive sensors. An important issue in this context is the real-time analysis of physiological signals performed on mobile devices, which requires optimized power consumption and short processing response time. The SAPhyRA framework performs real-time stream analysis for physiological risk assessment. To this aim, the framework evaluates people's health conditions by analyzing different clinical signals in a sliding time window. Given historical physiological measures, different models of patients and diseases are built. The most suitable model for the current monitored patient is exploited in the real time stream classification phase. Preliminary experiments performed on public physiological data show the effectiveness and flexibility of the proposed approach. Daniele Apiletti, Elena Baralis, Giulia Bruno, Tania Cerquitelli |
CBMS | 1 |