Francesco Folino

dblp:76/5134 · DBLP profile ↗
← Back
15ranked-venue papers in the field
7as first author
2since 2021 · last 2024
0000-0002-4952-1187ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 9 (4 first)Data Mining & Knowledge Discovery · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)Business Process & Enterprise Data · 2 (1 first)
YearPublicationVenuePosition
2024 Generating the Traces You Need: A Conditional Generative Model for Process Mining Data
abstract
In recent years, trace generation has emerged as a significant challenge within the Process Mining community. Deep Learning (DL) models have demonstrated accuracy in reproducing the features of the selected processes. However, current DL generative models are limited in their ability to adapt the learned distributions to generate data samples based on specific conditions or attributes. This limitation is particularly significant because the ability to control the type of generated data can be beneficial in various contexts, enabling a focus on specific behaviours, exploration of infrequent patterns, or simulation of alternative "what-if" scenarios.In this work, we address this challenge by introducing a conditional model for process data generation based on a conditional variational autoencoder (CVAE). Conditional models offer control over the generation process by tuning input conditional variables, enabling more targeted and controlled data generation. Unlike other domains, CVAE for process mining faces specific challenges due to the multiperspective nature of the data and the need to adhere to control-flow rules while ensuring data variability. Specifically, we focus on generating process executions conditioned on control flow and temporal features of the trace, allowing us to produce traces for specific, identified sub-processes. The generated traces are then evaluated using common metrics for generative model assessment, along with additional metrics to evaluate the quality of the conditional generation.
Riccardo Graziosi, Massimiliano Ronzani, Andrei Buliga 0001, Chiara Di Francescomarino, Francesco Folino, Chiara Ghidini, Francesca Meneghello 0002, Luigi Pontieri
ICPM5
2024 Data- & compute-efficient deviance mining via active learning and fast ensembles
abstract
Abstract Detecting deviant traces in business process logs is crucial for modern organizations, given the harmful impact of deviant behaviours (e.g., attacks or faults). However, training a Deviance Prediction Model (DPM) by solely using supervised learning methods is impractical in scenarios where only few examples are labelled. To address this challenge, we propose an Active-Learning-based approach that leverages multiple DPMs and a temporal ensembling method that can train and merge them in a few training epochs. Our method needs expert supervision only for a few unlabelled traces exhibiting high prediction uncertainty. Tests on real data (of either complete or ongoing process instances) confirm the effectiveness of the proposed approach.
Francesco Folino, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
J. Intell. Inf. Syst.1
2019 Predictive monitoring of temporally-aggregated performance indicators of business processes against low-level streaming events
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Inf. Syst.2
2018 A Predictive Learning Framework for Monitoring Aggregated Performance Indicators over Business Process Events
abstract
In many application contexts, a business process' executions are subject to performance constraints expressed in an aggregated form, usually over predefined time windows, and detecting a likely violation to such a constraint in advance could help undertake corrective measures for preventing it. This paper illustrates a prediction-aware event processing framework that addresses the problem of estimating whether the process instances of a given (unfinished) window w will violate an aggregate performance constraint, based on the continuous learning and application of an ensemble of models, capable each of making and integrating two kinds of predictions: single-instance predictions concerning the ongoing process instances of w, and time-series predictions concerning the "future" process instances of w (i.e. those that have not started yet, but will start by the end of w). Notably, the framework can continuously update the ensemble, fully exploiting the raw event data produced by the process under monitoring, suitably lifted to an adequate level of abstraction. The framework has been validated against historical event data coming from real-life business processes, showing promising results in terms of both accuracy and efficiency.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
IDEAS2
2016 A Robust and Versatile Multi-View Learning Framework for the Detection of Deviant Business Process Instances
abstract
Increasing attention has been paid to the detection and analysis of “deviant” instances of a business process that are connected with some kind of “hidden” undesired behavior (e.g. frauds and faults). In particular, several recent works faced the problem of inducing a binary classification model (here named deviance detection model ) that can discriminate between deviant traces and normal ones, based on a set of historical log traces (labeled as either deviant or normal). Current solutions rely on applying standard classifier-induction methods to a feature-based representation of the given traces, where the features include sequence-based patterns extracted from the corresponding sequences of activities. However, there is no consensus on which kinds of patterns are the most suitable for such a task. On the other hand, mixing multiple pattern families together may produce a heterogenous, redundant and sparse representation of the traces that likely leads to poor deviance detection models. In this paper, we propose an ensemble-learning method for solving this problem, where multiple base classifiers are trained on different feature-based views of the log (each obtained by mapping the traces onto a distinguished collection of patterns). A stacking procedure is used to combine the discovered base models into an overall probabilistic model that associates any new trace with an estimate of the probability that it reflects a deviant process instance. This helps the analyst prioritize the inspection of the cases that are more likely to be deviant. The method also takes advantage of all nonstructural data available in the log, and employs a resampling mechanism to deal with the rarity of deviances in the training log. It has been conceived as the core of a comprehensive framework for detecting and analyzing business process deviances. The framework supports the analyst to investigate suspect deviances, and provides some feedback to the learning method for improving the accuracy of the discovered deviance detection models. Tests on several real-life datasets proved the validity of the approach, as concerns its capability to discover an accurate deviance detection model, and to effectively exploit new (originally unlabeled) traces via active learning and self-training mechanisms.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Int. J. Cooperative Inf. Syst.2
2014 Mining Predictive Process Models out of Low-level Multidimensional Logs
Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
CAiSE1
2014 An Evolutionary Multiobjective Approach for Community Discovery in Dynamic Networks
abstract
The discovery of evolving communities in dynamic networks is an important research topic that poses challenging tasks. Evolutionary clustering is a recent framework for clustering dynamic networks that introduces the concept of temporal smoothness inside the community structure detection method. Evolutionary-based clustering approaches try to maximize cluster accuracy with respect to incoming data of the current time step, and minimize clustering drift from one time step to the successive one. In order to optimize both these two competing objectives, an input parameter that controls the preference degree of a user towards either the snapshot quality or the temporal quality is needed. In this paper the detection of communities with temporal smoothness is formulated as a multiobjective problem and a method based on genetic algorithms is proposed. The main advantage of the algorithm is that it automatically provides a solution representing the best trade-off between the accuracy of the clustering obtained, and the deviation from one time step to the successive. Experiments on synthetic data sets show the very good performance of the method when compared with state-of-the-art approaches.
Francesco Folino, Clara Pizzuti
IEEE Trans. Knowl. Data Eng.1
2013 DynamicNet: an effective and efficient algorithm for supporting community evolution detection in time-evolving information networks
abstract
DynamicNet, an effective and efficient algorithm for supporting community evolution detection in time-evolving information networks is presented and experimentally evaluated in this paper. DynamicNet introduces a graph-based model-theoretic approach to represent time-evolving information networks, and to capture how they change over time. A central feature of DynamicNet is represented by the ability of supporting matching-based community evolution detection, by identifying several classes of community transitions. Experimental results clearly demonstrate the reliability and the efficiency of our proposal.
Alfredo Cuzzocrea, Francesco Folino, Clara Pizzuti
IDEAS2
2011 Mining usage scenarios in business processes: Outlier-aware discovery and run-time prediction
Francesco Folino, Gianluigi Greco, Antonella Guzzo, Luigi Pontieri
Data Knowl. Eng.1
2010 A Multiobjective and Evolutionary Clustering Method for Dynamic Networks
abstract
The discovery of evolving communities in dynamic networks is an important research topic that poses challenging tasks. Previous evolutionary based clustering methods try to maximize cluster accuracy, with respect to incoming data of the current time step, and minimize clustering drift from one time step to the successive one. In order to optimize both these two competing objectives, an input parameter that controls the preference degree of a user towards either the snapshot quality or the temporal quality is needed. In this paper the detection of communities with temporal smoothness is formulated as a multiobjective problem and a method based on genetic algorithms is proposed. The main advantage of the algorithm is that it automatically provides a solution representing the best trade-off between the accuracy of the clustering obtained, and the deviation from one time step to the successive. Experiments on synthetic data sets show the very good performance of the method compared to state-of-the-art approaches.
Francesco Folino, Clara Pizzuti
ASONAM1
2009 Discovering expressive process models from noised log data
abstract
Process-oriented systems have been increasingly attracting data mining researchers, mainly due to the advantages that the application of inductive process mining techniques to log data could open to both the analysis of complex processes and the design of new process models.
Francesco Folino, Gianluigi Greco, Antonella Guzzo, Luigi Pontieri
IDEAS1
2008 Boosting text segmentation via progressive classification
Eugenio Cesario, Francesco Folino, Antonio Locane, Giuseppe Manco 0001, Riccardo Ortale
Knowl. Inf. Syst.2
2006 Effective Incremental Clustering for Duplicate Detection in Large Databases
abstract
We propose an incremental algorithm for discovering clusters of duplicate tuples in large databases. The core of the approach is the usage of an indexing technique which, for any newly arrived tuple mu, allows to efficiently retrieve a set of tuples in the database which are mostly similar to mu, and which are likely to refer to the same real-world entity which is associated with mu. The proposed index is based on a hashing approach which tends to assign similar objects to the same buckets. Empirical and analytical evaluation demonstrates that the proposed approach achieves satisfactory efficiency results, at the cost of low accuracy loss
Francesco Folino, Giuseppe Manco 0001, Luigi Pontieri
IDEAS1
2005 An Incremental Clustering Scheme for Duplicate Detection in Large Databases
abstract
We propose an incremental algorithm for clustering duplicate tuples in large databases, which allows to assign any new tuple t to the cluster containing the database tuples which are most similar to t (and hence are likely to refer to the same real-world entity t is associated with). The core of the approach is a hash-based indexing technique that tends to assign highly similar objects to the same buckets. Empirical evaluation proves that the proposed method allows to gain considerable efficiency improvement over a state-of-art index structure for proximity searches in metric spaces.
Eugenio Cesario, Francesco Folino, Giuseppe Manco 0001, Luigi Pontieri
IDEAS2
2004 Putting Enhanced Hypermedia Personalization into Practice via Web Mining
Eugenio Cesario, Francesco Folino, Riccardo Ortale
DEXA2