Massimo Guarascio 0001

dblp:41/7251 · DBLP profile ↗
← Back
38ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0001-7711-9833ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 14 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Boosting SSVEP Multi-subject Classification Through Mixture of ResNet-Based Experts
Francesco Capria, Franco Cicirelli, Alberto Falcone, Massimo Guarascio 0001, Antonio Guerrieri
ISMIS4
2026 MalARN: An Adversarial Reconstruction Network for Improving Detection of Evolving Malware
Francesco Pasqualatto, Luca Caviglione, Massimo Guarascio 0001, Angelica Liguori, Giuseppe Manco 0001, Ettore Ritacco, Antonino Rullo
ISMIS3
2026 Deobfuscation of JavaScript code and identification of security weaknesses through large language models
abstract
Advancements in Large Language Models (LLMs) allow solving many challenging tasks related to software security in an automatic manner, e.g., the generation of test cases. An important aspect concerns the deobfuscation of source code, especially for improving its readability or preventing the elusion of signature-based countermeasures. Although LLMs are increasingly deployed to reveal the presence of malicious payloads within obfuscated software components, a comprehensive understanding of their potential and limitations is still missing. In this work, we evaluate the effectiveness of deobfuscating JavaScript code through an LLM-based pipeline. In more detail, we investigate whether LLMs can preserve structural properties of the software, especially to enhance the identification of weaknesses. Compared to two standard tools (i.e., JSNice and js-deobfuscator ), our approach provides a more readable JavaScript prose according to several metrics, while retaining information on the Common Weaknesses Enumeration plaguing the software. To support the process of explaining issues within code, we performed tests on the use of two general-purpose LLMs, i.e., ChatGPT and Google Gemini. Results indicate that advancing the security of JavaScript through LLMs requires facing several challenges, which can be largely addressed via ad-hoc models.
Giacomo Benedetti, Luca Caviglione, Carmela Comito, Alberto Falcone, Massimo Guarascio 0001
Future Gener. Comput. Syst.5
2026 Seeing the invisible: Detection of stealth DoS attacks using variational U-Net-like models
abstract
The increasing sophistication of cyberattacks targeting companies and organizations continues to challenge the effectiveness of modern defense systems. Among these threats, slow Denial-of-Service (slow DoS) attacks are particularly difficult to detect, as they rely on evasion strategies that add significant complexity to cybersecurity efforts. Modern intrusion detection systems, especially those based on deep learning, have become essential tools in combating such attacks. However, their performance is often hindered by challenges such as limited data availability, noisy inputs, and the presence of out-of-distribution samples. Furthermore, their dependence on large labeled datasets makes detecting subtle or rare attack patterns particularly challenging. To overcome these limitations, this work proposes a novel unsupervised deep learning framework for detecting slow DoS attacks. The proposed approach incorporates a customized preprocessing pipeline to improve input data quality and leverages a sparse variational U-Net-like architecture for robust anomaly identification. Extensive experiments conducted on three real-world datasets demonstrate the ability of the framework to accurately and efficiently detect slow DoS attacks, highlighting its robustness, generalizability, and practical suitability for deployment in operational environments.
Enrico Cambiaso, Francesco Folino, Massimo Guarascio 0001, Angelica Liguori, Antonino Rullo
J. Inf. Secur. Appl.3
2026 A deep learning-based approach for stegomalware sanitization in digital images
abstract
Abstract Malware is increasingly endowed with steganographic mechanisms for concealing malicious data to avoid detection or bypass security measures. As a result, an emerging wave of threats named stegomalware has started to rise. Among the various approaches, real-world stegomalware primarily hides information within digital images, for instance, to retrieve additional payloads or configuration data. Unfortunately, developing attack-agnostic mitigation tools is difficult, especially due to the tight relation between the image format and the steganographic technique. Therefore, this paper presents an autoencoder-based approach to perform sanitization , i.e., to disrupt the malicious content hidden in images without altering their visual quality. For this purpose, we used an enhanced U-Net-like neural architecture, and we compared our idea against other mechanisms, including JPG transcoding and simple addition of Gaussian noise. Results obtained by considering different hiding patterns and realistic payloads showcased the effectiveness of our approach. Moreover, the U-Net-based sanitization solution prevents the recovery of the payload while preserving the original image quality and reducing risks arising from side-channel attacks.
Angelica Liguori, Marco Zuppelli, Daniela Gallo, Massimo Guarascio 0001, Luca Caviglione
J. Intell. Inf. Syst.4
2026 DALEK: combining deep active learning and explanations methods for fake news detection on COVID-19
Carmela Comito, Massimo Guarascio 0001, Angelica Liguori, Francesco Sergio Pisani
Neural Comput. Appl.2
2025 Learning Fast to Detect Slow: A Few-Shot Neural Approach to Slow DoS Attack Detection
abstract
Abstract The increasing frequency and sophistication of Denial of Service (DoS) and Distributed Denial of Service (DDoS) attacks, pose significant challenges to modern cybersecurity systems. These threats are further complicated by stealthy variants such as slow DoS attacks, which often evade timely detection. While Deep Learning (DL)-based Intrusion Detection Systems (IDSs) have shown promise in analyzing complex network traffic, their effectiveness is hindered by challenges like limited labeled data, noise, and the presence of Out-of-Distribution (OOD) samples. This paper proposes a hybrid DL-based IDS framework ( ENE4 ) that integrates unsupervised and supervised components to improve detection performance under label-scarce conditions. The unsupervised module extracts task-independent features from network traffic, while the supervised one learns task-specific representations. These complementary features are fused to enable robust detection even in few-shot learning settings. Additionally, the model incorporates an adaptation mechanism to leverage knowledge from more frequent and related attack types, enhancing generalization to rare patterns. Experimental results on two standard benchmark datasets demonstrate the effectiveness and robustness of the proposed approach in detecting evasive DoS attacks.
Francesco Scala, Massimo Guarascio 0001, Carlo Parrotta, Luigi Pontieri
DS2
2025 Breaking domain barriers: mixture of experts for cross-domain fake news detection
abstract
Social media have become a key tool for rapidly spreading information worldwide, amplifying the risks of misinformation and fake news. This is also intensified by the fact that fake news covers a wide range of topics across multiple domains. Machine learning, particularly language models, offers a promising solution for detecting fake news. However, a major limitation of existing methods is their inability to classify instances from new or unseen domains. To tackle this issue, we introduce MERMAID, a mixture of experts approach that leverages the knowledge from different specialized models to classify examples from unknown domains. Each expert is initially trained on a specific known domain and then fine-tuned using data from other known domains. A model merging procedure is then applied to combine related experts, reducing the number of models required for predicting instances from unknown domains. In addition, our approach can effectively be used in few-shot learning scenarios, where a small amount of data from the target/unknown domain is available during training. Experiments on five benchmark datasets demonstrate the effectiveness of our method in both zero-shot and few-shot learning settings.
Angelica Liguori, Francesco Sergio Pisani, Carmela Comito, Massimo Guarascio 0001, Giuseppe Manco 0001
Mach. Learn.4
2025 The force of few: boosting deviance detection in data scarcity scenarios through self-supervised learning and pattern-based encoding
abstract
Abstract In modern business environments, identifying anomalous or deviant instances in business process executions is a critical concern for enterprises and organizations. Recent advancements show that deep deviance detection models (DDMs), trained on process traces using (semi-)supervised learning techniques, outperform traditional machine learning methods. However, the effectiveness of these deep learning models often depends on large training datasets, which are not always available in practice, particularly in Green AI contexts, where data and computational resources are limited. To address these challenges, this paper presents a novel methodology for discovering deep DDMs that mitigates the impact of limited training data. Our approach incorporates an auxiliary self-supervised learning task that complements the primary deviance classification objective. In addition, we enhance the model with an autoencoder, using its reconstruction error as an additional self-supervisory signal. To promote interpretability, the model adopts a pattern-based encoding mechanism, on top of which two parallel feature-representation layers are efficiently and robustly learned through residual-like skip connections. Our method demonstrates its ability to handle the dual challenges of data efficiency and model explainability, as shown in a case study involving the execution traces of a real-world business process. The results highlight the potential of deep DDMs to achieve high performance in deviance detection, even when faced with limited data availability. Notably, our approach achieves an average performance gain (across all performance metrics) of over 15% while using only 5% of the labelled data, compared to a fully supervised baseline model, when evaluated on two publicly available logs from the current literature.
Francesco Folino, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
Soft Comput.3
2024 No Country for Leaking Containers: Detecting Exfiltration of Secrets Through AI and Syscalls
abstract
Containers offer lightweight execution environments for implementing microservices or cloud-native applications. Owing to their ubiquitous diffusion jointly with the complex interplay of hardware, computing, and network resources, effectively enforcing container security is a difficult task. Specifically, runtime detection of threats poses many challenges since container images are often immutable, and many malware deploys obfuscation or elusive mechanisms. Therefore, in this work, we propose a deep-learning-based approach for identifying the presence of two containers colluding to covertly leak secret information. In more detail, we consider a threat actor trying to exfiltrate a 4,096-bit private TLS key via five different covert channels. To decide whether containers are colluding for leaking data, the deep learning model is fed with statistical indicators of the syscalls, which are built starting from simple counters. Results indicate the effectiveness of our approach, even if some adjustments are needed to reduce the number of false positives.
Marco Zuppelli, Massimo Guarascio 0001, Luca Caviglione, Angelica Liguori
ARES2
2024 Beyond the Horizon: Using Mixture of Experts for Domain Agnostic Fake News Detection
Carmela Comito, Massimo Guarascio 0001, Angelica Liguori, Giuseppe Manco 0001, Francesco Sergio Pisani
DS (2)2
2024 Erasing the Shadow: Sanitization of Images with Malicious Payloads Using Deep Autoencoders
Angelica Liguori, Marco Zuppelli, Daniela Gallo, Massimo Guarascio 0001, Luca Caviglione
ISMIS4
2024 Mitigation of Covert Communications in MQTT Topics Through Small Language Models
abstract
Modern IoT ecosystems face many security issues. An aspect often neglected concerns covert channels, which allow for exfiltrating data or preventing detection. To this aim, the Message Queuing Telemetry Transport (MQTT) protocol can be abused to create various hidden communication paths, mainly due to its textual nature. Alas, simpler detection metrics could be ineffective and their optimization requires a vast number of test cases. Therefore, this paper proposes to use a small language model trained over real MQTT topics to automatically generate the required test cases. Results indicate the need for optimizations to make popular detection metrics usable “in the wild”.
Camilla Cespi Polisiani, Marco Zuppelli, Mariacarla Calzarossa, Luca Caviglione, Massimo Guarascio 0001
MASCOTS5
2024 Learning autoencoder ensembles for detecting malware hidden communications in IoT ecosystems
abstract
Abstract Modern IoT ecosystems are the preferred target of threat actors wanting to incorporate resource-constrained devices within a botnet or leak sensitive information. A major research effort is then devoted to create countermeasures for mitigating attacks, for instance, hardware-level verification mechanisms or effective network intrusion detection frameworks. Unfortunately, advanced malware is often endowed with the ability of cloaking communications within network traffic, e.g., to orchestrate compromised IoT nodes or exfiltrate data without being noticed. Therefore, this paper showcases how different autoencoder-based architectures can spot the presence of malicious communications hidden in conversations, especially in the TTL of IPv4 traffic. To conduct tests, this work considers IoT traffic traces gathered in a real setting and the presence of an attacker deploying two hiding schemes (i.e., naive and “elusive” approaches). Collected results showcase the effectiveness of our method as well as the feasibility of deploying autoencoders in production-quality IoT settings.
Nunzio Cassavia, Luca Caviglione, Massimo Guarascio 0001, Angelica Liguori, Marco Zuppelli
J. Intell. Inf. Syst.3
2024 Data- & compute-efficient deviance mining via active learning and fast ensembles
abstract
Abstract Detecting deviant traces in business process logs is crucial for modern organizations, given the harmful impact of deviant behaviours (e.g., attacks or faults). However, training a Deviance Prediction Model (DPM) by solely using supervised learning methods is impractical in scenarios where only few examples are labelled. To address this challenge, we propose an Active-Learning-based approach that leverages multiple DPMs and a temporal ensembling method that can train and merge them in a few training epochs. Our method needs expert supervision only for a few unlabelled traces exhibiting high prediction uncertainty. Tests on real data (of either complete or ongoing process instances) confirm the effectiveness of the proposed approach.
Francesco Folino, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
J. Intell. Inf. Syst.3
2024 Movie tag prediction: An extreme multi-label multi-modal transformer-based solution with explanation
Massimo Guarascio 0001, Marco Minici, Francesco Sergio Pisani, Erika De Francesco, Pasquale Lambardi
J. Intell. Inf. Syst.1
2023 Exploiting Deep Learning and Explanation Methods for Movie Tag Prediction
abstract
Indexing multimedia content with rich and accurate metadata allows for improving the quality of the search engines’ results and boosting the recommender systems performances, which can benefit from this information to yield more effective recommendation lists. Therefore, the adoption of tools able to automatically label multimedia content with informative tags represents an important task for all the companies offering streaming entertainment services. However, domain experts generally perform the tagging process manually, making it time-consuming and error-prone. In the last few years, Machine Learning techniques have been proposed as a promising solution to automate this type of task, but the lack of clean and labeled training data hinders the learning of robust classification models. To cope with the issues described above, in this work, we devised a Deep Learning based solution for semi-automatic multi-label classification integrating post-hoc explanation techniques. Specifically, model explanation methods are exploited to assist the operator in the labeling process by facilitating an understanding of the model predictions. The proposed approach has been validated on a real dataset, and the experimental results demonstrate its effectiveness.
Erica Coppolillo, Massimo Guarascio 0001, Marco Minici, Francesco Sergio Pisani
IDEAS2
2023 Learning ensembles of deep neural networks for extreme rainfall event detection
abstract
Abstract Accurate rainfall estimation is crucial to adequately assess the risk associated with extreme events capable of triggering floods and landslides. Data gathered from Rain Gauges (RGs), sensors devoted to measuring the intensity of the rain at individual points, are commonly used to feed interpolation methods (e.g., the Kriging geostatistical approach) and estimate the precipitation field over an area of interest. However, the information provided by RGs could be insufficient to model complex phenomena, and computationally expensive interpolation methods could not be used in real-time environments. Integrating additional data sources (e.g., radar and geostationary satellites) is an effective solution for improving the quality of the estimate, but it needs to cope with Big Data issues. To overcome all these issues, we propose a Rainfall Estimation Model (REM) based on an Ensemble of Deep Neural Networks (DeepEns-REM) that can automatically fuse heterogeneous data sources. The usage of Residual Blocks in the base models and the adoption of a Snapshot procedure to build the ensemble guarantees a fast convergence and scalability. Experimental results, conducted on a real dataset concerning a southern region in Italy, demonstrate the quality of the proposal in comparison with the Kriging interpolation technique and other machine learning techniques, especially in the case of exceptional rainfall events.
Gianluigi Folino, Massimo Guarascio 0001, Francesco Chiaravalloti
Neural Comput. Appl.2
2023 Neuro-Symbolic AI for Compliance Checking of Electrical Control Panels
abstract
Abstract Artificial Intelligence plays a main role in supporting and improving smart manufacturing and Industry 4.0, by enabling the automation of different types of tasks manually performed by domain experts. In particular, assessing the compliance of a product with the relative schematic is a time-consuming and prone-to-error process. In this paper, we address this problem in a specific industrial scenario. In particular, we define a Neuro-Symbolic approach for automating the compliance verification of the electrical control panels. Our approach is based on the combination of Deep Learning techniques with Answer Set Programming (ASP), and allows for identifying possible anomalies and errors in the final product even when a very limited amount of training data is available. The experiments conducted on a real test case provided by an Italian Company operating in electrical control panel production demonstrate the effectiveness of the proposed approach.
Vito Barbara, Massimo Guarascio 0001, Nicola Leone, Giuseppe Manco 0001, Alessandro Quarta, Francesco Ricca, Ettore Ritacco
Theory Pract. Log. Program.2
2022 Revealing MageCart-like Threats in Favicons via Artificial Intelligence
abstract
Modern malware increasingly takes advantage of information hiding to avoid detection, spread infections, and obfuscate code. A major offensive strategy exploits steganography to conceal scripts or URLs, which can be used to steal credentials or retrieve additional payloads. A recent example is the attack campaign against the Magento e-commerce platform, where a web skimmer has been cloaked in favicons to steal payment information of users.
Massimo Guarascio 0001, Marco Zuppelli, Nunzio Cassavia, Luca Caviglione, Giuseppe Manco 0001
ARES1
2022 Detecting DoS and DDoS Attacks through Sparse U-Net-like Autoencoders
abstract
In the last few years, we experienced exponential growth in the number of cyber-attacks performed against com-panies and organizations. In particular, because of their ability to mask themselves as legitimate traffic, DoS and DDoS have become two of the most common kinds of attacks on computer networks. Modern Intrusion Detection Systems (IDSs) represent a precious tool to mitigate the risk of unauthorized network access as they allow for accurately discriminating between benign and malicious traffic. Among the plethora of approaches proposed in the literature for detecting network intrusions, Deep Learning (DL)-based IDSs have been proved to be an effective solution because of their ability to analyze low-level data (e.g., flow and packet traffic) directly. However, many current solutions require large amounts of labeled data to yield reliable models. Unfortunately, in real scenarios, small portions of data carry label information due to the cost of manual labeling conducted by human experts. Labels can even be completely missing for some reason (e.g., privacy concerns). To cope with the lack of labeled data, we propose an unsupervised DL-based intrusion detection methodology, combining an ad-hoc preprocessing procedure on input data with a sparse U-Net-like autoencoder architecture. The experimentation on an IDS benchmark dataset substantiates our approach's ability to recognize malicious behaviors correctly.
Nunzio Cassavia, Francesco Folino, Massimo Guarascio 0001
ICTAI3
2022 Ensembling Sparse Autoencoders for Network Covert Channel Detection in IoT Ecosystems
Nunzio Cassavia, Luca Caviglione, Massimo Guarascio 0001, Angelica Liguori, Marco Zuppelli
ISMIS3
2022 Combining Active Learning and Fast DNN Ensembles for Process Deviance Discovery
Francesco Folino, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
ISMIS3
2022 Learning and Explanation of Extreme Multi-label Deep Classification Models for Media Content
Marco Minici, Francesco Sergio Pisani, Massimo Guarascio 0001, Erika De Francesco, Pasquale Lambardi
ISMIS3
2022 Combining deep ensemble learning and explanation for intelligent ticket management
Paolo Zicari, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
Expert Syst. Appl.3
2022 Boosting Cyber-Threat Intelligence via Collaborative Intrusion Detection
abstract
Sharing threat events and Indicators of Compromise (IoCs) enables quick and crucial decision making relative to effective countermeasures against cyberattacks. However, the current threat information sharing solutions do not allow easy communication and knowledge sharing among threat detection systems (in particular Intrusion Detection Systems (IDS)) exploiting Machine Learning (ML) techniques. Moreover, the interaction with the expert, which represents an important component to gather verified and reliable input data for the ML algorithms, is weakly supported. To address all these issues, ORISHA, a platform for ORchestrated Information SHaring and Awareness enabling the cooperation among threat detection systems and other information awareness components, is proposed here. ORISHA is backed by a distributed Threat Intelligence Platform based on a network of interconnected Malware Information Sharing Platform instances, which enables the communication with several Threat Detection layers belonging to different organizations. Within this ecosystem, Threat Detection Systems mutually benefit by sharing knowledge that allows them to refine the underlying predictive accuracy. Uncertain cases, i.e. examples with low anomaly scores, are proposed to the expert, who acts with the role of oracle in an Active Learning scheme. By interfacing with a honeynet, ORISHA allows for enriching the knowledge base with further positive attack instances and then yielding robust detection models. An experimentation conducted on a well-known Intrusion Detection benchmark demonstrates the validity of the proposed architecture.
Massimo Guarascio 0001, Nunzio Cassavia, Francesco Sergio Pisani, Giuseppe Manco 0001
Future Gener. Comput. Syst.1
2022 A Machine Learning Approach for Rainfall Estimation Integrating Heterogeneous Data Sources
abstract
Providing an accurate rainfall estimate at individual points is a challenging problem in order to mitigate risks derived from severe rainfall events, such as floods and landslides. Dense networks of sensors, named rain gauges (RGs), are typically used to obtain direct measurements of precipitation intensity in these points. These measurements are usually interpolated by using spatial interpolation methods for estimating the precipitation field over the entire area of interest. However, these methods are computationally expensive, and to improve the estimation of the variable of interest in unknown points, it is necessary to integrate further information. To overcome these issues, this work proposes a machine learning-based methodology that exploits a classifier based on ensemble methods for rainfall estimation and is able to integrate information from different remote sensing measurements. The proposed approach supplies an accurate estimate of the rainfall where RGs are not available, permits the integration of heterogeneous data sources exploiting both the high quantitative precision of RGs and the spatial pattern recognition ensured by radars and satellites, and is computationally less expensive than the interpolation methods. Experimental results, conducted on real data concerning an Italian region, Calabria, show a significant improvement in comparison with Kriging with external drift (KED), a well-recognized method in the field of rainfall estimation, both in terms of the probability of detection (0.58 versus 0.48) and mean-square error (0.11 versus 0.15).
Massimo Guarascio 0001, Gianluigi Folino, Francesco Chiaravalloti, Salvatore Gabriele, Antonio Procopio, Pietro Sabatino
IEEE Trans. Geosci. Remote. Sens.1
2020 Deep Autoencoder Ensembles for Anomaly Detection on Blockchain
Francesco Scicchitano, Angelica Liguori, Massimo Guarascio 0001, Ettore Ritacco, Giuseppe Manco 0001
ISMIS3
2019 Learning Effective Neural Nets for Outcome Prediction from Partially Labelled Log Data
abstract
The problem of inducing a model for forecasting the outcome of an ongoing process instance from historical log traces has attracted notable attention in the field of Process Mining. Approaches based on deep neural networks have become popular in this context, as a more effective alternative to previous feature-based outcome-prediction methods. However, these approaches rely on a pure supervised learning scheme, and unfit many real-life scenarios where the outcome of (fully unfolded) training traces must be provided by experts. Indeed, since in such a scenario only a small amount of labeled traces are usually given, there is a risk that an inaccurate or overfitting model is discovered. To overcome these issues, a novel outcome-discovery approach is proposed here, which leverages a fine-tuning strategy that learns general-enough trace representations from unlabelled log traces, which are then reused (and adapted) in the discovery of the outcome predictor. Results on real-life data confirmed that our proposal makes a more effective and robust solution for label-scarcity scenarios than current outcome-prediction methods.
Francesco Folino, Gianluigi Folino, Massimo Guarascio 0001, Luigi Pontieri
ICTAI3
2019 A Deep Learning based architecture for rainfall estimation integrating heterogeneous data sources
abstract
Rain gauges are sensors providing direct measurement of precipitation intensity at individual point sites, and, usually, spatial interpolation methods are used to obtain an estimate of the precipitation field over the entire area of interest. Among them, Kriging with External Drift (KED) is a largely used and well-recognized method in this field. However, interpolation methods need to work with real-time data, and therefore can be hardly used in real-time scenarios. To overcome this issue, we propose a general machine learning framework, which can be trained offline, based on a deep learning architecture, also integrating information derived from remote sensing measurements such as weather radars and satellites. The framework allows to provide accurate estimations of the rainfall in the areas where no rain gauge data is available. Experimental results, conducted on real data concerning a southern region in Italy, provided by the Department of Civil Protection (DCP), show significant improvement in comparison with KED and other machine learning techniques.
Gianluigi Folino, Massimo Guarascio 0001, Francesco Chiaravalloti, Salvatore Gabriele
IJCNN2
2019 Predictive monitoring of temporally-aggregated performance indicators of business processes against low-level streaming events
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Inf. Syst.3
2018 A Predictive Learning Framework for Monitoring Aggregated Performance Indicators over Business Process Events
abstract
In many application contexts, a business process' executions are subject to performance constraints expressed in an aggregated form, usually over predefined time windows, and detecting a likely violation to such a constraint in advance could help undertake corrective measures for preventing it. This paper illustrates a prediction-aware event processing framework that addresses the problem of estimating whether the process instances of a given (unfinished) window w will violate an aggregate performance constraint, based on the continuous learning and application of an ensemble of models, capable each of making and integrating two kinds of predictions: single-instance predictions concerning the ongoing process instances of w, and time-series predictions concerning the "future" process instances of w (i.e. those that have not started yet, but will start by the end of w). Notably, the framework can continuously update the ensemble, fully exploiting the raw event data produced by the process under monitoring, suitably lifted to an adequate level of abstraction. The framework has been validated against historical event data coming from real-life business processes, showing promising results in terms of both accuracy and efficiency.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
IDEAS3
2017 Deviance-Aware Discovery of High Quality Process Models
abstract
Despite performance-oriented process mining techniques have been successfully employed in numerous application contexts, they hardly produce models with a satisfactory level of accuracy, generality, and readability when applied to processes featuring complex and heterogeneous behaviors. In particular, the presence of deviant (i.e. anomalous/exceptional) traces often lead to cumbersome models with misleading performance statistics. Noise/outlier filtering solutions help alleviate this problem, and discover a better model for “normal“ executions, but do not provide insight on the nature and impact of deviant ones. The discovery approach proposed here tries to recognize and describe both a normal execution scenario and a number of deviant ones for a business process, by inducing two different kinds of models from a given execution log of the process: (i) a list of readable clustering rules defining the deviance scenarios; (ii) a performance model for each discovered deviance scenario, and a “distilled” one for the “normal” cases that do not fall in any deviant scenario. Technically, these models are found by mainly exploiting a conceptual clustering method, which greedily tries to identify groups of traces that maximally deviate from a current normality model. Tests on real-life logs confirmed the validity of the approach, and its ability to both find good performance models and support the analysis of deviant process instances.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
ICTAI3
2016 A multi-view multi-dimensional ensemble learning approach to mining business process deviances
abstract
The execution logs of a business process have been recently exploited to extract classification models for discriminating “deviant” instances of the process - i.e. instances diverging from normal/desired outcomes (e.g., frauds, faults, SLA violations). Regarding all log traces as sequences of task labels, current solutions essentially map each trace onto a vector space where the features correspond to sequence-oriented patterns, and any standard classifier-induction method can be applied to separate the two classes of instances. An ensemble-learning approach was also recently proposed to combine multiple base learners trained on heterogenous pattern-based log views. However, as these approaches simply abstract each event into an activity symbol, they disregard all the non structural event data that are typically stored in real-life logs, and which may well help improve the detection of deviances. Moreover, the usefulness of deviance models could be enhanced by equipping each prediction with a confidence measure, allowing the analyst to focus on (or prioritize) more suspicious cases. To overcome these limitations, we propose a multi-view ensemble learning approach, which: (i) fully exploits the multi-dimensional nature of log events, with the help of a clustering-based trace abstraction method; and (ii) implements a context- and probability-aware stacking method for combining base models' predictions. Tests on a real-life log confirmed the validity of the approach, and its capability to achieve compelling performances w.r.t. state-of-the-art methods.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
IJCNN3
2016 A Robust and Versatile Multi-View Learning Framework for the Detection of Deviant Business Process Instances
abstract
Increasing attention has been paid to the detection and analysis of “deviant” instances of a business process that are connected with some kind of “hidden” undesired behavior (e.g. frauds and faults). In particular, several recent works faced the problem of inducing a binary classification model (here named deviance detection model ) that can discriminate between deviant traces and normal ones, based on a set of historical log traces (labeled as either deviant or normal). Current solutions rely on applying standard classifier-induction methods to a feature-based representation of the given traces, where the features include sequence-based patterns extracted from the corresponding sequences of activities. However, there is no consensus on which kinds of patterns are the most suitable for such a task. On the other hand, mixing multiple pattern families together may produce a heterogenous, redundant and sparse representation of the traces that likely leads to poor deviance detection models. In this paper, we propose an ensemble-learning method for solving this problem, where multiple base classifiers are trained on different feature-based views of the log (each obtained by mapping the traces onto a distinguished collection of patterns). A stacking procedure is used to combine the discovered base models into an overall probabilistic model that associates any new trace with an estimate of the probability that it reflects a deviant process instance. This helps the analyst prioritize the inspection of the cases that are more likely to be deviant. The method also takes advantage of all nonstructural data available in the log, and employs a resampling mechanism to deal with the rarity of deviances in the training log. It has been conceived as the core of a comprehensive framework for detecting and analyzing business process deviances. The framework supports the analyst to investigate suspect deviances, and provides some feedback to the learning method for improving the accuracy of the discovered deviance detection models. Tests on several real-life datasets proved the validity of the approach, as concerns its capability to discover an accurate deviance detection model, and to effectively exploit new (originally unlabeled) traces via active learning and self-training mechanisms.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Int. J. Cooperative Inf. Syst.3
2015 A Prediction Framework for Proactively Monitoring Aggregate Process-Performance Indicators
abstract
Monitoring the performances of a business process is a key issue in many organizations, especially when predefined constraints exist on them, due to contracts or internal requirements. Several approaches were defined recently in the literature for predicting the performances of a single process instance. However, in many real situations, process-oriented performance metrics and associated constraints are defined in an aggregated form, on a time-window basis. This work right addresses the problem of predicting whether (the process instances in) each time window will infringe an aggregate performance constraint, at a series of checkpoints within the window. To this end, at each checkpoint, three kinds of measures are to be estimated: what performance outcome each ongoing process instance will yield, how many process instances will start in the rest of the window, and what their aggregate performance outcomes will be. The approach proposed is general (it can reuse a wide range of regression methods), and it can be embedded in a continuous monitoring-and-learning scheme. Tests on real-life logs showed its validity in terms of prediction accuracy.
Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
EDOC2
2014 Mining Predictive Process Models out of Low-level Multidimensional Logs
Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
CAiSE2
2009 Rule Learning with Probabilistic Smoothing
Gianni Costa, Massimo Guarascio 0001, Giuseppe Manco 0001, Riccardo Ortale, Ettore Ritacco
DaWaK2