EDBT 2026 Demo / reviewers in the wild / expert
Donato Malerba
dblp:m/DonatoMalerba
· DBLP profile ↗
66ranked-venue papers in the field
6as first author
14since 2021 · last 2026
0000-0001-8432-4608ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 19 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 17 (1 first)Database Systems & Data Management · 12 (1 first)Other / Interdisciplinary · 11 (2 first)Business Process & Enterprise Data · 4Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal predictive process monitoring and its application to explainable clinical pathwaysabstractThis paper presents one of the first contributions in the context of Multimodal Predictive Process Monitoring (MM-PPM) . In recent years, Predictive Process Monitoring (PPM) has evolved at the intersection of process mining, machine learning, and data science, as organizations seek to anticipate the future course of ongoing processes. Traditional PPM mainly relies on structured event log data, but many real-world scenarios generate richer information, including text, images, audio, and video. MM-PPM promises to start addressing this rich data scenario by integrating complementary knowledge from heterogeneous modalities through modality-specific representations and information fusion techniques. The growing digitization of healthcare systems, combined with advances in Artificial Intelligence (AI), has accelerated AI-based PPM for analyzing sequences of clinical events, supporting decision-making, enabling personalized care, and improving clinical facility management. Given these characteristics, clinical pathways represent an ideal domain for experimenting with MM-PPM, as they may naturally involve diverse modalities such as structured records, free-text notes, or medical images. To handle multimodal information available with clinical pathways, we introduce MEDUSA , an MM-PPM approach for outcome prediction, which jointly processes medical image information coupled with the storytelling of structural records and text notes collected during the clinical pathway of a patient until the acquisition of the considered image. The evaluation of MEDUSA is done in a COVID-19 case study, to assess the performance of the proposed approach and explain how specific information within each modality influences the decisions of the predictive model. Vincenzo Pasquadibisceglie, Ivan Donadello, Annalisa Appice, Oswald Lanz, Fabrizio Maria Maggi, Giuseppe Fiameni, Donato Malerba |
Inf. Syst. | 7 |
| 2025 | Leveraging a foundation deep neural embedding in process discovery under not-Pareto distributionabstractProcess discovery aims to automatically discover a process model to explain the behavior of event traces recorded in an event log during the execution of the activities of an underlying process. Several powerful process discovery algorithms are already formulated in process mining to identify regular control-flow structures in event logs and strike different trade-offs between the accuracy in capturing the behavior recorded in the event $\log$ and the complexity of the derived process model. This is commonly done under the assumption that log event traces are distributed according to the Pareto principle with a large portion of event traces held by a small fraction of top-frequent variants. However, the Pareto principle is not always satisfied in several real-life, complex processes. For example, the majority of log event traces produced in various healthcare or gaming processes is often spanned on a high number of top-frequent varianttraces. Various techniques (e.g. event filtering, trace extraction and trace abstraction) are already formulated in process mining to cope traditional process discovery algorithms also with event logs that do not conform the Pareto principle. Following this line of research, we explore the performance of a trace extraction method introduced to support the quest for Pareto-like event trace distribution during process discovery. The trace extraction is done resorting to a deep embedding representation of event traces, which sees traces at an abstraction level that removes noise and anomalous trace excerpt by enabling the discovery of simpler process models with higher accuracy. The deep embedding is used in combination with clustering. Experiments with several benchmark event logs show the effectiveness of the two proposed methods also compared to prior methods. Vincenzo Pasquadibisceglie, Annalisa Appice, Giovanni Discanno, Donato Malerba |
ICPM | 4 |
| 2025 | OLIVANDER: a counterfactual-based method to generate adversarial Windows PE malwareabstractAbstract Artificial Intelligence (AI) is transforming cybersecurity practices thanks to the amazing accuracy performance achieved with several AI-based malware detection systems. However, several recent studies have shown that AI decision models can be vulnerable to adversarial attacks. In malware detection scenarios, adversarial attacks are realistic manipulations of existing malware, which preserve the executable and malicious behaviour but evade the malware detection measures. In this study, we consider Windows Portable Executable (PE) malware, which is currently trending to prominent malware types, and we show that counterfactual explanations can be used to drive the generation of realistic adversarial Windows PE malware to evade AI-based detection. In particular, the proposed method OLIVANDER works in a black-box manner, which is the most restrictive attack option, as the evasion method interacts with the target decision system to evade by merely knowing the model input and output. The evaluation study explores the effectiveness of the proposed evasion method in terms of evasion ability, efficiency of computation, and attack transferability compared to two state-of-the-art evasion methods. In addition, the performed evaluation accounts for performances on commercial anti-malware systems. Luca De Rose, Giuseppina Andresini, Annalisa Appice, Donato Malerba |
Data Min. Knowl. Discov. | 4 |
| 2024 | LUPIN: A LLM Approach for Activity Suffix Prediction in Business Process Event LogsabstractForecasting future states of running process instances is one of the main challenges of Predictive Process Monitoring (PPM). Several deep learning approaches have recently achieved a valuable accuracy performance by addressing this task. On the other hand, with the recent boom of Large Language Models (LLMs) in multiple fields, LLMs have started attracting attention in PPM research also. In this study, we leverage the rich context of textual data to transform information recorded in event logs in smart textual data ready for boosting accurate PPM learning. In detail, we propose LUPIN, a LLM approach to predict the activity suffix of a running process instance. First it encodes historical running process instances in semantic text stories formulated according to narrative templates that account for information recorded in the event log. Then it fine tunes a pre-trained LLM model – medium BERT – on the text stories of historic running instances of a business process, to predict the activity suffix of any future running instance of the same business process. Finally, LUPIN integrates the XAI Integrated Gradient (IG) algorithm to explain how each part of the textual description of a running process instance has an effect on the prediction of its activity completion. The experimental evaluation explores the accuracy performance of LUPIN compared to that of several related methods and draws insights from the explanation retrieved through the IG algorithm. Vincenzo Pasquadibisceglie, Annalisa Appice, Donato Malerba |
ICPM | 3 |
| 2024 | Data-Centric AIabstractThe evolution of Artificial Intelligence (AI) has been driven by two core components: data and algorithms. Historically, AI research has predominantly followed the Model-Centric paradigm, which focuses on developing and refining models, while often treating data as static. This approach has led to the creation of increasingly sophisticated algorithms, which demand vast amounts of manually labeled and meticulously curated data. However, as data becomes central to AI development, it is also emerging as a significant bottleneck. The Data-Centric AI (DCAI) paradigm shifts the focus towards improving data quality, enabling the achievement of accuracy levels that are unattainable with Model-Centric approaches alone. This special issue presents recent advancements in DCAI, offering insights into the paradigm and exploring future research directions, aiming to contextualize the contributions included in this issue. Donato Malerba, Vincenzo Pasquadibisceglie |
J. Intell. Inf. Syst. | 1 |
| 2024 | TSUNAMI - an explainable PPM approach for customer churn prediction in evolving retail data environments
Vincenzo Pasquadibisceglie, Annalisa Appice, Giuseppe Ieva, Donato Malerba |
J. Intell. Inf. Syst. | 4 |
| 2023 | PANACEA: A Neural Model Ensemble for Cyber-Threat DetectionabstractThis study describes a new cyber-threat detection method, named PANACEA, that uses Ensemble Deep Learning coupled with Adversarial Training and XAI, to gain accuracy with neural models trained in cybersecurity problems. Malik Al-Essa, Giuseppina Andresini, Annalisa Appice, Donato Malerba |
DSAA | 4 |
| 2023 | An AI framework to support decisions on GDPR complianceabstractAbstract The Italian Public Administration (PA) relies on costly manual analyses to ensure the GDPR compliance of public documents and secure personal data. Despite recent advances in Artificial Intelligence (AI) have benefited many legal fields, the automation of workflows for data protection of public documents is still only marginally affected. The main aim of this work is to design a framework that can be effectively adopted to check whether PA documents written in Italian meet the GDPR requirements. The main outcome of our interdisciplinary research is INTREPID (art ficial i elligence for gdp complianc of ublic adm nistration ocuments), an AI-based framework that can help the Italian PA to ensure GDPR compliance of public documents. INTREPID is realized by tuning some linguistic resources for Italian language processing (i.e. SpaCy and Tint) to the GDPR intelligence. In addition, we set the foundations for a text classification methodology to recognise the public documents published by the Italian PA, which perform data breaches. We show the effectiveness of the framework over a text corpus of public documents that were published online by the Italian PA. We also perform an inter-annotator study and analyse the agreement of the annotation predictions of the proposed methodology with the annotations by domain experts. Finally, we evaluate the accuracy of the proposed text classification model in detecting breaches of security. Filippo Lorè, Pierpaolo Basile, Annalisa Appice, Marco de Gemmis, Donato Malerba, Giovanni Semeraro |
J. Intell. Inf. Syst. | 5 |
| 2022 | Anomaly Detection for Public Transport and Air Pollution AnalysisabstractAnomaly detection is a machine learning task that has been investigated within diverse research areas and application domains. In this paper, we performed anomaly detection for air pollution and public transport traffic analysis for the city of Oslo, Norway. To this aim, the state-of-the-art method SparkGHSOM was considered to learn predictive models for normal (i.e. regular) scenarios of air quality and traffic jams in a distributed fashion. Furthermore, we extended the main algorithm to make the detected anomalies explainable through an instance-based feature ranking approach. The results showed that SparkGHSOM is able to detect anomalies for both the real applications considered in this study, despite the fact it was designed for different tasks. Paolo Mignone, Donato Malerba, Michelangelo Ceci |
IEEE Big Data | 2 |
| 2022 | Leveraging autoencoders in change vector analysis of optical satellite imagesabstractAbstract Various applications in remote sensing demand automatic detection of changes in optical satellite images of the same scene acquired over time. This paper investigates how to leverage autoencoders in change vector analysis, in order to better delineate possible changes in a couple of co-registered, optical satellite images. Let us consider both a primary image and a secondary image acquired over time in the same scene. First an autoencoder artificial neural network is trained on the primary image. Then the reconstruction of both images is restored via the trained autoencoder so that the spectral angle distance can be computed pixelwise on the reconstructed data vectors. Finally, a threshold algorithm is used to automatically separate the foreground changed pixels from the unchanged background. The assessment of the proposed method is performed in three couples of benchmark hyperspectral images using different criteria, such as overall accuracy, missed alarms and false alarms. In addition, the method supplies promising results in the analysis of a couple of multispectral images of the burned area in the Majella National Park (Italy). Giuseppina Andresini, Annalisa Appice, Daniele Iaia, Donato Malerba, Nicolò Taggio, Antonello Aiello |
J. Intell. Inf. Syst. | 4 |
| 2021 | FOX: a neuro-Fuzzy model for process Outcome prediction and eXplanationabstractPredictive process monitoring (PPM) techniques have become a key element in both public and private organizations by enabling crucial operational support of their business processes. Thanks to the availability of large amounts of data, different solutions based on machine and deep learning have been proposed in the literature for the monitoring of process instances. These state-of-the-art approaches leverage accuracy as main objective of the predictive modeling, while they often neglect the interpretability of the model. Recent studies have addressed the problem of interpretability of predictive models leading to the emerging area of Explainable AI (XAI). In an attempt to bring XAI in PPM, in this paper we propose a fully interpretable model for outcome prediction. The proposed method is based on a set of fuzzy rules acquired from event data via the training of a neuro-fuzzy network. This solution provides a good trade-off between accuracy and interpretability of the predictive model. Experimental results on different benchmark event logs are encouraging and motivate the importance to develop explainable models for predictive process analytics. Vincenzo Pasquadibisceglie, Giovanna Castellano, Annalisa Appice, Donato Malerba |
ICPM | 4 |
| 2021 | Autoencoder-based deep metric learning for network intrusion detection
Giuseppina Andresini, Annalisa Appice, Donato Malerba |
Inf. Sci. | 3 |
| 2021 | Leveraging colour-based pseudo-labels to supervise saliency detection in hyperspectral image datasetsabstractAbstract Saliency detection mimics the natural visual attention mechanism that identifies an imagery region to be salient when it attracts visual attention more than the background. This image analysis task covers many important applications in several fields such as military science, ocean research, resources exploration, disaster and land-use monitoring tasks. Despite hundreds of models have been proposed for saliency detection in colour images, there is still a large room for improving saliency detection performances in hyperspectral imaging analysis. In the present study, an ensemble learning methodology for saliency detection in hyperspectral imagery datasets is presented. It enhances saliency assignments yielded through a robust colour-based technique with new saliency information extracted by taking advantage of the abundance of spectral information on multiple hyperspectral images. The experiments performed with the proposed methodology provide encouraging results, also compared to several competitors. Annalisa Appice, Angelo Cannarile, Antonella Falini, Donato Malerba, Francesca Mazzia, Cristiano Tamborrino |
J. Intell. Inf. Syst. | 4 |
| 2021 | Mining emotion-aware sequential rules at user-level from micro-blogs
Marjana Prifti Skenduli, Marenglen Biba, Corrado Loglisci, Michelangelo Ceci, Donato Malerba |
J. Intell. Inf. Syst. | 5 |
| 2020 | Clustering-Aided Multi-View Classification: A Case Study on Android Malware Detection
Annalisa Appice, Giuseppina Andresini, Donato Malerba |
J. Intell. Inf. Syst. | 3 |
| 2019 | Using Convolutional Neural Networks for Predictive Process AnalyticsabstractPredictive process monitoring has recently become one of the main enablers of data-driven insights in process mining. As an application of predictive analytics, process prediction is mainly concerned with predicting the evolution of running traces based on models extracted from historical event logs. This paper presents a process mining approach, which uses convolutional neural networks to equip the execution scenario of a business process with a means to predict the next activity in a running trace. The basic idea is to convert the temporal data enclosed in the historical event log of a business process into spatial data so as to treat them as images. To this purpose, every trace of the event log is first transformed into the set of its prefix traces (i.e. sequences of events that represent the prefix of a trace). These prefix traces are mapped into 2D image-like data structures. Created spatial data are finally used to train a Convolutional Neural Network, in order to learn a deep learning model capable to predict the next activity (i.e. the activity associated to the event occurring after the last event in the considered prefix trace). This predictive deep model can be employed as a powerful service to support participants in performing business processes since it guarantees a higher utilization by acting proactively in anticipation. Preliminary tests with two benchmark logs are carried out to investigate the viability of the proposed approach. Vincenzo Pasquadibisceglie, Annalisa Appice, Giovanna Castellano, Donato Malerba |
ICPM | 4 |
| 2019 | Spatial autocorrelation and entropy for renewable energy forecasting
Michelangelo Ceci, Roberto Corizzo, Donato Malerba, Aleksandra Rashkovska |
Data Min. Knowl. Discov. | 3 |
| 2018 | Distributed Learning of Process Models for Next Activity PredictionabstractProcess mining is a research discipline that aims to discover, monitor and improve real processing using event logs. In this paper we tackle the problem of next activity prediction/recommendation via "nested prediction model" learning, that is, we first identify recurrent and frequent sequences of activities and then we learn a prediction model for each frequent sequence. The key principle underlying the design of the proposed solution is in the ability to process massive logs by means of a parallel and distributed solution (by exploiting the Spark parallel computation framework) which can make reasonable decisions in the absence of perfect models. Indeed, given the classical threshold for minimum support and a user-specified error bound, our approach exploits the Chernoff bound to mine "approximate" frequent sequences with statistical error guarantees on their actual supports. Experiments on real-world log data prove the effectiveness of the proposed approach. Michelangelo Ceci, Michele Spagnoletta, Pasqua Fabiana Lanotte, Donato Malerba |
IDEAS | 4 |
| 2018 | Leveraging correlation across space and time to interpolate geophysical data via CoKrigingabstractManaging geophysical data generated by emerging spatiotemporal data sources (e.g. geosensor networks) presents a growing challenge to Geographic Information System science. The presence of correlation poses difficulties with respect to traditional spatial data analysis. This paper describes a novel spatiotemporal analytical scheme that allows us to yield a characterization of correlation in geophysical data along the spatial and temporal dimensions. We resort to a multivariate statistical model, namely CoKriging, in order to derive accurate spatiotemporal interpolation models. These predict unknown data by utilizing not only their own geosensor values at the same time, but also information from near past data. We use a window-based computation methodology that leverages the power of temporal correlation in a spatial modeling phase. This is done by also fitting the computed interpolation model to data which may change over time. In an assessment, using various geophysical data sets, we show that the presented algorithm is often able to deal with both spatial and temporal correlations. This helps to gain accuracy during the interpolation phase, compared to spatial and spatiotemporal competitors. Specifically, we evaluate the efficacy of the interpolation phase by using established machine-learning metrics (i.e. root mean squared error, Akaike information criterion and computation time). Sonja Pravilovic, Annalisa Appice, Donato Malerba |
Int. J. Geogr. Inf. Sci. | 3 |
| 2018 | Active learning via collective inference in network regression problems
Annalisa Appice, Corrado Loglisci, Donato Malerba |
Inf. Sci. | 3 |
| 2018 | Multi-type clustering and classification from heterogeneous networks
Gianvito Pio, Francesco Serafino 0002, Donato Malerba, Michelangelo Ceci |
Inf. Sci. | 3 |
| 2017 | Using multiple time series analysis for geosensor data forecasting
Sonja Pravilovic, Massimo Bilancia, Annalisa Appice, Donato Malerba |
Inf. Sci. | 4 |
| 2016 | Collective regression for handling autocorrelation of network data in a transductive setting
Corrado Loglisci, Annalisa Appice, Donato Malerba |
J. Intell. Inf. Syst. | 3 |
| 2016 | CloFAST: closed sequential pattern mining using sparse and vertical id-lists
Fabio Fumarola, Pasqua Fabiana Lanotte, Michelangelo Ceci, Donato Malerba |
Knowl. Inf. Syst. | 4 |
| 2015 | Big Data Techniques For Supporting Accurate Predictions of Energy Production From Renewable SourcesabstractPredicting the output power of renewable energy production plants distributed on a wide territory is a really valuable goal, both for marketing and energy management purposes. Vi-POC (Virtual Power Operating Center) project aims at designing and implementing a prototype which is able to achieve this goal. Due to the heterogeneity and the high volume of data, it is necessary to exploit suitable Big Data analysis techniques in order to perform a quick and secure access to data that cannot be obtained with traditional approaches for data management. In this paper, we describe Vi-POC -- a distributed system for storing huge amounts of data, gathered from energy production plants and weather prediction services. We use HBase over Hadoop framework on a cluster of commodity servers in order to provide a system that can be used as a basis for running machine learning algorithms. Indeed, we perform one-day ahead forecast of PV energy production based on Artificial Neural Networks in two learning settings, that is, structured and non-structured output prediction. Preliminary experimental results confirm the validity of the approach, also when compared with a baseline approach. Michelangelo Ceci, Roberto Corizzo, Fabio Fumarola, Michele Ianni, Donato Malerba, Gaspare Maria, Elio Masciari, Marco Oliverio, Aleksandra Rashkovska |
IDEAS | 5 |
| 2015 | Mining Multi-Relational Gradual PatternsabstractGradual patterns highlight covariations of attributes of the form “The more/less X, the more/less Y”. Their usefulness in several applications has recently stimulated the synthesis of several algorithms for their automated discovery from large datasets. However, existing techniques require all the interesting data to be in a single database relation or table. This paper extends the notion of gradual pattern to the case in which the co-variations are possibly expressed between attributes of different database relations. The interestingness measure for this class of “relational gradual patterns” is defined on the basis of both Kendall's τ and gradual supports. Moreover, this paper proposes two algorithms, named τRGP Miner and gRGP Miner, for the discovery of relational gradual rules. Three pruning strategies to reduce the search space are proposed. The efficiency of the algorithms is empirically validated, and the usefulness of relational gradual patterns is proved on some real-world databases. NhatHai Phan, Dino Ienco, Donato Malerba, Pascal Poncelet, Maguelonne Teisseire |
SDM | 3 |
| 2015 | Summarizing numeric spatial data streams by trend cluster discovery
Annalisa Appice, Anna Ciampi, Donato Malerba |
Data Min. Knowl. Discov. | 3 |
| 2015 | Effectively and efficiently supporting roll-up and drill-down OLAP operations over continuous dimensions via hierarchical clustering
Michelangelo Ceci, Alfredo Cuzzocrea, Donato Malerba |
J. Intell. Inf. Syst. | 3 |
| 2014 | Innovative power operating center management exploiting big data techniquesabstractThe problem of accurately predicting the energy production from renewable sources has recently received an increasing attention from both the industrial and the research communities. It presents several challenges, such as facing with the rate data are provided by sensors, the heterogeneity of the data collected, power plants efficiency, as well as uncontrollable factors, such as weather conditions and user consumption profiles. In this paper we describe Vi-POC (Virtual Power Operating Center), a project conceived to assist energy producers and decision makers in the energy market. In this paper we present the Vi-POC project and how we face with challenges posed by the specific application. The solutions we propose have roots both in big data management and in stream data mining. Michelangelo Ceci, Nunzio Cassavia, Roberto Corizzo, Pietro Dicosta, Donato Malerba, Gaspare Maria, Elio Masciari, Camillo Pastura |
IDEAS | 5 |
| 2014 | Network Reconstruction for the Identification of miRNA: mRNA Interaction Networks
Gianvito Pio, Michelangelo Ceci, Domenica D'Elia, Donato Malerba |
ECML/PKDD (3) | 4 |
| 2014 | Leveraging the power of local spatial autocorrelation in geophysical interpolative clustering
Annalisa Appice, Donato Malerba |
Data Min. Knowl. Discov. | 2 |
| 2014 | Dealing with temporal and spatial correlations to classify outliers in geophysical data streams
Annalisa Appice, Pietro Guccione, Donato Malerba, Anna Ciampi |
Inf. Sci. | 3 |
| 2012 | Continuously Mining Sliding Window Trend Clusters in a Sensor Network
Annalisa Appice, Donato Malerba, Anna Ciampi |
DEXA (2) | 2 |
| 2012 | An Unsupervised Framework for Topological Relations Extraction from Geographic Documents
Corrado Loglisci, Dino Ienco, Mathieu Roche, Maguelonne Teisseire, Donato Malerba |
DEXA (2) | 5 |
| 2012 | Learning to Rank from Concept-Drifting Network Data Streams
Lucrezia Macchia, Michelangelo Ceci, Donato Malerba |
DEXA (1) | 3 |
| 2012 | Guest Editors' Introduction: special issue of selected papers from ECML PKDD 2011
Dimitrios Gunopulos, Donato Malerba, Michalis Vazirgiannis |
Data Min. Knowl. Discov. | 2 |
| 2011 | Trend cluster based compression of geographically distributed data streamsabstractIn many real-time applications, such as wireless sensor network monitoring, traffic control or health monitoring systems, it is required to analyze continuous and unbounded geographically distributed streams of data (e.g. temperature or humidity measurements transmitted by sensors of weather stations). Storing and querying geo-referenced stream data poses specific challenges both in time (real-time processing) and in space (limited storage capacity). Summarization algorithms can be used to reduce the amount of data to be permanently stored into a data warehouse without losing information for further subsequent analysis. In this paper we present a framework in which data streams are seen as time-varying realizations of stochastic processes. Signal compression techniques, based on transformed domains, are applied and compared with a geometrical segmentation in terms of compression efficiency and accuracy in the subsequent reconstruction. Anna Ciampi, Annalisa Appice, Donato Malerba, Pietro Guccione |
CIDM | 3 |
| 2011 | Discovering process models through relational disjunctive patterns miningabstractThe automatic discovery of process models can help to gain insight into various perspectives (e.g., control flow or data perspective) of the process executions traced in an event log. Frequent patterns mining offers a means to build human understandable representations of these process models. This paper describes the application of a multi-relational method of frequent pattern discovery into process mining. Multi-relational data mining is demanded for the variety of activities and actors involved in the process executions traced in an event log which leads to a relational (or structural) representation of the process executions. Peculiarity of this work is in the integration of disjunctive forms into relational patterns discovered from event logs. The introduction of disjunctive forms enables relational patterns to express frequent variants of process models. The effectiveness of using relational patterns with disjunctions to describe process models with variants is assessed on real logs of process executions. Corrado Loglisci, Michelangelo Ceci, Annalisa Appice, Donato Malerba |
CIDM | 4 |
| 2011 | A Temporal Data Mining Framework for Analyzing Longitudinal Data
Corrado Loglisci, Michelangelo Ceci, Donato Malerba |
DEXA (2) | 3 |
| 2010 | Mapping web pages to database records via link pathsabstractIn this paper we propose a new knowledge management task which aims to map Web pages to their corresponding records in a structured database. For example, the DBLP database contains records for many computer scientists, and most of these persons have public Web pages; if we can map the database record with the appropriate Web page then the new information could be used to further describe the person's database record. To accomplish this goal we employ link paths which contain anchor texts from multiple paths through the Web ending at the Web page in question. We hypothesize that the information from these link paths can be used to generate an accurate Web page to database record mapping. Experiments on two large, real world data sets, DBLP and IMDB for the structured data and computer science faculty members' Web pages and official movie homepages for the Web page data, show that our method does provide an accurate mapping. Finally, we conclude by issuing a call for further research on this promising new task. Tim Weninger, Fabio Fumarola, Jiawei Han 0001, Donato Malerba |
CIKM | 4 |
| 2010 | A Relational Approach for Discovering Frequent Patterns with Disjunctions
Corrado Loglisci, Michelangelo Ceci, Donato Malerba |
DaWak | 3 |
| 2008 | Emerging Pattern Based Classification in Relational Data Mining
Michelangelo Ceci, Annalisa Appice, Donato Malerba |
DEXA | 3 |
| 2008 | A Grid-Based Multi-relational Approach to Process Mining
Antonio Turi, Annalisa Appice, Michelangelo Ceci, Donato Malerba |
DEXA | 4 |
| 2007 | A Data Mining Approach to Reading Order DetectionabstractDetermining the reading order for layout components ex- tracted from a document image can be a crucial problem for several applications. It enables the reconstruction of a single textual element from texts associated to multiple layout components and makes both information extraction and content-based retrieval of documents more effective. A common aspect for all methods reported in the literature is that they strongly depend on the specific domain and are scarcely reusable when the classes of documents or the task at hand changes. In this paper, we investigate the prob- lem of detecting the reading order of layout components by resorting to a data mining approach which acquires the do- main specific knowledge from a set of training examples. The input of the learning method is the description of the "chains" of layout components defined by the user. Only spatial information is exploited to describe a chain, thus making the proposed approach also applicable to the cases in which no text can be associated to a layout component. The method induces a probabilistic classifier based on the Bayesian framework which is used for reconstructing either single or multiple chains of layout components. It has been evaluated on a set of document images. Michelangelo Ceci, Margherita Berardi, G. Porcelli, Donato Malerba |
ICDAR | 4 |
| 2007 | Discovering Emerging Patterns in Spatial Databases: A Multi-relational Approach
Michelangelo Ceci, Annalisa Appice, Donato Malerba |
PKDD | 3 |
| 2007 | Classifying web documents in a hierarchy of categories: a comprehensive study
Michelangelo Ceci, Donato Malerba |
J. Intell. Inf. Syst. | 2 |
| 2006 | Mining spatio-temporal data
Gennady L. Andrienko, Donato Malerba, Michael May 0001, Maguelonne Teisseire |
J. Intell. Inf. Syst. | 2 |
| 2005 | A color-based layout analysis to process censorship cards of film archivesabstractProcessing censorship cards of the 20/sup th/ century in order to support annotation and retrieval processes, leads to a number of challenges for many DIA systems. Problems due to the low layout quality and standard of such a material can be reduced by exploiting information conveyed by color. In this paper, taking into account lessons learned in the context of the 1ST project Collate, we propose a new method for image segmentation and layout analysis that takes full advantage of color information. The method has been implemented in the DIA system WISDOM++ and tested on a corpus of multi-format documents concerning historic film censorships. Margherita Berardi, Oronzo Altamura, Michelangelo Ceci, Donato Malerba |
ICDAR | 4 |
| 2005 | Relational Learning techniques for Document Image Understanding: Comparing Statistical and Logical approachesabstractIn this paper, we evaluate and systematically compare two different (multi-)relational learning methods based on a statistical approach and a logical approach for the task of document image understanding. For a fair comparison, both methods are tested on the same real world dataset consisting of multipage articles published in an international journal. An analysis of pros and cons of both approaches is reported. Michelangelo Ceci, Margherita Berardi, Donato Malerba |
ICDAR | 3 |
| 2005 | Mining Model Trees from Spatial Data
Donato Malerba, Michelangelo Ceci, Annalisa Appice |
PKDD | 1 |
| 2004 | An Integrated Approach for Automatic Semantic Structure Extraction in Document Images
Margherita Berardi, Michele Lapi, Donato Malerba |
Document Analysis Systems | 3 |
| 2004 | Spatial Associative Classification at Different Levels of Granularity: A Probabilistic Approach
Michelangelo Ceci, Annalisa Appice, Donato Malerba |
PKDD | 3 |
| 2003 | XML and Knowledge Technologies for Semantic-Based Indexing of Paper Documents
Donato Malerba, Michelangelo Ceci, Margherita Berardi |
DEXA | 1 |
| 2003 | Hierarchical Classification of HTML Documents with WebClassII
Michelangelo Ceci, Donato Malerba |
ECIR | 2 |
| 2003 | Correcting the Document Layout: A Machine Learning ApproachabstractIn this paper, a machine learning approach to support the user during the correction of the layout analysis is proposed. Layout analysis is the process of extracting a hierarchical structure describing the layout of a page. In our approach, the layout analysis is performed in two steps: firstly, the global analysis determines possible areas containing paragraphs, sections, columns, figures and tables, and secondly, the local analysis groups together blocks that possibly fall within the same area. The result of the local analysis process strongly depends on the quality of the results of the first step. We investigate the possibility of supporting the user during the correction of the results of the global analysis. This is done by allowing the user to correct the results of the global analysis and then by learning rules for layout correction from the sequence of user actions. Experimental results on a set of multi-page documents are reported and commented. Donato Malerba, Floriana Esposito, Oronzo Altamura, Michelangelo Ceci, Margherita Berardi |
ICDAR | 1 |
| 2003 | Mr-SBC: A Multi-relational Naïve Bayes Classifier
Michelangelo Ceci, Annalisa Appice, Donato Malerba |
PKDD | 3 |
| 2001 | Automated Discovery of Dependencies Between Logical Components in Document Image UnderstandingabstractDocument image understanding denotes the recognition of semantically relevant components in the layout extracted from a document image. This recognition process is based on some visual models, whose manual specification can be a highly demanding task. In order to automatically acquire these models, we propose the application of machine learning techniques. Problems raised by possible dependencies between concepts to be learned are illustrated and solved with a computational strategy based on the separate-and-parallel-conquer search. The approach is tested on a set of real multi-page documents processed by the system WISDOM++. New results confirm the validity of the proposed strategy and show some limits of the learning system used in this work. Donato Malerba, Floriana Esposito, Francesca A. Lisi, Oronzo Altamura |
ICDAR | 1 |
| 2000 | Machine Learning for Intelligent Processing of Printed Documents
Floriana Esposito, Donato Malerba, Francesca A. Lisi |
J. Intell. Inf. Syst. | 2 |
| 1999 | WISDOM++: An Interactive and Adaptive Document Analysis SystemabstractWISDOM++ is a document analysis system whose main design requirements are real-time user interaction and adaptivity. This paper presents the two-phased skew estimation algorithm and the adaptive document block segmentation and classification techniques. An evaluation of the performance of some of these tasks is also conducted according to a benchmarking procedure. Oronzo Altamura, Floriana Esposito, Donato Malerba |
ICDAR | 3 |
| 1997 | Information Capture and Semantic Indexing of Digital Libraries through Machine Learning TechniquesabstractThis paper presents a prototypical digital library service. It integrates machine learning tools and techniques in order to make effective, efficient and economically feasible the process of capturing the information that should be stored and indexed by content in the digital library. Infact, information capture is one of the main bottleneck when building a digital library, since it involves complex pattern recognition problems, such as document analysis, classification and understanding. Experimental results show that learning systems can solve effectively and efficiently these problems. Floriana Esposito, Donato Malerba, Giovanni Semeraro, Cesare Daniele Antifora, Gioacchino de Gennaro |
ICDAR | 2 |
| 1995 | Simplifying Decision Trees by Pruning and Grafting: New Results (Extended Abstract)
Floriana Esposito, Donato Malerba, Giovanni Semeraro |
ECML | 2 |
| 1995 | A knowledge-based approach to the layout analysisabstractIn this paper, we present a hybrid approach to the problem of the document analysis in which the document image is segmented by means of a top-down technique and then basic blocks are grouped bottom-up in order to form complex layout components. In this latter process, called layout analysis, only generic knowledge on typesetting conventions is exploited. Such knowledge is independent of the particular class of processed documents and turns out to be valuable for a wide range of documents. Preliminary results of the layout analysis system LEX (Layout EXpert) show the methodological validity of this approach. Floriana Esposito, Donato Malerba, Giovanni Semeraro |
ICDAR | 2 |
| 1994 | An Analytic and Empirical Comparison of Two Methods for Discovering Probabilistic Causal Relationships
Donato Malerba, Giovanni Semeraro, Floriana Esposito |
ECML | 1 |
| 1993 | Decision Tree Pruning as a Search in the State Space
Floriana Esposito, Donato Malerba, Giovanni Semeraro |
ECML | 2 |
| 1993 | Automated acquisition of rules for document understandingabstractA study on the possibility of adopting a supervised inductive learning approach to the problem of document understanding is presented. A representation language used to describe a page layout is introduced and the opportunity of extending such a language by means of intentionally defined predicates is discussed. Experimental results obtained by using a well-known learning system, FOCL, are presented. They confirm the exigency of redefining the problem of document understanding in terms of a new strategy of supervised inductive learning, called contextual learning. Some experiments in which a dependence hierarchy between concepts is defined show that contextual rules increase predictive accuracy and decrease learning time for labeling problems like document understanding.> Floriana Esposito, Donato Malerba, Giovanni Semeraro |
ICDAR | 2 |
| 1990 | A Distance Measure for Decision Making in Uncertain Domains
Floriana Esposito, Donato Malerba, Giovanni Semeraro |
IPMU | 2 |