VLDB 2026 Research / reviewers in the wild / expert
Paolo Garza
dblp:g/PaoloGarza
· DBLP profile ↗
61ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-1263-7522ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 33 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Computer networks · 7 · 6 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Is Enough: Efficient Modeling of RTP Traffic for QoS Predictions in Real-Time CommunicationsabstractIn recent years, we have witnessed an unprecedented upsurge in popularity and advancement of Real-time Transport Protocol (RTP)-based real-time communication (RTC) applications. For the sake of their optimizations, Quality of Service (QoS) prediction serves as a viable venue for enhancing network monitoring and enabling preemptive solutions. However, existing methodologies are typically tailored and constrained to individual traffic flows and QoS metrics, lagging in correlation capturing and computational efficiency. In light of this, we argue that “one model is enough” to conquer these challenges, and propose a novel deep learning (DL) framework namelyOh, employing a teacher-student scheme with two training stages. The first (teacher) involves a sophisticated Long Short-Term Memory (LSTM) neural network (NN) empowered by a customized attention structure, and the second (student) comprises simple feedforward NNs to distill knowledge and reduce complexity. Specifically,Ohleverages a multi-task learning paradigm, mapping extracted features to four key QoS indicators. It is capable of simultaneously handling unlimited amount of concurrent RTP flows with packet-level information and performing end-to-end predictions of multiple QoS metrics in one single shot. Our work is based on massive traffic collected during real video-teleconferencing calls using various software, and benchmarked against multiple other machine learning (ML)/DL algorithms. As a result,Oh-teacher yields superior prediction performance, whereasOh-student achieves distinctly enhanced temporal efficiency with comparable forecasting outcomes. Tailai Song, Paolo Garza, Michela Meo, Maurizio M. Munafò |
IEEE Trans. Netw. | 2 |
| 2025 | KIEPrompter: Leveraging Lightweight Models' Predictions for Cost-Effective Key Information Extraction using Vision LLMsabstractKey information extraction (KIE) from visually rich documents, such as receipts and forms, involves a deep understanding of textual, visual, and layout feature information. Transformers fine-tuned for KIE achieve state-of-the-art performance but lack generality and portability across different domains. In contrast, vision large language models (VLLMs) offer higher flexibility and zero-shot capability but fall short with domain-specific layout relations unless performing a resource-demanding supervised fine-tuning. To reach the best compromise solution between lightweight models and VLLMs, we propose KIEPrompter, a cost-effective LLM-based KIE approach that leverages the predictions of lightweight models as external knowledge injected into VLLM prompts. By incorporating these auxiliary predictions, VLLMs are guided to attend relevant multimodal content without ad hoc training. The accuracy results achieved by KIEPrompter in three benchmark document collections are superior to those of VLLMs in both zero-shot and layout-sensitive scenarios. We compare various strategies for incorporating lightweight model predictions, ranging from coarse-grained predictions without explicit confidence scores to fine-grained per-element network logits. We also demonstrate that our approach is robust to the absence of specific classes in trained lightweight models, as the VLLMs' pre-training compensates for the limited generality of lightweight models. Lorenzo Vaiani, Yihao Ding, Luca Cagliero, Jean Lee, Paolo Garza, Josiah Poon, Soyeon Caren Han |
CIKM | 5 |
| 2025 | HydroChronos: Forecasting Decades of Surface Water ChangeabstractForecasting surface water dynamics is crucial for water resource management and climate change adaptation. However, the field lacks comprehensive datasets and standardized benchmarks. In this paper, we introduce HydroChronos, a large-scale, multi-modal spatiotemporal dataset for surface water dynamics forecasting designed to address this gap. We couple the dataset with three forecasting tasks. The dataset includes over three decades of aligned Landsat 5 and Sentinel-2 imagery, climate data, and Digital Elevation Models for diverse lakes and rivers across Europe, North America, and South America. We also propose AquaClimaTempo UNet, a novel spatiotemporal architecture with a dedicated climate data branch, as a strong benchmark baseline. Our model significantly outperforms a Persistence baseline for forecasting future water dynamics by +14% and +11% F1 across change detection and direction of change classification tasks, and by +0.1 MAE on the magnitude of change regression. Finally, we conduct an Explainable AI analysis to identify the key climate variables and input channels that influence surface water change, providing insights to inform and guide future modeling efforts. Daniele Rege Cambrin, Eleonora Poeta, Eliana Pastor, Isaac Corley, Tania Cerquitelli, Elena Baralis, Paolo Garza |
SIGSPATIAL/GIS | 7 |
| 2025 | Cross-modal consistency types in multimodal social dataabstractSocial media content, such as internet memes or tweets, are nowadays largely or mainly multimodal. Machine learning models often need to jointly process images and text to solve complex tasks such as hate speech detection or sentiment analysis. For example, the misogyny of a meme cannot be accurately predicted while considering the visual and textual modalities separately. Similarly, sentiment annotations for tweets’ images and text can be discordant. Detecting the samples with inconsistent modality contributions is particularly relevant to analyze machine learning model performance and explain classification errors. In this paper, we formalize the types of cross-modal consistency by differentiating between consistent cases and not. Cross-modal consistency denotes whether all modalities agree on the label (i.e., full consistency) or not (i.e., inconsistency). When the visual and textual modalities are discordant, we distinguish the cases in which a joint analysis of multimodal features is sufficient to solve the issue from those requiring a human agreement (i.e., NOR consistency). We also propose a CLIP-based architecture to predict the cross-modal consistency types and identify the modalities causing the inconsistency. The results achieved on benchmark datasets show that cross-modal consistency annotation is cost-effective, i.e., it provides relevant insights into model predictions while requiring a limited extra human effort. Lorenzo Vaiani, Luca Cagliero, Paolo Garza, Jason Ravagli |
Knowl. Based Syst. | 3 |
| 2025 | Discovering SpatioTemporally Invariant Event Patterns From Mobility DataabstractThe discovery of sequential patterns from spatiotemporal data is known to be a very complex data mining task. The relevance of spatiotemporal patterns to study event correlations in mobility data is established. Prior works addressed either the separate analysis of spatial and temporal dependencies among data, such as the co-location of events, or the study of the joint spatiotemporal properties of the trajectories observed over a region of interest. The aim of this paper is instead to overcome existing approaches by extracting sequences of discrete events showing spatiotemporally invariant properties. For example, ifan arbitrary bike sharing station becomes full (all its docks are used)thenwe will observe an increase in the occupancy level of the bike sharing stations in the surrounding area within ten minutes. We denote such a new pattern as a SpatioTemporally Invariant (STInv) event pattern because we observe several instances in the source data differing just in spatiotemporal shifts. We also propose a new algorithm to mine STInvs based on a prefix-projected sequential pattern growth approach and different quality metrics to quantify the contribution of the spatial invariance. The proposed approach is empirically evaluated on two mobility datasets related to a bike sharing system and traffic data. The results confirm the usability of the proposed solution in real-world scenarios. Luca Colomba, Luca Cagliero, Paolo Garza |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Packet Loss in Real-Time Communications: Can ML Tame Its Unpredictable Nature?abstractDue to the flourishing development of networks, and abetted by the Covid-19 pandemic, we have witnessed an exponential surge in the global proliferation of Real-Time Communications (RTC) applications in recent years. In light of this, the necessity for robust, scalable, and intelligent network infrastructures and technologies has become increasingly apparent. Among the principal challenges encountered in RTC lies the issue of packet loss. Indeed, the occurrence of losses leads to communication degradation and reallocation that adversely affect the Quality of Experience (QoE). In this paper, we investigate the feasibility of predicting packet loss phenomena through the utilization of machine learning techniques, solely based on statistics derived directly from packets. We provide different definitions of packet loss, subsequently focusing on the most critical scenario, which is defined as the first loss of a series. By delineating the concept of loss, we propose different problem formulations to determine whether there exists a mathematically advantageous scenario over others. To substantiate our analysis, we demonstrate that these phenomena can be correctly identified with a recall up to 66%, leveraging three ample datasets of RTC traffic, which were collected under distinct conditions at different times, further solidifying the validity of our findings. Tailai Song, Gianluca Perna, Paolo Garza, Michela Meo, Maurizio M. Munafò |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | BitFormer: Transformer-Based Neural Network for Bitrate Prediction in Real-Time CommunicationsabstractIn recent years, an exponential upsurge in the global proliferation of Real-Time Communications (RTC) applications has been witnessed, due to the prosperous development of networks and further fueled by the ramifications of the COVID-19 pandemic. Consequently, the imperative for development of intelligent, resilient, and scalable network infrastructures and technologies has grown significantly. Real-time bitrate prediction could play a crucial role, offering network observability and bolstering proactive system management. By accurately forecasting bitrate, it becomes possible to implement improvements at either application level or network level, such as swift and appropriate bandwidth adaptation. In this paper, we propose a novel Transformer-based deep learning framework called BitFormer designed to predict the short-term bitrate. Our work is based on extensive traffic data collected under various conditions using two prevalent RTC applications, and our model relies solely on packet-level information, which contains the fundamental traffic characteristics and facilitates effortless feature extraction. Through comprehensive evaluations and comparisons, we achieve a superior accuracy of 74% in identifying peak bitrates, while simultaneously ensuring commendable overall performance. Tailai Song, Gianluca Perna, Paolo Garza, Michela Meo, Maurizio M. Munafò |
CCNC | 3 |
| 2024 | DQNC2S: DQN-Based Cross-Stream Crisis Event Summarizer
Daniele Rege Cambrin, Luca Cagliero, Paolo Garza |
ECIR (3) | 3 |
| 2024 | Throughput Prediction in Real-Time Communications: Spotlight on Traffic ExtremesabstractAmidst the thriving advancement of networks, further catalyzed by the COVID-19 pandemic, we have witnessed a marked escalation in the worldwide adoption of Real-Time Communications (RTC) applications. In this context, there is a compelling necessity to cultivate intelligent and robust network infrastructures and technologies. Real-time throughput prediction emerges as a promising candidate for this purpose to foster network observability and provide preemptive functions, supporting advanced system management, e.g., bandwidth allocation and adaptive streaming. Nonetheless, contemporary solutions grapple with predicting extreme conditions in traffic throughput, notably peaks, valleys, and abrupt changes. To address the challenges, we propose a Transformer-based Deep Learning (DL) Neural Network (NN), leveraging solely packet-level information and adopting a multi-task learning paradigm, to predict short-term throughput, with an emphasis on critical values. In particular, our work is grounded in voluminous traffic traces procured from real video-teleconferencing sessions, and we formulate a time-series regression problem, comparing numerous technologies, from an adaptive filter to Machine Learning (ML) and DL approaches. Conclusively, our methodology exhibits superior efficacy, especially in forecasting traffic extremities. Tailai Song, Paolo Garza, Michela Meo, Maurizio M. Munafò |
ISCC | 2 |
| 2024 | Modelling Concurrent RTP Flows for End-to-end Predictions of QoS in Real Time CommunicationsabstractThe Real-time Transport Protocol (RTP)-based real-time communications (RTC) applications, exemplified by video conferencing, have experienced an unparalleled surge in popularity and development in recent years. In pursuit of optimizing their performance, the prediction of Quality of Service (QoS) metrics emerges as a pivotal endeavor, bolstering network monitoring and proactive solutions. However, contemporary approaches are confined to individual RTP flows and metrics, falling short in relationship capture and computational efficiency. To this end, we propose Packet-to-Prediction (P2P), a novel deep learning (DL) framework that hinges on raw packets to simultaneously process concurrent RTP flows and perform end-to-end prediction of multiple QoS metrics. Specifically, we implement a streamlined architecture, namely length-free Transformer with cross and neighbourhood attention, capable of handling an unlimited number of RTP flows, and employ a multi-task learning paradigm to forecast four key metrics in a single shot. Our work is based on extensive traffic collected during real video calls, and conclusively, P2P excels comparative models in both prediction performance and temporal efficiency. Tailai Song, Paolo Garza, Michela Meo, Maurizio M. Munafò |
ISM | 2 |
| 2024 | Towards the Detection of Unobservable Losses in Real-Time CommunicationsabstractPacket loss, an omnipresent issue that degrades the QoE in Real-time Transport Protocol (RTP)-based real-time communications (RTC) applications, serves as a pivotal indicator for gauging network performance. Conventionally, loss detection hinges on sequence number irregularities. However, many contemporary applications incorporate customized mechanisms that diverge from the standard, confounding loss identification. Although the actual losses are transparent to applications themselves, they remain unobservable to other entities such as network operators, hampering the prospect of overall network management and performance optimization. To address this challenge, we investigate multitudinous RTC traffic gathered across various locations and times. Consequently, we uncover two types of anomalous patterns pertaining to sequence numbers. To discern between factual losses and aberrations in RTP flows, i.e., to detect the unobservable losses, we curate three distinct datasets, aggregating packets into time bins and calculating multiple traffic statistics. Subsequently, we leverage Machine Learning (ML) technologies, training the algorithm on one dataset while testing the remaining two, to classify the loss presence in a bin. Despite the inherent hurdles posed by class imbalance and intricate traffic dynamics, we achieve decent outcomes (0.64 Fl-score), effectively identifying the majority of lossy bins (0.64 recall) while guaranteeing the performance for lossless scenarios (0.94 recall). Tailai Song, Paolo Garza, Michela Meo, Maurizio M. Munafò |
LANMAN | 2 |
| 2024 | DeX: Deep learning-based throughput prediction for real-time communications with emphasis on traffic eXtremesabstractRecent years have witnessed a remarkable upsurge in the global proliferation of Real-Time Communications (RTC) applications, a trend propelled by the flourishing advancement of network technologies and further amplified by the COVID-19 pandemic. Within this context, there is a burgeoning interest in the innovation of sophisticated and intelligent network infrastructures and technologies. Positioned as a promising candidate for this purpose, real-time throughput prediction emerges as a key enabler to foster network observability and offer proactive functions, upholding advanced system management, including but not limited to, bandwidth allocation and adaptive streaming. Nonetheless, existing methodologies struggle with predicting extreme conditions of throughput, notably peaks, valleys, and abrupt changes, that are critical in RTC traffic. To surmount these obstacles, we introduce DeX, a Deep Learning (DL)-based framework, designed to predict short-term throughput, with a dexterous proficiency and dedicated focus on navigating the complexities of traffic eXtremes. In particular, DeX leverages solely packet-level information as features and is composed of three integral components: a packet selection module that opts for an optimal subset of input features, a feature extraction block that partially incorporates the Transformer architecture, and a multi-task learning pipeline that improves the proficiency in handling traffic extremes. Moreover, our work is anchored in extensive traffic traces garnered during actual video-teleconferencing calls, and we formulate a time-series regression problem, rigorously evaluating a spectrum of technologies ranging from an adaptive filter to diverse Machine Learning (ML) and DL approaches. Initially, we aim at predicting throughput within 500-ms time windows using historical 1024 packets out of 2048, and consequently, our methodology exhibits exceptional efficacy, especially in forecasting traffic extremities. Conclusively, we conduct a series of ablation experiments and thorough analyses to showcase the enhanced performance of various scenarios, further validating the effectiveness and robustness of DeX. Tailai Song, Paolo Garza, Michela Meo, Maurizio M. Munafò |
Comput. Networks | 2 |
| 2023 | Shortlisting machine learning-based stock trading recommendations using candlestick pattern recognition
Luca Cagliero, Jacopo Fior, Paolo Garza |
Expert Syst. Appl. | 3 |
| 2023 | Density-Based Clustering by Means of Bridge Point IdentificationabstractDensity-based clustering focuses on defining clusters consisting of contiguous regions characterized by similar densities of points. Traditional approaches identify core points first, whereas more recent ones initially identify the cluster borders and then propagate cluster labels within the delimited regions. Both strategies encounter issues in presence of multi-density regions or when clusters are characterized by noisy borders. To overcome the above issues, we present a new clustering algorithm that relies on the concept of bridge point. A bridge point is a point whose neighborhood includes points of different clusters. The key idea is to use bridge points, rather than border points, to partition points into clusters. We have proved that a correct bridge point identification yields a cluster separation consistent with the expectation. To correctly identify bridge points in absence of a priori cluster information we leverage an established unsupervised outlier detection algorithm. Specifically, we empirically show that, in most cases, the detected outliers are actually a superset of the bridge point set. Therefore, to define clusters we spread cluster labels like a wildfire until an outlier, acting as a candidate bridge point, is reached. The proposed algorithm performs statistically better than state-of-the-art methods on a large set of benchmark datasets and is particularly robust to the presence of intra-cluster multiple densities and noisy borders. Luca Colomba, Luca Cagliero, Paolo Garza |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | A Dataset for Burned Area Delineation and Severity Estimation from Satellite ImageryabstractThe ability to correctly identify areas damaged by forest wildfires is essential to plan and monitor the restoration process and estimate the environmental damages after such catastrophic events. The wide availability of satellite data, combined with the recent development of machine learning and deep learning methodologies applied to the computer vision field, makes it extremely interesting to apply the aforementioned techniques to the field of automatic burned area detection. One of the main issues in such a context is the limited amount of labeled data, especially in the context of semantic segmentation. In this paper, we introduce a publicly available dataset for the burned area detection problem for semantic segmentation. The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel. Data were collected from the Sentinel-2 L2A satellite mission and the target labels were generated from the Copernicus Emergency Management Service (EMS) annotations, with five different severity levels, ranging from undamaged to completely destroyed. Finally, we report the benchmark values obtained by applying a Convolutional Neural Network on the proposed dataset to address the burned area identification problem. Luca Colomba, Alessandro Farasin, Simone Monaco, Salvatore Greco, Paolo Garza, Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
CIKM | 5 |
| 2022 | Mining spatiotemporally invariant patternsabstractDiscovering patterns that represent key spatial or temporal dependencies among data is a well-known exploratory data mining task. However, prior works either separately analyze spatial and temporal dependencies or discover joint spatiotemporal properties of specific trajectories observed over a region of interest. With the goal of generalizing the information provided by spatiotemporal patterns, in this paper we extract sequences of discrete events showing spatiotemporally invariant properties. We seek patterns whose corresponding instances in the source data differ only due to an invariant spatiotemporal transformation. We denote such a new type of patterns as SpatioTemporally Invariant. We also propose an efficient algorithm to mine STInvs and validate its efficiency and effectiveness on real data. Luca Colomba, Luca Cagliero, Paolo Garza |
SIGSPATIAL/GIS | 3 |
| 2022 | How Much Attention Should we Pay to Mosquitoes?abstractMosquitoes are a major global health problem. They are responsible for the transmission of diseases and can have a large impact on local economies. Monitoring mosquitoes is therefore helpful in preventing the outbreak of mosquito-borne diseases. In this paper, we propose a novel data-driven approach that leverages Transformer-based models for the identification of mosquitoes in audio recordings. The task aims at detecting the time intervals corresponding to the acoustic mosquito events in an audio signal. We formulate the problem as a sequence tagging task and train a Transformer-based model using a real-world dataset collecting mosquito recordings. By leveraging the sequential nature of mosquito recordings, we formulate the training objective so that the input recordings do not require fine-grained annotations. We show that our approach is able to outperform baseline methods using standard evaluation metrics, albeit suffering from unexpectedly high false negatives detection rates. In view of the achieved results, we propose future directions for the design of more effective mosquito detection models. Moreno La Quatra, Lorenzo Vaiani, Alkis Koudounas, Luca Cagliero, Paolo Garza, Elena Baralis |
ACM Multimedia | 5 |
| 2022 | Retina: An open-source tool for flexible analysis of RTC traffic
Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò |
Comput. Networks | 4 |
| 2022 | Complementing Location-Based Social Network Data With Mobility Data: A Pattern-Based ApproachabstractLocation-Based Social Networks can be profitably exploited to characterize citizens’ activities in urban environments. However, collecting LBSN is potentially challenging due to privacy concerns, connectivity issues, and potential imbalances in LBSN service usage. We propose to complement LBSN data with mobility data in the analysis of citizens’ activities in urban areas. Unlike the explicit insights provided by LBSN users, mobility data give implicit feedback on citizens’ habits. This paper explores the spatial and temporal conditions under which user habits are coherent according to both sources and reports the most reliable common sequences of visited categories of Points-Of-Interests. To this aim, it relies on a multidimensional model in which recurrent citizens’ activities are described by a new pattern type, namely the generalized activity pattern. It also detects the eventual presence of bias between LBSN and mobility user activities by customizing the established Statistical Parity metric. The motivations behind the detected bias are explained in terms of combinations of POI categories that are most likely to be the main causes. We evaluate the proposed approach on real-world data achieved from Foursquare check-ins, taxi service, and free-floating car sharing. The results highlight not only the complementarity of the data sources regarding specific POI categories, but also their interchangeability in many spatio-temporal conditions. Elena Daraio, Luca Cagliero, Silvia Chiusano, Paolo Garza |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Real-Time Classification of Real-Time CommunicationsabstractReal-time communication (RTC) applications have become largely popular in the last decade with the spread of broadband and mobile Internet access. Nowadays, these platforms are a fundamental means for connecting people and supporting businesses that increasingly rely on forms of remote work. In this context, it is of paramount importance to operate at the network level to ensure adequate Quality of Experience (QoE) for users, and appropriate traffic management policies are essential to prioritize RTC traffic. This in turn requires the network to be able to identify RTC streams and the type of content they carry. In this paper, we propose a machine learning-based application to classify media streams generated by RTC applications encapsulated in Secure Real-Time Protocol (SRTP) flows in real-time. Using carefully tuned features extracted from packet characteristics, we train models to classify streams into a variety of classes, including media type (audio/video), video quality, and redundant streams. We validate our approach using traffic from over 62 hours of multi-party meetings conducted using two popular RTC applications, namely Cisco Webex Teams and Jitsi Meet. We achieve an overall accuracy of 96% for Webex and 95% for Jitsi, using a lightweight decision tree model that makes decisions based solely on 1 second of real-time traffic. Our results show that models trained for a particular meeting software have difficulty when used with another one, although domain adaptation techniques facilitate the transfer of pre-trained models. Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò, Giovanna Carofiglio |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2021 | Online Classification of RTC TrafficabstractReal-time communication (RTC) platforms have become increasingly popular in the last decade, together with the spread of broadband Internet access. They are nowadays a fundamental means for connecting people and supporting the economy, which relies more and more on forms of remote working. In this context, it is particularly important to act at the network level to ensure adequate Quality of Experience (QoE) to users, where proper traffic management policies are essential to prioritize RTC traffic. This, in turn, requires in-network devices to identify RTC streams and the type of content they carry. In this paper, we propose a machine learning-based application to classify, in real-time, the media streams generated by RTC applications encapsulated in Secure Real Time Protocol (SRTP) flows. Using carefully tuned features extracted from packet characteristics, we train a model to classify streams into an ample set of classes, including media type (audio/video), video quality and redundant streams. To validate our approach, we use traffic from more than 88 hours of multi-party meeting calls made using the Cisco Webex Teams application. We reach an overall accuracy of 97% with a light-weight decision tree model, which makes decisions using only 1 second of traffic. Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò, Giovanna Carofiglio |
CCNC | 4 |
| 2021 | DBSCOUT: A Density-based Method for Scalable Outlier Detection in Very Large DatasetsabstractRecent technological advancements have enabled generating and collecting huge amounts of data in a daily manner. This data is used for different purposes that may impact us on an unprecedented scale. Understanding the data, including detecting its outliers, is a critical step before utilizing it.Outlier detection has been studied well in the literature but the existing approaches fail to scale to these very large settings. In this paper, we propose DBSCOUT, an efficient exact algorithm for outlier detection with a linear complexity that can run in parallel over multiple independent machines, making it a fit for the settings with billions of tuples. Besides the theoretical analysis, our experiment results confirm orders of magnitude improvement over the existing work, proving the efficiency, scalability, and effectiveness of our approach. Matteo Corain, Paolo Garza, Abolfazl Asudeh |
ICDE | 2 |
| 2020 | Improving Wildfire Severity Classification of Deep Learning U-Nets from Satellite ImagesabstractUncontrolled wildfires are dangerous events capable of harming people safety. To contrast their increasing impact in recent years, a key task is an accurate detection of the affected areas and their damage assessment from satellite images. Current state-of-the-art solutions address such problem through a double convolutional neural network able to automatically detect wildfires in satellite acquisitions and associate a damage index from a defined scale. However, such deep-learning model performance is strongly dependent on many factors. In this work, we specifically focus on a key parameter, i.e., the loss function, exploited in the underlying neural networks. Besides the state-of-the-art solutions based on the Dice-MSE, among the many loss functions proposed in literature, we focus on the Binary Cross-Entropy (BCE) and the Intersection over Union (IoU), as two representatives of the distribution-based and region-based categories, respectively. Experiments show that the BCE loss function coupled with a double-step U-Net architecture provides better results than current state-of-the-art solutions on a public labeled dataset of European wildfires. Simone Monaco, Andrea Pasini, Daniele Apiletti, Luca Colomba, Paolo Garza, Elena Baralis |
IEEE BigData | 5 |
| 2020 | DSLE: A Smart Platform for Designing Data Science CompetitionsabstractDuring the last years an increasing number of university-level and post-graduation courses on Data Science have been offered. Practices and assessments need specific learning environments where learners could play with data samples and run machine learning and data mining algorithms. To foster learner engagement many closed-and open-source platforms support the design of data science competitions. However, they show limitations on the ability to handle private data, customize the analytics and evaluation processes, and visualize learners' activities and outcomes. This paper presents Data Science Lab Environment (DSLE, in short), a new open-source platform to design and monitor data science competitions. DSLE offers a easily configurable interface to share training and test data, design group works or individual sessions, evaluate the competition runs according to customizable metrics, manage public and private leaderboards, monitor participants' activities and their progress over time. The paper describes also a real experience of usage of DSLE in the context of a 1st-year M.Sc. course, which has involved around 160 students. Giuseppe Attanasio, Flavio Giobergia, Andrea Pasini, Francesco Ventura, Elena Baralis, Luca Cagliero, Paolo Garza, Daniele Apiletti, Tania Cerquitelli, Silvia Chiusano |
COMPSAC | 7 |
| 2020 | An explainable data-driven approach to web directory taxonomy mappingabstractThe spread of e-commerce and web applications has fostered the integration of cross-domain business activities. To efficiently retrieve products and services, web directories allow customers to browse multiple-level taxonomies to find specific products or services according to a predefined categorization. Providers need to periodically update web directory lists by aligning in-house taxonomies to domain-specific hierarchies coming from external sources. However, such taxonomy mapping procedures are often semi-automatic and rely on traditional word disambiguation techniques to capture the semantics behind categories and products descriptions. Hence, the flexibility and explainability of the underlying models are quite limited. This paper proposes an automated, explainable approach to web directory taxonomy mapping based on text categorization. It exploits two complementary word-based text representations: a frequency-based representation, which captures syntactic text similarities, and an embedding one, which highlights the underlying semantic relationships among words. Since the proposed solution is purely data-driven, it can be successfully applied to business domains where there is a lack of semantic models. The frequency-based text representation has shown to be particularly suitable for driving the automated taxonomy mapping procedure, whereas the embedding space has been profitably used to provide local explanations of the category assignments. Elena Daraio, Luca Cagliero, Silvia Chiusano, Paolo Garza, Giuseppe Ricupero |
KES | 4 |
| 2019 | ELSA: A Multilingual Document Summarization Algorithm Based on Frequent Itemsets and Latent Semantic AnalysisabstractSentence-based summarization aims at extracting concise summaries of collections of textual documents. Summaries consist of a worthwhile subset of document sentences. The most effective multilingual strategies rely on Latent Semantic Analysis (LSA) and on frequent itemset mining, respectively. LSA-based summarizers pick the document sentences that cover the most important concepts. Concepts are modeled as combinations of single-document terms and are derived from a term-by-sentence matrix by exploiting Singular Value Decomposition (SVD). Itemset-based summarizers pick the sentences that contain the largest number of frequent itemsets, which represent combinations of frequently co-occurring terms. The main drawbacks of existing approaches are (i) the inability of LSA to consider the correlation between combinations of multiple-document terms and the underlying concepts, (ii) the inherent redundancy of frequent itemsets because similar itemsets may be related to the same concept, and (iii) the inability of itemset-based summarizers to correlate itemsets with the underlying document concepts. To overcome the issues of both of the abovementioned algorithms, we propose a new summarization approach that exploits frequent itemsets to describe all of the latent concepts covered by the documents under analysis and LSA to reduce the potentially redundant set of itemsets to a compact set of uncorrelated concepts. The summarizer selects the sentences that cover the latent concepts with minimal redundancy. We tested the summarization algorithm on both multilingual and English-language benchmark document collections. The proposed approach performed significantly better than both itemset- and LSA-based summarizers, and better than most of the other state-of-the-art approaches. Luca Cagliero, Paolo Garza, Elena Baralis |
ACM Trans. Inf. Syst. | 2 |
| 2018 | A Density-based Preprocessing Technique to Scale Out ClusteringabstractClustering big data is a challenging task, because the majority of high-quality clustering algorithms do not scale well with respect to the data set cardinality. To tackle the scalability problem, we propose a general-purpose density-based preprocessing technique, called SCOUT, implemented in the Spark framework. It allows compacting the original data by means of a set of representative points, while still preserving the original data distribution and density information. This small set of representative points may become the input to almost any clustering algorithm. Thus, also complex, high-quality in-memory algorithms can be applied. A thorough experimental evaluation shows that the proposed approach is efficient and at the same time effective. Elena Baralis, Paolo Garza, Eliana Pastor |
IEEE BigData | 2 |
| 2018 | Characterizing unpredictable patterns in Wireless Sensor Network data
Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Antonio Attanasio |
Inf. Sci. | 4 |
| 2017 | SQL versus NoSQL databases for geospatial applicationsabstractIn the last years, we are witnessing an increasing availability of geolocated data, ranging from satellite images to user generated content (e.g., tweets). This big amount of data is exploited by several cloud-based applications to deliver effective and customized services to end users. In order to provide a good user experience, a low-latency response time is needed, both when data are retrieved and provided. To achieve this goal, current geospatial applications need to exploit efficient and scalable geospatial databases, the choice of which has a high impact on the overall performance of the deployed applications. In this paper, we compare, from a qualitative point of view, four state-of-the-art SQL and NoSQL databases with geospatial features, and then we analyze the performances of two of them, selecting the ones based on the Database-as-a-service (DBaaS) model: Azure SQL Database and Azure DocumentDB (i.e., an SQL database versus a NoSQL one). The empirical evaluation shows pros and cons of both solutions and it is performed on a real use case related to an emergency management application. Elena Baralis, Andrea Dalla Valle, Paolo Garza, Claudio Rossi 0003, Francesco Scullino |
IEEE BigData | 3 |
| 2017 | Planning stock portfolios by means of weighted frequent itemsets
Elena Baralis, Luca Cagliero, Paolo Garza |
Expert Syst. Appl. | 3 |
| 2017 | Discovering profitable stocks for intraday trading
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Fabio Pulvirenti |
Inf. Sci. | 4 |
| 2016 | SaFe-NeC: A scalable and flexible system for network data characterizationabstractNowadays, large volumes of data and measurements are being continuously generated by computer and telecommunication networks, but such volumes make it difficult to extract meaningful knowledge from them. This paper presents SaFe-NeC, an innovative methodology for analyzing network traffic by exploiting data mining techniques, i.e. clustering and classification algorithms, focusing on self-learning capabilities of state-of-the-art scalable approaches. Self-learning algorithms, coupled with self-assessment indicators and domain-driven semantics enriching data mining results, are able to build a model of the data with minimal user intervention and highlight possibly meaningful interpretations to domain experts. Furthermore, a self-evolving model evaluation phase is included to continuously track the quality degradation of the model itself, whose rebuilding is triggered as soon as quality indicators fall below a threshold of tolerance. The proposed methodology can exploit the computational advantages of distributed computing frameworks, as the current implementation runs on Apache Spark. Preliminary experimental results on a real traffic dataset show the full potential of the proposed methodology to characterize network traffic data. Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Luca Venturini |
NOMS | 4 |
| 2016 | Modeling Correlations among Air Pollution-Related Data through Generalized Association RulesabstractToday's citizens and city administrations have an increasing interest in monitoring the air quality in urban areas. Studying the causes of air pollution entails analyzing the correlations between heterogeneous data, among which pollutant concentrations, traffic flow measurements, and meteorological data. To this end, innovative data analytics solutions able to acquire, integrate, and analyze very large amounts of data are needed. This paper presents a new data mining system, named GEneralized Correlation analyzer of pOllution data (GECKO), to discover interesting and multiple-level correlations among a large variety of open air pollution-related data. Specifically, correlations among pollutant levels and traffic and climate conditions are discovered and analyzed at different abstraction levels. The knowledge extraction process is driven by a taxonomy to generalize low-level measurement values as the corresponding categories. To ease the manual inspection of the result, the extracted correlations are classified into few classes based on the semantics of underlying data. The experiments, performed on real data acquired in a major Italian Smart City, demonstrate the effectiveness of the proposed analytics engine in discovering correlations among pollutant data that are potentially useful for supporting city administrators in decision-making. Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Giuseppe Ricupero, Xin Xiao 0002 |
SMARTCOMP | 4 |
| 2016 | Characterization and search of web services through intensional knowledge
Devis Bianchini, Paolo Garza, Elisa Quintarelli |
J. Intell. Inf. Syst. | 2 |
| 2016 | SeLINA: A Self-Learning Insightful Network AnalyzerabstractUnderstanding the behavior of a network from a large scale traffic dataset is a challenging problem. Big data frameworks offer scalable algorithms to extract information from raw data, but often require a sophisticated fine-tuning and a detailed knowledge of machine learning algorithms. To streamline this process, we propose self-learning insightful network analyzer (SeLINA), a generic, self-tuning, simple tool to extract knowledge from network traffic measurements. SeLINA includes different data analytics techniques providing self-learning capabilities to state-of-the-art scalable approaches, jointly with parameter auto-selection to off-load the network expert from parameter tuning. We combine both unsupervised and supervised approaches to mine data with a scalable approach. SeLINA embeds mechanisms to check if the new data fits the model, to detect possible changes in the traffic, and to, possibly automatically, trigger model rebuilding. The result is a system that offers human-readable models of the data with minimal user intervention, supporting domain experts in extracting actionable knowledge and highlighting possibly meaningful interpretations. SeLINA's current implementation runs on Apache Spark. We tested it on large collections of real-world passive network measurements from a nationwide ISP, investigating YouTube, and P2P traffic. The experimental results confirmed the ability of SeLINA to provide insights and detect changes in the data that suggest further analyses. Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Danilo Giordano, Marco Mellia, Luca Venturini |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2015 | Digging deep into weighted patient data through multiple-level patterns
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza |
Inf. Sci. | 5 |
| 2015 | Pattern set mining with schema-based constraint
Luca Cagliero, Silvia Chiusano, Paolo Garza, Giulia Bruno |
Knowl. Based Syst. | 3 |
| 2015 | An Expert CAD Flow for Incremental Functional Diagnosis of Complex Electronic BoardsabstractFunctional diagnosis for complex systems can be a very time-consuming and expensive task, trying to identify the source of an observed misbehavior. We propose an automatic incremental diagnostic methodology and CAD flow, based on data mining (DM). It is a model-based approach that incrementally determines the tests to be executed to isolate the faulty component, aiming at minimizing the total number of executed tests, without compromising 100% diagnostic accuracy. The DM engine allows for shorter test sequences with respect to other reasoning-based solutions (e.g., Bayesian belief networks), not requiring complex pre and post-conditions management. Experimental results on a large set of synthetic examples and on three industrial boards substantiate the quality of the proposed approach. Cristiana Bolchini, Luca Cassano, Paolo Garza, Elisa Quintarelli, Fabio Salice |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | MeTA: Characterization of Medical Treatments at Different Abstraction LevelsabstractPhysicians and health care organizations always collect large amounts of data during patient care. These large and high-dimensional datasets are usually characterized by an inherent sparseness. Hence, analyzing these datasets to figure out interesting and hidden knowledge is a challenging task. This article proposes a new data mining framework based on generalized association rules to discover multiple-level correlations among patient data. Specifically, correlations among prescribed examinations, drugs, and patient profiles are discovered and analyzed at different abstraction levels. The rule extraction process is driven by a taxonomy to generalize examinations and drugs into their corresponding categories. To ease the manual inspection of the result, a worthwhile subset of rules (i.e., nonredundant generalized rules) is considered. Furthermore, rules are classified according to the involved data features (medical treatments or patient profiles) and then explored in a top-down fashion: from the small subset of high-level rules, a drill-down is performed to target more specific rules. The experiments, performed on a real diabetic patient dataset, demonstrate the effectiveness of the proposed approach in discovering interesting rule groups at different abstraction levels. Dario Antonelli, Elena Baralis, Giulia Bruno, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Naeem Ahmed Mahoto |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2015 | MWI-Sum: A Multilingual Summarizer Based on Frequent Weighted ItemsetsabstractMultidocument summarization addresses the selection of a compact subset of highly informative sentences, i.e., the summary, from a collection of textual documents. To perform sentence selection, two parallel strategies have been proposed: (a) apply general-purpose techniques relying on data mining or information retrieval techniques, and/or (b) perform advanced linguistic analysis relying on semantics-based models (e.g., ontologies) to capture the actual sentence meaning. Since there is an increasing need for processing documents written in different languages, the attention of the research community has recently focused on summarizers based on strategy (a). This article presents a novel multilingual summarizer, namely MWI-Sum (Multilingual Weighted Itemset-based Summarizer), that exploits an itemset-based model to summarize collections of documents ranging over the same topic. Unlike previous approaches, it extracts frequent weighted itemsets tailored to the analyzed collection and uses them to drive the sentence selection process. Weighted itemsets represent correlations among multiple highly relevant terms that are neglected by previous approaches. The proposed approach makes minimal use of language-dependent analyses. Thus, it is easily applicable to document collections written in different languages. Experiments performed on benchmark and real-life collections, English-written and not, demonstrate that the proposed approach performs better than state-of-the-art multilingual document summarizers. Elena Baralis, Luca Cagliero, Alessandro Fiori, Paolo Garza |
ACM Trans. Inf. Syst. | 4 |
| 2014 | Misleading Generalized Itemset Mining in the CloudabstractIn the era of smart cities huge data volumes are continuously generated and collected, thus prompting the need for efficient and distributed data mining approaches. Generalized itemset mining is an established data mining technique, which entails the discovery of multiple-level patterns hidden in the analyzed data by exploiting analyst-provided taxonomies. Among the generalized itemsets, the most peculiar high-level patterns are those with many contrasting correlations among items at different abstraction levels. They represent misleading situations that are worth analyzing separately by experts during manual inspection. This paper proposes a novel cloud-based service, named MGI-CLOUD, to efficiently mine misleading multiple-level patterns, i.e., the Misleading Generalized Itemsets, on a distributed computing environment. MGI-CLOUD consists of a set of distributed MapReduce jobs running in the cloud. As a case study, the system has been contextualized in a real-life scenario, i.e., the analysis of traffic law infractions committed in a smart city environment. The experiments, performed on real datasets, demonstrate the efficiency and effectiveness of MGI-CLOUD. Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Luigi Grimaudo, Fabio Pulvirenti |
ISPA | 5 |
| 2014 | Misleading Generalized Itemset discovery
Luca Cagliero, Tania Cerquitelli, Paolo Garza, Luigi Grimaudo |
Expert Syst. Appl. | 3 |
| 2014 | Expressive generalized itemsets
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Vincenzo D'Elia, Paolo Garza |
Inf. Sci. | 5 |
| 2014 | Twitter data analysis by means of Strong Flipping Generalized Itemsets
Luca Cagliero, Tania Cerquitelli, Paolo Garza, Luigi Grimaudo |
J. Syst. Softw. | 3 |
| 2014 | Infrequent Weighted Itemset Mining Using Frequent Pattern GrowthabstractFrequent weighted itemsets represent correlations frequently holding in data in which items may weight differently. However, in some contexts, e.g., when the need is to minimize a certain cost function, discovering rare data correlations is more interesting than mining frequent ones. This paper tackles the issue of discovering rare and weighted itemsets, i.e., the infrequent weighted itemset (IWI) mining problem. Two novel quality measures are proposed to drive the IWI mining process. Furthermore, two algorithms that perform IWI and Minimal IWI mining efficiently, driven by the proposed measures, are presented. Experimental results show efficiency and effectiveness of the proposed approach. Luca Cagliero, Paolo Garza |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Hadoop on a Low-Budget General Purpose HPC Cluster in Academia
Paolo Garza, Paolo Margara, Nicolò Nepote, Luigi Grimaudo, Elio Piccolo |
ADBIS (2) | 1 |
| 2013 | Frequent weighted itemset mining from gene expression dataabstractGene Expression Datasets (GEDs) usually consist of the expression values of thousands of genes within hundreds of samples. Frequent itemset and association rule mining algorithms have been applied to discover significant co-expressions among multiple genes from GEDs. To perform these data analyses, gene expression values are commonly discretized into a predefined number of bins. Such an expert-driven and not trivial preprocessing step could bias the quality of the mining result. This paper presents a novel approach to discovering gene correlations from GEDs which does not require data discretization. By representing per-sample gene expression values as item weights, frequent weighted itemsets can be extracted. The discovery of weighted itemsets instead of traditional (not weighted) ones prevents experts from discretizing GEDs before analyzing them and thus improves the effectiveness of the knowledge discovery process. Experiments performed on real GEDs demonstrate the effectiveness of the proposed approach. Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza |
BIBE | 5 |
| 2013 | Improving classification models with taxonomy information
Luca Cagliero, Paolo Garza |
Data Knowl. Eng. | 2 |
| 2013 | Itemset generalization with cardinality-based constraints
Luca Cagliero, Paolo Garza |
Inf. Sci. | 2 |
| 2013 | EnBay: A Novel Pattern-Based Bayesian ClassifierabstractA promising approach to Bayesian classification is based on exploiting frequent patterns, i.e., patterns that frequently occur in the training data set, to estimate the Bayesian probability. Pattern-based Bayesian classification focuses on building and evaluating reliable probability approximations by exploiting a subset of frequent patterns tailored to a given test case. This paper proposes a novel and effective approach to estimate the Bayesian probability. Differently from previous approaches, the Entropy-based Bayesian classifier, namely EnBay, focuses on selecting the minimal set of long and not overlapped patterns that best complies with a conditional-independence model, based on an entropy-based evaluator. Furthermore, the probability approximation is separately tailored to each class. An extensive experimental evaluation, performed on both real and synthetic data sets, shows that EnBay is significantly more accurate than most state-of-the-art classifiers, Bayesian and not. Elena Baralis, Luca Cagliero, Paolo Garza |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | I-prune: Item selection for associative classificationabstractAssociative classification is characterized by accurate models and high model generation time. Most time is spent in extracting and postprocessing a large set of irrelevant rules, which are eventually pruned. We propose I-prune, an item-pruning approach that selects uninteresting items by means of an interestingness measure and prunes them as soon as they are detected. Thus, the number of extracted rules is reduced and model generation time decreases correspondingly. A wide set of experiments on real and synthetic data sets has been performed to evaluate I-prune and select the appropriate interestingness measure. The experimental results show that I-prune allows a significant reduction in model generation time, while increasing (or at worst preserving) model accuracy. Experimental evaluation also points to the chi-square measure as the most effective interestingness measure for item pruning. © 2012 Wiley Periodicals, Inc. Elena Baralis, Paolo Garza |
Int. J. Intell. Syst. | 2 |
| 2012 | Generalized association rule mining with constraints
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza |
Inf. Sci. | 4 |
| 2011 | Structured data classification by means of matrix factorizationabstractSingular Value Decomposition (SVD) has been extensively used in the classification context as a preprocessing step aiming to reduce the number of features of the input space. Traditional classification algorithms are then applied on the new space to generate accurate models. In this paper, we propose a different use of SVD. In our approach SVD is the building block of a new classification algorithm, called CMF, and not that of a feature reduction algorithm. In particular, we propose a new classification algorithm where the classification model corresponds to the k largest right singular vectors of the factorization of the training dataset obtained by applying SVD. The selected singular vectors allows representing the main "characteristics" of the training data and can be used to provide accurate predictions. The experiments performed on 15 structured UCI datasets show that CMF is efficient and, despite its simplicity, it is more accurate than many state of the art classification algorithms. Paolo Garza |
CIKM | 1 |
| 2011 | Semi-Automatic Ontology Construction by Exploiting Functional Dependencies and Association RulesabstractThis paper presents a novel semi-automatic approach to construct conceptual ontologies over structured data by exploiting both the schema and content of the input dataset. It effectively combines two well-founded database and data mining techniques, i.e., functional dependency discovery and association rule mining, to support domain experts in the construction of meaningful ontologies, tailored to the analyzed data, by using Description Logic (DL). To this aim, functional dependencies are first discovered to highlight valuable conceptual relationships among attributes of the data schema (i.e., among concepts). The set of discovered correlations effectively support analysts in the assertion of the Tbox ontological statements (i.e., the statements involving shared data conceptualizations and their relationships). Then, the analyst-validated dependencies are exploited to drive the association rule mining process. Association rules represent relevant and hidden correlations among data content and they are used to provide valuable knowledge at the instance level. The pushing of functional dependency constraints into the rule mining process allows analysts to look into and exploit only the most significant data item recurrences in the assertion of the Abox ontological statements (i.e., the statements involving concept instances and their relationships). Luca Cagliero, Tania Cerquitelli, Paolo Garza |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2011 | CAS-Mine: providing personalized services in context-aware applications by means of generalized rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti |
Knowl. Inf. Syst. | 4 |
| 2010 | TOD: Temporal outlier detection by using quasi-functional temporal dependencies
Giulia Bruno, Paolo Garza |
Data Knowl. Eng. | 2 |
| 2009 | Context-Aware User and Service Profiling by Means of Generalized Association Rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti |
KES (2) | 4 |
| 2008 | A Lazy Approach to Associative ClassificationabstractAssociative classification is a promising technique to build accurate classifiers. However, in large or correlated datasets, association rule mining may yield huge rule sets. Hence, several pruning techniques have been proposed to select a small subset of high quality rules. We argue that rule pruning should be reduced to a minimum, since the availability of a "rich" rule set may improve the accuracy of the classifier. The L^3 associative classifier is built by means of a lazy pruning technique which discards exclusively rules that only misclassify training data. Classification of unlabeled data is performed in two steps. A small subset of high quality rules is first considered. When this set is not able to classify the data, a larger rule set is exploited. This second set includes rules usually discarded by previous approaches. To cope with the need of mining large rule sets and efficiently use them for classification, a compact form is proposed to represent a complete rule set in a space-efficient way and without information loss. An extensive experimental evaluation on real and synthetic datasets shows that L^3 improves the classification accuracy with respect to previous approaches. Elena Baralis, Silvia Chiusano, Paolo Garza |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2007 | Answering XML queries by means of data summariesabstractXML is a rather verbose representation of semistructured data, which may require huge amounts of storage space. We propose a summarized representation of XML data, based on the concept of instance pattern, which can both provide succinct information and be directly queried. The physical representation of instance patterns exploits itemsets or association rules to summarize the content of XML datasets. Instance patterns may be used for (possibly partially) answering queries, either when fast and approximate answers are required, or when the actual dataset is not available, for example, it is currently unreachable. Experiments on large XML documents show that instance patterns allow a significant reduction in storage space, while preserving almost entirely the completeness of the query result. Furthermore, they provide fast query answers and show good scalability on the size of the dataset, thus overcoming the document size limitation of most current XQuery engines. Elena Baralis, Paolo Garza, Elisa Quintarelli, Letizia Tanca |
ACM Trans. Inf. Syst. | 2 |
| 2003 | Majority Classification by Means of Association Rules
Elena Baralis, Paolo Garza |
PKDD | 2 |
| 2002 | A Lazy Approach to Pruning Classification RulesabstractAssociative classification is a promising technique for the generation of highly precise classifiers. Previous works propose several clever techniques to prune the huge set of generated rules, with the twofold aim of selecting a small set of high quality rules, and reducing the chance of overfitting. In this paper, we argue that pruning should be reduced to a minimum and that the availability of a large rule base may improve the precision of the classifier without affecting its performance. In L/sup 3/ (Live and Let Live), a new algorithm for associative classification, a lazy pruning technique iteratively discards all rules that only yield wrong case classifications. Classification is performed in two steps. Initially, rules which have already correctly classified at least one training case, sorted by confidence, are considered If the case is still unclassified, the remaining rules (unused during the training phase) are considered, again sorted by confidence. Extensive experiments on 26 databases from the UCI machine learning database repository show that L/sup 3/ improves the classification precision with respect to previous approaches. Elena Baralis, Paolo Garza |
ICDM | 2 |