Elena Baralis

dblp:77/5283 · also Elena Maria Baralis · DBLP profile ↗
← Back
107ranked-venue papers
50as first author
30since 2021 · last 2026
0000-0001-9231-467XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 54 · 34 first-author · 12 since 2021Artificial intelligence and machine learning · 44 · 16 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 10 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Computer networks · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 On the Evaluation of Machine Unlearning Methods: A Multi-domain Classification Benchmark
abstract
Abstract Machine Unlearning (MU), the process of removing specific data influences from trained machine learning models, is critical for regulatory compliance (e.g., GDPR’s right to be forgotten) and for addressing copyright and privacy concerns in large-scale models. While a wide range of methods and metrics have been proposed, systematic evaluations remain fragmented, typically limited in scope by modality, metric coverage, or the number of methods considered. Moreover, the lack of standardized benchmarks leaves several gaps in evaluation protocols, including how to efficiently compare methods, identify optimal hyperparameters, and determine which experimental settings are appropriate for fair and meaningful benchmarking. To address these gaps, we present the most comprehensive MU benchmark to date, evaluating 12 unlearning methods across 8 classification datasets, 4 modalities, several hyperparameters and settings. Based on previous literature and our empirical results, we formalize evaluation protocol desiderata to guide future MU benchmarking. Following these guidelines, we report benchmark results highlighting the best methods within and across domains. To help with method comparison, we also introduce LUMA, a unified metric that aggregates core unlearning dimensions into a single score. Our code is reproducible and extensible to serve as a benchmark for MU research.
Andrea D'Angelo, Claudio Savelli, Flavio Giobergia, Elena Baralis, Giovanni Stilo
Mach. Learn.4
2025 ERASURE: A Modular and Extensible Framework for Machine Unlearning
abstract
Machine Unlearning (MU) is an emerging research area that enables models to selectively forget specific data, a critical requirement for privacy compliance (e.g., GDPR, CCPA) and security. However, the lack of standardized benchmarks makes evaluating and developing unlearning methods difficult. To address this gap, we introduce ERASURE, a benchmarking and development framework designed to systematically assess MU techniques. ERASURE provides a modular, extensible, open-source environment with real-world datasets and standardized unlearning measures. The framework is designed with configuration-driven workflows and an inversion of control architecture, allowing integration of new datasets, models, and evaluation measures. ERASURE advances trustworthy AI research as a tool for researchers to develop and benchmark new MU methods.
Andrea D'Angelo, Claudio Savelli, Gabriele Tagliente, Flavio Giobergia, Elena Baralis, Giovanni Stilo
CIKM5
2025 HydroChronos: Forecasting Decades of Surface Water Change
abstract
Forecasting surface water dynamics is crucial for water resource management and climate change adaptation. However, the field lacks comprehensive datasets and standardized benchmarks. In this paper, we introduce HydroChronos, a large-scale, multi-modal spatiotemporal dataset for surface water dynamics forecasting designed to address this gap. We couple the dataset with three forecasting tasks. The dataset includes over three decades of aligned Landsat 5 and Sentinel-2 imagery, climate data, and Digital Elevation Models for diverse lakes and rivers across Europe, North America, and South America. We also propose AquaClimaTempo UNet, a novel spatiotemporal architecture with a dedicated climate data branch, as a strong benchmark baseline. Our model significantly outperforms a Persistence baseline for forecasting future water dynamics by +14% and +11% F1 across change detection and direction of change classification tasks, and by +0.1 MAE on the magnitude of change regression. Finally, we conduct an Explainable AI analysis to identify the key climate variables and input channels that influence surface water change, providing insights to inform and guide future modeling efforts.
Daniele Rege Cambrin, Eleonora Poeta, Eliana Pastor, Isaac Corley, Tania Cerquitelli, Elena Baralis, Paolo Garza
SIGSPATIAL/GIS6
2025 voc2vec: A Foundation Model for Non-Verbal Vocalization
abstract
Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various real-world applications. Audio foundation models well handle non-speech data but also fail to capture the nuanced features of non-verbal human sounds. In this work, we aim to overcome the above shortcoming and propose a novel foundation model, termed voc2vec, specifically designed for non-verbal human data leveraging exclusively open-soruce non-verbal audio datasets. We employ a collection of 10 datasets covering around 125 hours of non-verbal audio. Experimental results prove that voc2vec is effective in non-verbal vocalization classification, and it outperforms conventional speech and audio foundation models. Moreover, voc2vec consistently outperforms strong baselines, namely OpenSmile and emotion2vec, on six different benchmark datasets. To the best of the authors’ knowledge, voc2vec is the first universal representation model for vocalization tasks.
Alkis Koudounas, Moreno La Quatra, Sabato Marco Siniscalchi, Elena Baralis
ICASSP4
2025 How to Make Reproducible Research in Machine Unlearning with ERASURE
abstract
Machine unlearning, the process of removing specific data influences from Machine Learning models, is critical for complying with regulations like the GDPR's right to be forgotten and addressing copyright disputes in large models. Despite its rising importance, the field still lacks standardized tools, hindering reproducibility and evaluation. Here, we present, in an extensive way, ERASURE, a unified framework enabling reproducibility by implementing common unlearning techniques, evaluation metrics, and dedicated datasets. ERASURE advances research, ensures solution comparability, and facilitates reproducibility, addressing future legal and ethical challenges in data management.
Andrea D'Angelo, Claudio Savelli, Gabriele Tagliente, Flavio Giobergia, Elena Baralis, Giovanni Stilo
IJCAI5
2025 MVP: Multi-source Voice Pathology detection
abstract
Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces MVP (Multi-source Voice Pathology detection), a novel approach that leverages transformers operating directly on raw voice signals. We explore three fusion strategies to combine sentence reading and sustained vowel recordings: waveform concatenation, intermediate feature fusion, and decision-level combination. Empirical validation across the German, Portuguese, and Italian languages shows that intermediate feature fusion using transformers best captures the complementary characteristics of both recording types. Our approach achieves up to +13% AUC improvement over single-source methods.
Alkis Koudounas, Moreno La Quatra, Gabriele Ciravegna, Marco Fantini, Erika Crosetti, Giovanni Succo, Tania Cerquitelli, Sabato Marco Siniscalchi, Elena Baralis
INTERSPEECH9
2025 "KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
abstract
Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional neural architectures, yet their application to speech processing remains under explored. This work presents the first investigation of KANs for Spoken Language Understanding (SLU) tasks. We experiment with 2D-CNN models on two datasets, integrating KAN layers in five different configurations within the dense block. The best-performing setup, which places a KAN layer between two linear layers, is directly applied to transformer-based models and evaluated on five SLU datasets with increasing complexity. Our results show that KAN layers can effectively replace the linear layers, achieving comparable or superior performance in most cases. Finally, we provide insights into how KAN and linear layers on top of transformers differently attend to input regions of the raw waveforms.
Alkis Koudounas, Moreno La Quatra, Eliana Pastor, Sabato Marco Siniscalchi, Elena Baralis
INTERSPEECH5
2025 "Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
Alkis Koudounas, Claudio Savelli, Flavio Giobergia, Elena Baralis
INTERSPEECH4
2025 Detecting Interpretable Subgroup Drifts
Flavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena Baralis
KDD (1)4
2025 MAD: Multicriteria Anomaly Detection of Suspicious Financial Accounts from Billions of Cash Transactions
abstract
This paper presents a real-world deployment case study on using unsupervised anomaly detection for Anti-Money Laundering (AML).Using more than 2 billion anonymized bank transactions that Intesa Sanpaolo, a primary Italian financial institution, registered over 8 months, we developed, tuned and deployed a machine learning pipeline in production.Experts from Intesa Sanpaolo validated the performance of our approach against the institution's traditional rule-based system and checked new real-world cases the system allowed them to identify.Besides increasing both precision and recall by a factor of 6 in the detection of high-risk cases, our pipeline raises 200+ additional alerts during the 8-month period, manually identified by branch managers, but missed by the rulebased system.More importantly, a manual inspection of 100 new unseen cases revealed 28 significant previously unreported cases.The pipeline, now fully deployed in Intesa Sanpaolo's Transaction Monitoring system, highlights the advantages of machine learning over traditional approaches typically adopted in this traditionally very conservative sector.
Giordano Paoletti, Flavio Giobergia, Danilo Giordano, Luca Cagliero, Silvia Ronchiadin, Dario Moncalvo, Marco Mellia, Elena Baralis
KDD (2)8
2024 Speech Analysis of Language Varieties in Italy
abstract
Italy exhibits rich linguistic diversity across its territory due to the distinct regional languages spoken in different areas. Recent advances in self-supervised learning provide new opportunities to analyze Italy’s linguistic varieties using speech data alone. This includes the potential to leverage representations learned from large amounts of data to better examine nuances between closely related linguistic varieties. In this study, we focus on automatically identifying the geographic region of origin of speech samples drawn from Italy’s diverse language varieties. We leverage self-supervised learning models to tackle this task and analyze differences and similarities between Italy’s regional languages. In doing so, we also seek to uncover new insights into the relationships among these diverse yet closely related varieties, which may help linguists understand their interconnected evolution and regional development over time and space. To improve the discriminative ability of learned representations, we evaluate several supervised contrastive learning objectives, both as pre-training steps and additional fine-tuning objectives. Experimental evidence shows that pre-trained self-supervised models can effectively identify regions from speech recording. Additionally, incorporating contrastive objectives during fine-tuning improves classification accuracy and yields embeddings that distinctly separate regional varieties, demonstrating the value of combining self-supervised pre-training and contrastive learning for this task.
Moreno La Quatra, Alkis Koudounas, Elena Baralis, Sabato Marco Siniscalchi
LREC/COLING3
2024 Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic Features
abstract
Eliana Pastor, Alkis Koudounas, Giuseppe Attanasio, Dirk Hovy, Elena Baralis. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Eliana Pastor, Alkis Koudounas, Giuseppe Attanasio, Dirk Hovy, Elena Baralis
EACL (1)5
2024 Ainur: Harmonizing Speed and Quality in Deep Music Generation Through Lyrics-Audio Embeddings
abstract
In the domain of music generation, prevailing methods focus on text-to-music tasks, predominantly relying on diffusion models. However, they fail to achieve good vocal quality in synthetic music compositions.To tackle this critical challenge, we present Ainur, a hierarchical diffusion model that concentrates on the lyrics-to-music generation task. Through its use of multimodal Lyrics-Audio Spectrogram Pre-training (CLASP) embeddings, Ainur distinguishes itself from past approaches by specifically enhancing the vocal quality of synthetically produced music. Notably, Ainur’s training and testing processes are highly efficient, requiring only a single GPU. According to experimental results, Ainur meets or exceeds the quality of other state-of-the-art models like MusicGen, MusicLM, and AudioLDM2 in both objective and subjective evaluations. Additionally, Ainur offers near real-time inference speed, which facilitate its use in practical, real-world applications.
Giuseppe Concialdi, Alkis Koudounas, Eliana Pastor, Barbara Di Eugenio, Elena Baralis
ICASSP5
2024 Prioritizing Data Acquisition for end-to-end Speech Model Improvement
abstract
As speech processing moves toward more data-hungry models, data selection and acquisition become crucial to building better systems. Recent efforts have championed quantity over quality, following the mantra "The more data, the better." However, not every data brings the same benefit. This paper proposes a data acquisition solution that yields better models with less data – and lower cost. Given a model, a task, and an objective to maximize, we propose a process with three steps. First, we assess the model’s baseline performance on the task. Second, we use efficient mining techniques to identify subgroups that maximize the target objective if acquired first as new samples. Being the subgroups interpretable, we can determine which samples to acquire. Third, we run incremental training sampling from those subgroups. Experiments with two state-of-the-art speech models for Intent Classification across two datasets in English and Italian show that our method is significantly better than random or complete acquisition and clustering-based techniques.
Alkis Koudounas, Eliana Pastor, Giuseppe Attanasio, Luca de Alfaro, Elena Baralis
ICASSP5
2024 Voice Disorder Analysis: a Transformer-based Approach
Alkis Koudounas, Gabriele Ciravegna, Marco Fantini, Erika Crosetti, Giovanni Succo, Tania Cerquitelli, Elena Baralis
INTERSPEECH7
2024 A Contrastive Learning Approach to Mitigate Bias in Speech Models
Alkis Koudounas, Flavio Giobergia, Eliana Pastor, Elena Baralis
INTERSPEECH4
2024 Towards Comprehensive Subgroup Performance Analysis in Speech Models
abstract
The evaluation of spoken language understanding (SLU) systems is often restricted to assessing their global performance or examining predefined subgroups of interest. However, a more detailed analysis at the subgroup level has the potential to uncover valuable insights into how speech system performance differs across various subgroups. In this work, we identify biased data subgroups and describe them at the level of user demographics, recording conditions, and speech targets. We propose a new task-, model- and dataset-agnostic approach to detect significant intra- and cross-model performance gaps. We detect problematic data subgroups in SLU models by leveraging the notion of subgroup divergence. We also compare the outcome of different SLU models on the same dataset and task at the subgroup level. We identify significant gaps in subgroup performance between models different in size, architecture, or pre-training objectives, including multi-lingual and mono-lingual models, yet comparable to each other in overall performance. The results, obtained on two SLU models, four datasets, and three different tasks–intent classification, automatic speech recognition, and emotion recognition–confirm the effectiveness of the proposed approach in providing a nuanced SLU model assessment.
Alkis Koudounas, Eliana Pastor, Giuseppe Attanasio, Vittorio Mazzia, Manuel Giollo, Thomas Gueudré, Elisa Reale, Luca Cagliero, Sandro Cumani, Luca de Alfaro, Elena Baralis, Daniele Amberti
IEEE ACM Trans. Audio Speech Lang. Process.11
2023 Late Fusion-based Distributed Multimodal Learning
abstract
Multimodal artificial intelligence promises deeper insights by analyzing data from diverse sources such as text, images, audio and more. However, efficiently processing and fusing large multimodal datasets remains an open challenge. This paper presents a Spark-based approach to parallelize multimodal encoding tasks. A key aspect is the use of late fusion with frozen backbone encoders, allowing encodings to be processed independently across cluster nodes. The encoded vectors can then be used for a variety of supervised and unsupervised tasks, regardless of whether they are gradient-based or not. Experimental results on image, text and audio datasets show that Spark clusters can offer competitive performance compared to GPUs, especially for I/O-intensive modalities. While GPUs outperform the Spark cluster when sufficient CPU cores are available, Spark makes it possible to use already available commodity hardware. The presented architecture demonstrates how distributed computing platforms like Spark can be effectively “repurposed” for multimodal AI, enhancing scalability and making such systems more accessible.
Flavio Giobergia, Elena Baralis
IEEE Big Data2
2023 Exploring Subgroup Performance in End-to-End Speech Models
abstract
End-to-End Spoken Language Understanding models are generally evaluated according to their overall accuracy, or separately on (a priori defined) data subgroups of interest. We propose a technique for analyzing model performance at the subgroup level, which considers all subgroups that can be defined via a given set of metadata and are above a specified minimum size. The metadata can represent user characteristics, recording conditions, and speech targets. Our technique is based on advances in model bias analysis, enabling efficient exploration of resulting subgroups. A fine-grained analysis reveals how model performance varies across sub-groups, identifying modeling issues or bias towards specific subgroups.We compare the subgroup-level performance of models based on wav2vec 2.0 and HuBERT on the Fluent Speech Commands dataset. The experimental results illustrate how subgroup-level analysis reveals a finer and more complete picture of performance changes when models are replaced, automatically identifying the subgroups that most benefit or fail to benefit from the change.
Alkis Koudounas, Eliana Pastor, Giuseppe Attanasio, Vittorio Mazzia, Manuel Giollo, Thomas Gueudré, Luca Cagliero, Luca de Alfaro, Elena Baralis, Daniele Amberti
ICASSP9
2023 A Hierarchical Approach to Anomalous Subgroup Discovery
abstract
Understanding peculiar and anomalous behavior of machine learning models for specific data subgroups is a fundamental building block of model performance and fairness evaluation. The analysis of these data subgroups can provide useful insights into model inner working and highlight its potentially discriminatory behavior. Current approaches to subgroup exploration ignore the presence of hierarchies in the data, and can only be applied to discretized attributes. The discretization process required for continuous attributes may significantly affect the identification of relevant subgroups.We propose a hierarchical subgroup exploration technique to identify anomalous subgroup behavior at multiple granularity levels, along with a technique for the hierarchical discretization of data attributes. The hierarchical discretization produces, for each continuous attribute, a hierarchy of intervals. The subsequent hierarchical exploration can exploit data hierarchies, selecting for each attribute the optimal granularity to identify subgroups that are both anomalous, and with enough elements to be statistically and practically significant. Compared to non- hierarchical approaches, we show that our hierarchical approach is more powerful in identifying anomalous subgroups and more stable with respect to discretization and exploration parameters.
Eliana Pastor, Elena Baralis, Luca de Alfaro
ICDE2
2023 On Computing Paradigms - Where Will Large Language Models Be Going
abstract
Computing generates intelligence. With this statement we do not mean computing’s capabilities of manipulating numbers, shapes, symbols, and even logics. What we mean is the ingenious design of computing structures which serve as the basis of intelligence generation during program running. In this panel discussion, we consider how to obtain such capabilities through some computing paradigms as examples, including principal computing, logic computing, discriminative computing, and generative computing. The panelists express their thoughts about the inherent advantages and disadvantages of each of these paradigms, in terms of their adaptivity, interpretability, generality and specificity, and dives into detailed discussions about Large Language Models (LLMs), a mainstream generative paradigm which leverages the strengths of large pre-trained models and downstream prompt tuning to deliver combined intelligence, superior to most existing frameworks in natural language processing. The panel outlines potential challenges of the generative paradigm, with a strong focus on LLMs, and emphasizes that future directions of such models will need to address (1) tackling bias, discrimination, and transparency challenges; (2) delivering logical answers with high specificity; (3) enabling personalized, lightweight, and rapid updating mechanisms; (4) assessing accreditation, tracing, and misusages; and (5) ensuring sustainable LLMs.
Xindong Wu 0001, Xingquan Zhu 0001, Elena Baralis, Ruqian Lu, Vipin Kumar 0001, Leszek Rutkowski
ICDM3
2023 ITALIC: An Italian Intent Classification Dataset
abstract
Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects.We introduce ITALIC, the first largescale speech dataset designed for intent classification in Italian.The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata.We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models.Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks.We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.
Alkis Koudounas, Moreno La Quatra, Lorenzo Vaiani, Luca Colomba, Giuseppe Attanasio, Eliana Pastor, Luca Cagliero, Elena Baralis
INTERSPEECH8
2022 Legal Entity Disambiguation for Financial Crime Detection
abstract
Transaction Monitoring is one of the main labor-intensive tasks of anti-financial crime and it requires to scrutinise billions of transactions per month against possible crimes. The first step in the process is the correct identification of the involved parties. This foundational step defines the focal entities on which transaction monitoring algorithms rely to spot suspicious events. Unfortunately, the loose syntax of protocols and the free text fields of inter-banking communications make party disambiguation particularly challenging. The first step of a fully automated data-driven strategy is thus the detection of the actual entity owning or using a given account.In this paper, we leverage data-driven techniques to identify and disambiguate the owners of accounts involved in cross-border international transactions when a Financial Institution only knows a minority fraction of such parties as its own customers. For this, we propose a data science pipeline relying on hierarchical clustering to capture similarities among names of parties involved in actual transactions. We test and tune the proposed approach using a large, real-world, multi-language, proprietary dataset of actual international transactions. Our highly parallel implementation completes the identification of parties that share an account and identifies all accounts owned by a party with f-score higher than 0.8.
Jacopo Fior, Thomas Favale, Luca Cagliero, Danilo Giordano, Marco Mellia, Elena Baralis, Silvia Ronchiadin, Paolo Baracco, Dario Moncalvo
IEEE Big Data6
2022 A Dataset for Burned Area Delineation and Severity Estimation from Satellite Imagery
abstract
The ability to correctly identify areas damaged by forest wildfires is essential to plan and monitor the restoration process and estimate the environmental damages after such catastrophic events. The wide availability of satellite data, combined with the recent development of machine learning and deep learning methodologies applied to the computer vision field, makes it extremely interesting to apply the aforementioned techniques to the field of automatic burned area detection. One of the main issues in such a context is the limited amount of labeled data, especially in the context of semantic segmentation. In this paper, we introduce a publicly available dataset for the burned area detection problem for semantic segmentation. The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel. Data were collected from the Sentinel-2 L2A satellite mission and the target labels were generated from the Copernicus Emergency Management Service (EMS) annotations, with five different severity levels, ranging from undamaged to completely destroyed. Finally, we report the benchmark values obtained by applying a Convolutional Neural Network on the proposed dataset to address the burned area identification problem.
Luca Colomba, Alessandro Farasin, Simone Monaco, Salvatore Greco, Paolo Garza, Daniele Apiletti, Elena Baralis, Tania Cerquitelli
CIKM7
2022 How Much Attention Should we Pay to Mosquitoes?
abstract
Mosquitoes are a major global health problem. They are responsible for the transmission of diseases and can have a large impact on local economies. Monitoring mosquitoes is therefore helpful in preventing the outbreak of mosquito-borne diseases. In this paper, we propose a novel data-driven approach that leverages Transformer-based models for the identification of mosquitoes in audio recordings. The task aims at detecting the time intervals corresponding to the acoustic mosquito events in an audio signal. We formulate the problem as a sequence tagging task and train a Transformer-based model using a real-world dataset collecting mosquito recordings. By leveraging the sequential nature of mosquito recordings, we formulate the training objective so that the input recordings do not require fine-grained annotations. We show that our approach is able to outperform baseline methods using standard evaluation metrics, albeit suffering from unexpectedly high false negatives detection rates. In view of the achieved results, we propose future directions for the design of more effective mosquito detection models.
Moreno La Quatra, Lorenzo Vaiani, Alkis Koudounas, Luca Cagliero, Paolo Garza, Elena Baralis
ACM Multimedia6
2021 Summarize Dates First: A Paradigm Shift in Timeline Summarization
abstract
Timeline summarization aims at presenting long news stories in a compact manner. State-of-the-art approaches first select the most relevant dates from the original event timeline then produce per-date news summaries. Date selection is driven by either per-date news content or date-level references. When coping with complex event data, characterized by inherent news flow redundancy, this pipeline may encounter relevant issues in both date selection and summarization due to a limited use of news content in date selection and no use of high-level temporal references (e.g., the past month). This paper proposes a paradigm shift in timeline summarization aimed at overcoming the above issues. It presents a new approach, namely Summarize Date First, which focuses on first generating date-level summaries then selecting the most relevant dates on top of summarized knowledge. In the latter stage, it performs date aggregations to consider high-level temporal references as well. The proposed pipeline also supports frequent incremental timeline updates more efficiently than previous approaches. We tested our unsupervised approach both on existing benchmark datasets and on a newly proposed benchmark dataset describing the COVID-19 news timeline. The achieved results were superior to state-of-the-art unsupervised methods and competitive against supervised ones.
Moreno La Quatra, Luca Cagliero, Elena Baralis, Alberto Messina, Maurizio Montagnuolo
SIGIR3
2021 Looking for Trouble: Analyzing Classifier Behavior via Pattern Divergence
abstract
Machine learning models may perform differently on different data subgroups, which we represent as itemsets (i.e., conjunctions of simple predicates). The identification of these critical data subgroups plays an important role in many applications, for example model validation and testing, or evaluation of model fairness. Typically, domain expert help is required to identify relevant (or sensitive) subgroups.
Eliana Pastor, Luca de Alfaro, Elena Baralis
SIGMOD Conference3
2021 Enhancing manufacturing intelligence through an unsupervised data-driven methodology for cyclic industrial processes
Tania Cerquitelli, Francesco Ventura, Daniele Apiletti, Elena Baralis, Enrico Macii, Massimo Poncino
Expert Syst. Appl.4
2021 Dissecting a data-driven prognostic pipeline: A powertrain use case
Danilo Giordano, Eliana Pastor, Flavio Giobergia, Tania Cerquitelli, Elena Baralis, Marco Mellia, Alessandra Neri, Davide Tricarico
Expert Syst. Appl.5
2021 How Divergent Is Your Data?
abstract
We present DivExplorer, a tool that enables users to explore datasets and find subgroups of data for which a classifier behaves in an anomalous manner. These subgroups, denoted as divergent subgroups, may exhibit, for example, higher-than-normal false positive or negative rates. DivExplorer can be used to analyze and debug classifiers. If the data has ethical or social implications, DivExplorer can be also used to identify bias in classifiers.
Eliana Pastor, Andrew Gavgavian, Elena Baralis, Luca de Alfaro
Proc. VLDB Endow.3
2020 Improving Wildfire Severity Classification of Deep Learning U-Nets from Satellite Images
abstract
Uncontrolled wildfires are dangerous events capable of harming people safety. To contrast their increasing impact in recent years, a key task is an accurate detection of the affected areas and their damage assessment from satellite images. Current state-of-the-art solutions address such problem through a double convolutional neural network able to automatically detect wildfires in satellite acquisitions and associate a damage index from a defined scale. However, such deep-learning model performance is strongly dependent on many factors. In this work, we specifically focus on a key parameter, i.e., the loss function, exploited in the underlying neural networks. Besides the state-of-the-art solutions based on the Dice-MSE, among the many loss functions proposed in literature, we focus on the Binary Cross-Entropy (BCE) and the Intersection over Union (IoU), as two representatives of the distribution-based and region-based categories, respectively. Experiments show that the BCE loss function coupled with a double-step U-Net architecture provides better results than current state-of-the-art solutions on a public labeled dataset of European wildfires.
Simone Monaco, Andrea Pasini, Daniele Apiletti, Luca Colomba, Paolo Garza, Elena Baralis
IEEE BigData6
2020 DSLE: A Smart Platform for Designing Data Science Competitions
abstract
During the last years an increasing number of university-level and post-graduation courses on Data Science have been offered. Practices and assessments need specific learning environments where learners could play with data samples and run machine learning and data mining algorithms. To foster learner engagement many closed-and open-source platforms support the design of data science competitions. However, they show limitations on the ability to handle private data, customize the analytics and evaluation processes, and visualize learners' activities and outcomes. This paper presents Data Science Lab Environment (DSLE, in short), a new open-source platform to design and monitor data science competitions. DSLE offers a easily configurable interface to share training and test data, design group works or individual sessions, evaluate the competition runs according to customizable metrics, manage public and private leaderboards, monitor participants' activities and their progress over time. The paper describes also a real experience of usage of DSLE in the context of a 1st-year M.Sc. course, which has involved around 160 students.
Giuseppe Attanasio, Flavio Giobergia, Andrea Pasini, Francesco Ventura, Elena Baralis, Luca Cagliero, Paolo Garza, Daniele Apiletti, Tania Cerquitelli, Silvia Chiusano
COMPSAC5
2020 Bring Your Own Data to X-PLAIN
abstract
Exploring and understanding the motivations behind black-box model predictions is becoming essential in many different applications. X-PLAIN is an interactive tool that allows human-in-the-loop inspection of the reasons behind model predictions. Its support for the local analysis of individual predictions enables users to inspect the local behavior of different classifiers and compare the knowledge different classifiers are exploiting for their prediction. The interactive exploration of prediction explanation provides actionable insights for both trusting and validating model predictions and, in case of unexpected behaviors, for debugging and improving the model itself.
Eliana Pastor, Elena Baralis
SIGMOD Conference2
2020 entity2rec: Property-specific knowledge graph embeddings for item recommendation
Enrico Palumbo, Diego Monti, Giuseppe Rizzo 0002, Raphaël Troncy, Elena Baralis
Expert Syst. Appl.5
2019 Fast Self-Organizing Maps Training
abstract
Self-organizing maps are an unsupervised machine learning technique that offers interpretable results by identifying topological properties in high-dimensional datasets and projecting them on a 2-dimensional grid. An important problem of self-organizing maps is the computational expensiveness of their training phase. In this paper, we propose a fast approach to train self-organizing maps. The approach consists of 2 steps. First, a small map identifies the most relevant areas from the entire high-dimensional input space. Then a larger map (initialized from the small one) is fine-tuned to further explore the local areas identified in the first step. The resulting map has performance (measured in terms of accuracy and quantization error) on par with self-organizing maps trained with the standard approach, but with a significantly reduced training time.
Flavio Giobergia, Elena Baralis
IEEE BigData2
2019 Tinderbook: Fall in Love with Culture
abstract
More than 2 millions of new books are published every year and choosing a good book among the huge amount of available options can be a challenging endeavor. Recommender systems help in choosing books by providing personalized suggestions based on the user reading history. However, most book recommender systems are based on collaborative filtering, involving a long onboarding process that requires to rate many books before providing good recommendations. Tinderbook provides book recommendations, given a single book that the user likes, through a card-based playful user interface that does not require an account creation. Tinderbook is strongly rooted in semantic technologies, using the DBpedia knowledge graph to enrich book descriptions and extending a hybrid state-of-the-art knowledge graph embeddings algorithm to derive an item relatedness measure for cold start recommendations. Tinderbook is publicly available ( http://www.tinderbook.it ) and has already generated interest in the public, involving passionate readers, students, librarians, and researchers. The online evaluation shows that Tinderbook achieves almost 50% of precision of the recommendations.
Enrico Palumbo, Alberto Buzio, Andrea Gaiardo 0001, Giuseppe Rizzo 0002, Raphaël Troncy, Elena Baralis
ESWC6
2019 ELSA: A Multilingual Document Summarization Algorithm Based on Frequent Itemsets and Latent Semantic Analysis
abstract
Sentence-based summarization aims at extracting concise summaries of collections of textual documents. Summaries consist of a worthwhile subset of document sentences. The most effective multilingual strategies rely on Latent Semantic Analysis (LSA) and on frequent itemset mining, respectively. LSA-based summarizers pick the document sentences that cover the most important concepts. Concepts are modeled as combinations of single-document terms and are derived from a term-by-sentence matrix by exploiting Singular Value Decomposition (SVD). Itemset-based summarizers pick the sentences that contain the largest number of frequent itemsets, which represent combinations of frequently co-occurring terms. The main drawbacks of existing approaches are (i) the inability of LSA to consider the correlation between combinations of multiple-document terms and the underlying concepts, (ii) the inherent redundancy of frequent itemsets because similar itemsets may be related to the same concept, and (iii) the inability of itemset-based summarizers to correlate itemsets with the underlying document concepts. To overcome the issues of both of the abovementioned algorithms, we propose a new summarization approach that exploits frequent itemsets to describe all of the latent concepts covered by the documents under analysis and LSA to reduce the potentially redundant set of itemsets to a compact set of uncorrelated concepts. The summarizer selects the sentences that cover the latent concepts with minimal redundancy. We tested the summarization algorithm on both multilingual and English-language benchmark document collections. The proposed approach performed significantly better than both itemset- and LSA-based summarizers, and better than most of the other state-of-the-art approaches.
Luca Cagliero, Paolo Garza, Elena Baralis
ACM Trans. Inf. Syst.3
2018 A Density-based Preprocessing Technique to Scale Out Clustering
abstract
Clustering big data is a challenging task, because the majority of high-quality clustering algorithms do not scale well with respect to the data set cardinality. To tackle the scalability problem, we propose a general-purpose density-based preprocessing technique, called SCOUT, implemented in the Spark framework. It allows compacting the original data by means of a set of representative points, while still preserving the original data distribution and density information. This small set of representative points may become the input to almost any clustering algorithm. Thus, also complex, high-quality in-memory algorithms can be applied. A thorough experimental evaluation shows that the proposed approach is efficient and at the same time effective.
Elena Baralis, Paolo Garza, Eliana Pastor
IEEE BigData1
2018 Mining Sensor Data for Predictive Maintenance in the Automotive Industry
abstract
Predictive maintenance is an ever-growing area of interest, spanning different fields and approaches. In the automotive industry faulty behaviors of the oxygen sensor are a key challenge to address. This paper presents OxyClog, a data-driven framework that, given a large number of time series collected from a vehicle's ECU (engine control unit), builds a model to predict if the oxygen sensor is currently unclogged, almost clogged (since the clogging of the sensor happens gradually), or clogged. OxyClog is characterized by a tailored preprocessing, which includes a custom and interpretable feature selection algorithm, along with a summarization strategy to transform a time-dependent problem into a time-independent one. Furthermore, a semi-supervised labeling methodology has been devised to use different data sources with different characteristics to define meaningful clogging labels. OxyClog integrates state-of-the-art classification algorithms - both interpretable and non-interpretable - to process real ECU data with good prediction performance.
Flavio Giobergia, Elena Baralis, Maria Camuglia, Tania Cerquitelli, Marco Mellia, Alessandra Neri, Davide Tricarico, Alessia Tuninetti
DSAA2
2017 SQL versus NoSQL databases for geospatial applications
abstract
In the last years, we are witnessing an increasing availability of geolocated data, ranging from satellite images to user generated content (e.g., tweets). This big amount of data is exploited by several cloud-based applications to deliver effective and customized services to end users. In order to provide a good user experience, a low-latency response time is needed, both when data are retrieved and provided. To achieve this goal, current geospatial applications need to exploit efficient and scalable geospatial databases, the choice of which has a high impact on the overall performance of the deployed applications. In this paper, we compare, from a qualitative point of view, four state-of-the-art SQL and NoSQL databases with geospatial features, and then we analyze the performances of two of them, selecting the ones based on the Database-as-a-service (DBaaS) model: Azure SQL Database and Azure DocumentDB (i.e., an SQL database versus a NoSQL one). The empirical evaluation shows pros and cons of both solutions and it is performed on a real use case related to an emergency management application.
Elena Baralis, Andrea Dalla Valle, Paolo Garza, Claudio Rossi 0003, Francesco Scullino
IEEE BigData1
2017 Experimental Validation of a Massive Educational Service in a Blended Learning Environment
abstract
New information and communication technologies offer today many opportunities to improve the quality of educational services in universities and in particular they allow to design and implement innovative learning models. This paper describes and validates our university blended learning model, and specifically the massive educational video service that we offer to our students since 2010. In these years, we have gathered a huge amount of detailed data about the students' access to the service, and the paper describes a number of analyses that we carried out with these data. The common goal was to find out experimentally whether the main objectives of the educational video service we had in our mind when we designed it, namely appreciation, effectiveness and flexibility, were reflected by the users' behavior. We analyzed how many students used the service, for how many courses, and how many videos they accessed within a course (appreciation of the service). We analyzed the correlation between the use of the service and the performance of the students in terms of successful examination rate and average mark (effectiveness of the service). Finally, by using data mining techniques we profiled users according to their behavior while accessing the educational video service. We found out six different patterns that reflect different uses of the services matching different learning goals (flexibility of the service). The results of these analyses show the quality of the proposed blended learning model and the coherency of its implementation with respect to the design goals.
Elena Baralis, Luca Cagliero, Laura Farinetti, Marco Mezzalama, Enrico Venuto
COMPSAC (1)1
2017 Test-Driven Summarization: Combining Formative Assessment with Teaching Document Summarization
abstract
The diffusion of learning technologies has fostered the use of mobile and Web-based applications to assess the knowledge level of learners. In parallel, an increasing research interest has been devoted to studying new learning analytics tools able to summarize the content of large sets of learning documents. To bridge the gap between formative assessment tools and document summarization systems, this paper addresses the problem of recommending short summaries of large sets of learning documents based on the outcomes of multiple-choice tests. Specifically, it presents a new methodology for integrating formative assessment through mobile applications and summarization of learning documents in textual form. The content of the multiple-choice tests is exploited to drive the generation of document summaries tailored to specific topics. Furthermore, the outcomes of the tests are used to automatically recommend the generated summaries to learners based on their actual needs. As a case study, we performed an evaluation experience of students' progresses, which was conducted in the context of a university-level course. The achieved results show the applicability of the proposed methodology.
Luca Cagliero, Laura Farinetti, Elena Baralis
COMPSAC (1)3
2017 Educational video services in universities: A systematic effectiveness analysis
abstract
Our university has offered a massive educational video service since 2010, as part of a blended learning model that allows students to balance active participation in the classroom with remote access to video-recorded lectures. In these years, we have collected a huge amount of very detailed data about the students' access to the service. Together with additional information that characterize a university system (e.g. students' performance or course population), these data represent a precious ground set to assess the educational model. The paper describes an experimental set to profile the use of the educational video service, whose results will contribute to improve the model. Specifically the paper analyzes the students' service use relatively to different transversal course characteristics, such as level, main topic, population, success rate. As a result, it outlines the profile of the “ideal” courses for which students highly appreciate the service. This information will help educational designers to select the future courses to be included in the service, but it will also give directions on the sectors where improvements are necessary. Finally, the paper experimentally demonstrates a positive impact of the educational video service on students' performance, and specifically on the exam success rate.
Luca Cagliero, Laura Farinetti, Marco Mezzalama, Enrico Venuto, Elena Baralis
FIE5
2017 Planning stock portfolios by means of weighted frequent itemsets
Elena Baralis, Luca Cagliero, Paolo Garza
Expert Syst. Appl.1
2017 Discovering profitable stocks for intraday trading
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Fabio Pulvirenti
Inf. Sci.1
2016 SaFe-NeC: A scalable and flexible system for network data characterization
abstract
Nowadays, large volumes of data and measurements are being continuously generated by computer and telecommunication networks, but such volumes make it difficult to extract meaningful knowledge from them. This paper presents SaFe-NeC, an innovative methodology for analyzing network traffic by exploiting data mining techniques, i.e. clustering and classification algorithms, focusing on self-learning capabilities of state-of-the-art scalable approaches. Self-learning algorithms, coupled with self-assessment indicators and domain-driven semantics enriching data mining results, are able to build a model of the data with minimal user intervention and highlight possibly meaningful interpretations to domain experts. Furthermore, a self-evolving model evaluation phase is included to continuously track the quality degradation of the model itself, whose rebuilding is triggered as soon as quality indicators fall below a threshold of tolerance. The proposed methodology can exploit the computational advantages of distributed computing frameworks, as the current implementation runs on Apache Spark. Preliminary experimental results on a real traffic dataset show the full potential of the proposed methodology to characterize network traffic data.
Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Luca Venturini
NOMS2
2016 MAGMA network behavior classifier for malware traffic
Enrico Bocchi, Luigi Grimaudo, Marco Mellia, Elena Baralis, Sabyasachi Saha, Stanislav Miskovic, Gaspar Modelo-Howard, Sung-Ju Lee 0001
Comput. Networks4
2016 SeLINA: A Self-Learning Insightful Network Analyzer
abstract
Understanding the behavior of a network from a large scale traffic dataset is a challenging problem. Big data frameworks offer scalable algorithms to extract information from raw data, but often require a sophisticated fine-tuning and a detailed knowledge of machine learning algorithms. To streamline this process, we propose self-learning insightful network analyzer (SeLINA), a generic, self-tuning, simple tool to extract knowledge from network traffic measurements. SeLINA includes different data analytics techniques providing self-learning capabilities to state-of-the-art scalable approaches, jointly with parameter auto-selection to off-load the network expert from parameter tuning. We combine both unsupervised and supervised approaches to mine data with a scalable approach. SeLINA embeds mechanisms to check if the new data fits the model, to detect possible changes in the traffic, and to, possibly automatically, trigger model rebuilding. The result is a system that offers human-readable models of the data with minimal user intervention, supporting domain experts in extracting actionable knowledge and highlighting possibly meaningful interpretations. SeLINA's current implementation runs on Apache Spark. We tested it on large collections of real-world passive network measurements from a nationwide ISP, investigating YouTube, and P2P traffic. The experimental results confirmed the ability of SeLINA to provide insights and detect changes in the data that suggest further analyses.
Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Paolo Garza, Danilo Giordano, Marco Mellia, Luca Venturini
IEEE Trans. Netw. Serv. Manag.2
2015 Predicting Cardiopulmonary Response to Incremental Exercise Test
abstract
Cardiopulmonary exercise testing is a non-invasive method widely used to monitor various physiological signals, describing the cardiac and respiratory response of the patient to increasing workload. Since this method is physically very demanding, innovative data analysis techniques are needed to predict patient response thus lowering body stress and avoiding cardiopulmonary overload. This paper proposes the Cardiopulmonary Response Prediction (CRP) framework for early predicting the physiological signal values that can be reached during an incremental exercise test. The learning phase creates different models tailored to specific conditions (i.e., single-test and multiple-test models). Each model can be exploited in the real-time stream prediction phase to periodically predict, during the test execution, signal values achievable by the patient. Experimental results on a real dataset showed that CRP prediction is performed with a limited and acceptable error.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Alessandro Mezzani, Davide Susta, Xin Xiao 0002
CBMS1
2015 Macroscopic view of malware in home networks
abstract
Malicious activities on the Web are increasingly threatening users in the Internet. Home networks are one of the prime targets of the attackers to host malware, commonly exploited as a stepping stone to further launch a variety of attacks. Due to diversification, existing security solutions often fail to detect malicious activities that remain hidden and pose threats to users' security and privacy. Characterizing behavioral patterns of known malware can help to improve the classification accuracy of threats. More importantly, as different malware might share commonalities, studying the behavior of known malware could help the detection of previously unknown malicious activities. We pose the research question if it is possible to characterize such behavioral patterns analyzing the traffic from known infected clients. We present our quest to discover such characterizations. Results show that commonalities arise but their identification may require some ingenuity. We also present our discovery of malicious activities that were left undetected by commercial IDS.
Alessandro Finamore, Sabyasachi Saha, Gaspar Modelo-Howard, Sung-Ju Lee 0001, Enrico Bocchi, Luigi Grimaudo, Marco Mellia, Elena Baralis
CCNC8
2015 Generation and Evaluation of Summaries of Academic Teaching Materials
abstract
E-learning systems commonly rely on advanced ICT technologies to enable users to access and browse electronic resources. Document summarization is an established text mining technique which focuses on extracting succinct summaries of potentially long textual documents. The application of summarization algorithms in the e-learning context is particularly appealing, because readers may want to pinpoint the key concepts by reading short summaries instead of the whole document content. This paper investigates the application of a state-of-the-art summarization algorithm to English-written academic teaching material. The summarizer produces an ordered sequence of key phrases extracted from learning material organized in different sections. The generated summaries are provided to students as additional material for study and revision. A crowd-sourcing experience of evaluation of the generated summaries was conducted by involving the students of a B.S. Course given by a technical university. The results show that the automatically generated summaries reflect, to a large extent, the student's expectations and therefore they can be useful for supporting individual and collective learning activities.
Elena Baralis, Luca Cagliero, Laura Farinetti
COMPSAC1
2015 Network Connectivity Graph for Malicious Traffic Dissection
abstract
Malware is a major threat to security and privacy of network users. A huge variety of malware typically spreads over the Internet, evolving every day, and challenging the research community and security practitioners to improve the effectiveness of countermeasures. In this paper, we present a system that automatically extracts patterns of network activity related to a specific malicious event, i.e., a seed. Our system is based on a methodology that correlates network events of hosts normally connected to the Internet over (i) time (i.e., analyzing different samples of traffic from the same host), (ii) space (i.e., correlating patterns across different hosts), and (iii) network layers (e.g., HTTP, DNS, etc.). The result is a Network Connectivity Graph that captures the overall "network behavior" of the seed. That is a focused and enriched representation of the malicious pattern infected hosts exhibit, purified from ordinary network activities and background traffic. We applied our approach on a large dataset collected in a real commercial ISP where the aggregated traffic produced by more than 20,000 households has been monitored. A commercial IDS has been used to complement network data with alerts related to malicious activities. We use such alerts to trigger our processing system. Results shows that the richness of the Network Connectivity Graph provides a much more detailed picture of malicious activities, considerably enhancing our understanding.
Enrico Bocchi, Luigi Grimaudo, Marco Mellia, Elena Baralis, Sabyasachi Saha, Stanislav Miskovic, Gaspar Modelo-Howard, Sung-Ju Lee 0001
ICCCN4
2015 Digging deep into weighted patient data through multiple-level patterns
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza
Inf. Sci.1
2015 Scalable out-of-core itemset mining
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Alberto Grand
Inf. Sci.1
2015 MeTA: Characterization of Medical Treatments at Different Abstraction Levels
abstract
Physicians and health care organizations always collect large amounts of data during patient care. These large and high-dimensional datasets are usually characterized by an inherent sparseness. Hence, analyzing these datasets to figure out interesting and hidden knowledge is a challenging task. This article proposes a new data mining framework based on generalized association rules to discover multiple-level correlations among patient data. Specifically, correlations among prescribed examinations, drugs, and patient profiles are discovered and analyzed at different abstraction levels. The rule extraction process is driven by a taxonomy to generalize examinations and drugs into their corresponding categories. To ease the manual inspection of the result, a worthwhile subset of rules (i.e., nonredundant generalized rules) is considered. Furthermore, rules are classified according to the involved data features (medical treatments or patient profiles) and then explored in a top-down fashion: from the small subset of high-level rules, a drill-down is performed to target more specific rules. The experiments, performed on a real diabetic patient dataset, demonstrate the effectiveness of the proposed approach in discovering interesting rule groups at different abstraction levels.
Dario Antonelli, Elena Baralis, Giulia Bruno, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Naeem Ahmed Mahoto
ACM Trans. Intell. Syst. Technol.2
2015 MWI-Sum: A Multilingual Summarizer Based on Frequent Weighted Itemsets
abstract
Multidocument summarization addresses the selection of a compact subset of highly informative sentences, i.e., the summary, from a collection of textual documents. To perform sentence selection, two parallel strategies have been proposed: (a) apply general-purpose techniques relying on data mining or information retrieval techniques, and/or (b) perform advanced linguistic analysis relying on semantics-based models (e.g., ontologies) to capture the actual sentence meaning. Since there is an increasing need for processing documents written in different languages, the attention of the research community has recently focused on summarizers based on strategy (a). This article presents a novel multilingual summarizer, namely MWI-Sum (Multilingual Weighted Itemset-based Summarizer), that exploits an itemset-based model to summarize collections of documents ranging over the same topic. Unlike previous approaches, it extracts frequent weighted itemsets tailored to the analyzed collection and uses them to drive the sentence selection process. Weighted itemsets represent correlations among multiple highly relevant terms that are neglected by previous approaches. The proposed approach makes minimal use of language-dependent analyses. Thus, it is easily applicable to document collections written in different languages. Experiments performed on benchmark and real-life collections, English-written and not, demonstrate that the proposed approach performs better than state-of-the-art multilingual document summarizers.
Elena Baralis, Luca Cagliero, Alessandro Fiori, Paolo Garza
ACM Trans. Inf. Syst.1
2014 Misleading Generalized Itemset Mining in the Cloud
abstract
In the era of smart cities huge data volumes are continuously generated and collected, thus prompting the need for efficient and distributed data mining approaches. Generalized itemset mining is an established data mining technique, which entails the discovery of multiple-level patterns hidden in the analyzed data by exploiting analyst-provided taxonomies. Among the generalized itemsets, the most peculiar high-level patterns are those with many contrasting correlations among items at different abstraction levels. They represent misleading situations that are worth analyzing separately by experts during manual inspection. This paper proposes a novel cloud-based service, named MGI-CLOUD, to efficiently mine misleading multiple-level patterns, i.e., the Misleading Generalized Itemsets, on a distributed computing environment. MGI-CLOUD consists of a set of distributed MapReduce jobs running in the cloud. As a case study, the system has been contextualized in a real-life scenario, i.e., the analysis of traffic law infractions committed in a smart city environment. The experiments, performed on real datasets, demonstrate the efficiency and effectiveness of MGI-CLOUD.
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Luigi Grimaudo, Fabio Pulvirenti
ISPA1
2014 Expressive generalized itemsets
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Vincenzo D'Elia, Paolo Garza
Inf. Sci.1
2014 RIB: A Robust Itemset-based Bayesian approach to classification
Elena Baralis, Luca Cagliero
Knowl. Based Syst.1
2014 SeLeCT: Self-Learning Classifier for Internet Traffic
abstract
Network visibility is a critical part of traffic engineering, network management, and security. The most popular current solutions - Deep Packet Inspection (DPI) and statistical classification, deeply rely on the availability of a training set. Besides the cumbersome need to regularly update the signatures, their visibility is limited to classes the classifier has been trained for. Unsupervised algorithms have been envisioned as a viable alternative to automatically identify classes of traffic. However, the accuracy achieved so far does not allow to use them for traffic classification in practical scenario. To address the above issues, we propose SeLeCT, a Self-Learning Classifier for Internet Traffic. It uses unsupervised algorithms along with an adaptive seeding approach to automatically let classes of traffic emerge, being identified and labeled. Unlike traditional classifiers, it requires neither a-priori knowledge of signatures nor a training set to extract the signatures. Instead, SeLeCT automatically groups flows into pure (or homogeneous) clusters using simple statistical features. SeLeCT simplifies label assignment (which is still based on some manual intervention) so that proper class labels can be easily discovered. Furthermore, SeLeCT uses an iterative seeding approach to boost its ability to cope with new protocols and applications. We evaluate the performance of SeLeCT using traffic traces collected in different years from various ISPs located in 3 different continents. Our experiments show that SeLeCT achieves excellent precision and recall, with overall accuracy close to 98%. Unlike state-of-art classifiers, the biggest advantage of SeLeCT is its ability to discover new protocols and applications in an almost automated fashion.
Luigi Grimaudo, Marco Mellia, Elena Baralis, Ram Keralapura
IEEE Trans. Netw. Serv. Manag.3
2013 Frequent weighted itemset mining from gene expression data
abstract
Gene Expression Datasets (GEDs) usually consist of the expression values of thousands of genes within hundreds of samples. Frequent itemset and association rule mining algorithms have been applied to discover significant co-expressions among multiple genes from GEDs. To perform these data analyses, gene expression values are commonly discretized into a predefined number of bins. Such an expert-driven and not trivial preprocessing step could bias the quality of the mining result. This paper presents a novel approach to discovering gene correlations from GEDs which does not require data discretization. By representing per-sample gene expression values as item weights, frequent weighted itemsets can be extracted. The discovery of weighted itemsets instead of traditional (not weighted) ones prevents experts from discretizing GEDs before analyzing them and thus improves the effectiveness of the knowledge discovery process. Experiments performed on real GEDs demonstrate the effectiveness of the proposed approach.
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza
BIBE1
2013 Self-learning classifier for Internet traffic
abstract
Network visibility is a critical part of traffic engineering, network management, and security. Recently, unsupervised algorithms have been envisioned as a viable alternative to automatically identify classes of traffic. However, the accuracy achieved so far does not allow to use them for traffic classification in practical scenario. In this paper, we propose SeLeCT, a Self-Learning Classifier for Internet traffic. It uses unsupervised algorithms along with an adaptive learning approach to automatically let classes of traffic emerge, being identified and (easily) labeled. SeLeCT automatically groups flows into pure (or homogeneous) clusters using alternating simple clustering and filtering phases to remove outliers. SeLeCT uses an adaptive learning approach to boost its ability to spot new protocols and applications. Finally, SeLeCT also simplifies label assignment (which is still based on some manual intervention) so that proper class labels can be easily discovered. We evaluate the performance of SeLeCT using traffic traces collected in different years from various ISPs located in 3 different continents. Our experiments show that SeLeCT achieves overall accuracy close to 98%. Unlike state-of-art classifiers, the biggest advantage of SeLeCT is its ability to help discovering new protocols and applications in an almost automated fashion.
Luigi Grimaudo, Marco Mellia, Elena Baralis, Ram Keralapura
INFOCOM3
2013 Analysis of Twitter Data Using a Multiple-level Clustering Strategy
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Luigi Grimaudo, Xin Xiao 0002
MEDI1
2013 NetCluster: A clustering-based framework to analyze internet passive measurements data
Elena Baralis, Andrea Bianco, Tania Cerquitelli, Luca Chiaraviglio, Marco Mellia
Comput. Networks1
2013 Analysis of diabetic patients through their examination history
Dario Antonelli, Elena Baralis, Giulia Bruno, Tania Cerquitelli, Silvia Chiusano, Naeem Ahmed Mahoto
Expert Syst. Appl.2
2013 Multi-document summarization based on the Yago ontology
Elena Baralis, Luca Cagliero, Saima Jabeen, Alessandro Fiori, Sajid Shah
Expert Syst. Appl.1
2013 GraphSum: Discovering correlations among multiple terms for graph-based summarization
Elena Baralis, Luca Cagliero, Naeem Ahmed Mahoto, Alessandro Fiori
Inf. Sci.1
2013 Early prediction of the highest workload in incremental cardiopulmonary tests
abstract
Incremental tests are widely used in cardiopulmonary exercise testing, both in the clinical domain and in sport sciences. The highest workload (denoted Wpeak) reached in the test is key information for assessing the individual body response to the test and for analyzing possible cardiac failures and planning rehabilitation, and training sessions. Being physically very demanding, incremental tests can significantly increase the body stress on monitored individuals and may cause cardiopulmonary overload. This article presents a new approach to cardiopulmonary testing that addresses these drawbacks. During the test, our approach analyzes the individual body response to the exercise and predicts the Wpeakvalue that will be reached in the test and an evaluation of its accuracy. When the accuracy of the prediction becomes satisfactory, the test can be prematurely stopped, thus avoiding its entire execution. To predict Wpeak, we introduce a new index, the CardioPulmonary Efficiency Index (CPE), summarizing the cardiopulmonary response of the individual to the test. Our approach analyzes the CPE trend during the test, together with the characteristics of the individual, and predicts Wpeak. A K-nearest-neighbor-based classifier and an ANN-based classier are exploited for the prediction. The experimental evaluation showed that the Wpeakvalue can be predicted with a limited error from the first steps of the test.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Vincenzo D'Elia, Riccardo Molinari, Davide Susta
ACM Trans. Intell. Syst. Technol.1
2013 EnBay: A Novel Pattern-Based Bayesian Classifier
abstract
A promising approach to Bayesian classification is based on exploiting frequent patterns, i.e., patterns that frequently occur in the training data set, to estimate the Bayesian probability. Pattern-based Bayesian classification focuses on building and evaluating reliable probability approximations by exploiting a subset of frequent patterns tailored to a given test case. This paper proposes a novel and effective approach to estimate the Bayesian probability. Differently from previous approaches, the Entropy-based Bayesian classifier, namely EnBay, focuses on selecting the minimal set of long and not overlapped patterns that best complies with a conditional-independence model, based on an entropy-based evaluator. Furthermore, the probability approximation is separately tailored to each class. An extensive experimental evaluation, performed on both real and synthetic data sets, shows that EnBay is significantly more accurate than most state-of-the-art classifiers, Bayesian and not.
Elena Baralis, Luca Cagliero, Paolo Garza
IEEE Trans. Knowl. Data Eng.1
2012 Hierarchical learning for fine grained internet traffic classification
abstract
Traffic classification is still today a challenging problem given the ever evolving nature of the Internet in which new protocols and applications arise at a constant pace. In the past, so called behavioral approaches have been successfully proposed as valid alternatives to traditional DPI based tools to properly classify traffic into few and coarse classes. In this paper we push forward the adoption of behavioral classifiers by engineering a Hierarchical classifier that allows proper classification of traffic into more than twenty fine grained classes. Thorough engineering has been followed which considers both proper feature selection and testing seven different classification algorithms. Results obtained over actual and large data sets show that the proposed Hierarchical classifier outperforms off-the-shelf non hierarchical classification algorithms by exhibiting average accuracy higher than 90%, with precision and recall that are higher than 95% for most popular classes of traffic.
Luigi Grimaudo, Marco Mellia, Elena Baralis
IWCMC3
2012 MaskedPainter: Feature selection for microarray data analysis
abstract
Selecting a small number of discriminative genes from thousands is a fundamental task in microarray data analysis. An effective feature selection allows biologists to investigate only a subset of genes instead of the entire set, thus avoiding insignificant, noisy, and redundant features. This paper presents the MaskedPainter feature selection method for gene expression data. The proposed method measures the ability of each gene to classify samples belonging to different classes and ranks genes by computing an overlap score. A density based technique is exploited to smooth the effects of outliers in the overlap score computation. Analogously to other approaches, the number of selected genes can be set by the user. However, our algorithm may automatically detect the minimum set of genes that yields the best classification coverage of training set samples. The effectiveness of our approach has been demonstrated through an empirical study on public microarray datasets with different characteristics. Experimental results show that the proposed approach yields a higher classification accuracy with respect to widely used feature selection techniques.
Daniele Apiletti, Elena Baralis, Giulia Bruno, Alessandro Fiori
Intell. Data Anal.2
2012 I-prune: Item selection for associative classification
abstract
Associative classification is characterized by accurate models and high model generation time. Most time is spent in extracting and postprocessing a large set of irrelevant rules, which are eventually pruned. We propose I-prune, an item-pruning approach that selects uninteresting items by means of an interestingness measure and prunes them as soon as they are detected. Thus, the number of extracted rules is reduced and model generation time decreases correspondingly. A wide set of experiments on real and synthetic data sets has been performed to evaluate I-prune and select the appropriate interestingness measure. The experimental results show that I-prune allows a significant reduction in model generation time, while increasing (or at worst preserving) model accuracy. Experimental evaluation also points to the chi-square measure as the most effective interestingness measure for item pruning. © 2012 Wiley Periodicals, Inc.
Elena Baralis, Paolo Garza
Int. J. Intell. Syst.1
2012 Generalized association rule mining with constraints
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza
Inf. Sci.1
2011 An Efficient Itemset Mining Approach for Data Streams
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Alberto Grand, Luigi Grimaudo
KES (2)1
2011 Energy-saving models for wireless sensor networks
Daniele Apiletti, Elena Baralis, Tania Cerquitelli
Knowl. Inf. Syst.2
2011 Measuring gene similarity by means of the classification distance
Elena Baralis, Giulia Bruno, Alessandro Fiori
Knowl. Inf. Syst.1
2011 CAS-Mine: providing personalized services in context-aware applications by means of generalized rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti
Knowl. Inf. Syst.1
2010 Predicting the highest workload in cardiopulmonary test
abstract
Cardiopulmonary exercise testing is an objective method to evaluate both the cardiac and pulmonary functions. It is used in different application domains, ranging from the clinical domain to sport sciences, to assess possible cardiac failures as well as athete performance. The highest workload reached in the test is a key information to evaluate the individual's physiological characteristics, to plan rehabilitation and/or training sessions. However, these tests are physically very demanding and may expose the tested individual to cardiopulmonary overload. This paper presents a new approach that allows an early prediction of the highest workload that will be reached in the cardiopulmonary test. The test can be prematurely stopped, avoiding its entire execution. The proposed approach relies on a new index, the CardioPulmonary Efficiency Index, which describes the cardiopulmonary response of an individual by summarizing the physiological signals monitored during the test. A k-Nearest Neighbor based classifier analyzes the index trend during the test, together with the characteristics of the individual, and predicts the highest workload. Preliminary experiments, performed on a real dataset provided by the CSA Sport Training Center, showed that the proposed approach is able to effectively predict the highest workload with a limited error since the first steps of the test.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Vincenzo D'Elia, Riccardo Molinari, Davide Susta
CBMS1
2010 Summarizing biological literature with BioSumm
abstract
BioSumm is a summarization environment that supports user queries on online repositories of scientific publications by providing abstract descriptions of focused document groups. The summarization approach is driven by a grading function which evaluates the occurrences of domain dictionary terms.
Elena Baralis, Alessandro Fiori
CIKM1
2010 Analysis of Medical Pathways by Means of Frequent Closed Sequences
Elena Baralis, Giulia Bruno, Silvia Chiusano, Virna C. Domenici, Naeem Ahmed Mahoto, Caterina Petrigni
KES (3)1
2010 Constrained itemset mining on a sequence of incoming data blocks
abstract
Many real-life databases are updated by means of incoming business information. In these databases (e.g., transactional data from large retail chains, call-detail records), the content evolves through periodical insertions (or deletions) of data blocks. Since data evolve over time, algorithms have to be devised to incrementally update data mining models. This paper presents a novel index, called I-Forest, to support itemset mining on incoming data blocks, where new blocks are inserted periodically, or old blocks are discarded. The I-Forest structure provides a complete data representation and allows different kind of analyses (e.g., investigate quarterly data), besides supporting user-defined time and support constraints. The I-Forest index has been implemented into the PostgreSQL open source DBMS and exploits its physical level access methods. Experiments, run for both sparse and dense data distributions, show the effectiveness of the I-Forest-based approach to perform itemset mining with both time and support constraints. The execution time of the I-Forest-based itemset mining technique is often faster than the Prefix-Tree algorithm accessing static data on flat files. © 2010 Wiley Periodicals, Inc.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano
Int. J. Intell. Syst.1
2009 NetCluster: A Clustering-Based Framework for Internet Tomography
abstract
In this paper, Internet data collected via passive measurement are analyzed to obtain localization information on nodes by clustering (i.e., grouping together) nodes that exhibit similar network path properties. Since traditional clustering algorithms fail to correctly identify clusters of homogeneous nodes, we propose a novel framework, named "NetCluster", suited to analyze Internet measurement datasets. We show that the proposed framework correctly analyzes synthetically generated traces. Finally, we apply it to real traces collected at the access link of our campus LAN and discuss the network characteristics as seen at the vantage point.
Elena Baralis, Andrea Bianco, Tania Cerquitelli, Luca Chiaraviglio, Marco Mellia
ICC1
2009 Context-Aware User and Service Profiling by Means of Generalized Association Rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti
KES (2)1
2009 Characterizing network traffic by means of the NetMine framework
Daniele Apiletti, Elena Baralis, Tania Cerquitelli, Vincenzo D'Elia
Comput. Networks2
2009 Real-Time Analysis of Physiological Data to Support Medical Applications
abstract
This paper presents a flexible framework that performs real-time analysis of physiological data to monitor people's health conditions in any context (e.g., during daily activities, in hospital environments). Given historical physiological data, different behavioral models tailored to specific conditions (e.g., a particular disease, a specific patient) are automatically learnt. A suitable model for the currently monitored patient is exploited in the real-time stream classification phase. The framework has been designed to perform both instantaneous evaluation and stream analysis over a sliding time window. To allow ubiquitous monitoring, real-time analysis could also be executed on mobile devices. As a case study, the framework has been validated in the intensive care scenario. Experimental validation, performed on 64 patients affected by different critical illnesses, demonstrates the effectiveness and the flexibility of the proposed framework in detecting different severity levels of monitored people's clinical situations.
Daniele Apiletti, Elena Baralis, Giulia Bruno, Tania Cerquitelli
IEEE Trans. Inf. Technol. Biomed.2
2009 IMine: Index Support for Item Set Mining
abstract
This paper presents the IMine index, a general and compact structure which provides tight integration of item set extraction in a relational DBMS. Since no constraint is enforced during the index creation phase, IMine provides a complete representation of the original database. To reduce the I/O cost, data accessed together during the same extraction phase are clustered on the same disk block. The IMine index structure can be efficiently exploited by different item set extraction algorithms. In particular, IMine data access methods currently support the FP-growth and LCM v.2 algorithms, but they can straightforwardly support the enforcement of various constraint categories. The IMine index has been integrated into the PostgreSQL DBMS and exploits its physical level access methods. Experiments, run for both sparse and dense data distributions, show the efficiency of the proposed index and its linear scalability also for large datasets. Item set mining supported by the IMine index shows performance always comparable with, and sometimes better than, state of the art algorithms accessing data on flat file.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano
IEEE Trans. Knowl. Data Eng.1
2008 BioSumm: A novel summarizer oriented to biological information
abstract
The availability of increasingly wider repositories of biomedical and biological texts requires effective techniques to manage the huge mass of unstructured information there contained. The availability of ad-hoc document summaries, targeted to specific topics, may assist researchers in inferring previously undisclosed knowledge and in performing the biological validation of the results of data mining analysis. This paper presents BioSumm, a flexible framework which analyzes large collections of unclassified biomedical texts and produces ad-hoc summaries oriented to inferring knowledge of gene/protein relationships. Summary generation is driven by a novel grading function, which biases sentence selection by means of an appropriate domain dictionary.
Elena Baralis, Alessandro Fiori, Lorenzo Montrucchio
BIBE1
2008 A Lazy Approach to Associative Classification
abstract
Associative classification is a promising technique to build accurate classifiers. However, in large or correlated datasets, association rule mining may yield huge rule sets. Hence, several pruning techniques have been proposed to select a small subset of high quality rules. We argue that rule pruning should be reduced to a minimum, since the availability of a "rich" rule set may improve the accuracy of the classifier. The L^3 associative classifier is built by means of a lazy pruning technique which discards exclusively rules that only misclassify training data. Classification of unlabeled data is performed in two steps. A small subset of high quality rules is first considered. When this set is not able to classify the data, a larger rule set is exploited. This second set includes rules usually discarded by previous approaches. To cope with the need of mining large rule sets and efficiently use them for classification, a compact form is proposed to represent a complete rule set in a space-efficient way and without information loss. An extensive experimental evaluation on real and synthetic datasets shows that L^3 improves the classification accuracy with respect to previous approaches.
Elena Baralis, Silvia Chiusano, Paolo Garza
IEEE Trans. Knowl. Data Eng.1
2007 Gene-Markers Representation for Microarray Data Integration
abstract
When analyzing the relationship between genes under different scenarios, the integration of different microarray experiments becomes a relevant task. This paper presents a framework to address some intrinsic problems of integration, due for instance to scaling issues, error bias, different experimental conditions or technology and protocols. Our approach projects original microarray data in a common transformed space to create a common representation of different microarray datasets. This approach allows us to integrate data from various microarray platforms or microarrays based on different experimental conditions. We validate our framework with experiments on real microarray datasets. The results suggest that our approach can be a profitably exploited for microarray data integration and further gene expression analysis applications.
Elena Baralis, Elisa Ficarra, Alessandro Fiori, Enrico Macii
BIBE1
2007 SAPhyRA: Stream Analysis for Physiological Risk Assessment
abstract
Advances in technology allow the continuous physiological monitoring of people using noninvasive sensors. An important issue in this context is the real-time analysis of physiological signals performed on mobile devices, which requires optimized power consumption and short processing response time. The SAPhyRA framework performs real-time stream analysis for physiological risk assessment. To this aim, the framework evaluates people's health conditions by analyzing different clinical signals in a sliding time window. Given historical physiological measures, different models of patients and diseases are built. The most suitable model for the current monitored patient is exploited in the real time stream classification phase. Preliminary experiments performed on public physiological data show the effectiveness and flexibility of the proposed approach.
Daniele Apiletti, Elena Baralis, Giulia Bruno, Tania Cerquitelli
CBMS2
2007 Topic 5 Parallel and Distributed Databases
Marta Patiño-Martínez, Genoveva Vargas-Solar, Elena Baralis, Bettina Kemme
Euro-Par3
2007 Answering XML queries by means of data summaries
abstract
XML is a rather verbose representation of semistructured data, which may require huge amounts of storage space. We propose a summarized representation of XML data, based on the concept of instance pattern, which can both provide succinct information and be directly queried. The physical representation of instance patterns exploits itemsets or association rules to summarize the content of XML datasets. Instance patterns may be used for (possibly partially) answering queries, either when fast and approximate answers are required, or when the actual dataset is not available, for example, it is currently unreachable. Experiments on large XML documents show that instance patterns allow a significant reduction in storage space, while preserving almost entirely the completeness of the query result. Furthermore, they provide fast query answers and show good scalability on the size of the dataset, thus overcoming the document size limitation of most current XQuery engines.
Elena Baralis, Paolo Garza, Elisa Quintarelli, Letizia Tanca
ACM Trans. Inf. Syst.1
2005 Index Support for Frequent Itemset Mining in a Relational DBMS
abstract
Many efforts have been devoted to couple data mining activities with relational DBMSs, but a true integration into the relational DBMS kernel has been rarely achieved. This paper presents a novel indexing technique, which represents transactions in a succinct form, appropriate for tightly integrating frequent itemset mining in a relational DBMS. The data representation is complete, i.e., no support threshold is enforced, in order to allow reusing the index for mining itemsets with any support threshold. Furthermore, an appropriate structure of the stored information has been devised, in order to allow a selective access of the index blocks necessary for the current extraction phase. The index has been implemented into the PostgreSQL open source DBMS and exploits its physical level access methods. Experiments have been run for various datasets, characterized by different data distributions. The execution time of the frequent itemset extraction task exploiting the index is always comparable with and sometime faster than a C++ implementation of the FP-growth algorithm accessing data stored on a flat file.
Elena Baralis, Tania Cerquitelli, Silvia Chiusano
ICDE1
2005 Data mining techniques for effective and scalable traffic analysis
abstract
This paper describes a novel approach to traffic analysis in high speed networks based on data mining techniques. Data mining techniques are here applied as a means to effectively process the significant amount of captured data. The paper provides a first evaluation of the proposed approach in terms of its ability of extracting relevant information and its computational requirements. Such evaluation is based on experiments run on a prototypal implementation of the proposed approach.
Mario Baldi, Elena Baralis, Fulvio Risso
Integrated Network Management2
2004 Essential classification rule sets
abstract
Given a class model built from a dataset including labeled data, classification assigns a new data object to the appropriate class. In associative classification the class model (i.e., the classifier) is a set of association rules. Associative classification is a promising technique for the generation of highly accurate classifiers. In this article, we present a compact form which encodes without information loss the classification knowledge available in a classification rule set. This form includes the rules that are essential for classification purposes, and thus it can replace the complete rule set. The proposed form is particularly effective in dense datasets, where traditional extraction techniques may generate huge rule sets. The reduction in size of the rule set allows decreasing the complexity of both the rule generation step and the rule pruning step. Hence, classification rule extraction can be performed also with low support, in order to extract more, possibly useful, rules.
Elena Baralis, Silvia Chiusano
ACM Trans. Database Syst.1
2003 Majority Classification by Means of Association Rules
Elena Baralis, Paolo Garza
PKDD1
2002 A Lazy Approach to Pruning Classification Rules
abstract
Associative classification is a promising technique for the generation of highly precise classifiers. Previous works propose several clever techniques to prune the huge set of generated rules, with the twofold aim of selecting a small set of high quality rules, and reducing the chance of overfitting. In this paper, we argue that pruning should be reduced to a minimum and that the availability of a large rule base may improve the precision of the classifier without affecting its performance. In L/sup 3/ (Live and Let Live), a new algorithm for associative classification, a lazy pruning technique iteratively discards all rules that only yield wrong case classifications. Classification is performed in two steps. Initially, rules which have already correctly classified at least one training case, sorted by confidence, are considered If the case is still unclassified, the remaining rules (unused during the training phase) are considered, again sorted by confidence. Extensive experiments on 26 databases from the UCI machine learning database repository show that L/sup 3/ improves the classification precision with respect to previous approaches.
Elena Baralis, Paolo Garza
ICDM1
2000 An algebraic approach to static analysis of active database rules
abstract
Rules in active database systems can be very difficult to program due to the unstructured and unpredictable nature of rule processing. We provide static analysis techniques for predicting whether a given rule set is guaranteed to terminate and whether rule execution is confluent (guaranteed to have a unique final state). Our methods are based on previous techniques for analyzing rules in active database systems. We improve considerably on the previous techniques by providing analysis criteria that are much less conservative: our methods often determine that a rule set will terminate or is confluent when previous methods could not make this determination. Our improved analysis is based on a “propagation” algorithm, which uses an extended relational algebra to accurately determine when the action of one rule can affect the condition of another, and determine when rule actions commute. We consider both conditon-action rules and event-condition-action-rules, making our approach widely applicable to relational active database rule languages and to the trigger language in the SQL:1999 standard.
Elena Baralis, Jennifer Widom
ACM Trans. Database Syst.1
1999 Incremental Refinement of Mining Queries
Elena Baralis, Giuseppe Psaila
DaWaK1
1998 Compile-Time and Runtime Analysis of Active Behaviors
abstract
Active rules may interact in complex and sometimes unpredictable ways, thus possibly yielding infinite rule executions by triggering each other indefinitely. This paper presents analysis techniques focused on detecting termination of rule execution. We describe an approach which combines static analysis of a rule set at compile-time and detection of endless loops during rule processing at runtime. The compile-time analysis technique is based on the distinction between mutual triggering and mutual activation of rules. This distinction motivates the introduction of two graphs defining rule interaction, called Triggering and Activation Graphs, respectively. This analysis technique allows us to identify reactive behaviors which are guaranteed to terminate and reactive behaviors which may lead to infinite rule processing. When termination cannot be guaranteed at compile-time, it is crucial to detect infinite rule executions at runtime. We propose a technique for identifying loops which is based on recognizing that a given situation has already occurred in the past and, therefore, will occur an infinite number of times in the future. This technique is potentially very expensive, therefore, we explain how it can be implemented in practice with limited computational effort. A particular use of this technique allows us to develop cycle monitors, which check that critical rule sequences, detected at compile time, do not repeat forever We bridge compile-time analysis to runtime monitoring by showing techniques, based on the result of rule analysis, for the identification of rule sets that can be independently monitored and for the optimal selection of cycle monitors.
Elena Baralis, Stefano Ceri, Stefano Paraboschi
IEEE Trans. Knowl. Data Eng.1
1997 Performance Evaluation of Rule Semantics in Active Databases
abstract
Different rule execution semantics may be available in the same active database system. We perform several simulation experiments to evaluate the performance trade-offs yielded by different execution semantics in various operating conditions. In particular, we evaluate the effect of executing transaction and rule statements that affect a varying number of data instances, and applications with different rule triggering breadth and depth. Since references to data changed by the database operation triggering the rules are commonly used in active rule programming, we also analyze the impact of its management on overall performance.
Elena Baralis, Andrea Bianco
ICDE1
1997 Materialized Views Selection in a Multidimensional Database
Elena Baralis, Stefano Paraboschi, Ernest Teniente
VLDB1
1997 Designing Templates for Mining Association Rules
Elena Baralis, Giuseppe Psaila
J. Intell. Inf. Syst.1
1996 Support Environment for Active Rule Design
Elena Baralis, Stefano Ceri, Piero Fraternali, Stefano Paraboschi
J. Intell. Inf. Syst.1
1996 Modularization Techniques for Active Rules Design
abstract
Active database systems can be used to establish and enforce data management policies. A large amount of the semantics that normally needs to be coded in application programs can be abstracted and assigned to active rules. This trend is sometimes called “knowledge independence” a nice consequence of achieving full knowledge independence is that data management policies can then effectively evolve just by modifying rules instead of application programs. Active rules, however, may be quite complex to understand and manage: rules react to arbitrary event sequences, they trigger each other, and sometimes the outcome of rule processing may depend on the order in which events occur or rules are scheduled. Although reasoning on a large collection of rules is very difficult, the task becomes more manageable when the rules are few. Therefore, we are convinced that modularization, similar to what happens in any software development process, is the key principle for designing active rules; however, this important notion has not been addressed so far. This article introduces a modularization technique for active rules called stratification; it presents a theory of stratification and indicates how stratification can be practically applied. The emphasis of this article is on providing a solution to a very concrete and practical problem; therefore, our approach is illustrated by several examples.
Elena Baralis, Stefano Ceri, Stefano Paraboschi
ACM Trans. Database Syst.1
1994 Declarative Specification of Constraint Maintenance
Elena Baralis, Stefano Ceri, Stefano Paraboschi
ER1
1994 An Algebraic Approach to Rule Analysis in Expert Database Systems
Elena Baralis, Jennifer Widom
VLDB1