VLDB 2026 Research / reviewers in the wild / expert
Tania Cerquitelli
dblp:20/5348
· DBLP profile ↗
39ranked-venue papers in the field
1as first author
20since 2021 · last 2026
0000-0002-9039-6226ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 11 (1 first)Database Systems & Data Management · 9Knowledge Engineering, Semantic Web & Information Systems · 9Big Data, Cloud & Distributed Data Systems · 8Information Retrieval & Web Search · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prime convolutional model: Breaking the ground for theoretical explainability
Francesco Panelli, Doaa Almhaithawi, Tania Cerquitelli, Alessandro Bellini |
Inf. Sci. | 3 |
| 2025 | HydroChronos: Forecasting Decades of Surface Water ChangeabstractForecasting surface water dynamics is crucial for water resource management and climate change adaptation. However, the field lacks comprehensive datasets and standardized benchmarks. In this paper, we introduce HydroChronos, a large-scale, multi-modal spatiotemporal dataset for surface water dynamics forecasting designed to address this gap. We couple the dataset with three forecasting tasks. The dataset includes over three decades of aligned Landsat 5 and Sentinel-2 imagery, climate data, and Digital Elevation Models for diverse lakes and rivers across Europe, North America, and South America. We also propose AquaClimaTempo UNet, a novel spatiotemporal architecture with a dedicated climate data branch, as a strong benchmark baseline. Our model significantly outperforms a Persistence baseline for forecasting future water dynamics by +14% and +11% F1 across change detection and direction of change classification tasks, and by +0.1 MAE on the magnitude of change regression. Finally, we conduct an Explainable AI analysis to identify the key climate variables and input channels that influence surface water change, providing insights to inform and guide future modeling efforts. Daniele Rege Cambrin, Eleonora Poeta, Eliana Pastor, Isaac Corley, Tania Cerquitelli, Elena Baralis, Paolo Garza |
SIGSPATIAL/GIS | 5 |
| 2025 | Towards Better Generalization and Interpretability in Unsupervised Concept-Based Models
Francesco De Santis, Philippe Bich, Gabriele Ciravegna, Pietro Barbiero, Tania Cerquitelli, Danilo Giordano |
ECML/PKDD (3) | 5 |
| 2025 | Advances on data management systems
Ladjel Bellatreche, Marlon Dumas, Panagiotis Karras, Raimundas Matulevicius, Silvia Chiusano, Tania Cerquitelli, Robert Wrembel |
Inf. Syst. | 6 |
| 2025 | Towards AI-Assisted Inclusive Language Writing in Italian Formal CommunicationsabstractFormal communications such as public calls, announcements, or regulations are supposed to exhibit respect for diversity in terms of gender, race, age, and disability. However, human writers often lack adequate inclusive writing skills. For instance, they tend to overuse the masculine as a neutral form, mainly because they are self-trained on biased text examples. To overcome this issue, we propose to leverage Generative Artificial Intelligence to support inclusive language writing. Focusing on formal Italian communications, we have designed and developed an AI-assisted tool for non-inclusive text detection and reformulation. Thanks to the joint work with a team of linguistic experts, we first define a set of linguistic criteria necessary to model inclusive writing forms in Italian. Based on these criteria, we collect and annotate a dataset of Italian administrative documents enriched with fine-grained inclusive annotations. Finally, we train deep learning models on the collected data for non-inclusive language detection and inclusive language reformulation tasks. We perform quantitative and human-driven evaluations on the trained models. The best detection model correctly classifies 89% of the sentences, whereas the best reformulation model produces 73% fully correct reformulations. Both models have been integrated into a writing assistance tool acting as a text proofreader and self-learning tool for non-expert writers, namely Inclusively . Once a non-inclusive piece of text is detected, the proposed approach suggests inclusive reformulations. The tool also provides explanations of the models’ outputs to increase system transparency. Furthermore, it allows expert end-users to provide further annotations for system fine-tuning. The trained models and the writing assistance tool are publicly available for research purposes. Salvatore Greco, Moreno La Quatra, Luca Cagliero, Tania Cerquitelli |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | Unsupervised Concept Drift Detection From Deep Learning Representations in Real-TimeabstractConcept drift is the phenomenon in which the underlying data distributions and statistical properties of a target domain change over time, leading to a degradation in model performance. Consequently, production models require continuous drift detection monitoring. Most drift detection methods to date are supervised, relying on ground-truth labels. However, they are inapplicable in many real-world scenarios, as true labels are often unavailable. Although recent efforts have proposed unsupervised drift detectors, many lack the accuracy required for reliable detection or are too computationally intensive for real-time use in high-dimensional, large-scale production environments. Moreover, they often fail to characterize or explain drift effectively. To address these limitations, we proposeDRIFTLENS, an unsupervised framework for real-time concept drift detection and characterization. Designed for deep learning classifiers handling unstructured data,DRIFTLENSleverages distribution distances in deep learning representations to enable efficient and accurate detection. Additionally, it characterizes drift by analyzing and explaining its impact on each label. Our evaluation across classifiers and data-types demonstrates thatDRIFTLENS(i) outperforms previous methods in detecting drift in 15/17 use cases; (ii) runs at least 5 times faster; (iii) produces drift curves that align closely with actual drift (correlation$\geq 0.85$); (iv) effectively identifies representative drift samples as explanations. Salvatore Greco, Bartolomeo Vacchetti, Daniele Apiletti, Tania Cerquitelli |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Decoding Narratives: Towards a Classification Analysis for Stereotypical Patterns in Italian News HeadlinesabstractMedia headlines shape our initial interpretation of news, framing narratives that influence societal engagement with political and social issues. Yet, they often rely on sensationalism and bias to capture readers’ attention.In this paper, we aim to uncover distinct patterns in Italian headline composition, examining how language and framing vary across political leanings. We analyze a dataset of daily Italian newspaper articles from two outlets with opposing political perspectives, anonymized as Newspaper A and Newspaper B. Our study encompasses the entire set of news and a subset of topics (n = 8) likely to contain stereotypes or clickbait headlines identified using a Large Language Model. Our methodology combines (1) a lexicometric analysis to identify characteristic words of each newspaper, and (2) the training of an accurate deep learning classifier (F 1 = 0.84) to learn specific patterns for categorizing headlines into these two perspectives and leveraging explainability techniques to extract and interpret these patterns.Our analysis reveals distinct tonal differences between the two newspapers: Newspaper A generally adopts a more balanced and nuanced approach, while Newspaper B often favors a more direct and sometimes provocative style, especially regarding topics like immigration and social justice. Additionally, Newspaper B’s headlines tend to be brief and punchy, in contrast to the longer, more detailed ones from Newspaper A. Despite these tonal differences, both outlets exhibit similar stereotypical patterns in their coverage, such as consistently emphasizing nationality and group distinctions in ways that can reinforce social stereotypes. This shared tendency suggests that, although their narrative strategies differ, both outlets could contribute to a broader pattern of stereotype reinforcement. Matteo Berta, Salvatore Greco, Giuseppe Tipaldo, Tania Cerquitelli |
IEEE Big Data | 4 |
| 2024 | DriftLens: A Concept Drift Detection Tool
Salvatore Greco, Bartolomeo Vacchetti, Daniele Apiletti, Tania Cerquitelli |
EDBT | 4 |
| 2024 | Equity, Diversity & Inclusion (EDI): Special Day at ACM KDD 2024abstractThe Equity, Diversity & Inclusion event is a special day organized in conjunction with KDD '24, the 30 𝑡ℎ ACM SIGKDD Conference on Knowledge Discovery and Data Mining, which will take place from Sunday, August 25 to Thursday, August 29, 2024 at the Center de Convencions Internacional de Barcelona in Barcelona, Spain.This special day, scheduled for August 28, 2024, promotes equity, diversity, and inclusion (EDI) in data science, artificial intelligence, and beyond.It will bring together academics, researchers, practitioners, and human resources professionals (i) to present algorithms, techniques, methodologies, and projects in data science that enable responsible data processing and modeling; (ii) to discuss policies, best practices, and guidelines to promote an inclusive work environment and effective collaboration; (iii) to share personal stories to encourage young researchers, including those from groups unrepresented in the research community, to develop strong careers in data science; and (iv) to collaboratively develop and discuss an EDI Manifesto to promote an inclusive workplace environment and guiding principles in the development of research activities. Tania Cerquitelli, Amin Mantrach |
KDD | 1 |
| 2024 | Workshop on Human-Interpretable AIabstractThis workshop aims to spearhead research on Human-Interpretable Artificial Intelligence (HI-AI) by providing: (i) a general overview of the key aspects of HI-AI, in order to equip all researchers with the necessary background and set of definitions; (ii) novel and interesting ideas coming from both invited talks and top paper contributions; (iii) the chance to engage in dialogue with prominent scientists during poster presentations and coffee breaks. The workshop welcomes contributions covering novel interpretable-by-design or post-hoc approaches, as well as theoretical analysis of existing works. Additionally, we accept visionary contributions speculating on the future potential of this field. Finally, we welcome contributions from related fields such as Ethical AI, Knowledge-driven Machine learning, Human-machine Interaction, but also applications in Medicine and Industry, and analyses from Regulatory experts. Gabriele Ciravegna, Mateo Espinosa Zarlenga, Pietro Barbiero, Francesco Giannini, Zohreh Shams, Damien Garreau, Mateja Jamnik, Tania Cerquitelli |
KDD | 8 |
| 2024 | Explaining deep convolutional models by measuring the influence of interpretable features in image classificationabstractAbstract The accuracy and flexibility of Deep Convolutional Neural Networks (DCNNs) have been highly validated over the past years. However, their intrinsic opaqueness is still affecting their reliability and limiting their application in critical production systems, where the black-box behavior is difficult to be accepted. This work proposes EBAnO, an innovative explanation framework able to analyze the decision-making process of DCNNs in image classification by providing prediction-local and class-based model-wise explanations through the unsupervised mining of knowledge contained in multiple convolutional layers. EBAnO provides detailed visual and numerical explanations thanks to two specific indexes that measure the features’ influence and their influence precision in the decision-making process. The framework has been experimentally evaluated, both quantitatively and qualitatively, by (i) analyzing its explanations with four state-of-the-art DCNN architectures, (ii) comparing its results with three state-of-the-art explanation strategies and (iii) assessing its effectiveness and easiness of understanding through human judgment, by means of an online survey. EBAnO has been released as open-source code and it is freely available online. Francesco Ventura, Salvatore Greco, Daniele Apiletti, Tania Cerquitelli |
Data Min. Knowl. Discov. | 4 |
| 2023 | GINN: Towards Gender InclusioNeural NetworkabstractToday’s data-driven systems and official statistics often oversimplify the concept of gender, reducing it to binary data, with far-reaching implications for policy development and equitable access to services. This simplification can lead to misclassification and discrimination against individuals who identify as non-binary.We are working to advance our research in this area to develop new, more equitable approaches that can avoid discrimination based on gender identity. Within this research framework, our primary focus is on mitigating the problem of underrepresentation and, in some cases, the complete absence of non-binary individuals in data collection.With this goal in mind, we present the GINN Gender InclusioNeural Network. This is our first attempt to develop an equitable neural network that accurately identifies gender in a multiclass context and includes individuals whose gender identity does not fall on the binary spectrum. To achieve this goal, we conducted a comprehensive comparative analysis of several fine-tuned neural network models. Our goal was to gain a deep understanding of the crucial distinguishing features in gender identify classification and to highlight the limitation of current methods using explainable AI techniques.The initial results are promising and demonstrate the effectiveness of a fine-tuned EfficientNetB0 model in accurately categorizing images of individuals into their self-reported gender, but we are skeptical about the application in a real-world scenario because of the amount of data available about non-binary people at the moment. Matteo Berta, Bartolomeo Vacchetti, Tania Cerquitelli |
IEEE Big Data | 3 |
| 2023 | Multi-View Latent DiffusionabstractMulti-view observations potentially offer a more comprehensive understanding of real-world phenomena compared to observations acquired from a single viewpoint. Existing models that utilize multi-view data often consider that all views are available during inference, but this assumption may not hold in practical scenarios. To address this limitation, we introduce MVLD, a novel method that, by employing a deterministic autoencoder and a score-based diffusion model, is capable of imputing missing views. We finally envision MVLD being used in a communication system for image transmission. Giuseppe Di Giacomo, Giulio Franzese, Tania Cerquitelli, Carla Fabiana Chiasserini, Pietro Michiardi |
IEEE Big Data | 3 |
| 2022 | Physical and mental health of university staff during the Covid-19 pandemicabstractThe 2020 Covid-19 pandemic caused a sudden and massive change in work organizations. One of the major consequences of the crisis was the acceleration towards teleworking, through the specific phenomenon of Mandatory Work From Home: the situations in which workers overnight found themselves to work seven days a week from their home environment, constantly online, often without adequate equipment and little to no preparation. Different workers reacted in different way to this important change, depending on age, gender, family characteristics and other impacting factors. Mandatory work from home and these other variables impacted employees’ physical and mental health, triggering or increasing symptoms of overwork and emotional exhaustion among others. This paper contributes to the literature on the impact of the pandemic on workers’ health by giving an overview of the effects of MWFH on university staff, using Politecnico di Torino as a case study. Alessandra Colombelli, Greta Temporin, Francesco Serraino, Tania Cerquitelli |
IEEE Big Data | 4 |
| 2022 | Promoting equity, diversity and inclusion: policies, strategies and future directions in higher education, research communities and businessabstractThis paper provides a multi-perspective vision of diversity and inclusion (D&I) projects aiming to promote equity in organisations seeking to build virtuous contexts where people can achieve positive professional and personal objectives. It introduces the understanding of D&I, best practices and outcomes of projects promoted in multicultural organisations, including academia, universities and research centres (Politecnico di Torino, university education in France and the French CNRS) and in leading international companies, namely Accenture and Nestlé. The paper gathers and extends the discussion and ideas exchanged in the D&I panel of the conference ADBIS-2022. Genoveva Vargas-Solar, Tania Cerquitelli, Arianna Montorsi, Stefania Salvai, Maria Teresa Sangineti, Jérôme Darmont, Cécile Favre |
IEEE Big Data | 2 |
| 2022 | A Dataset for Burned Area Delineation and Severity Estimation from Satellite ImageryabstractThe ability to correctly identify areas damaged by forest wildfires is essential to plan and monitor the restoration process and estimate the environmental damages after such catastrophic events. The wide availability of satellite data, combined with the recent development of machine learning and deep learning methodologies applied to the computer vision field, makes it extremely interesting to apply the aforementioned techniques to the field of automatic burned area detection. One of the main issues in such a context is the limited amount of labeled data, especially in the context of semantic segmentation. In this paper, we introduce a publicly available dataset for the burned area detection problem for semantic segmentation. The dataset contains 73 satellite images of different forests damaged by wildfires across Europe with a resolution of up to 10m per pixel. Data were collected from the Sentinel-2 L2A satellite mission and the target labels were generated from the Copernicus Emergency Management Service (EMS) annotations, with five different severity levels, ranging from undamaged to completely destroyed. Finally, we report the benchmark values obtained by applying a Convolutional Neural Network on the proposed dataset to address the burned area identification problem. Luca Colomba, Alessandro Farasin, Simone Monaco, Salvatore Greco, Paolo Garza, Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
CIKM | 8 |
| 2022 | Trusting deep learning natural-language models via local and global explanationsabstractAbstract Despite the high accuracy offered by state-of-the-art deep natural-language models (e.g., LSTM, BERT), their application in real-life settings is still widely limited, as they behave like a black-box to the end-user. Hence, explainability is rapidly becoming a fundamental requirement of future-generation data-driven systems based on deep-learning approaches. Several attempts to fulfill the existing gap between accuracy and interpretability have been made. However, robust and specialized eXplainable Artificial Intelligence solutions, tailored to deep natural-language models, are still missing. We propose a new framework, named T-EBAnO, which provides innovative prediction-local and class-based model-global explanation strategies tailored to deep learning natural-language models. Given a deep NLP model and the textual input data, T-EBAnO provides an objective, human-readable, domain-specific assessment of the reasons behind the automatic decision-making process. Specifically, the framework extracts sets of interpretable features mining the inner knowledge of the model. Then, it quantifies the influence of each feature during the prediction process by exploiting the normalized Perturbation Influence Relation index at the local level and the novel Global Absolute Influence and Global Relative Influence indexes at the global level. The effectiveness and the quality of the local and global explanations obtained with T-EBAnO are proved on an extensive set of experiments addressing different tasks, such as a sentiment-analysis task performed by a fine-tuned BERT model and a toxic-comment classification task performed by an LSTM model. The quality of the explanations proposed by T-EBAnO, and, specifically, the correlation between the influence index and human judgment, has been evaluated by humans in a survey with more than 4000 judgments. To prove the generality of T-EBAnO and its model/task-independent methodology, experiments with other models (ALBERT, ULMFit) on popular public datasets (Ag News and Cola) are also discussed in detail. Francesco Ventura, Salvatore Greco, Daniele Apiletti, Tania Cerquitelli |
Knowl. Inf. Syst. | 4 |
| 2021 | E-MIMIC: Empowering Multilingual Inclusive CommunicationabstractPreserving diversity and inclusion is becoming a compelling need in both industry and academia. The ability to use appropriate forms of writing, speaking, and gestures is not widespread even in formal communications such as public calls, public announcements, official reports, and legal documents. The improper use of linguistic expressions can foment unacceptable forms of exclusion, stereotypes as well as forms of verbal violence against minorities, including women. Furthermore, existing machine translation tools are not designed to generate inclusive content.The present paper investigates a joint effort of the research communities of linguistics and Deep Learning Natural Language Understanding in fighting against non-inclusive, prejudiced language forms. It presents a methodology aimed at tackling the improper use of language in formal communication, with a particular attention paid to Romanic languages (Italian, in particular). State-of-the-art Deep Language Modeling architectures are exploited to automatically identify non-inclusive text snippets, suggest alternative forms, and produce inclusive text rephrasing. A preliminary evaluation conducted on a benchmark dataset shows promising results, i.e., 85% accuracy in predicting inclusive/non-inclusive communications. Giuseppe Attanasio, Salvatore Greco, Moreno La Quatra, Luca Cagliero, Michela Tonti, Tania Cerquitelli, Rachele Raus |
IEEE BigData | 6 |
| 2021 | DS4ALL: All you need for democratizing data exploration and analysisabstractToday, large amounts of data are collected in various domains, presenting unprecedented economic and societal opportunities. Yet, at present, the exploitation of these data sets through data science methods is primarily dominated by AI-savvy users. From an inclusive perspective, there is a need for solutions that can democratise data science that can guide non-specialists intuitively to explore data collections and extract knowledge out of them. This paper introduces the vision of a new data science engine, called DS4ALL (Data Science for ALL), that empowers users who are neither computer nor AI experts to perform sophisticated data exploration and analysis tasks. Therefore, DS4ALL is based on a conversational and intuitive approach that insulates users from the complexity of AI algorithms. DS4ALL allows a dialogue-based approach that gives the user greater freedom of expression. It will enable them to communicate using natural language without requiring a high level of expertise on data-driven algorithms. User requests are interpreted and handled internally by the system in an automated manner, providing the user with the required output by masking the complexity of the data science workflow. The system can also collect feedback on the displayed results, leveraging these comments to address personalized data analysis sessions. The benefits of the envisioned system are discussed, and a use case is also presented to describe the innovative aspects. Paolo Bethaz, Khalid Belhajjame, Genoveva Vargas-Solar, Tania Cerquitelli |
IEEE BigData | 4 |
| 2021 | Estimating the job's pending time on a High-Performance Computing cluster through a hierarchical data-driven methodology
Fabio Carfi, Enrica Capitelli, Vladi Massimo Nosenzo, Tania Cerquitelli |
DOLAP | 4 |
| 2018 | Mining Sensor Data for Predictive Maintenance in the Automotive IndustryabstractPredictive maintenance is an ever-growing area of interest, spanning different fields and approaches. In the automotive industry faulty behaviors of the oxygen sensor are a key challenge to address. This paper presents OxyClog, a data-driven framework that, given a large number of time series collected from a vehicle's ECU (engine control unit), builds a model to predict if the oxygen sensor is currently unclogged, almost clogged (since the clogging of the sensor happens gradually), or clogged. OxyClog is characterized by a tailored preprocessing, which includes a custom and interpretable feature selection algorithm, along with a summarization strategy to transform a time-dependent problem into a time-independent one. Furthermore, a semi-supervised labeling methodology has been devised to use different data sources with different characteristics to define meaningful clogging labels. OxyClog integrates state-of-the-art classification algorithms - both interpretable and non-interpretable - to process real ECU data with good prediction performance. Flavio Giobergia, Elena Baralis, Maria Camuglia, Tania Cerquitelli, Marco Mellia, Alessandra Neri, Davide Tricarico, Alessia Tuninetti |
DSAA | 4 |
| 2018 | Characterizing unpredictable patterns in Wireless Sensor Network data
Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Antonio Attanasio |
Inf. Sci. | 2 |
| 2017 | All in a twitter: Self-tuning strategies for a deeper understanding of a crisis tweet collectionabstractNatural disasters have become more frequent during the past 20 years due to significant climate changes. These natural events are hotly debated on social networks like Twitter and a huge amount of short text messages are continuously and promptly exchanged with personal opinions, descriptions of the natural events and their corresponding consequences. The analysis of these large and complex data could help policy-makers to better understand the event as well as to set priorities. However, the correct configuration of the tweet mining process is still challenging due to variable data distribution and the availability of a large number of algorithms with different specific parameters. The analyst need to perform a large number of experiments to identify the best configuration for the overall knowledge discovery process. Innovative, scalable, and parameter-free solutions need to be explored to streamline the analytics process. This paper presents an enhanced version of PASTA (a distributed self-tuning engine) applied to a crisis tweet collection to group a corpus of tweets into cohesive and well-separated clusters with minimal analyst intervention. Experimental results performed on real data collected during natural disasters show the effectiveness of PASTA in discovering interesting groups of correlated tweets without selecting neither the algorithms nor their parameters. Evelina Di Corso, Francesco Ventura, Tania Cerquitelli |
IEEE BigData | 3 |
| 2017 | Discovering profitable stocks for intraday trading
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Fabio Pulvirenti |
Inf. Sci. | 3 |
| 2015 | Reducing the search space in ontology alignment using clustering techniques and topic identificationabstractOne of the current challenges in ontology alignment is scalability and one technique to deal with this issue is to reduce the search space for the generation of mapping suggestions. In this paper we develop a method to prune that search space by using clustering techniques and topic identification. Further, we provide experiments showing that we are able to generate partitions that allow for high quality alignments with a highly reduced effort for computation and validation of mapping suggestions for the parts of the ontologies in the partition. Other techniques will still be needed for finding mappings that are not in the partition. Agnese Chiatti, Zlatan Dragisic, Tania Cerquitelli, Patrick Lambrix |
K-CAP | 3 |
| 2015 | Digging deep into weighted patient data through multiple-level patterns
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza |
Inf. Sci. | 3 |
| 2015 | Scalable out-of-core itemset mining
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Alberto Grand |
Inf. Sci. | 2 |
| 2015 | MeTA: Characterization of Medical Treatments at Different Abstraction LevelsabstractPhysicians and health care organizations always collect large amounts of data during patient care. These large and high-dimensional datasets are usually characterized by an inherent sparseness. Hence, analyzing these datasets to figure out interesting and hidden knowledge is a challenging task. This article proposes a new data mining framework based on generalized association rules to discover multiple-level correlations among patient data. Specifically, correlations among prescribed examinations, drugs, and patient profiles are discovered and analyzed at different abstraction levels. The rule extraction process is driven by a taxonomy to generalize examinations and drugs into their corresponding categories. To ease the manual inspection of the result, a worthwhile subset of rules (i.e., nonredundant generalized rules) is considered. Furthermore, rules are classified according to the involved data features (medical treatments or patient profiles) and then explored in a top-down fashion: from the small subset of high-level rules, a drill-down is performed to target more specific rules. The experiments, performed on a real diabetic patient dataset, demonstrate the effectiveness of the proposed approach in discovering interesting rule groups at different abstraction levels. Dario Antonelli, Elena Baralis, Giulia Bruno, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Naeem Ahmed Mahoto |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | Expressive generalized itemsets
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Vincenzo D'Elia, Paolo Garza |
Inf. Sci. | 3 |
| 2013 | New Trends in Databases and Information Systems: Contributions from ADBIS 2013
Yamine Aït-Ameur, Witold Andrzejewski, Ladjel Bellatreche, Barbara Catania, Tania Cerquitelli, Silvia Chiusano, Matteo Golfarelli, Giovanna Guerrini, Krzysztof Kaczmarski, Mirko Kämpf, Alfons Kemper, Tobias Lauer, Boris Novikov 0001, Themis Palpanas, Jaroslav Pokorný, Stefano Rizzi, Athena Vakali |
ADBIS (2) | 5 |
| 2013 | Analysis of Twitter Data Using a Multiple-level Clustering Strategy
Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Luigi Grimaudo, Xin Xiao 0002 |
MEDI | 2 |
| 2013 | Early prediction of the highest workload in incremental cardiopulmonary testsabstractIncremental tests are widely used in cardiopulmonary exercise testing, both in the clinical domain and in sport sciences. The highest workload (denoted Wpeak) reached in the test is key information for assessing the individual body response to the test and for analyzing possible cardiac failures and planning rehabilitation, and training sessions. Being physically very demanding, incremental tests can significantly increase the body stress on monitored individuals and may cause cardiopulmonary overload. This article presents a new approach to cardiopulmonary testing that addresses these drawbacks. During the test, our approach analyzes the individual body response to the exercise and predicts the Wpeakvalue that will be reached in the test and an evaluation of its accuracy. When the accuracy of the prediction becomes satisfactory, the test can be prematurely stopped, thus avoiding its entire execution. To predict Wpeak, we introduce a new index, the CardioPulmonary Efficiency Index (CPE), summarizing the cardiopulmonary response of the individual to the test. Our approach analyzes the CPE trend during the test, together with the characteristics of the individual, and predicts Wpeak. A K-nearest-neighbor-based classifier and an ANN-based classier are exploited for the prediction. The experimental evaluation showed that the Wpeakvalue can be predicted with a limited error from the first steps of the test. Elena Baralis, Tania Cerquitelli, Silvia Chiusano, Vincenzo D'Elia, Riccardo Molinari, Davide Susta |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Generalized association rule mining with constraints
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza |
Inf. Sci. | 3 |
| 2011 | Semi-Automatic Ontology Construction by Exploiting Functional Dependencies and Association RulesabstractThis paper presents a novel semi-automatic approach to construct conceptual ontologies over structured data by exploiting both the schema and content of the input dataset. It effectively combines two well-founded database and data mining techniques, i.e., functional dependency discovery and association rule mining, to support domain experts in the construction of meaningful ontologies, tailored to the analyzed data, by using Description Logic (DL). To this aim, functional dependencies are first discovered to highlight valuable conceptual relationships among attributes of the data schema (i.e., among concepts). The set of discovered correlations effectively support analysts in the assertion of the Tbox ontological statements (i.e., the statements involving shared data conceptualizations and their relationships). Then, the analyst-validated dependencies are exploited to drive the association rule mining process. Association rules represent relevant and hidden correlations among data content and they are used to provide valuable knowledge at the instance level. The pushing of functional dependency constraints into the rule mining process allows analysts to look into and exploit only the most significant data item recurrences in the assertion of the Abox ontological statements (i.e., the statements involving concept instances and their relationships). Luca Cagliero, Tania Cerquitelli, Paolo Garza |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2011 | Energy-saving models for wireless sensor networks
Daniele Apiletti, Elena Baralis, Tania Cerquitelli |
Knowl. Inf. Syst. | 3 |
| 2011 | CAS-Mine: providing personalized services in context-aware applications by means of generalized rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti |
Knowl. Inf. Syst. | 3 |
| 2010 | Constrained itemset mining on a sequence of incoming data blocksabstractMany real-life databases are updated by means of incoming business information. In these databases (e.g., transactional data from large retail chains, call-detail records), the content evolves through periodical insertions (or deletions) of data blocks. Since data evolve over time, algorithms have to be devised to incrementally update data mining models. This paper presents a novel index, called I-Forest, to support itemset mining on incoming data blocks, where new blocks are inserted periodically, or old blocks are discarded. The I-Forest structure provides a complete data representation and allows different kind of analyses (e.g., investigate quarterly data), besides supporting user-defined time and support constraints. The I-Forest index has been implemented into the PostgreSQL open source DBMS and exploits its physical level access methods. Experiments, run for both sparse and dense data distributions, show the effectiveness of the I-Forest-based approach to perform itemset mining with both time and support constraints. The execution time of the I-Forest-based itemset mining technique is often faster than the Prefix-Tree algorithm accessing static data on flat files. © 2010 Wiley Periodicals, Inc. Elena Baralis, Tania Cerquitelli, Silvia Chiusano |
Int. J. Intell. Syst. | 2 |
| 2009 | IMine: Index Support for Item Set MiningabstractThis paper presents the IMine index, a general and compact structure which provides tight integration of item set extraction in a relational DBMS. Since no constraint is enforced during the index creation phase, IMine provides a complete representation of the original database. To reduce the I/O cost, data accessed together during the same extraction phase are clustered on the same disk block. The IMine index structure can be efficiently exploited by different item set extraction algorithms. In particular, IMine data access methods currently support the FP-growth and LCM v.2 algorithms, but they can straightforwardly support the enforcement of various constraint categories. The IMine index has been integrated into the PostgreSQL DBMS and exploits its physical level access methods. Experiments, run for both sparse and dense data distributions, show the efficiency of the proposed index and its linear scalability also for large datasets. Item set mining supported by the IMine index shows performance always comparable with, and sometimes better than, state of the art algorithms accessing data on flat file. Elena Baralis, Tania Cerquitelli, Silvia Chiusano |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2005 | Index Support for Frequent Itemset Mining in a Relational DBMSabstractMany efforts have been devoted to couple data mining activities with relational DBMSs, but a true integration into the relational DBMS kernel has been rarely achieved. This paper presents a novel indexing technique, which represents transactions in a succinct form, appropriate for tightly integrating frequent itemset mining in a relational DBMS. The data representation is complete, i.e., no support threshold is enforced, in order to allow reusing the index for mining itemsets with any support threshold. Furthermore, an appropriate structure of the stored information has been devised, in order to allow a selective access of the index blocks necessary for the current extraction phase. The index has been implemented into the PostgreSQL open source DBMS and exploits its physical level access methods. Experiments have been run for various datasets, characterized by different data distributions. The execution time of the frequent itemset extraction task exploiting the index is always comparable with and sometime faster than a C++ implementation of the FP-growth algorithm accessing data stored on a flat file. Elena Baralis, Tania Cerquitelli, Silvia Chiusano |
ICDE | 2 |