Luca Cagliero

dblp:60/7439 · DBLP profile ↗
← Back
37ranked-venue papers in the field
10as first author
19since 2021 · last 2026
0000-0002-7185-5247ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 9 (3 first)Data Mining & Knowledge Discovery · 8 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 8 (3 first)Information Retrieval & Web Search · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 5 (2 first)
YearPublicationVenuePosition
2026 Lightweight Strategies to Mitigate Small VideoLLM's Challenges in Video Summarization
Lorenzo Vaiani, Luca Cagliero, Wiktoria Woronko
DEXA (1)2
2026 From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architectures
Chiara Bonfanti, Davide Colaiacomo, Luca Cagliero, Cataldo Basile
ICWE3
2026 A Unified Approach for Sexism Detection in Social Media Memes Under Hard and Soft Evaluation Settings
Lorenzo Calogiuri, Elöd Egyed-Zsigmond, Nathan Nowakowski, Luca Cagliero
ICWE4
2025 KIEPrompter: Leveraging Lightweight Models' Predictions for Cost-Effective Key Information Extraction using Vision LLMs
abstract
Key information extraction (KIE) from visually rich documents, such as receipts and forms, involves a deep understanding of textual, visual, and layout feature information. Transformers fine-tuned for KIE achieve state-of-the-art performance but lack generality and portability across different domains. In contrast, vision large language models (VLLMs) offer higher flexibility and zero-shot capability but fall short with domain-specific layout relations unless performing a resource-demanding supervised fine-tuning. To reach the best compromise solution between lightweight models and VLLMs, we propose KIEPrompter, a cost-effective LLM-based KIE approach that leverages the predictions of lightweight models as external knowledge injected into VLLM prompts. By incorporating these auxiliary predictions, VLLMs are guided to attend relevant multimodal content without ad hoc training. The accuracy results achieved by KIEPrompter in three benchmark document collections are superior to those of VLLMs in both zero-shot and layout-sensitive scenarios. We compare various strategies for incorporating lightweight model predictions, ranging from coarse-grained predictions without explicit confidence scores to fine-grained per-element network logits. We also demonstrate that our approach is robust to the absence of specific classes in trained lightweight models, as the VLLMs' pre-training compensates for the limited generality of lightweight models.
Lorenzo Vaiani, Yihao Ding, Luca Cagliero, Jean Lee, Paolo Garza, Josiah Poon, Soyeon Caren Han
CIKM3
2025 MAD: Multicriteria Anomaly Detection of Suspicious Financial Accounts from Billions of Cash Transactions
abstract
This paper presents a real-world deployment case study on using unsupervised anomaly detection for Anti-Money Laundering (AML).Using more than 2 billion anonymized bank transactions that Intesa Sanpaolo, a primary Italian financial institution, registered over 8 months, we developed, tuned and deployed a machine learning pipeline in production.Experts from Intesa Sanpaolo validated the performance of our approach against the institution's traditional rule-based system and checked new real-world cases the system allowed them to identify.Besides increasing both precision and recall by a factor of 6 in the detection of high-risk cases, our pipeline raises 200+ additional alerts during the 8-month period, manually identified by branch managers, but missed by the rulebased system.More importantly, a manual inspection of 100 new unseen cases revealed 28 significant previously unreported cases.The pipeline, now fully deployed in Intesa Sanpaolo's Transaction Monitoring system, highlights the advantages of machine learning over traditional approaches typically adopted in this traditionally very conservative sector.
Giordano Paoletti, Flavio Giobergia, Danilo Giordano, Luca Cagliero, Silvia Ronchiadin, Dario Moncalvo, Marco Mellia, Elena Baralis
KDD (2)4
2025 Towards AI-Assisted Inclusive Language Writing in Italian Formal Communications
abstract
Formal communications such as public calls, announcements, or regulations are supposed to exhibit respect for diversity in terms of gender, race, age, and disability. However, human writers often lack adequate inclusive writing skills. For instance, they tend to overuse the masculine as a neutral form, mainly because they are self-trained on biased text examples. To overcome this issue, we propose to leverage Generative Artificial Intelligence to support inclusive language writing. Focusing on formal Italian communications, we have designed and developed an AI-assisted tool for non-inclusive text detection and reformulation. Thanks to the joint work with a team of linguistic experts, we first define a set of linguistic criteria necessary to model inclusive writing forms in Italian. Based on these criteria, we collect and annotate a dataset of Italian administrative documents enriched with fine-grained inclusive annotations. Finally, we train deep learning models on the collected data for non-inclusive language detection and inclusive language reformulation tasks. We perform quantitative and human-driven evaluations on the trained models. The best detection model correctly classifies 89% of the sentences, whereas the best reformulation model produces 73% fully correct reformulations. Both models have been integrated into a writing assistance tool acting as a text proofreader and self-learning tool for non-expert writers, namely Inclusively . Once a non-inclusive piece of text is detected, the proposed approach suggests inclusive reformulations. The tool also provides explanations of the models’ outputs to increase system transparency. Furthermore, it allows expert end-users to provide further annotations for system fine-tuning. The trained models and the writing assistance tool are publicly available for research purposes.
Salvatore Greco, Moreno La Quatra, Luca Cagliero, Tania Cerquitelli
ACM Trans. Intell. Syst. Technol.3
2025 QATCH: Automatic Evaluation of SQL-Centric Tasks on Proprietary Data
abstract
Tabular Representation Learning (TRL) and Large Language Models (LLMs) have become established for tackling Question Answering (QA) and Semantic Parsing (SP) tasks on tabular data. State-of-the-art models are pre-trained and evaluated on large open-domain datasets. However, the performance on existing QA and SP benchmarks is not necessarily representative of that achieved on proprietary data as the characteristics of the input and the complexity of the posed queries show high variability. To tackle this challenge, our goal is to allow end-users to evaluate TRL and LLM performance on their own proprietary data. We present Query-Aided TRL CHecklist (QATCH), a toolbox to automatically generate a testing checklist tailored to QA and SP. QATCH provides a testing suite highlighting models’ strengths and weaknesses on relational tables unseen at training time. The proposed toolbox relies on a SQL query generator that crafts tests of varying types and complexity including, amongst others, tests on null values, projection, selections, joins, group by, and having clauses. QATCH also supports a set of general cross-task performance metrics providing more insights into SQL-related model capabilities than currently used metrics. The empirical results, achieved by state-of-the-art TRL models and LLMs, show substantial performance differences (1) between existing benchmarks and proprietary data, (2) across queries of different complexity.
Simone Papicchio, Paolo Papotti, Luca Cagliero
ACM Trans. Intell. Syst. Technol.3
2024 Empowering University-level Help Desk for International Applicants with AI Chatbots
abstract
In this paper, we present the design, development, and testing of AI chatbots to support the help desk service provided by the Recruitment and Admissions Unit at Politecnico di Torino, a technical university located in the north-west of Italy. We explore the use of data-driven, AI-based conversational agents providing targeted responses to applicants’ queries based on both the past student-office interactions through a ticketing system and a collection of Frequently Asked Questions (FAQs). With an ever-increasing number of requests from international applicants (20k+ requests from 100+ countries since January 2024), the adoption of AI-based solutions allows a significant reduction of the average waiting time per request, fostering the application, enrollment, and integration of foreign students coming from a variety of different countries. We develop separate chatbot systems to handle FAQs and manage inquiries submitted through the university’s ticketing system. We explore the use of intent-based and generative approaches as well as a combination of the two. Our findings indicate that the intent-based approach excels in handling FAQs and well-defined tickets, whereas the generative-only strategy is more suitable for open-ended requests.
Zahra Karimi, Giuseppe Gallipoli, Luca Cagliero, Francesca Chicco, Carola Rosa, Alessandra Sechi
IEEE Big Data3
2024 DQNC2S: DQN-Based Cross-Stream Crisis Event Summarizer
Daniele Rege Cambrin, Luca Cagliero, Paolo Garza
ECIR (3)2
2024 Self-supervised Text Style Transfer Using Cycle-Consistent Adversarial Networks
abstract
Text Style Transfer (TST) is a relevant branch of natural language processing that aims to control the style attributes of a piece of text while preserving its original content. To address TST in the absence of parallel data, Cycle-consistent Generative Adversarial Networks (CycleGANs) have recently emerged as promising solutions. Existing CycleGAN-based TST approaches suffer from the following limitations: (1) They apply self-supervision, based on the cycle-consistency principle, in the latent space. This approach turns out to be less robust to mixed-style inputs, i.e., when the source text is partly in the original and partly in the target style; (2) Generators and discriminators rely on recurrent networks, which are exposed to known issues with long-term text dependencies; (3) The target style is weakly enforced, as the discriminator distinguishes real from fake sentences without explicitly accounting for the generated text's style. We propose a new CycleGAN-based TST approach that applies self-supervision directly at the sequence level to effectively handle mixed-style inputs and employs Transformers to leverage the attention mechanism for both text encoding and decoding. We also employ a pre-trained style classifier to guide the generation of text in the target style while maintaining the original content's meaning. The experimental results achieved on the formality and sentiment transfer tasks show that our approach outperforms existing ones, both CycleGAN-based and not (including an open-source Large Language Model), on benchmark data and shows better robustness to mixed-style inputs.
Moreno La Quatra, Giuseppe Gallipoli, Luca Cagliero
ACM Trans. Intell. Syst. Technol.3
2024 Efficient Neural Network-Based Estimation of Interval Shapley Values
abstract
The use of Shapley Values (SVs) to explain machine learning model predictions is established. Recent research efforts have been devoted to generating efficient Neural Network-based SVs estimates. However, the variability of the generated estimates, which depend on the selected data sampling, model, and training parameters, brings the reliability of such estimates into question. By leveraging the concept of Interval SVs, we propose to incorporate SVs uncertainty directly into the learning process. Specifically, we explain ensemble models composed of multiple predictors, each one generating potentially different outcomes. Unlike all existing approaches, the explainer design is tailored to Interval SVs learning instead of SVs only. We present three new Network-based explainers relying on different ISV paradigms, i.e., a Multi-Task Learning network inspired by the Shapley value's weighted least squares characterization and two Interval Shapley-Like Value Neural estimators. The experiments thoroughly evaluate the new approaches on ten benchmark datasets, looking for the best compromise between intervals’ accuracy and explainers’ efficiency.
Davide Napolitano, Lorenzo Vaiani, Luca Cagliero
IEEE Trans. Knowl. Data Eng.3
2023 Detecting industrial vehicles' duty levels using contrastive learning
abstract
Industrial vehicles equipped with CAN bus devices transmit large volumes of IoT signals. The analysis of CAN bus data can be helpful to monitor the current vehicles’ workload, namely the vehicle duty levels, in an automated fashion. Despite the use of machine learning techniques to automatically detect vehicle duty levels is particularly appealing, existing approaches are challenged by the high cost of human data annotation and by the high heterogeneity of the analyzed vehicles types and models. In this paper, we present a self-supervised approach to automatically detect vehicles’ duty levels based on contrastive learning. The multivariate CAN Bus signals are first divided into fixed-sized segments and then embedded into a vector space shared by all vehicles of the same model by leveraging a constrastive approach with a mixup augmentation strategy. The key idea is to embed similar segments in close proximity by self-learning a model-specific clustering, which allows automatic duty level assignment with minimal human supervision. We validate the proposed approach in a real industrial use case, analyzing CAN Bus data acquired from test heavy-duty vehicles. Data were provided by a multinational Internet-of-Things company specialized in telematics solution. The experiments show clustering performance superior to state-of-the-art models and as well as an higher ability to differentiate between Moving and Working duty levels.
Luca Cagliero, Silvia Buccafusco, Francesco Vaccarino, Lucia Salvatori, Riccardo Loti
IEEE Big Data1
2023 Early portfolio pruning: a scalable approach to hybrid portfolio selection
abstract
Abstract Driving the decisions of stock market investors is among the most challenging financial research problems. Markowitz’s approach to portfolio selection models stock profitability and risk level through a mean–variance model, which involves estimating a very large number of parameters. In addition to requiring considerable computational effort, this raises serious concerns about the reliability of the model in real-world scenarios. This paper presents a hybrid approach that combines itemset extraction with portfolio selection. We propose to adapt Markowitz’s model logic to deal with sets of candidate portfolios rather than with single stocks. We overcome some of the known issues of the Markovitz model as follows: (i) Complexity: we reduce the model complexity, in terms of parameter estimation, by studying the interactions among stocks within a shortlist of candidate stock portfolios previously selected by an itemset mining algorithm. (ii) Portfolio-level constraints: we not only perform stock-level selection, but also support the enforcement of arbitrary constraints at the portfolio level, including the properties of diversification and the fundamental indicators. (iii) Usability: we simplify the decision-maker’s work by proposing a decision support system that enables flexible use of domain knowledge and human-in-the-loop feedback. The experimental results, achieved on the US stock market, confirm the proposed approach’s flexibility, effectiveness, and scalability.
Daniele Giovanni Gioia, Jacopo Fior, Luca Cagliero
Knowl. Inf. Syst.3
2023 Density-Based Clustering by Means of Bridge Point Identification
abstract
Density-based clustering focuses on defining clusters consisting of contiguous regions characterized by similar densities of points. Traditional approaches identify core points first, whereas more recent ones initially identify the cluster borders and then propagate cluster labels within the delimited regions. Both strategies encounter issues in presence of multi-density regions or when clusters are characterized by noisy borders. To overcome the above issues, we present a new clustering algorithm that relies on the concept of bridge point. A bridge point is a point whose neighborhood includes points of different clusters. The key idea is to use bridge points, rather than border points, to partition points into clusters. We have proved that a correct bridge point identification yields a cluster separation consistent with the expectation. To correctly identify bridge points in absence of a priori cluster information we leverage an established unsupervised outlier detection algorithm. Specifically, we empirically show that, in most cases, the detected outliers are actually a superset of the bridge point set. Therefore, to define clusters we spread cluster labels like a wildfire until an outlier, acting as a candidate bridge point, is reached. The proposed algorithm performs statistically better than state-of-the-art methods on a large set of benchmark datasets and is particularly robust to the presence of intra-cluster multiple densities and noisy borders.
Luca Colomba, Luca Cagliero, Paolo Garza
IEEE Trans. Knowl. Data Eng.2
2022 Generating Comparative Explanations of Financial Time Series
Jacopo Fior, Luca Cagliero, Tommaso Calò
ADBIS2
2022 Legal Entity Disambiguation for Financial Crime Detection
abstract
Transaction Monitoring is one of the main labor-intensive tasks of anti-financial crime and it requires to scrutinise billions of transactions per month against possible crimes. The first step in the process is the correct identification of the involved parties. This foundational step defines the focal entities on which transaction monitoring algorithms rely to spot suspicious events. Unfortunately, the loose syntax of protocols and the free text fields of inter-banking communications make party disambiguation particularly challenging. The first step of a fully automated data-driven strategy is thus the detection of the actual entity owning or using a given account.In this paper, we leverage data-driven techniques to identify and disambiguate the owners of accounts involved in cross-border international transactions when a Financial Institution only knows a minority fraction of such parties as its own customers. For this, we propose a data science pipeline relying on hierarchical clustering to capture similarities among names of parties involved in actual transactions. We test and tune the proposed approach using a large, real-world, multi-language, proprietary dataset of actual international transactions. Our highly parallel implementation completes the identification of parties that share an account and identifies all accounts owned by a party with f-score higher than 0.8.
Jacopo Fior, Thomas Favale, Luca Cagliero, Danilo Giordano, Marco Mellia, Elena Baralis, Silvia Ronchiadin, Paolo Baracco, Dario Moncalvo
IEEE Big Data3
2022 Mining spatiotemporally invariant patterns
abstract
Discovering patterns that represent key spatial or temporal dependencies among data is a well-known exploratory data mining task. However, prior works either separately analyze spatial and temporal dependencies or discover joint spatiotemporal properties of specific trajectories observed over a region of interest. With the goal of generalizing the information provided by spatiotemporal patterns, in this paper we extract sequences of discrete events showing spatiotemporally invariant properties. We seek patterns whose corresponding instances in the source data differ only due to an invariant spatiotemporal transformation. We denote such a new type of patterns as SpatioTemporally Invariant. We also propose an efficient algorithm to mine STInvs and validate its efficiency and effectiveness on real data.
Luca Colomba, Luca Cagliero, Paolo Garza
SIGSPATIAL/GIS2
2021 E-MIMIC: Empowering Multilingual Inclusive Communication
abstract
Preserving diversity and inclusion is becoming a compelling need in both industry and academia. The ability to use appropriate forms of writing, speaking, and gestures is not widespread even in formal communications such as public calls, public announcements, official reports, and legal documents. The improper use of linguistic expressions can foment unacceptable forms of exclusion, stereotypes as well as forms of verbal violence against minorities, including women. Furthermore, existing machine translation tools are not designed to generate inclusive content.The present paper investigates a joint effort of the research communities of linguistics and Deep Learning Natural Language Understanding in fighting against non-inclusive, prejudiced language forms. It presents a methodology aimed at tackling the improper use of language in formal communication, with a particular attention paid to Romanic languages (Italian, in particular). State-of-the-art Deep Language Modeling architectures are exploited to automatically identify non-inclusive text snippets, suggest alternative forms, and produce inclusive text rephrasing. A preliminary evaluation conducted on a benchmark dataset shows promising results, i.e., 85% accuracy in predicting inclusive/non-inclusive communications.
Giuseppe Attanasio, Salvatore Greco, Moreno La Quatra, Luca Cagliero, Michela Tonti, Tania Cerquitelli, Rachele Raus
IEEE BigData4
2021 Summarize Dates First: A Paradigm Shift in Timeline Summarization
abstract
Timeline summarization aims at presenting long news stories in a compact manner. State-of-the-art approaches first select the most relevant dates from the original event timeline then produce per-date news summaries. Date selection is driven by either per-date news content or date-level references. When coping with complex event data, characterized by inherent news flow redundancy, this pipeline may encounter relevant issues in both date selection and summarization due to a limited use of news content in date selection and no use of high-level temporal references (e.g., the past month). This paper proposes a paradigm shift in timeline summarization aimed at overcoming the above issues. It presents a new approach, namely Summarize Date First, which focuses on first generating date-level summaries then selecting the most relevant dates on top of summarized knowledge. In the latter stage, it performs date aggregations to consider high-level temporal references as well. The proposed pipeline also supports frequent incremental timeline updates more efficiently than previous approaches. We tested our unsupervised approach both on existing benchmark datasets and on a newly proposed benchmark dataset describing the COVID-19 news timeline. The achieved results were superior to state-of-the-art unsupervised methods and competitive against supervised ones.
Moreno La Quatra, Luca Cagliero, Elena Baralis, Alberto Messina, Maurizio Montagnuolo
SIGIR2
2019 ELSA: A Multilingual Document Summarization Algorithm Based on Frequent Itemsets and Latent Semantic Analysis
abstract
Sentence-based summarization aims at extracting concise summaries of collections of textual documents. Summaries consist of a worthwhile subset of document sentences. The most effective multilingual strategies rely on Latent Semantic Analysis (LSA) and on frequent itemset mining, respectively. LSA-based summarizers pick the document sentences that cover the most important concepts. Concepts are modeled as combinations of single-document terms and are derived from a term-by-sentence matrix by exploiting Singular Value Decomposition (SVD). Itemset-based summarizers pick the sentences that contain the largest number of frequent itemsets, which represent combinations of frequently co-occurring terms. The main drawbacks of existing approaches are (i) the inability of LSA to consider the correlation between combinations of multiple-document terms and the underlying concepts, (ii) the inherent redundancy of frequent itemsets because similar itemsets may be related to the same concept, and (iii) the inability of itemset-based summarizers to correlate itemsets with the underlying document concepts. To overcome the issues of both of the abovementioned algorithms, we propose a new summarization approach that exploits frequent itemsets to describe all of the latent concepts covered by the documents under analysis and LSA to reduce the potentially redundant set of itemsets to a compact set of uncorrelated concepts. The summarizer selects the sentences that cover the latent concepts with minimal redundancy. We tested the summarization algorithm on both multilingual and English-language benchmark document collections. The proposed approach performed significantly better than both itemset- and LSA-based summarizers, and better than most of the other state-of-the-art approaches.
Luca Cagliero, Paolo Garza, Elena Baralis
ACM Trans. Inf. Syst.1
2018 Characterizing unpredictable patterns in Wireless Sensor Network data
Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Antonio Attanasio
Inf. Sci.1
2017 Summarization of emergency news articles driven by relevance feedback
abstract
Many articles on the same news are daily published by online newspapers and by various social media. To ease news article exploration sentence-based summarization algorithms aim at automatically generating for each news a summary consisting of the most salient sentences in the original articles. However, since sentence selection is error-prone, the automatically generated summaries are still subject to manual validation by domain experts. If the validation step not only focuses on pruning less relevant content but also on enriching summaries with missing yet relevant sentences this activity may become extremely time consuming. The paper focuses on summarizing news articles by means of an itemset-based technique. To tune summarizer performance a relevance feedback given on sentences is exploited to drive the generation of a new, more targeted summary. The feedback indicates the pertinence of the sentences that are already in the summary. Among the words or the word combinations selected by the summarization model, those occurring in sentences with high feedback score represent concepts that may be deemed as particularly relevant. Therefore, they are exploited to drive the new sentence selection process. The proposed approach was tested on collections of news articles reporting emergency situations. The results show the effectiveness of the proposed approach.
Luca Cagliero
IEEE BigData1
2017 Discovering profitable stocks for intraday trading
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Fabio Pulvirenti
Inf. Sci.2
2015 Digging deep into weighted patient data through multiple-level patterns
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza
Inf. Sci.2
2015 MeTA: Characterization of Medical Treatments at Different Abstraction Levels
abstract
Physicians and health care organizations always collect large amounts of data during patient care. These large and high-dimensional datasets are usually characterized by an inherent sparseness. Hence, analyzing these datasets to figure out interesting and hidden knowledge is a challenging task. This article proposes a new data mining framework based on generalized association rules to discover multiple-level correlations among patient data. Specifically, correlations among prescribed examinations, drugs, and patient profiles are discovered and analyzed at different abstraction levels. The rule extraction process is driven by a taxonomy to generalize examinations and drugs into their corresponding categories. To ease the manual inspection of the result, a worthwhile subset of rules (i.e., nonredundant generalized rules) is considered. Furthermore, rules are classified according to the involved data features (medical treatments or patient profiles) and then explored in a top-down fashion: from the small subset of high-level rules, a drill-down is performed to target more specific rules. The experiments, performed on a real diabetic patient dataset, demonstrate the effectiveness of the proposed approach in discovering interesting rule groups at different abstraction levels.
Dario Antonelli, Elena Baralis, Giulia Bruno, Luca Cagliero, Tania Cerquitelli, Silvia Chiusano, Paolo Garza, Naeem Ahmed Mahoto
ACM Trans. Intell. Syst. Technol.4
2015 MWI-Sum: A Multilingual Summarizer Based on Frequent Weighted Itemsets
abstract
Multidocument summarization addresses the selection of a compact subset of highly informative sentences, i.e., the summary, from a collection of textual documents. To perform sentence selection, two parallel strategies have been proposed: (a) apply general-purpose techniques relying on data mining or information retrieval techniques, and/or (b) perform advanced linguistic analysis relying on semantics-based models (e.g., ontologies) to capture the actual sentence meaning. Since there is an increasing need for processing documents written in different languages, the attention of the research community has recently focused on summarizers based on strategy (a). This article presents a novel multilingual summarizer, namely MWI-Sum (Multilingual Weighted Itemset-based Summarizer), that exploits an itemset-based model to summarize collections of documents ranging over the same topic. Unlike previous approaches, it extracts frequent weighted itemsets tailored to the analyzed collection and uses them to drive the sentence selection process. Weighted itemsets represent correlations among multiple highly relevant terms that are neglected by previous approaches. The proposed approach makes minimal use of language-dependent analyses. Thus, it is easily applicable to document collections written in different languages. Experiments performed on benchmark and real-life collections, English-written and not, demonstrate that the proposed approach performs better than state-of-the-art multilingual document summarizers.
Elena Baralis, Luca Cagliero, Alessandro Fiori, Paolo Garza
ACM Trans. Inf. Syst.2
2014 Expressive generalized itemsets
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Vincenzo D'Elia, Paolo Garza
Inf. Sci.2
2014 Infrequent Weighted Itemset Mining Using Frequent Pattern Growth
abstract
Frequent weighted itemsets represent correlations frequently holding in data in which items may weight differently. However, in some contexts, e.g., when the need is to minimize a certain cost function, discovering rare data correlations is more interesting than mining frequent ones. This paper tackles the issue of discovering rare and weighted itemsets, i.e., the infrequent weighted itemset (IWI) mining problem. Two novel quality measures are proposed to drive the IWI mining process. Furthermore, two algorithms that perform IWI and Minimal IWI mining efficiently, driven by the proposed measures, are presented. Experimental results show efficiency and effectiveness of the proposed approach.
Luca Cagliero, Paolo Garza
IEEE Trans. Knowl. Data Eng.1
2013 Improving classification models with taxonomy information
Luca Cagliero, Paolo Garza
Data Knowl. Eng.1
2013 GraphSum: Discovering correlations among multiple terms for graph-based summarization
Elena Baralis, Luca Cagliero, Naeem Ahmed Mahoto, Alessandro Fiori
Inf. Sci.2
2013 Itemset generalization with cardinality-based constraints
Luca Cagliero, Paolo Garza
Inf. Sci.1
2013 Personalized tag recommendation based on generalized rules
abstract
Tag recommendation is focused on recommending useful tags to a user who is annotating a Web resource. A relevant research issue is the recommendation of additional tags to partially annotated resources, which may be based on either personalized or collective knowledge. However, since the annotation process is usually not driven by any controlled vocabulary, the collections of user-specific and collective annotations are often very sparse. Indeed, the discovery of the most significant associations among tags becomes a challenging task. This article presents a novel personalized tag recommendation system that discovers and exploits generalized association rules, that is, tag correlations holding at different abstraction levels, to identify additional pertinent tags to suggest. The use of generalized rules relevantly improves the effectiveness of traditional rule-based systems in coping with sparse tag collections, because: (i) correlations hidden at the level of individual tags may be anyhow figured out at higher abstraction levels and (ii) low-level tag associations discovered from collective data may be exploited to specialize high-level associations discovered in the user-specific context. The effectiveness of the proposed system has been validated against other personalized approaches on real-life and benchmark collections retrieved from the popular photo-sharing system Flickr.
Luca Cagliero, Alessandro Fiori, Luigi Grimaudo
ACM Trans. Intell. Syst. Technol.1
2013 EnBay: A Novel Pattern-Based Bayesian Classifier
abstract
A promising approach to Bayesian classification is based on exploiting frequent patterns, i.e., patterns that frequently occur in the training data set, to estimate the Bayesian probability. Pattern-based Bayesian classification focuses on building and evaluating reliable probability approximations by exploiting a subset of frequent patterns tailored to a given test case. This paper proposes a novel and effective approach to estimate the Bayesian probability. Differently from previous approaches, the Entropy-based Bayesian classifier, namely EnBay, focuses on selecting the minimal set of long and not overlapped patterns that best complies with a conditional-independence model, based on an entropy-based evaluator. Furthermore, the probability approximation is separately tailored to each class. An extensive experimental evaluation, performed on both real and synthetic data sets, shows that EnBay is significantly more accurate than most state-of-the-art classifiers, Bayesian and not.
Elena Baralis, Luca Cagliero, Paolo Garza
IEEE Trans. Knowl. Data Eng.2
2013 Discovering Temporal Change Patterns in the Presence of Taxonomies
abstract
Frequent itemset mining is a widely exploratory technique that focuses on discovering recurrent correlations among data. The steadfast evolution of markets and business environments prompts the need of data mining algorithms to discover significant correlation changes in order to reactively suit product and service provision to customer needs. Change mining, in the context of frequent itemsets, focuses on detecting and reporting significant changes in the set of mined itemsets from one time period to another. The discovery of frequent generalized itemsets, i.e., itemsets that 1) frequently occur in the source data, and 2) provide a high-level abstraction of the mined knowledge, issues new challenges in the analysis of itemsets that become rare, and thus are no longer extracted, from a certain point. This paper proposes a novel kind of dynamic pattern, namely the History GENeralized Pattern (HIGEN), that represents the evolution of an itemset in consecutive time periods, by reporting the information about its frequent generalizations characterized by minimal redundancy (i.e., minimum level of abstraction) in case it becomes infrequent in a certain time period. To address HIGEN mining, it proposes HIGEN MINER, an algorithm that focuses on avoiding itemset mining followed by postprocessing by exploiting a support-driven itemset generalization approach. To focus the attention on the minimally redundant frequent generalizations and thus reduce the amount of the generated patterns, the discovery of a smart subset of HIGENs, namely the NONREDUNDANT HIGENs, is addressed as well. Experiments performed on both real and synthetic datasets show the efficiency and the effectiveness of the proposed approach as well as its usefulness in a real application context.
Luca Cagliero
IEEE Trans. Knowl. Data Eng.1
2012 Generalized association rule mining with constraints
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza
Inf. Sci.2
2011 Semi-Automatic Ontology Construction by Exploiting Functional Dependencies and Association Rules
abstract
This paper presents a novel semi-automatic approach to construct conceptual ontologies over structured data by exploiting both the schema and content of the input dataset. It effectively combines two well-founded database and data mining techniques, i.e., functional dependency discovery and association rule mining, to support domain experts in the construction of meaningful ontologies, tailored to the analyzed data, by using Description Logic (DL). To this aim, functional dependencies are first discovered to highlight valuable conceptual relationships among attributes of the data schema (i.e., among concepts). The set of discovered correlations effectively support analysts in the assertion of the Tbox ontological statements (i.e., the statements involving shared data conceptualizations and their relationships). Then, the analyst-validated dependencies are exploited to drive the association rule mining process. Association rules represent relevant and hidden correlations among data content and they are used to provide valuable knowledge at the instance level. The pushing of functional dependency constraints into the rule mining process allows analysts to look into and exploit only the most significant data item recurrences in the assertion of the Abox ontological statements (i.e., the statements involving concept instances and their relationships).
Luca Cagliero, Tania Cerquitelli, Paolo Garza
Int. J. Semantic Web Inf. Syst.1
2011 CAS-Mine: providing personalized services in context-aware applications by means of generalized rules
Elena Baralis, Luca Cagliero, Tania Cerquitelli, Paolo Garza, Marco Marchetti
Knowl. Inf. Syst.2