Fosca Giannotti

dblp:g/FoscaGiannotti · DBLP profile ↗
← Back
79ranked-venue papers in the field
14as first author
8since 2021 · last 2025
0000-0003-3099-3835ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 52 (6 first)Database Systems & Data Management · 21 (7 first)Big Data, Cloud & Distributed Data Systems · 3Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 Explanations Go Linear: Post-Hoc Explainability for Tabular Data with Interpretable Meta-Encoding
abstract
Post-hoc explainability is essential for understanding black-box machine learning models. Surrogate-based techniques are widely used for local and global model-agnostic explanations but have significant limitations. Local surrogates capture non-linearities but are computationally expensive and sensitive to parameters, while global surrogates are more efficient but struggle with complex local behaviors. In this paper, we present ILLUME, a flexible and interpretable framework grounded in representation learning, that can be integrated with various surrogate models to provide explanations for any black-box classifier. Specifically, our approach combines a globally trained surrogate with instance-specific linear transformations learned with a meta-encoder to generate both local and global explanations. Through extensive empirical evaluations, we demonstrate the effectiveness of ILLUME in producing feature attributions and decision rules that are not only accurate but also robust and computationally efficient, thus providing a unified explanation framework that effectively addresses the limitations of traditional surrogate methods.
Simone Piaggesi, Riccardo Guidotti, Fosca Giannotti, Dino Pedreschi
ICDM3
2025 MAINLE: A Multi-Agent, Interactive, Natural Language Local Explainer of Classification Tasks
Paulo Serafim, Rômulo Férrer Filho, Stenio Freitas, Gizem Gezici, Fosca Giannotti, Franco Raimondi, Alexandre Santos
ECML/PKDD (4)5
2024 FLocalX - Local to Global Fuzzy Explanations for Black Box Classifiers
Guillermo Fernández 0005, Riccardo Guidotti, Fosca Giannotti, Mattia Setzu, Juan A. Aledo, José A. Gámez 0001, José M. Puerta
IDA (2)3
2024 Interpretable and Fair Mechanisms for Abstaining Classifiers
Daphne Lenders, Andrea Pugnana, Roberto Pellungrini, Toon Calders, Dino Pedreschi, Fosca Giannotti
ECML/PKDD (7)6
2024 Stable and actionable explanations of black-box models through factual and counterfactual rules
abstract
Abstract Recent years have witnessed the rise of accurate but obscure classification models that hide the logic of their internal decision processes. Explaining the decision taken by a black-box classifier on a specific input instance is therefore of striking interest. We propose a local rule-based model-agnostic explanation method providing stable and actionable explanations. An explanation consists of a factual logic rule, stating the reasons for the black-box decision, and a set of actionable counterfactual logic rules, proactively suggesting the changes in the instance that lead to a different outcome. Explanations are computed from a decision tree that mimics the behavior of the black-box locally to the instance to explain. The decision tree is obtained through a bagging-like approach that favors stability and fidelity: first, an ensemble of decision trees is learned from neighborhoods of the instance under investigation; then, the ensemble is merged into a single decision tree. Neighbor instances are synthetically generated through a genetic algorithm whose fitness function is driven by the black-box behavior. Experiments show that the proposed method advances the state-of-the-art towards a comprehensive approach that successfully covers stability and actionability of factual and counterfactual explanations.
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Francesca Naretto, Franco Turini, Dino Pedreschi, Fosca Giannotti
Data Min. Knowl. Discov.7
2024 Understanding Any Time Series Classifier with a Subsequence-based Explainer
abstract
The growing availability of time series data has increased the usage of classifiers for this data type. Unfortunately, state-of-the-art time series classifiers are black-box models and, therefore, not usable in critical domains such as healthcare or finance, where explainability can be a crucial requirement. This paper presents a framework to explain the predictions of any black-box classifier for univariate and multivariate time series. The provided explanation is composed of three parts. First, a saliency map highlighting the most important parts of the time series for the classification. Second, an instance-based explanation exemplifies the black-box’s decision by providing a set of prototypical and counterfactual time series. Third, a factual and counterfactual rule-based explanation, revealing the reasons for the classification through logical conditions based on subsequences that must, or must not, be contained in the time series. Experiments and benchmarks show that the proposed method provides faithful, meaningful, stable, and interpretable explanations.
Francesco Spinnato, Riccardo Guidotti, Anna Monreale, Mirco Nanni, Dino Pedreschi, Fosca Giannotti
ACM Trans. Knowl. Discov. Data6
2023 Benchmarking and survey of explanation methods for black box models
abstract
Abstract The rise of sophisticated black-box machine learning models in Artificial Intelligence systems has prompted the need for explanation methods that reveal how these models work in an understandable way to users and decision makers. Unsurprisingly, the state-of-the-art exhibits currently a plethora of explainers providing many different types of explanations. With the aim of providing a compass for researchers and practitioners, this paper proposes a categorization of explanation methods from the perspective of the type of explanation they return, also considering the different input data formats. The paper accounts for the most representative explainers to date, also discussing similarities and discrepancies of returned explanations through their visual appearance. A companion website to the paper is provided as a continuous update to new explainers as they appear. Moreover, a subset of the most robust and widely adopted explainers, are benchmarked with respect to a repertoire of quantitative metrics.
Francesco Bodria, Fosca Giannotti, Riccardo Guidotti, Francesca Naretto, Dino Pedreschi, Salvatore Rinzivillo
Data Min. Knowl. Discov.2
2022 Transparent Latent Space Counterfactual Explanations for Tabular Data
abstract
Artificial Intelligence decision-making systems have dramatically increased their predictive performance in recent years, beating humans in many different specific tasks. However, with increased performance has come an increase in the complexity of the black-box models adopted by the AI systems, making them entirely obscure for the decision process adopted. Explainable AI is a field that seeks to make AI decisions more transparent by producing explanations. In this paper, we propose T-LACE, an approach able to retrieve post-hoc counterfactual explanations for a given pre-trained black-box model. T-LACE exploits the similarity and linearity proprieties of a custom-created transparent latent space to build reliable counterfactual explanations. We tested T-LACE on several tabular datasets and provided qualitative evaluations of the generated explanations in terms of similarity, robustness, and diversity. Comparative analysis against various state-of-the-art counterfactual explanation methods shows the higher effectiveness of our approach.
Francesco Bodria, Riccardo Guidotti, Fosca Giannotti, Dino Pedreschi
DSAA3
2020 Estimating countries' peace index through the lens of the world news as monitored by GDELT
abstract
Peacefulness is a principal dimension of well-being, and its measurement has lately drawn the attention of researchers and policy-makers. During the last years, novel digital data streams have drastically changed research in this field. In the current study, we exploit information extracted from Global Data on Events, Location, and Tone (GDELT) digital news database, to capture peacefulness through the Global Peace Index (GPI). Applying machine learning techniques, we demonstrate that news media attention, sentiment, and social stability from GDELT can be used as proxies for measuring GPI at a monthly level. Additionally, through the variable importance analysis, we show that each country's socio-economic, political, and military profile emerges. This could bring added value to researchers interested in "Data Science for Social Good", to policy-makers, and peacekeeping organizations since they could monitor peacefulness almost real-time, and therefore facilitate timely and more efficient policy-making.
Vasiliki Voukelatou, Luca Pappalardo, Ioanna Miliou, Lorenzo Gabrielli, Fosca Giannotti
DSAA5
2020 Digital Footprints of International Migration on Twitter
abstract
Studying migration using traditional data has some limitations. To date, there have been several studies proposing innovative methodologies to measure migration stocks and flows from social big data. Nevertheless, a uniform definition of a migrant is difficult to find as it varies from one work to another depending on the purpose of the study and nature of the dataset used. In this work, a generic methodology is developed to identify migrants within the Twitter population. This describes a migrant as a person who has the current residence different from the nationality. The residence is defined as the location where a user spends most of his/her time in a certain year. The nationality is inferred from linguistic and social connections to a migrant’s country of origin. This methodology is validated first with an internal gold standard dataset and second with two official statistics, and shows strong performance scores and correlation coefficients. Our method has the advantage that it can identify both immigrants and emigrants, regardless of the origin/destination countries. The new methodology can be used to study various aspects of migration, including opinions, integration, attachment, stocks and flows, motivations for migration, etc. Here, we exemplify how trending topics across and throughout different migrant communities can be observed.
Alina Sîrbu, Fosca Giannotti, Lorenzo Gabrielli
IDA3
2020 PRIMULE: Privacy risk mitigation for user profiles
Francesca Pratesi, Lorenzo Gabrielli, Paolo Cintia, Anna Monreale, Fosca Giannotti
Data Knowl. Eng.5
2019 PlayeRank: Data-driven Performance Evaluation and Player Ranking in Soccer via a Machine Learning Approach
abstract
The problem of evaluating the performance of soccer players is attracting the interest of many companies and the scientific community, thanks to the availability of massive data capturing all the events generated during a match (e.g., tackles, passes, shots, etc.). Unfortunately, there is no consolidated and widely accepted metric for measuring performance quality in all of its facets. In this article, we design and implement PlayeRank, a data-driven framework that offers a principled multi-dimensional and role-aware evaluation of the performance of soccer players. We build our framework by deploying a massive dataset of soccer-logs and consisting of millions of match events pertaining to four seasons of 18 prominent soccer competitions. By comparing PlayeRank to known algorithms for performance evaluation in soccer, and by exploiting a dataset of players’ evaluations made by professional soccer scouts, we show that PlayeRank significantly outperforms the competitors. We also explore the ratings produced by PlayeRank and discover interesting patterns about the nature of excellent performances and what distinguishes the top players from the others. At the end, we explore some applications of PlayeRank—i.e. searching players and player versatility—showing its flexibility and efficiency, which makes it worth to be used in the design of a scalable platform for soccer analytics.
Luca Pappalardo, Paolo Cintia, Paolo Ferragina, Emanuele Massucco, Dino Pedreschi, Fosca Giannotti
ACM Trans. Intell. Syst. Technol.6
2019 Personalized Market Basket Prediction with Temporal Annotated Recurring Sequences
abstract
Nowadays, a hot challenge for supermarket chains is to offer personalized services to their customers. Market basket prediction, i.e., supplying the customer a shopping list for the next purchase according to her current needs, is one of these services. Current approaches are not capable of capturing at the same time the different factors influencing the customer's decision process: co-occurrence, sequentuality, periodicity, and recurrency of the purchased items. To this aim, we define a pattern TemporalAnnotated Recurring Sequence (TARS) able to capture simultaneously and adaptively all these factors. We define the method to extract TARS and develop a predictor for next basket named TBP (TARS Based Predictor) that, on top of TARS, is able to understand the level of the customer's stocks and recommend the set of most necessary items. By adopting the TBP the supermarket chains could crop tailored suggestions for each individual customer which in turn could effectively speed up their shopping sessions. A deep experimentation shows that TARS are able to explain the customer purchase behavior, and that TBP outperforms the state-of-the-art competitors.
Riccardo Guidotti, Giulio Rossetti, Luca Pappalardo, Fosca Giannotti, Dino Pedreschi
IEEE Trans. Knowl. Data Eng.4
2017 There's a Path for Everyone: A Data-Driven Personal Model Reproducing Mobility Agendas
abstract
The avalanche of mobility data like GPS and GSM daily produced by each user through mobile devices enables personalized mobility-services improving everyday life. The base for these mobility-services lies in the predictability of human behavior. In this paper we propose an approach for reproducing the user's personal mobility agenda that is able to predict the user's positions for the whole day. We reproduce the agenda by exploiting a data-driven personal mobility model able to capture and summarize different aspects of the systematic mobility behavior of a user. We show how the proposed approach outperforms typical methodologies adopted in the literature on four different real GPS datasets. Moreover, we analyze some features of the mobility models and we discuss how they can be employed as agents of a simulator for what-if mobility analysis.
Riccardo Guidotti, Roberto Trasarti, Mirco Nanni, Fosca Giannotti, Dino Pedreschi
DSAA4
2017 NDlib: Studying Network Diffusion Dynamics
abstract
Nowadays the analysis of diffusive phenomena occurring on top of complex networks represents a hot topic in the Social Network Analysis playground. In order to support students, teachers, developers and researchers in this work we introduce a novel simulation framework, NDlib. NDlib is designed to be a multi-level ecosystem that can be fruitfully used by different user segments. Upon the diffusion library, we designed a simulation server that allows remote execution of experiments and an online visualization tool that abstract the programmatic interface and makes available the simulation platform to non-technicians.
Giulio Rossetti, Letizia Milli, Salvatore Rinzivillo, Alina Sîrbu, Dino Pedreschi, Fosca Giannotti
DSAA6
2017 Market Basket Prediction Using User-Centric Temporal Annotated Recurring Sequences
abstract
Nowadays, a hot challenge for supermarket chains is to offer personalized services to their customers. Market basket prediction, i.e., supplying the customer a shopping list for the next purchase according to her current needs, is one of these services. Current approaches are not capable of capturing at the same time the different factors influencing the customer's decision process: co-occurrence, sequentuality, periodicity and recurrency of the purchased items. To this aim, we define a pattern named Temporal Annotated Recurring Sequence (TARS). We define the method to extract TARS and develop a predictor for next basket named TBP (TARS Based Predictor) that, on top of TARS, is able to understand the level of the customer's stocks and recommend the set of most necessary items. A deep experimentation shows that TARS can explain the customers' purchase behavior, and that TBP outperforms the state-of-the-art competitors.
Riccardo Guidotti, Giulio Rossetti, Luca Pappalardo, Fosca Giannotti, Dino Pedreschi
ICDM4
2017 Clustering Individual Transactional Data for Masses of Users
abstract
Mining a large number of datasets recording human activities for making sense of individual data is the key enabler of a new wave of personalized knowledge-based services. In this paper we focus on the problem of clustering individual transactional data for a large mass of users. Transactional data is a very pervasive kind of information that is collected by several services, often involving huge pools of users. We propose txmeans, a parameter-free clustering algorithm able to efficiently partitioning transactional data in a completely automatic way. Txmeans is designed for the case where clustering must be applied on a massive number of different datasets, for instance when a large set of users need to be analyzed individually and each of them has generated a long history of transactions. A deep experimentation on both real and synthetic datasets shows the practical effectiveness of txmeans for the mass clustering of different personal datasets, and suggests that txmeans outperforms existing methods in terms of quality and efficiency. Finally, we present a personal cart assistant application based on txmeans
Riccardo Guidotti, Anna Monreale, Mirco Nanni, Fosca Giannotti, Dino Pedreschi
KDD4
2017 Never drive alone: Boosting carpooling with network analysis
Riccardo Guidotti, Mirco Nanni, Salvatore Rinzivillo, Dino Pedreschi, Fosca Giannotti
Inf. Syst.5
2017 MyWay: Location prediction via mobility profiling
Roberto Trasarti, Riccardo Guidotti, Anna Monreale, Fosca Giannotti
Inf. Syst.4
2015 Interaction Prediction in Dynamic Networks exploiting Community Discovery
abstract
Due to the growing availability of online social services, interactions between people became more and more easy to establish and track. Online social human activities generate digital footprints, that describe complex, rapidly evolving, dynamic networks. In such scenario one of the most challenging task to address involves the prediction of future interactions between couples of actors. In this study, we want to leverage networks dynamics and community structure to predict which are the future interactions more likely to appear. To this extent, we propose a supervised learning approach which exploit features computed by time-aware forecasts of topological measures calculated between pair of nodes belonging to the same community. Our experiments on real dynamic networks show that the designed analytical process is able to achieve interesting results.
Giulio Rossetti, Riccardo Guidotti, Diego Pennacchioli, Dino Pedreschi, Fosca Giannotti
ASONAM5
2015 Community-centric analysis of user engagement in Skype social network
abstract
Traditional approaches to user engagement analysis focus on individual users. In this paper we address user engagement analysis at the level of groups of users (social communities). From the entire Skype social network we extract communities by means of representative community detection methods each one providing node partitions having their own peculiarities. We then examine user engagement in the extracted communities putting into evidence clear relations between topological and geographic features of communities and their mean user engagement. In particular we show that user engagement can be to a great extent predicted from such features. Moreover, from the analysis it clearly emerges that the choice of community definition and granularity deeply affect the predictive performance.
Giulio Rossetti, Luca Pappalardo, Riivo Kikas, Dino Pedreschi, Fosca Giannotti, Marlon Dumas
ASONAM5
2015 City users' classification with mobile phone data
abstract
Nowadays mobile phone data are an actual proxy for studying the users' social life and urban dynamics. In this paper we present the Sociometer, and analytical framework aimed at classifying mobile phone users into behavioral categories by means of their call habits. The analytical process starts from spatio-temporal profiles, learns the different behaviors, and returns annotated profiles. After the description of the methodology and its evaluation, we present an application of the Sociometer for studying city users of one small and one big city, evaluating the impact of big events in these cities.
Lorenzo Gabrielli, Barbara Furletti, Roberto Trasarti, Fosca Giannotti, Dino Pedreschi
IEEE BigData4
2015 Using big data to study the link between human mobility and socio-economic development
abstract
Big Data offer nowadays the potential capability of creating a digital nervous system of our society, enabling the measurement, monitoring and prediction of relevant aspects of socio-economic phenomena in quasi real time. This potential has fueled, in the last few years, a growing interest around the usage of Big Data to support official statistics in the measurement of individual and collective economic well-being. In this work we study the relations between human mobility patterns and socioeconomic development. Starting from nation-wide mobile phone data we extract a measure of mobility volume and a measure of mobility diversity for each individual. We then aggregate the mobility measures at municipality level and investigate the correlations with external socio-economic indicators independently surveyed by an official statistics institute. We find three main results. First, aggregated human mobility patterns are correlated with these socio-economic indicators. Second, the diversity of mobility, defined in terms of entropy of the individual users' trajectories, exhibits the strongest correlation with the external socio-economic indicators. Third, the volume of mobility and the diversity of mobility show opposite correlations with the socioeconomic indicators. Our results, validated against a null model, open an interesting perspective to study human behavior through Big Data by means of new statistical indicators that quantify and possibly "nowcast" the socio-economic development of our society.
Luca Pappalardo, Dino Pedreschi, Zbigniew Smoreda, Fosca Giannotti
IEEE BigData4
2015 The harsh rule of the goals: Data-driven performance indicators for football teams
abstract
Sports analytics in general, and football (soccer in USA) analytics in particular, have evolved in recent years in an amazing way, thanks to automated or semi-automated sensing technologies that provide high-fidelity data streams extracted from every game. In this paper we propose a data-driven approach and show that there is a large potential to boost the understanding of football team performance. From observational data of football games we extract a set of pass-based performance indicators and summarize them in the H indicator. We observe a strong correlation among the proposed indicator and the success of a team, and therefore perform a simulation on the four major European championships (78 teams, almost 1500 games). The outcome of each game in the championship was replaced by a synthetic outcome (win, loss or draw) based on the performance indicators computed for each team. We found that the final rankings in the simulated championships are very close to the actual rankings in the real championships, and show that teams with high ranking error show extreme values of a defense/attack efficiency measure, the Pezzali score. Our results are surprising given the simplicity of the proposed indicators, suggesting that a complex systems' view on football data has the potential of revealing hidden patterns and behavior of superior quality.
Paolo Cintia, Fosca Giannotti, Luca Pappalardo, Dino Pedreschi, Marco Malvaldi
DSAA2
2015 Quantification in social networks
abstract
In many real-world applications there is a need to monitor the distribution of a population across different classes, and to track changes in this distribution over time. As an example, an important task is to monitor the percentage of unemployed adults in a given region. When the membership of an individual in a class cannot be established deterministically, a typical solution is the classification task. However, in the above applications the final goal is not determining which class the individuals belong to, but estimating the prevalence of each class in the unlabeled data. This task is called quantification. Most of the work in the literature addressed the quantification problem considering data presented in conventional attribute format. Since the ever-growing availability of web and social media we have a flourish of network data representing a new important source of information and by using quantification network techniques we could quantify collective behavior, i.e., the number of users that are involved in certain type of activities, preferences, or behaviors. In this paper we exploit the homophily effect observed in many social networks in order to construct a quantifier for networked data. Our experiments show the effectiveness of the proposed approaches and the comparison with the existing state-of-the-art quantification methods shows that they are more accurate.
Letizia Milli, Anna Monreale, Giulio Rossetti, Dino Pedreschi, Fosca Giannotti, Fabrizio Sebastiani 0001
DSAA5
2015 Discrimination- and privacy-aware patterns
Sara Hajian, Josep Domingo-Ferrer, Anna Monreale, Dino Pedreschi, Fosca Giannotti
Data Min. Knowl. Discov.5
2014 The purpose of motion: Learning activities from Individual Mobility Networks
abstract
The large availability of mobility data allows us to investigate complex phenomena about human movement. However this adundance of data comes with few information about the purpose of movement. In this work we address the issue of activity recognition by introducing Activity-Based Cascading (ABC) classification. Such approach departs completely from probabilistic approaches for two main reasons. First, it exploits a set of structural features extracted from the Individual Mobility Network (IMN), a model able to capture the salient aspects of individual mobility. Second, it uses a cascading classification as a way to tackle the highly skewed frequency of activity classes. We show that our approach outperforms existing state-of-the-art probabilistic methods. Since it reaches high precision, ABC classification represents a very reliable semantic amplifier for Big Data.
Salvatore Rinzivillo, Lorenzo Gabrielli, Mirco Nanni, Luca Pappalardo, Dino Pedreschi, Fosca Giannotti
DSAA6
2014 Uncovering Hierarchical and Overlapping Communities with a Local-First Approach
abstract
Community discovery in complex networks is the task of organizing a network’s structure by grouping together nodes related to each other. Traditional approaches are based on the assumption that there is a global-level organization in the network. However, in many scenarios, each node is the bearer of complex information and cannot be classified in disjoint clusters. The top-down global view of the partition approach is not designed for this. Here, we represent this complex information as multiple latent labels, and we postulate that edges in the networks are created among nodes carrying similar labels. The latent labels are the communities a node belongs to and we discover them with a simple local-first approach to community discovery. This is achieved by democratically letting each node vote for the communities it sees surrounding it in its limited view of the global system, its ego neighborhood, using a label propagation algorithm, assuming that each node is aware of the label it shares with each of its connections. The local communities are merged hierarchically, unveiling the modular organization of the network at the global level and identifying overlapping groups and groups of groups. We tested this intuition against the state-of-the-art overlapping community discovery and found that our new method advances in the chosen scenarios in the quality of the obtained communities. We perform a test on benchmark and on real-world networks, evaluating the quality of the community coverage by using the extracted communities to predict the metadata attached to the nodes, which we consider external information about the latent labels. We also provide an explanation about why real-world networks contain overlapping communities and how our logic is able to capture them. Finally, we show how our method is deterministic, is incremental, and has a limited time complexity, so that it can be used on real-world scale networks.
Michele Coscia, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi
ACM Trans. Knowl. Discov. Data3
2013 "You know because I know": a multidimensional network approach to human resources problem
abstract
Finding talents, often among the people already hired, is an endemic challenge for organizations. The social networking revolution, with online tools like Linkedin, made possible to make explicit and accessible what we perceived, but not used, for thousands of years: the exact position and ranking of a person in a network of professional and personal connections. To search and mine where and how an employee is positioned on a global skill network will enable organizations to find unpredictable sources of knowledge, innovation and know-how. This data richness and hidden knowledge demands for a multidimensional and multiskill approach to the network ranking problem. Multidimensional networks are networks with multiple kinds of relations. To the best of our knowledge, no network-based ranking algorithm is able to handle multidimensional networks and multiple rankings over multiple attributes at the same time. In this paper we propose such an algorithm, whose aim is to address the node multi-ranking problem in multidimensional networks. We test our algorithm over several real world networks, extracted from DBLP and the Enron email corpus, and we show its usefulness in providing less trivial and more flexible rankings than the current state of the art algorithms.
Michele Coscia, Giulio Rossetti, Diego Pennacchioli, Damiano Ceccarelli, Fosca Giannotti
ASONAM5
2013 Explaining the product range effect in purchase data
abstract
In our market society, buyers are considered rational entities, driven by two utility functions: i) the amount of money spent, a universal quantity to be minimized; and ii) the individual needs to satisfy, a personal quantity, varying from person to person, to be maximized. In this paper, we propose an analytic framework based on big data to measure the personal utility function and we prove that this function has a stronger effect on customer behavior than the price. By focusing on the purchases in an Italian supermarket chain, we discover and describe a range effect of products: the more sophisticated the needs they satisfy, the more cost the customers are willing to pay to buy them, in terms of distance to travel more than in terms of the price of the item itself. We exhibit a striking empirical evidence of this theory by tracking the geographical information about points of sale and customers, in a large dataset containing tens of thousands of customers and thousands of products. We create a data mining framework able to scale to possibly hundreds of thousands, or millions, of customers and to let emerge from the data the knowledge about the actual range of each product. As an application of this finding, we show how it is possible to accurately predict how long a customer will travel (or which shop she will choose) to buy a product, as a function of the product's sophistication.
Diego Pennacchioli, Michele Coscia, Salvatore Rinzivillo, Dino Pedreschi, Fosca Giannotti
IEEE BigData5
2013 Quantification Trees
abstract
In many applications there is a need to monitor how a population is distributed across different classes, and to track the changes in this distribution that derive from varying circumstances, an example such application is monitoring the percentage (or "prevalence") of unemployed people in a given region, or in a given age range, or at different time periods. When the membership of an individual in a class cannot be established deterministically, this monitoring activity requires classification. However, in the above applications the final goal is not determining which class each individual belongs to, but simply estimating the prevalence of each class in the unlabeled data. This task is called quantification. In a supervised learning framework we may estimate the distribution across the classes in a test set from a training set of labeled individuals. However, this may be sub optimal, since the distribution in the test set may be substantially different from that in the training set (a phenomenon called distribution drift). So far, quantification has mostly been addressed by learning a classifier optimized for individual classification and later adjusting the distribution it computes to compensate for its tendency to either under-or over-estimate the prevalence of the class. In this paper we propose instead to use a type of decision trees (quantification trees) optimized not for individual classification, but directly for quantification. Our experiments show that quantification trees are more accurate than existing state-of-the-art quantification methods, while retaining at the same time the simplicity and understandability of the decision tree framework.
Letizia Milli, Anna Monreale, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi, Fabrizio Sebastiani 0001
ICDM4
2012 Optimal Spatial Resolution for the Analysis of Human Mobility
abstract
The availability of massive network and mobility data from diverse domains has fostered the analysis of human behaviors and interactions. This data availability leads to challenges in the knowledge discovery community. Several different analyses have been performed on the traces of human trajectories, such as understanding the real borders of human mobility or mining social interactions derived from mobility and vice versa. However, the data quality of the digital traces of human mobility has a dramatic impact over the knowledge that it is possible to mine, and this issue has not been thoroughly tackled so far in literature. In this paper, we mine and analyze with complex network techniques a large dataset of human trajectories, a GPS dataset from more than 150k vehicles in Italy. We build a multi resolution grid and we map the trajectories with several complex networks, by connecting the different areas of our region of interest. Then we analyze the structural properties of these networks and the quality of the borders it is possible to infer from them. The result is a significant advancement in our understanding of the data transformation process that is needed to connect mobility with social network analysis and mining.
Michele Coscia, Salvatore Rinzivillo, Fosca Giannotti, Dino Pedreschi
ASONAM3
2012 DEMON: a local-first discovery method for overlapping communities
abstract
Community discovery in complex networks is an interesting problem with a number of applications, especially in the knowledge extraction task in social and information networks. However, many large networks often lack a particular community organization at a global level. In these cases, traditional graph partitioning algorithms fail to let the latent knowledge embedded in modular structure emerge, because they impose a top-down global view of a network. We propose here a simple local-first approach to community discovery, able to unveil the modular organization of real complex networks. This is achieved by democratically letting each node vote for the communities it sees surrounding it in its limited view of the global system, i.e. its ego neighborhood, using a label propagation algorithm; finally, the local communities are merged into a global collection. We tested this intuition against the state-of-the-art overlapping and non-overlapping community discovery methods, and found that our new method clearly outperforms the others in the quality of the obtained communities, evaluated by using the extracted communities to predict the metadata about the nodes of several real world networks. We also show how our method is deterministic, fully incremental, and has a limited time complexity, so that it can be used on web-scale real networks.
Michele Coscia, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi
KDD3
2012 AUDIO: An Integrity Auditing Framework of Outlier-Mining-as-a-Service Systems
Wendy Hui Wang, Anna Monreale, Dino Pedreschi, Fosca Giannotti, Wenge Guo
ECML/PKDD (2)5
2011 Finding and Characterizing Communities in Multidimensional Networks
abstract
Complex networks have been receiving increasing attention by the scientific community, also due to the availability of massive network data from diverse domains. One problem studied so far in complex network analysis is Community Discovery, i.e. the detection of group of nodes densely connected, or highly related. However, one aspect of such networks has been disregarded so far: real networks are often multidimensional, i.e. many connections may reside between any two nodes, either to reflect different kinds of relationships, or to connect nodes by different values of the same type of tie. In this context, the problem of Community Discovery has to be redefined, taking into account multidimensionality. In this paper, we attempt to do so, by defining the problem in the multidimensional context, and by introducing also a new measure able to characterize the communities found. We then provide a complete framework for finding and characterizing multidimensional communities. Our experiments on real world multidimensional networks support the methodology proposed in this paper, and open the way for a new class of algorithms, aimed at capturing the multifaceted complexity of connections among nodes in a network.
Michele Berlingerio, Michele Coscia, Fosca Giannotti
ASONAM3
2011 Foundations of Multidimensional Network Analysis
abstract
Complex networks have been receiving increasing attention by the scientific community, thanks also to the increasing availability of real-world network data. In the last years, the multidimensional nature of many real world networks has been pointed out, i.e. many networks containing multiple connections between any pair of nodes have been analyzed. Despite the importance of analyzing this kind of networks was recognized by previous works, a complete framework for multidimensional network analysis is still missing. Such a framework would enable the analysts to study different phenomena, that can be either the generalization to the multidimensional setting of what happens inmonodimensional network, or a new class of phenomena induced by the additional degree of complexity that multidimensionality provides in real networks. The aim of this paper is then to give the basis for multidimensional network analysis: we develop a solid repertoire of basic concepts and analytical measures, which takes into account the general structure of multidimensional networks. We tested our framework on a real world multidimensional network, showing the validity and the meaningfulness of the measures introduced, that are able to extract important, nonrandom, information about complex phenomena.
Michele Berlingerio, Michele Coscia, Fosca Giannotti, Anna Monreale, Dino Pedreschi
ASONAM3
2011 Finding redundant and complementary communities in multidimensional networks
abstract
Community Discovery in networks is the problem of detecting, for each node, its membership to one of more groups of nodes, the communities, that are densely connected, or highly interactive. We define the community discovery problem in multidimensional networks, where more than one connection may reside between any two nodes. We also introduce two measures able to characterize the communities found. Our experiments on real world multidimensional networks support the methodology proposed in this paper, and open the way for a new class of algorithms, aimed at capturing the multifaceted complexity of connections among nodes in a network.
Michele Berlingerio, Michele Coscia, Fosca Giannotti
CIKM3
2011 Mining mobility user profiles for car pooling
abstract
In this paper we introduce a methodology for extracting mobility profiles of individuals from raw digital traces (in particular, GPS traces), and study criteria to match individuals based on profiles. We instantiate the profile matching problem to a specific application context, namely proactive car pooling services, and therefore develop a matching criterion that satisfies various basic constraints obtained from the background knowledge of the application domain. In order to evaluate the impact and robustness of the methods introduced, two experiments are reported, which were performed on a massive dataset containing GPS traces of private cars: (i) the impact of the car pooling application based on profile matching is measured, in terms of percentage shareable traffic; (ii) the approach is adapted to coarser-grained mobility data sources that are nowadays commonly available from telecom operators. In addition the ensuing loss in precision and coverage of profile matches is measured.
Roberto Trasarti, Fabio Pinelli, Mirco Nanni, Fosca Giannotti
KDD4
2011 Human mobility, social ties, and link prediction
abstract
Our understanding of how individual mobility patterns shape and impact the social network is limited, but is essential for a deeper understanding of network dynamics and evolution. This question is largely unexplored, partly due to the difficulty in obtaining large-scale society-wide data that simultaneously capture the dynamical information on individual movements and social interactions. Here we address this challenge for the first time by tracking the trajectories and communication records of 6 Million mobile phone users. We find that the similarity between two individuals' movements strongly correlates with their proximity in the social network. We further investigate how the predictive power hidden in such correlations can be exploited to address a challenging problem: which new links will develop in a social network. We show that mobility measures alone yield surprising predictive power, comparable to traditional network-based measures. Furthermore, the prediction accuracy can be significantly improved by learning a supervised classifier based on combined mobility and network measures. We believe our findings on the interplay of mobility patterns and social ties offer new perspectives on not only link prediction but also network dynamics.
Dashun Wang, Dino Pedreschi, Chaoming Song, Fosca Giannotti, Albert-László Barabási
KDD4
2011 Mobility, Data Mining and Privacy Understanding Human Movement Patterns from Trajectory Data
abstract
The technologies of mobile communications and ubiquitous computing pervade our society, and wireless net works sense the movement of people and vehicles, generating large volumes of mobility data. This is a scenario of great opportunities and risks. On one side, mining this data can produce useful knowledge. On the other side, individual privacy is at risk, as the mobility data contain sensitive personal information.
Fosca Giannotti
Mobile Data Management (1)1
2011 Traffic Jams Detection Using Flock Mining
Rebecca Ong, Fabio Pinelli, Roberto Trasarti, Mirco Nanni, Chiara Renso, Salvatore Rinzivillo, Fosca Giannotti
ECML/PKDD (3)7
2011 Unveiling the complexity of human mobility by querying and mining massive trajectory data
Fosca Giannotti, Mirco Nanni, Dino Pedreschi, Fabio Pinelli, Chiara Renso, Salvatore Rinzivillo, Roberto Trasarti
VLDB J.1
2010 Advanced knowledge discovery on movement data with the GeoPKDD system
abstract
The growing availability of mobile devices produces an enor- mous quantity of personal tracks which calls for advanced analysis methods capable of extracting knowledge out of massive trajectories datasets. In this paper we present an experiment on a real world scenario that demonstrates the strong analytical power of massive, raw trajectory data made available as a by-product of telecom services, in unveiling the complexity of urban mobility. The experiment has been made possible by the GeoPKDD system, an integrated plat- form for complex analysis of mobility data. The system com- bines spatio-temporal querying capabilities with data min- ing and semantic technologies, thus providing a full support for the Mobility Knowledge Discovery process.
Mirco Nanni, Roberto Trasarti, Chiara Renso, Fosca Giannotti, Dino Pedreschi
EDBT4
2010 As Time Goes by: Discovering Eras in Evolving Social Networks
Michele Berlingerio, Michele Coscia, Fosca Giannotti, Anna Monreale, Dino Pedreschi
PAKDD (1)3
2010 Exploring Real Mobility Data with M-Atlas
Roberto Trasarti, Salvatore Rinzivillo, Fabio Pinelli, Mirco Nanni, Anna Monreale, Chiara Renso, Dino Pedreschi, Fosca Giannotti
ECML/PKDD (3)8
2010 Hiding Sequential and Spatiotemporal Patterns
abstract
The process of discovering relevant patterns holding in a database was first indicated as a threat to database security by O'Leary in. Since then, many different approaches for knowledge hiding have emerged over the years, mainly in the context of association rules and frequent item sets mining. Following many real-world data and application demands, in this paper, we shift the problem of knowledge hiding to contexts where both the data and the extracted knowledge have a sequential structure. We define the problem of hiding sequential patterns and show its NP-hardness. Thus, we devise heuristics and a polynomial sanitization algorithm. Starting from this framework, we specialize it to the more complex case of spatiotemporal patterns extracted from moving objects databases. Finally, we discuss a possible kind of attack to our model, which exploits the knowledge of the underlying road network, and enhance our model to protect from this kind of attack. An exhaustive experiential analysis on real-world data sets shows the effectiveness of our proposal.
Osman Abul, Francesco Bonchi, Fosca Giannotti
IEEE Trans. Knowl. Data Eng.3
2009 Social Network Analysis as Knowledge Discovery Process: A Case Study on Digital Bibliography
abstract
Today digital bibliographies are a powerful instrument that collects a great amount of data about scientific publications. Digital bibliographies have been used as basis of many studies focused on the knowledge extraction in databases. Here we present anew methodology for mining knowledge in this field. Our approach aims to apply the potential of social network analysis techniques to accomplish this task, using a network representation of bibliography data. Besides we use some data mining techniques applied on social network representations in order to enrich this new point of view and to evolve our methodology towards a comprehensive local and global bibliography analysis workflow seen as a knowledge discovery process.
Michele Coscia, Fosca Giannotti, Ruggero G. Pensa
ASONAM2
2009 Geographic privacy-aware knowledge discovery and delivery
abstract
A flood of data pertinent to moving objects is available today, and will be more in the near future, particularly due to the automated collection of privacy-sensitive telecom data from mobile phones and other location-aware devices. Such wealth of data, referenced both in space and time, may enable novel classes of applications of high societal and economic impact, provided that the discovery of consumable and concise knowledge out of these raw data is made possible. Recent research activities have developed theory, techniques and systems for geographic knowledge discovery and delivery, some of them based on privacy-preserving methods for extracting knowledge from large amounts of raw data referenced in space and time. All these efforts aim at devising knowledge discovery and analysis methods for trajectories of moving objects.The fundamental hypothesis is that it is possible, in principle, to aid citizens in their mobile activities by analysing the traces of their past activities by means of data mining techniques. For instance, behavioural patterns derived from mobile trajectories may allow inducing traffic flow information, capable to help people travel efficiently, to help public administrations in traffic-related decision making for sustainable mobility and security management, as well as to help mobile operators in optimising bandwidth and power allocation on the network. On the other hand, it is clear that the use of personal sensitive data arouses concerns about citizen's privacy rights.In this tutorial, we establish a framework for the challenges and the mining solutions for the geographic information collected by Moving Object Database (MOD) engines. We first discuss the challenges of collecting mobility data, and elaborate on the impact of trajectory data analysis in several modern applications. We then discuss methodologies and techniques to collect raw data, reconstruct trajectory information, and efficiently store it in MODs. We continue with an overview of knowledge discovery approaches for movement data. Finally, we propose a research agenda and identify areas where interdisciplinary studies are needed.
Fosca Giannotti, Dino Pedreschi, Yannis Theodoridis
EDBT1
2009 Mining the Temporal Dimension of the Information Propagation
Michele Berlingerio, Michele Coscia, Fosca Giannotti
IDA3
2009 Temporal mining for interactive workflow data analysis
abstract
In the past few years there has been an increasing interest in the analysis of process logs. Several proposed techniques, such as workflow mining, are aimed at automatically deriving the underlying workflow models. However, current approaches only pay little attention on an important piece of information contained in process logs: the timestamps, which are used to define a sequential ordering of the performed tasks. In this work we try to overcome these limitations by explicitly including time in the extracted knowledge, thus making the temporal information a first-class citizen of the analysis process. This makes it possible to discern between apparently identical process executions that are performed with different transition times between consecutive tasks.
Michele Berlingerio, Fabio Pinelli, Mirco Nanni, Fosca Giannotti
KDD4
2009 WhereNext: a location predictor on trajectory pattern mining
abstract
The pervasiveness of mobile devices and location based services is leading to an increasing volume of mobility data.This side eect provides the opportunity for innovative methods that analyse the behaviors of movements. In this paper we propose WhereNext, which is a method aimed at predicting with a certain level of accuracy the next location of a moving object. The prediction uses previously extracted movement patterns named Trajectory Patterns, which are a concise representation of behaviors of moving objects as sequences of regions frequently visited with a typical travel time. A decision tree, named T-pattern Tree, is built and evaluated with a formal training and test process. The tree is learned from the Trajectory Patterns that hold a certain area and it may be used as a predictor of the next location of a new trajectory finding the best matching path in the tree. Three dierent best matching methods to classify a new moving object are proposed and their impact on the quality of prediction is studied extensively. Using Trajectory Patterns as predictive rules has the following implications: (I) the learning depends on the movement of all available objects in a certain area instead of on the individual history of an object; (II) the prediction tree intrinsically contains the spatio-temporal properties that have emerged from the data and this allows us to define matching methods that striclty depend on the properties of such movements. In addition, we propose a set of other measures, that evaluate a priori the predictive power of a set of Trajectory Patterns. This measures were tuned on a real life case study. Finally, an exhaustive set of experiments and results on the real dataset are presented.
Anna Monreale, Fabio Pinelli, Roberto Trasarti, Fosca Giannotti
KDD4
2009 A constraint-based querying system for exploratory pattern discovery
Francesco Bonchi, Fosca Giannotti, Claudio Lucchese, Salvatore Orlando 0001, Raffaele Perego 0001, Roberto Trasarti
Inf. Syst.2
2008 The DAEDALUS framework: progressive querying and mining of movement data
abstract
In this work we propose DAEDALUS, a formal framework and system, specifically focussed on progressive combination of mining and querying operators. The core component of DAEDALUS is the MO-DMQL query language that extends SQL in two respects, namely a pattern definition operator and the capability to uniform manipulating both raw data and unveiled patterns. DAEDALUS system is specifically focussed on movement data and has been implemented as a query execution layer on top of the Hermes Moving Object Database. The expressiveness and usefulness of the MODMQL language as well as the computational capabilities of DAEDALUS are qualitatively evaluated by means of a case study.
Riccardo Ortale, Ettore Ritacco, Nikos Pelekis, Roberto Trasarti, Gianni Costa, Fosca Giannotti, Giuseppe Manco 0001, Chiara Renso, Yannis Theodoridis
GIS6
2008 Clustering of German municipalities based on mobility characteristics: an overview of results
abstract
This paper presents a clustering approach which groups German municipalities according to mobility characteristics. As the number of measurements for nationwide mobility studies is usually restricted, this clustering provides a means to infer mobility information for locations without measurements based on values of their respective cluster representatives. Our approach considers local and global information, i.e. characteristics of municipalities as well as relationships between municipalities. We realize previous findings in urban geography by using techniques from graph theory and computer vision. Our clustering consists of a two-step model, which first extracts and condenses single mobility characteristics and subsequently combines the various features. We apply our model to all German municipalities between 10,000 and 50,000 inhabitants. The clustering has been successfully applied in practice for the inference of traffic frequencies.
Andrea Zanda, Christine Kopp, Fosca Giannotti, Daniel Schulz, Michael May 0001
GIS3
2008 An Application of Advanced Spatio-Temporal Formalisms to Behavioural Ecology
Alessandra Raffaetà, Tommaso Ceccarelli, Dominique Centeno, Fosca Giannotti, Alessandro Massolo, Christine Parent, Chiara Renso, Stefano Spaccapietra, Franco Turini
GeoInformatica4
2008 Anonymity preserving pattern discovery
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi
VLDB J.3
2007 Trajectory pattern mining
abstract
The increasing pervasiveness of location-acquisition technologies (GPS, GSM networks, etc.) is leading to the collection of large spatio-temporal datasets and to the opportunity of discovering usable knowledge about movement behaviour, which fosters novel applications and services. In this paper, we move towards this direction and develop an extension of the sequential pattern mining paradigm that analyzes the trajectories of moving objects. We introduce trajectory patterns as concise descriptions of frequent behaviours, in terms of both space (i.e., the regions of space visited during movements) and time (i.e., the duration of movements). In this setting, we provide a general formal statement of the novel mining problem and then study several different instantiations of different complexity. The various approaches are then empirically evaluated over real data and synthetic benchmarks, comparing their strengths and weaknesses.
Fosca Giannotti, Mirco Nanni, Fabio Pinelli, Dino Pedreschi
KDD1
2007 Privacy-Aware Knowledge Discovery from Location Data
abstract
Spatio-temporal, geo-referenced datasets are growing rapidly, and will be more in the near future. This phenomenon is mostly due to the daily collection of telecommunication data from mobile phones and other location-aware devices and is expected to enable novel classes of applications based on the extraction of behavioral patterns from mobility data. Such patterns could be used for instance in traffic and sustainable mobility management (e.g., to study the accessibility to services), urban planning, environmental monitoring, and collaborative location-based services. Clearly, in these applications privacy is a concern, since some knowledge may be sensitive, or an over-specific pattern may reveal the behaviour of groups of few individual. In this paper we focus on automated privacy-preserving methods we developed for extracting and sharing user- consumable forms of knowledge from large amounts of raw data referenced in space and in time.
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi, Osman Abul
MDM3
2006 ConQueSt: a Constraint-based Querying System for Exploratory Pattern Discovery
abstract
ConQueSt is a constraint-based querying system devised with the aim of supporting the intrinsically exploratory nature of pattern discovery. It provides users with an expressive constraint-based query language which allows the discovery process to be effectively driven toward potentially interesting patterns. Constraints are also exploited to reduce the cost of pattern mining. The system is built around an efficient constraint-based mining engine which entails several data and search space reduction techniques, and allows new user-defined constraints to be easily added.
Francesco Bonchi, Fosca Giannotti, Claudio Lucchese, Salvatore Orlando 0001, Raffaele Perego 0001, Roberto Trasarti
ICDE2
2006 Efficient Mining of Temporally Annotated Sequences
abstract
Sequential patterns mining received much attention in recent years, thanks to its various potential application domains. A large part of them represent data as collections of time-stamped itemsets, e.g., customers' purchases, logged web accesses, etc. Most approaches to sequence mining focus on sequentiality of data, using time-stamps only to order items and, in some cases, to constrain the temporal gap between items. In this paper, we propose an efficient algorithm for computing (temporally-)annotated sequential patterns, i.e., sequential patterns where each transition is annotated with a typical transition time derived from the source data. The algorithm adopts a prefix-projection approach to mine candidate sequences, and it is tightly integrated with an annotation mining process that associates sequences with temporal annotations. The pruning capabilities of the two steps sum together, yielding significant improvements in performances, as demonstrated by a set of experiments performed on synthetic datasets.
Fosca Giannotti, Mirco Nanni, Dino Pedreschi
SDM1
2005 Blocking Anonymity Threats Raised by Frequent Itemset Mining
abstract
In this paper we study when the disclosure of data mining results represents, per se, a threat to the anonymity of the individuals recorded in the analyzed database. The novelty of our approach is that we focus on an objective definition of privacy compliance of patterns without any reference to a preconceived knowledge of what is sensitive and what is not, on the basis of the rather intuitive and realistic constraint that the anonymity of individuals should be guaranteed. In particular, the problem addressed here arises from the possibility of inferring from the output of frequent itemset mining (i.e., a set of item-sets with support larger than a threshold a), the existence of patterns with very low support (smaller than an anonymity threshold k)[M. Atzori et. al, 2005]. In the following we develop a simple methodology to block such inference opportunities by introducing distortion on the dangerous patterns.
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi
ICDM3
2005 k-Anonymous Patterns
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi
PKDD3
2005 Efficient breadth-first mining of frequent pattern with monotone constraints
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi
Knowl. Inf. Syst.2
2004 Specifying Mining Algorithms with Iterative User-Defined Aggregates
abstract
We present a way of exploiting domain knowledge in the design and implementation of data mining algorithms, with special attention to frequent patterns discovery, within a deductive framework. In our framework, domain knowledge is represented by way of deductive rules, and data mining algorithms are specified by means of iterative user-defined aggregates and implemented by means of user-defined predicates. This choice allows us to exploit the full expressive power of deductive rules without loosing in performance. Iterative user-defined aggregates have a fixed scheme, in which user-defined predicates are to be added. This feature allows the modularization of data mining algorithms, thus providing a way to integrate the proper domain knowledge exploitation in the right point. As a case study, we present how user-defined aggregates can be exploited to specify and implement a version of the a priori algorithm. Some performance analyzes and comparisons are discussed in order to show the effectiveness of the approach.
Fosca Giannotti, Giuseppe Manco 0001, Franco Turini
IEEE Trans. Knowl. Data Eng.1
2003 ExAMiner: Optimized Level-wise Frequent Pattern Mining with Monotone Constraint
abstract
The key point is that, in frequent pattern mining, the most appropriate way of exploiting monotone constraints in conjunction with frequency is to use them in order to reduce the problem input together with the search space. Following this intuition, we introduce ExAMiner, a level-wise algorithm which exploits the real synergy of antimonotone and monotone constraints: the total benefit is greater than the sum of the two individual benefits. ExAMiner generalizes the basic idea of the preprocessing algorithm ExAnte [F. Bonchi et al., (2003)], embedding such ideas at all levels of an Apriori-like computation. The resulting algorithm is the generalization of the Apriori algorithm when a conjunction of monotone constraints is conjoined to the frequency antimonotone constraint. Experimental results confirm that this is, so far, the most efficient way of attacking the computational problem in analysis.
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi
ICDM2
2003 Adaptive Constraint Pushing in Frequent Pattern Mining
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi
PKDD2
2003 ExAnte: Anticipated Data Reduction in Constrained Pattern Mining
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi
PKDD2
2002 Clustering Transactional Data
Fosca Giannotti, Cristian Gozzi, Giuseppe Manco 0001
PKDD1
2001 Specifying Mining Algorithms with Iterative User-Defined Aggregates: A Case Study
Fosca Giannotti, Giuseppe Manco 0001, Franco Turini
PKDD1
2001 Web log data warehousing and mining for intelligent web caching
Francesco Bonchi, Fosca Giannotti, Cristian Gozzi, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi, Chiara Renso, Salvatore Ruggieri
Data Knowl. Eng.2
2001 Nondeterministic, Nonmonotonic Logic Databases
abstract
We consider an extension of Datalog with mechanisms for temporal, nonmonotonic, and nondeterministic reasoning, which we refer to as Datalog++. We show, by means of examples, its flexibility in expressing queries concerning aggregates and data cube. Also, we show how iterated fixpoint and stable model semantics can be combined to the purpose of clarifying the semantics of Datalog++ programs and supporting their efficient execution. Finally, we provide a more concrete implementation strategy on which basis the design of optimization techniques tailored for Datalog++ is addressed.
Fosca Giannotti, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi
IEEE Trans. Knowl. Data Eng.1
2000 Logic-Based Knowledge Discovery in Databases
Fosca Giannotti, Mirco Nanni, Dino Pedreschi
EJC1
2000 Declarative Knowledge Extraction with Interactive User-Defined Aggregates
abstract
We present the notion of Iterative User-Defined Aggregates as an extension of the notion of user-defined aggregates in deductive databases. Such an extension provides a versative mechanism for defining complex aggregation functions, that are not definable as distributive aggregates. As a result, we show how such a mechanism can be applied to the specification of complex data mining tasks as user-defined aggregates. The resulting formalism provides a flexible way to customize, tune and reason on both the evaluation functions and the extracted knowledge. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Fosca Giannotti, Giuseppe Manco 0001
FQAS1
2000 Making Knowledge Extraction and Reasoning Closer
Fosca Giannotti, Giuseppe Manco 0001
PAKDD1
1999 Using Data Mining Techniques in Fiscal Fraud Detection
Francesco Bonchi, Fosca Giannotti, Gianni Mainetto, Dino Pedreschi
DaWaK2
1999 A Classification-Based Methodology for Planning Audit Strategies in Fraud Detection
abstract
Planning adequate audit strategies is a key success factor in a posterion' fraud detection, e.g., in the fiscal and insurance domains, where audits are intended to detect tax evasion and fraudulent claims.A case study is presented in this paper, which illustrates how techniques based on classification can be used to support the task of planning audit strategies.The proposed approach is sensible to some conflicting issues of audit planning, e.g., the trade-off between maximizing audit benefits vs. minimizing audit costs.A methodological scenario, common to a whole class of similar applications, is then abstracted away from the case study.The limitations of available systems to support the identified overall KDD process, bring us to point out the key aspects of a logic-based database language, integrated with mining mechanisms, which is used to provide a uniform, highly expressive environment for the various steps in the construction of the considered case-study.
Francesco Bonchi, Fosca Giannotti, Gianni Mainetto, Dino Pedreschi
KDD2
1999 Querying Inductive Databases via Logic-Based User-Defined Aggregates
Fosca Giannotti, Giuseppe Manco 0001
PKDD1
1998 Query Answering in Nondeterministic, Nonmonotonic Logic Databases
Fosca Giannotti, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi
FQAS1
1993 Data Sharing Analysis for a Database Programming Lanaguage via Abstract Interpretation
Giuseppe Amato 0001, Fosca Giannotti, Gianni Mainetto
VLDB2