EDBT 2026 Demo / reviewers in the wild / expert
Dino Pedreschi
dblp:p/DPedreschi
· DBLP profile ↗
67ranked-venue papers in the field
3as first author
8since 2021 · last 2025
0000-0003-4801-3225ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 44 (2 first)Database Systems & Data Management · 14Big Data, Cloud & Distributed Data Systems · 3Knowledge Engineering, Semantic Web & Information Systems · 3Information Retrieval & Web Search · 1Business Process & Enterprise Data · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explanations Go Linear: Post-Hoc Explainability for Tabular Data with Interpretable Meta-EncodingabstractPost-hoc explainability is essential for understanding black-box machine learning models. Surrogate-based techniques are widely used for local and global model-agnostic explanations but have significant limitations. Local surrogates capture non-linearities but are computationally expensive and sensitive to parameters, while global surrogates are more efficient but struggle with complex local behaviors. In this paper, we present ILLUME, a flexible and interpretable framework grounded in representation learning, that can be integrated with various surrogate models to provide explanations for any black-box classifier. Specifically, our approach combines a globally trained surrogate with instance-specific linear transformations learned with a meta-encoder to generate both local and global explanations. Through extensive empirical evaluations, we demonstrate the effectiveness of ILLUME in producing feature attributions and decision rules that are not only accurate but also robust and computationally efficient, thus providing a unified explanation framework that effectively addresses the limitations of traditional surrogate methods. Simone Piaggesi, Riccardo Guidotti, Fosca Giannotti, Dino Pedreschi |
ICDM | 4 |
| 2024 | Interpretable and Fair Mechanisms for Abstaining Classifiers
Daphne Lenders, Andrea Pugnana, Roberto Pellungrini, Toon Calders, Dino Pedreschi, Fosca Giannotti |
ECML/PKDD (7) | 5 |
| 2024 | Stable and actionable explanations of black-box models through factual and counterfactual rulesabstractAbstract Recent years have witnessed the rise of accurate but obscure classification models that hide the logic of their internal decision processes. Explaining the decision taken by a black-box classifier on a specific input instance is therefore of striking interest. We propose a local rule-based model-agnostic explanation method providing stable and actionable explanations. An explanation consists of a factual logic rule, stating the reasons for the black-box decision, and a set of actionable counterfactual logic rules, proactively suggesting the changes in the instance that lead to a different outcome. Explanations are computed from a decision tree that mimics the behavior of the black-box locally to the instance to explain. The decision tree is obtained through a bagging-like approach that favors stability and fidelity: first, an ensemble of decision trees is learned from neighborhoods of the instance under investigation; then, the ensemble is merged into a single decision tree. Neighbor instances are synthetically generated through a genetic algorithm whose fitness function is driven by the black-box behavior. Experiments show that the proposed method advances the state-of-the-art towards a comprehensive approach that successfully covers stability and actionability of factual and counterfactual explanations. Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Francesca Naretto, Franco Turini, Dino Pedreschi, Fosca Giannotti |
Data Min. Knowl. Discov. | 6 |
| 2024 | Understanding Any Time Series Classifier with a Subsequence-based ExplainerabstractThe growing availability of time series data has increased the usage of classifiers for this data type. Unfortunately, state-of-the-art time series classifiers are black-box models and, therefore, not usable in critical domains such as healthcare or finance, where explainability can be a crucial requirement. This paper presents a framework to explain the predictions of any black-box classifier for univariate and multivariate time series. The provided explanation is composed of three parts. First, a saliency map highlighting the most important parts of the time series for the classification. Second, an instance-based explanation exemplifies the black-box’s decision by providing a set of prototypical and counterfactual time series. Third, a factual and counterfactual rule-based explanation, revealing the reasons for the classification through logical conditions based on subsequences that must, or must not, be contained in the time series. Experiments and benchmarks show that the proposed method provides faithful, meaningful, stable, and interpretable explanations. Francesco Spinnato, Riccardo Guidotti, Anna Monreale, Mirco Nanni, Dino Pedreschi, Fosca Giannotti |
ACM Trans. Knowl. Discov. Data | 5 |
| 2023 | Benchmarking and survey of explanation methods for black box modelsabstractAbstract The rise of sophisticated black-box machine learning models in Artificial Intelligence systems has prompted the need for explanation methods that reveal how these models work in an understandable way to users and decision makers. Unsurprisingly, the state-of-the-art exhibits currently a plethora of explainers providing many different types of explanations. With the aim of providing a compass for researchers and practitioners, this paper proposes a categorization of explanation methods from the perspective of the type of explanation they return, also considering the different input data formats. The paper accounts for the most representative explainers to date, also discussing similarities and discrepancies of returned explanations through their visual appearance. A companion website to the paper is provided as a continuous update to new explainers as they appear. Moreover, a subset of the most robust and widely adopted explainers, are benchmarked with respect to a repertoire of quantitative metrics. Francesco Bodria, Fosca Giannotti, Riccardo Guidotti, Francesca Naretto, Dino Pedreschi, Salvatore Rinzivillo |
Data Min. Knowl. Discov. | 5 |
| 2022 | Transparent Latent Space Counterfactual Explanations for Tabular DataabstractArtificial Intelligence decision-making systems have dramatically increased their predictive performance in recent years, beating humans in many different specific tasks. However, with increased performance has come an increase in the complexity of the black-box models adopted by the AI systems, making them entirely obscure for the decision process adopted. Explainable AI is a field that seeks to make AI decisions more transparent by producing explanations. In this paper, we propose T-LACE, an approach able to retrieve post-hoc counterfactual explanations for a given pre-trained black-box model. T-LACE exploits the similarity and linearity proprieties of a custom-created transparent latent space to build reliable counterfactual explanations. We tested T-LACE on several tabular datasets and provided qualitative evaluations of the generated explanations in terms of similarity, robustness, and diversity. Comparative analysis against various state-of-the-art counterfactual explanation methods shows the higher effectiveness of our approach. Francesco Bodria, Riccardo Guidotti, Fosca Giannotti, Dino Pedreschi |
DSAA | 4 |
| 2022 | How routing strategies impact urban emissionsabstractNavigation apps use routing algorithms to suggest the best path to reach a user's desired destination. Although undoubtedly useful, navigation apps' impact on the urban environment (e.g., CO2 emissions and pollution) is still largely unclear. In this work, we design a simulation framework to assess the impact of routing algorithms on carbon dioxide emissions within an urban environment. Using APIs from TomTom and OpenStreetMap, we find that settings in which either all vehicles or none of them follow a navigation app's suggestion lead to the worst impact in terms of CO2 emissions. In contrast, when just a portion (around half) of vehicles follow these suggestions, and some degree of randomness is added to the remaining vehicles' paths, we observe a reduction in the overall CO2 emissions over the road network. Our work is a first step towards designing next-generation routing principles that may increase urban well-being while satisfying individual needs. Giuliano Cornacchia, Matteo Böhm, Giovanni Mauro, Mirco Nanni, Dino Pedreschi, Luca Pappalardo |
SIGSPATIAL/GIS | 5 |
| 2021 | FairLens: Auditing black-box clinical decision support systemsabstractThe pervasive application of algorithmic decision-making is raising concerns on the risk of unintended bias in AI systems deployed in critical settings such as healthcare. The detection and mitigation of model bias is a very delicate task that should be tackled with care and involving domain experts in the loop. In this paper we introduce FairLens, a methodology for discovering and explaining biases. We show how this tool can audit a fictional commercial black-box model acting as a clinical decision support system (DSS). In this scenario, the healthcare facility experts can use FairLens on their historical data to discover the biases of the model before incorporating it into the clinical decision flow. FairLens first stratifies the available patient data according to demographic attributes such as age, ethnicity, gender and healthcare insurance; it then assesses the model performance on such groups highlighting the most common misclassifications. Finally, FairLens allows the expert to examine one misclassification of interest by explaining which elements of the affected patients’ clinical history drive the model error in the problematic group. We validate FairLens’ ability to highlight bias in multilabel clinical DSSs introducing a multilabel-appropriate metric of disparity and proving its efficacy against other standard metrics. Cecilia Panigutti, Alan Perotti, André Panisson, Paolo Bajardi, Dino Pedreschi |
Inf. Process. Manag. | 5 |
| 2020 | Causal inference for social discrimination reasoning
Bilal Qureshi, Faisal Kamiran, Asim Karim, Salvatore Ruggieri, Dino Pedreschi |
J. Intell. Inf. Syst. | 5 |
| 2019 | Black Box Explanation by Learning Image Exemplars in the Latent Feature Space
Riccardo Guidotti, Anna Monreale, Stan Matwin, Dino Pedreschi |
ECML/PKDD (1) | 4 |
| 2019 | PlayeRank: Data-driven Performance Evaluation and Player Ranking in Soccer via a Machine Learning ApproachabstractThe problem of evaluating the performance of soccer players is attracting the interest of many companies and the scientific community, thanks to the availability of massive data capturing all the events generated during a match (e.g., tackles, passes, shots, etc.). Unfortunately, there is no consolidated and widely accepted metric for measuring performance quality in all of its facets. In this article, we design and implement PlayeRank, a data-driven framework that offers a principled multi-dimensional and role-aware evaluation of the performance of soccer players. We build our framework by deploying a massive dataset of soccer-logs and consisting of millions of match events pertaining to four seasons of 18 prominent soccer competitions. By comparing PlayeRank to known algorithms for performance evaluation in soccer, and by exploiting a dataset of players’ evaluations made by professional soccer scouts, we show that PlayeRank significantly outperforms the competitors. We also explore the ratings produced by PlayeRank and discover interesting patterns about the nature of excellent performances and what distinguishes the top players from the others. At the end, we explore some applications of PlayeRank—i.e. searching players and player versatility—showing its flexibility and efficiency, which makes it worth to be used in the design of a scalable platform for soccer analytics. Luca Pappalardo, Paolo Cintia, Paolo Ferragina, Emanuele Massucco, Dino Pedreschi, Fosca Giannotti |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | Personalized Market Basket Prediction with Temporal Annotated Recurring SequencesabstractNowadays, a hot challenge for supermarket chains is to offer personalized services to their customers. Market basket prediction, i.e., supplying the customer a shopping list for the next purchase according to her current needs, is one of these services. Current approaches are not capable of capturing at the same time the different factors influencing the customer's decision process: co-occurrence, sequentuality, periodicity, and recurrency of the purchased items. To this aim, we define a pattern TemporalAnnotated Recurring Sequence (TARS) able to capture simultaneously and adaptively all these factors. We define the method to extract TARS and develop a predictor for next basket named TBP (TARS Based Predictor) that, on top of TARS, is able to understand the level of the customer's stocks and recommend the set of most necessary items. By adopting the TBP the supermarket chains could crop tailored suggestions for each individual customer which in turn could effectively speed up their shopping sessions. A deep experimentation shows that TARS are able to explain the customer purchase behavior, and that TBP outperforms the state-of-the-art competitors. Riccardo Guidotti, Giulio Rossetti, Luca Pappalardo, Fosca Giannotti, Dino Pedreschi |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Helping Your Docker Images to Spread Based on Explainable Models
Riccardo Guidotti, Jacopo Soldani, Davide Neri, Antonio Brogi, Dino Pedreschi |
ECML/PKDD (3) | 5 |
| 2017 | There's a Path for Everyone: A Data-Driven Personal Model Reproducing Mobility AgendasabstractThe avalanche of mobility data like GPS and GSM daily produced by each user through mobile devices enables personalized mobility-services improving everyday life. The base for these mobility-services lies in the predictability of human behavior. In this paper we propose an approach for reproducing the user's personal mobility agenda that is able to predict the user's positions for the whole day. We reproduce the agenda by exploiting a data-driven personal mobility model able to capture and summarize different aspects of the systematic mobility behavior of a user. We show how the proposed approach outperforms typical methodologies adopted in the literature on four different real GPS datasets. Moreover, we analyze some features of the mobility models and we discuss how they can be employed as agents of a simulator for what-if mobility analysis. Riccardo Guidotti, Roberto Trasarti, Mirco Nanni, Fosca Giannotti, Dino Pedreschi |
DSAA | 5 |
| 2017 | NDlib: Studying Network Diffusion DynamicsabstractNowadays the analysis of diffusive phenomena occurring on top of complex networks represents a hot topic in the Social Network Analysis playground. In order to support students, teachers, developers and researchers in this work we introduce a novel simulation framework, NDlib. NDlib is designed to be a multi-level ecosystem that can be fruitfully used by different user segments. Upon the diffusion library, we designed a simulation server that allows remote execution of experiments and an online visualization tool that abstract the programmatic interface and makes available the simulation platform to non-technicians. Giulio Rossetti, Letizia Milli, Salvatore Rinzivillo, Alina Sîrbu, Dino Pedreschi, Fosca Giannotti |
DSAA | 5 |
| 2017 | Market Basket Prediction Using User-Centric Temporal Annotated Recurring SequencesabstractNowadays, a hot challenge for supermarket chains is to offer personalized services to their customers. Market basket prediction, i.e., supplying the customer a shopping list for the next purchase according to her current needs, is one of these services. Current approaches are not capable of capturing at the same time the different factors influencing the customer's decision process: co-occurrence, sequentuality, periodicity and recurrency of the purchased items. To this aim, we define a pattern named Temporal Annotated Recurring Sequence (TARS). We define the method to extract TARS and develop a predictor for next basket named TBP (TARS Based Predictor) that, on top of TARS, is able to understand the level of the customer's stocks and recommend the set of most necessary items. A deep experimentation shows that TARS can explain the customers' purchase behavior, and that TBP outperforms the state-of-the-art competitors. Riccardo Guidotti, Giulio Rossetti, Luca Pappalardo, Fosca Giannotti, Dino Pedreschi |
ICDM | 5 |
| 2017 | Clustering Individual Transactional Data for Masses of UsersabstractMining a large number of datasets recording human activities for making sense of individual data is the key enabler of a new wave of personalized knowledge-based services. In this paper we focus on the problem of clustering individual transactional data for a large mass of users. Transactional data is a very pervasive kind of information that is collected by several services, often involving huge pools of users. We propose txmeans, a parameter-free clustering algorithm able to efficiently partitioning transactional data in a completely automatic way. Txmeans is designed for the case where clustering must be applied on a massive number of different datasets, for instance when a large set of users need to be analyzed individually and each of them has generated a long history of transactions. A deep experimentation on both real and synthetic datasets shows the practical effectiveness of txmeans for the mass clustering of different personal datasets, and suggests that txmeans outperforms existing methods in terms of quality and efficiency. Finally, we present a personal cart assistant application based on txmeans Riccardo Guidotti, Anna Monreale, Mirco Nanni, Fosca Giannotti, Dino Pedreschi |
KDD | 5 |
| 2017 | Never drive alone: Boosting carpooling with network analysis
Riccardo Guidotti, Mirco Nanni, Salvatore Rinzivillo, Dino Pedreschi, Fosca Giannotti |
Inf. Syst. | 4 |
| 2016 | Driving Profiles Computation and Monitoring for Car Insurance CRMabstractCustomer segmentation is one of the most traditional and valued tasks in customer relationship management (CRM). In this article, we explore the problem in the context of the car insurance industry, where the mobility behavior of customers plays a key role: Different mobility needs, driving habits, and skills imply also different requirements (level of coverage provided by the insurance) and risks (of accidents). In the present work, we describe a methodology to extract several indicators describing the driving profile of customers, and we provide a clustering-oriented instantiation of the segmentation problem based on such indicators. Then, we consider the availability of a continuous flow of fresh mobility data sent by the circulating vehicles, aiming at keeping our segments constantly up to date. We tackle a major scalability issue that emerges in this context when the number of customers is large—namely, the communication bottleneck—by proposing and implementing a sophisticated distributed monitoring solution that reduces communications between vehicles and company servers to the essential. We validate the framework on a large database of real mobility data coming from GPS devices on private cars. Finally, we analyze the privacy risks that the proposed approach might involve for the users, providing and evaluating a countermeasure based on data perturbation. Mirco Nanni, Roberto Trasarti, Anna Monreale, Valerio Grossi, Dino Pedreschi |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2015 | Interaction Prediction in Dynamic Networks exploiting Community DiscoveryabstractDue to the growing availability of online social services, interactions between people became more and more easy to establish and track. Online social human activities generate digital footprints, that describe complex, rapidly evolving, dynamic networks. In such scenario one of the most challenging task to address involves the prediction of future interactions between couples of actors. In this study, we want to leverage networks dynamics and community structure to predict which are the future interactions more likely to appear. To this extent, we propose a supervised learning approach which exploit features computed by time-aware forecasts of topological measures calculated between pair of nodes belonging to the same community. Our experiments on real dynamic networks show that the designed analytical process is able to achieve interesting results. Giulio Rossetti, Riccardo Guidotti, Diego Pennacchioli, Dino Pedreschi, Fosca Giannotti |
ASONAM | 4 |
| 2015 | Community-centric analysis of user engagement in Skype social networkabstractTraditional approaches to user engagement analysis focus on individual users. In this paper we address user engagement analysis at the level of groups of users (social communities). From the entire Skype social network we extract communities by means of representative community detection methods each one providing node partitions having their own peculiarities. We then examine user engagement in the extracted communities putting into evidence clear relations between topological and geographic features of communities and their mean user engagement. In particular we show that user engagement can be to a great extent predicted from such features. Moreover, from the analysis it clearly emerges that the choice of community definition and granularity deeply affect the predictive performance. Giulio Rossetti, Luca Pappalardo, Riivo Kikas, Dino Pedreschi, Fosca Giannotti, Marlon Dumas |
ASONAM | 4 |
| 2015 | City users' classification with mobile phone dataabstractNowadays mobile phone data are an actual proxy for studying the users' social life and urban dynamics. In this paper we present the Sociometer, and analytical framework aimed at classifying mobile phone users into behavioral categories by means of their call habits. The analytical process starts from spatio-temporal profiles, learns the different behaviors, and returns annotated profiles. After the description of the methodology and its evaluation, we present an application of the Sociometer for studying city users of one small and one big city, evaluating the impact of big events in these cities. Lorenzo Gabrielli, Barbara Furletti, Roberto Trasarti, Fosca Giannotti, Dino Pedreschi |
IEEE BigData | 5 |
| 2015 | Using big data to study the link between human mobility and socio-economic developmentabstractBig Data offer nowadays the potential capability of creating a digital nervous system of our society, enabling the measurement, monitoring and prediction of relevant aspects of socio-economic phenomena in quasi real time. This potential has fueled, in the last few years, a growing interest around the usage of Big Data to support official statistics in the measurement of individual and collective economic well-being. In this work we study the relations between human mobility patterns and socioeconomic development. Starting from nation-wide mobile phone data we extract a measure of mobility volume and a measure of mobility diversity for each individual. We then aggregate the mobility measures at municipality level and investigate the correlations with external socio-economic indicators independently surveyed by an official statistics institute. We find three main results. First, aggregated human mobility patterns are correlated with these socio-economic indicators. Second, the diversity of mobility, defined in terms of entropy of the individual users' trajectories, exhibits the strongest correlation with the external socio-economic indicators. Third, the volume of mobility and the diversity of mobility show opposite correlations with the socioeconomic indicators. Our results, validated against a null model, open an interesting perspective to study human behavior through Big Data by means of new statistical indicators that quantify and possibly "nowcast" the socio-economic development of our society. Luca Pappalardo, Dino Pedreschi, Zbigniew Smoreda, Fosca Giannotti |
IEEE BigData | 2 |
| 2015 | The harsh rule of the goals: Data-driven performance indicators for football teamsabstractSports analytics in general, and football (soccer in USA) analytics in particular, have evolved in recent years in an amazing way, thanks to automated or semi-automated sensing technologies that provide high-fidelity data streams extracted from every game. In this paper we propose a data-driven approach and show that there is a large potential to boost the understanding of football team performance. From observational data of football games we extract a set of pass-based performance indicators and summarize them in the H indicator. We observe a strong correlation among the proposed indicator and the success of a team, and therefore perform a simulation on the four major European championships (78 teams, almost 1500 games). The outcome of each game in the championship was replaced by a synthetic outcome (win, loss or draw) based on the performance indicators computed for each team. We found that the final rankings in the simulated championships are very close to the actual rankings in the real championships, and show that teams with high ranking error show extreme values of a defense/attack efficiency measure, the Pezzali score. Our results are surprising given the simplicity of the proposed indicators, suggesting that a complex systems' view on football data has the potential of revealing hidden patterns and behavior of superior quality. Paolo Cintia, Fosca Giannotti, Luca Pappalardo, Dino Pedreschi, Marco Malvaldi |
DSAA | 4 |
| 2015 | Behavioral entropy and profitability in retailabstractHuman behavior is predictable in principle: people are systematic in their everyday choices. This predictability can be used to plan events and infrastructure, both for the public good and for private gains. In this paper we investigate the largely unexplored relationship between the systematic behavior of a customer and its profitability for a retail company. We estimate a customer's behavioral entropy over two dimensions: the basket entropy is the variety of what customers buy, and the spatio-temporal entropy is the spatial and temporal variety of their shopping sessions. To estimate the basket and the spatio-temporal entropy we use data mining and information theoretic techniques. We find that predictable systematic customers are more profitable for a supermarket: their average per capita expenditures are higher than non systematic customers and they visit the shops more often. However, this higher individual profitability is masked by its overall level. The highly systematic customers are a minority of the customer set. As a consequence, the total amount of revenues they generate is small. We suggest that favoring a systematic behavior in their customers might be a good strategy for supermarkets to increase revenue. These results are based on data coming from a large Italian supermarket chain, including more than 50 thousand customers visiting 23 shops to purchase more than 80 thousand distinct products. Riccardo Guidotti, Michele Coscia, Dino Pedreschi, Diego Pennacchioli |
DSAA | 3 |
| 2015 | Quantification in social networksabstractIn many real-world applications there is a need to monitor the distribution of a population across different classes, and to track changes in this distribution over time. As an example, an important task is to monitor the percentage of unemployed adults in a given region. When the membership of an individual in a class cannot be established deterministically, a typical solution is the classification task. However, in the above applications the final goal is not determining which class the individuals belong to, but estimating the prevalence of each class in the unlabeled data. This task is called quantification. Most of the work in the literature addressed the quantification problem considering data presented in conventional attribute format. Since the ever-growing availability of web and social media we have a flourish of network data representing a new important source of information and by using quantification network techniques we could quantify collective behavior, i.e., the number of users that are involved in certain type of activities, preferences, or behaviors. In this paper we exploit the homophily effect observed in many social networks in order to construct a quantifier for networked data. Our experiments show the effectiveness of the proposed approaches and the comparison with the existing state-of-the-art quantification methods shows that they are more accurate. Letizia Milli, Anna Monreale, Giulio Rossetti, Dino Pedreschi, Fosca Giannotti, Fabrizio Sebastiani 0001 |
DSAA | 4 |
| 2015 | Discrimination- and privacy-aware patterns
Sara Hajian, Josep Domingo-Ferrer, Anna Monreale, Dino Pedreschi, Fosca Giannotti |
Data Min. Knowl. Discov. | 4 |
| 2014 | The purpose of motion: Learning activities from Individual Mobility NetworksabstractThe large availability of mobility data allows us to investigate complex phenomena about human movement. However this adundance of data comes with few information about the purpose of movement. In this work we address the issue of activity recognition by introducing Activity-Based Cascading (ABC) classification. Such approach departs completely from probabilistic approaches for two main reasons. First, it exploits a set of structural features extracted from the Individual Mobility Network (IMN), a model able to capture the salient aspects of individual mobility. Second, it uses a cascading classification as a way to tackle the highly skewed frequency of activity classes. We show that our approach outperforms existing state-of-the-art probabilistic methods. Since it reaches high precision, ABC classification represents a very reliable semantic amplifier for Big Data. Salvatore Rinzivillo, Lorenzo Gabrielli, Mirco Nanni, Luca Pappalardo, Dino Pedreschi, Fosca Giannotti |
DSAA | 5 |
| 2014 | Uncovering Hierarchical and Overlapping Communities with a Local-First ApproachabstractCommunity discovery in complex networks is the task of organizing a network’s structure by grouping together nodes related to each other. Traditional approaches are based on the assumption that there is a global-level organization in the network. However, in many scenarios, each node is the bearer of complex information and cannot be classified in disjoint clusters. The top-down global view of the partition approach is not designed for this. Here, we represent this complex information as multiple latent labels, and we postulate that edges in the networks are created among nodes carrying similar labels. The latent labels are the communities a node belongs to and we discover them with a simple local-first approach to community discovery. This is achieved by democratically letting each node vote for the communities it sees surrounding it in its limited view of the global system, its ego neighborhood, using a label propagation algorithm, assuming that each node is aware of the label it shares with each of its connections. The local communities are merged hierarchically, unveiling the modular organization of the network at the global level and identifying overlapping groups and groups of groups. We tested this intuition against the state-of-the-art overlapping community discovery and found that our new method advances in the chosen scenarios in the quality of the obtained communities. We perform a test on benchmark and on real-world networks, evaluating the quality of the community coverage by using the extracted communities to predict the metadata attached to the nodes, which we consider external information about the latent labels. We also provide an explanation about why real-world networks contain overlapping communities and how our logic is able to capture them. Finally, we show how our method is deterministic, is incremental, and has a limited time complexity, so that it can be used on real-world scale networks. Michele Coscia, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi |
ACM Trans. Knowl. Discov. Data | 4 |
| 2013 | Explaining the product range effect in purchase dataabstractIn our market society, buyers are considered rational entities, driven by two utility functions: i) the amount of money spent, a universal quantity to be minimized; and ii) the individual needs to satisfy, a personal quantity, varying from person to person, to be maximized. In this paper, we propose an analytic framework based on big data to measure the personal utility function and we prove that this function has a stronger effect on customer behavior than the price. By focusing on the purchases in an Italian supermarket chain, we discover and describe a range effect of products: the more sophisticated the needs they satisfy, the more cost the customers are willing to pay to buy them, in terms of distance to travel more than in terms of the price of the item itself. We exhibit a striking empirical evidence of this theory by tracking the geographical information about points of sale and customers, in a large dataset containing tens of thousands of customers and thousands of products. We create a data mining framework able to scale to possibly hundreds of thousands, or millions, of customers and to let emerge from the data the knowledge about the actual range of each product. As an application of this finding, we show how it is possible to accurately predict how long a customer will travel (or which shop she will choose) to buy a product, as a function of the product's sophistication. Diego Pennacchioli, Michele Coscia, Salvatore Rinzivillo, Dino Pedreschi, Fosca Giannotti |
IEEE BigData | 4 |
| 2013 | Quantification TreesabstractIn many applications there is a need to monitor how a population is distributed across different classes, and to track the changes in this distribution that derive from varying circumstances, an example such application is monitoring the percentage (or "prevalence") of unemployed people in a given region, or in a given age range, or at different time periods. When the membership of an individual in a class cannot be established deterministically, this monitoring activity requires classification. However, in the above applications the final goal is not determining which class each individual belongs to, but simply estimating the prevalence of each class in the unlabeled data. This task is called quantification. In a supervised learning framework we may estimate the distribution across the classes in a test set from a training set of labeled individuals. However, this may be sub optimal, since the distribution in the test set may be substantially different from that in the training set (a phenomenon called distribution drift). So far, quantification has mostly been addressed by learning a classifier optimized for individual classification and later adjusting the distribution it computes to compensate for its tendency to either under-or over-estimate the prevalence of the class. In this paper we propose instead to use a type of decision trees (quantification trees) optimized not for individual classification, but directly for quantification. Our experiments show that quantification trees are more accurate than existing state-of-the-art quantification methods, while retaining at the same time the simplicity and understandability of the decision tree framework. Letizia Milli, Anna Monreale, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi, Fabrizio Sebastiani 0001 |
ICDM | 5 |
| 2012 | Optimal Spatial Resolution for the Analysis of Human MobilityabstractThe availability of massive network and mobility data from diverse domains has fostered the analysis of human behaviors and interactions. This data availability leads to challenges in the knowledge discovery community. Several different analyses have been performed on the traces of human trajectories, such as understanding the real borders of human mobility or mining social interactions derived from mobility and vice versa. However, the data quality of the digital traces of human mobility has a dramatic impact over the knowledge that it is possible to mine, and this issue has not been thoroughly tackled so far in literature. In this paper, we mine and analyze with complex network techniques a large dataset of human trajectories, a GPS dataset from more than 150k vehicles in Italy. We build a multi resolution grid and we map the trajectories with several complex networks, by connecting the different areas of our region of interest. Then we analyze the structural properties of these networks and the quality of the borders it is possible to infer from them. The result is a significant advancement in our understanding of the data transformation process that is needed to connect mobility with social network analysis and mining. Michele Coscia, Salvatore Rinzivillo, Fosca Giannotti, Dino Pedreschi |
ASONAM | 4 |
| 2012 | "How Well Do We Know Each Other?" Detecting Tie Strength in Multidimensional Social NetworksabstractThe advent of social media have allowed us to build massive networks of weak ties: acquaintances and nonintimate ties we use all the time to spread information and thoughts. Conversely, strong ties are the people we really trust, people whose social circles tightly overlap with our own and, often, they are also the people most like us. Unfortunately, the majority of social media do not incorporate explicitly tie strength information in the creation and management of relationships, and treat all users the same: friend or stranger, with little or nothing in between. In the current work, we address the challenging issue of detecting on online social networks the strong and intimate ties from the huge mass of such mere social contacts. In order to do so, we propose a novel multidimensional definition of tie strength which exploits the existence of multiple online social links between two individuals. We test our definition on a multidimensional network constructed over users in Foursquare, Twitter and Facebook, analyzing the structural role of strong and weak links, and the correlations with the most common similarity measures. Luca Pappalardo, Giulio Rossetti, Dino Pedreschi |
ASONAM | 3 |
| 2012 | Mega-modeling for Big Data Analytics
Stefano Ceri, Emanuele Della Valle, Dino Pedreschi, Roberto Trasarti |
ER | 3 |
| 2012 | DEMON: a local-first discovery method for overlapping communitiesabstractCommunity discovery in complex networks is an interesting problem with a number of applications, especially in the knowledge extraction task in social and information networks. However, many large networks often lack a particular community organization at a global level. In these cases, traditional graph partitioning algorithms fail to let the latent knowledge embedded in modular structure emerge, because they impose a top-down global view of a network. We propose here a simple local-first approach to community discovery, able to unveil the modular organization of real complex networks. This is achieved by democratically letting each node vote for the communities it sees surrounding it in its limited view of the global system, i.e. its ego neighborhood, using a label propagation algorithm; finally, the local communities are merged into a global collection. We tested this intuition against the state-of-the-art overlapping and non-overlapping community discovery methods, and found that our new method clearly outperforms the others in the quality of the obtained communities, evaluated by using the extracted communities to predict the metadata about the nodes of several real world networks. We also show how our method is deterministic, fully incremental, and has a limited time complexity, so that it can be used on web-scale real networks. Michele Coscia, Giulio Rossetti, Fosca Giannotti, Dino Pedreschi |
KDD | 4 |
| 2012 | AUDIO: An Integrity Auditing Framework of Outlier-Mining-as-a-Service Systems
Wendy Hui Wang, Anna Monreale, Dino Pedreschi, Fosca Giannotti, Wenge Guo |
ECML/PKDD (2) | 4 |
| 2011 | Foundations of Multidimensional Network AnalysisabstractComplex networks have been receiving increasing attention by the scientific community, thanks also to the increasing availability of real-world network data. In the last years, the multidimensional nature of many real world networks has been pointed out, i.e. many networks containing multiple connections between any pair of nodes have been analyzed. Despite the importance of analyzing this kind of networks was recognized by previous works, a complete framework for multidimensional network analysis is still missing. Such a framework would enable the analysts to study different phenomena, that can be either the generalization to the multidimensional setting of what happens inmonodimensional network, or a new class of phenomena induced by the additional degree of complexity that multidimensionality provides in real networks. The aim of this paper is then to give the basis for multidimensional network analysis: we develop a solid repertoire of basic concepts and analytical measures, which takes into account the general structure of multidimensional networks. We tested our framework on a real world multidimensional network, showing the validity and the meaningfulness of the measures introduced, that are able to extract important, nonrandom, information about complex phenomena. Michele Berlingerio, Michele Coscia, Fosca Giannotti, Anna Monreale, Dino Pedreschi |
ASONAM | 5 |
| 2011 | Human mobility, social ties, and link predictionabstractOur understanding of how individual mobility patterns shape and impact the social network is limited, but is essential for a deeper understanding of network dynamics and evolution. This question is largely unexplored, partly due to the difficulty in obtaining large-scale society-wide data that simultaneously capture the dynamical information on individual movements and social interactions. Here we address this challenge for the first time by tracking the trajectories and communication records of 6 Million mobile phone users. We find that the similarity between two individuals' movements strongly correlates with their proximity in the social network. We further investigate how the predictive power hidden in such correlations can be exploited to address a challenging problem: which new links will develop in a social network. We show that mobility measures alone yield surprising predictive power, comparable to traditional network-based measures. Furthermore, the prediction accuracy can be significantly improved by learning a supervised classifier based on combined mobility and network measures. We believe our findings on the interplay of mobility patterns and social ties offer new perspectives on not only link prediction but also network dynamics. Dashun Wang, Dino Pedreschi, Chaoming Song, Fosca Giannotti, Albert-László Barabási |
KDD | 2 |
| 2011 | Challenges for Mobile Data Management in the Era of Cloud and Social ComputingabstractThe mobile data management community is experiencing a rapid evolutionary change due to the worldwide diffusion of always-on mobile devices and to the increased popularity of location and context-aware mobile applications. Accordingly to recent studies, in two years from now one fourth of the total mobile data will come from audio and video streaming and nearly all the rest from other Internet services. A large part of the increase in mobile data will come from cloud computing applications that are massively used for storing personal data, for sharing data, as well as for utility software (such as maps) and productivity tools. Social networking will strongly influence the way mobile users choose, share and use content from mobile devices. On the other side mobile devices are changing the way social networks have been used till now introducing geo-tagging, location sharing, and many innovative location based services. Chatschik Bisdikian, Bernhard Mitschang, Dino Pedreschi, Vincent S. Tseng, Claudio Bettini |
Mobile Data Management (1) | 3 |
| 2011 | Unveiling the complexity of human mobility by querying and mining massive trajectory data
Fosca Giannotti, Mirco Nanni, Dino Pedreschi, Fabio Pinelli, Chiara Renso, Salvatore Rinzivillo, Roberto Trasarti |
VLDB J. | 3 |
| 2010 | Advanced knowledge discovery on movement data with the GeoPKDD systemabstractThe growing availability of mobile devices produces an enor- mous quantity of personal tracks which calls for advanced analysis methods capable of extracting knowledge out of massive trajectories datasets. In this paper we present an experiment on a real world scenario that demonstrates the strong analytical power of massive, raw trajectory data made available as a by-product of telecom services, in unveiling the complexity of urban mobility. The experiment has been made possible by the GeoPKDD system, an integrated plat- form for complex analysis of mobility data. The system com- bines spatio-temporal querying capabilities with data min- ing and semantic technologies, thus providing a full support for the Mobility Knowledge Discovery process. Mirco Nanni, Roberto Trasarti, Chiara Renso, Fosca Giannotti, Dino Pedreschi |
EDBT | 5 |
| 2010 | As Time Goes by: Discovering Eras in Evolving Social Networks
Michele Berlingerio, Michele Coscia, Fosca Giannotti, Anna Monreale, Dino Pedreschi |
PAKDD (1) | 5 |
| 2010 | Exploring Real Mobility Data with M-Atlas
Roberto Trasarti, Salvatore Rinzivillo, Fabio Pinelli, Mirco Nanni, Anna Monreale, Chiara Renso, Dino Pedreschi, Fosca Giannotti |
ECML/PKDD (3) | 7 |
| 2010 | DCUBE: discrimination discovery in databasesabstractDiscrimination discovery in databases consists in finding unfair practices against minorities which are hidden in a dataset of historical decisions. The DCUBE system implements the approach of [5], which is based on classification rule extraction and analysis, by centering the analysis phase around an Oracle database. The proposed demonstration guides the audience through the legal issues about discrimination hidden in data, and through several legally-grounded analyses to unveil discriminatory situations. The SIGMOD attendees will freely pose complex discrimination analysis queries over the database of extracted classification rules, once they are presented with the database relational schema, a few ad-hoc functions and procedures, and several snippets of SQL queries for discrimination discovery. Salvatore Ruggieri, Dino Pedreschi, Franco Turini |
SIGMOD Conference | 2 |
| 2010 | Data mining for discrimination discoveryabstractIn the context of civil rights law, discrimination refers to unfair or unequal treatment of people based on membership to a category or a minority, without regard to individual merit. Discrimination in credit, mortgage, insurance, labor market, and education has been investigated by researchers in economics and human sciences. With the advent of automatic decision support systems, such as credit scoring systems, the ease of data collection opens several challenges to data analysts for the fight against discrimination. In this article, we introduce the problem of discovering discrimination through data mining in a dataset of historical decision records, taken by humans or by automatic systems. We formalize the processes of direct and indirect discrimination discovery by modelling protected-by-law groups and contexts where discrimination occurs in a classification rule based syntax. Basically, classification rules extracted from the dataset allow for unveiling contexts of unlawful discrimination, where the degree of burden over protected-by-law groups is formalized by an extension of the lift measure of a classification rule. In direct discrimination, the extracted rules can be directly mined in search of discriminatory contexts. In indirect discrimination, the mining process needs some background knowledge as a further input, for example, census data, that combined with the extracted rules might allow for unveiling contexts of discriminatory decisions. A strategy adopted for combining extracted classification rules with background knowledge is called an inference model. In this article, we propose two inference models and provide automatic procedures for their implementation. An empirical assessment of our results is provided on the German credit dataset and on the PKDD Discovery Challenge 1999 financial dataset. Salvatore Ruggieri, Dino Pedreschi, Franco Turini |
ACM Trans. Knowl. Discov. Data | 2 |
| 2009 | Geographic privacy-aware knowledge discovery and deliveryabstractA flood of data pertinent to moving objects is available today, and will be more in the near future, particularly due to the automated collection of privacy-sensitive telecom data from mobile phones and other location-aware devices. Such wealth of data, referenced both in space and time, may enable novel classes of applications of high societal and economic impact, provided that the discovery of consumable and concise knowledge out of these raw data is made possible. Recent research activities have developed theory, techniques and systems for geographic knowledge discovery and delivery, some of them based on privacy-preserving methods for extracting knowledge from large amounts of raw data referenced in space and time. All these efforts aim at devising knowledge discovery and analysis methods for trajectories of moving objects.The fundamental hypothesis is that it is possible, in principle, to aid citizens in their mobile activities by analysing the traces of their past activities by means of data mining techniques. For instance, behavioural patterns derived from mobile trajectories may allow inducing traffic flow information, capable to help people travel efficiently, to help public administrations in traffic-related decision making for sustainable mobility and security management, as well as to help mobile operators in optimising bandwidth and power allocation on the network. On the other hand, it is clear that the use of personal sensitive data arouses concerns about citizen's privacy rights.In this tutorial, we establish a framework for the challenges and the mining solutions for the geographic information collected by Moving Object Database (MOD) engines. We first discuss the challenges of collecting mobility data, and elaborate on the impact of trajectory data analysis in several modern applications. We then discuss methodologies and techniques to collect raw data, reconstruct trajectory information, and efficiently store it in MODs. We continue with an overview of knowledge discovery approaches for movement data. Finally, we propose a research agenda and identify areas where interdisciplinary studies are needed. Fosca Giannotti, Dino Pedreschi, Yannis Theodoridis |
EDBT | 2 |
| 2009 | Measuring Discrimination in Socially-Sensitive Decision RecordsabstractDiscrimination in social sense (e.g., against minorities and disadvantaged groups) is the subject of many laws worldwide, and it has been extensively studied in the social and economic sciences. We tackle the problem of determining, given a dataset of historical decision records, a precise measure of the degree of discrimination suffered by a given group (e.g., an etnic minority) in a given context (e.g., a geographic area) with respect to the decision (e.g. credit denial). In our approach, this problem is rephrased in a classification rule based setting, and a collection of quantitative measures of discrimination is introduced, on the basis of existing norms and regulations. The measures are defined as functions of the contingency table of a classification rule, and their statistical significance is assessed, relying on a large body of statistical inference methods for proportions. Based on this basic method, we are then able to address the more general problems of: (1) unveiling all discriminatory decision patterns hidden in the historical data, combining discrimination analysis with association rule mining, (2) unveiling discrimination in classifiers that learn over training data biased by discriminatory decisions, and (3) in the case of rule-based classifiers, sanitizing discriminatory rules by correcting their confidence. Our approach is validated on the German credit dataset and on the CPAR classifier. Dino Pedreschi, Salvatore Ruggieri, Franco Turini |
SDM | 1 |
| 2009 | A Visual Analytics Toolkit for Cluster-Based Classification of Mobility Data
Gennady L. Andrienko, Natalia V. Andrienko, Salvatore Rinzivillo, Mirco Nanni, Dino Pedreschi |
SSTD | 5 |
| 2008 | Discrimination-aware data miningabstractIn the context of civil rights law, discrimination refers to unfair or unequal treatment of people based on membership to a category or a minority, without regard to individual merit. Rules extracted from databases by data mining techniques, such as classification or association rules, when used for decision tasks such as benefit or credit approval, can be discriminatory in the above sense. In this paper, the notion of discriminatory classification rules is introduced and studied. Providing a guarantee of non-discrimination is shown to be a non trivial task. A naive approach, like taking away all discriminatory attributes, is shown to be not enough when other background knowledge is available. Our approach leads to a precise formulation of the redlining problem along with a formal result relating discriminatory rules with apparently safe ones by means of background knowledge. An empirical assessment of the results on the German credit dataset is also provided. Dino Pedreschi, Salvatore Ruggieri, Franco Turini |
KDD | 1 |
| 2008 | Anonymity preserving pattern discovery
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi |
VLDB J. | 4 |
| 2007 | Trajectory pattern miningabstractThe increasing pervasiveness of location-acquisition technologies (GPS, GSM networks, etc.) is leading to the collection of large spatio-temporal datasets and to the opportunity of discovering usable knowledge about movement behaviour, which fosters novel applications and services. In this paper, we move towards this direction and develop an extension of the sequential pattern mining paradigm that analyzes the trajectories of moving objects. We introduce trajectory patterns as concise descriptions of frequent behaviours, in terms of both space (i.e., the regions of space visited during movements) and time (i.e., the duration of movements). In this setting, we provide a general formal statement of the novel mining problem and then study several different instantiations of different complexity. The various approaches are then empirically evaluated over real data and synthetic benchmarks, comparing their strengths and weaknesses. Fosca Giannotti, Mirco Nanni, Fabio Pinelli, Dino Pedreschi |
KDD | 4 |
| 2007 | Privacy-Aware Knowledge Discovery from Location DataabstractSpatio-temporal, geo-referenced datasets are growing rapidly, and will be more in the near future. This phenomenon is mostly due to the daily collection of telecommunication data from mobile phones and other location-aware devices and is expected to enable novel classes of applications based on the extraction of behavioral patterns from mobility data. Such patterns could be used for instance in traffic and sustainable mobility management (e.g., to study the accessibility to services), urban planning, environmental monitoring, and collaborative location-based services. Clearly, in these applications privacy is a concern, since some knowledge may be sensitive, or an over-specific pattern may reveal the behaviour of groups of few individual. In this paper we focus on automated privacy-preserving methods we developed for extracting and sharing user- consumable forms of knowledge from large amounts of raw data referenced in space and in time. Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi, Osman Abul |
MDM | 4 |
| 2006 | Efficient Mining of Temporally Annotated SequencesabstractSequential patterns mining received much attention in recent years, thanks to its various potential application domains. A large part of them represent data as collections of time-stamped itemsets, e.g., customers' purchases, logged web accesses, etc. Most approaches to sequence mining focus on sequentiality of data, using time-stamps only to order items and, in some cases, to constrain the temporal gap between items. In this paper, we propose an efficient algorithm for computing (temporally-)annotated sequential patterns, i.e., sequential patterns where each transition is annotated with a typical transition time derived from the source data. The algorithm adopts a prefix-projection approach to mine candidate sequences, and it is tightly integrated with an annotation mining process that associates sequences with temporal annotations. The pruning capabilities of the two steps sum together, yielding significant improvements in performances, as demonstrated by a set of experiments performed on synthetic datasets. Fosca Giannotti, Mirco Nanni, Dino Pedreschi |
SDM | 3 |
| 2006 | Time-focused clustering of trajectories of moving objects
Mirco Nanni, Dino Pedreschi |
J. Intell. Inf. Syst. | 2 |
| 2005 | Blocking Anonymity Threats Raised by Frequent Itemset MiningabstractIn this paper we study when the disclosure of data mining results represents, per se, a threat to the anonymity of the individuals recorded in the analyzed database. The novelty of our approach is that we focus on an objective definition of privacy compliance of patterns without any reference to a preconceived knowledge of what is sensitive and what is not, on the basis of the rather intuitive and realistic constraint that the anonymity of individuals should be guaranteed. In particular, the problem addressed here arises from the possibility of inferring from the output of frequent itemset mining (i.e., a set of item-sets with support larger than a threshold a), the existence of patterns with very low support (smaller than an anonymity threshold k)[M. Atzori et. al, 2005]. In the following we develop a simple methodology to block such inference opportunities by introducing distortion on the dangerous patterns. Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi |
ICDM | 4 |
| 2005 | k-Anonymous Patterns
Maurizio Atzori, Francesco Bonchi, Fosca Giannotti, Dino Pedreschi |
PKDD | 4 |
| 2005 | Efficient breadth-first mining of frequent pattern with monotone constraints
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi |
Knowl. Inf. Syst. | 4 |
| 2003 | ExAMiner: Optimized Level-wise Frequent Pattern Mining with Monotone ConstraintabstractThe key point is that, in frequent pattern mining, the most appropriate way of exploiting monotone constraints in conjunction with frequency is to use them in order to reduce the problem input together with the search space. Following this intuition, we introduce ExAMiner, a level-wise algorithm which exploits the real synergy of antimonotone and monotone constraints: the total benefit is greater than the sum of the two individual benefits. ExAMiner generalizes the basic idea of the preprocessing algorithm ExAnte [F. Bonchi et al., (2003)], embedding such ideas at all levels of an Apriori-like computation. The resulting algorithm is the generalization of the Apriori algorithm when a conjunction of monotone constraints is conjoined to the frequency antimonotone constraint. Experimental results confirm that this is, so far, the most efficient way of attacking the computational problem in analysis. Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi |
ICDM | 4 |
| 2003 | Adaptive Constraint Pushing in Frequent Pattern Mining
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi |
PKDD | 4 |
| 2003 | ExAnte: Anticipated Data Reduction in Constrained Pattern Mining
Francesco Bonchi, Fosca Giannotti, Alessio Mazzanti, Dino Pedreschi |
PKDD | 4 |
| 2001 | Web log data warehousing and mining for intelligent web caching
Francesco Bonchi, Fosca Giannotti, Cristian Gozzi, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi, Chiara Renso, Salvatore Ruggieri |
Data Knowl. Eng. | 6 |
| 2001 | Nondeterministic, Nonmonotonic Logic DatabasesabstractWe consider an extension of Datalog with mechanisms for temporal, nonmonotonic, and nondeterministic reasoning, which we refer to as Datalog++. We show, by means of examples, its flexibility in expressing queries concerning aggregates and data cube. Also, we show how iterated fixpoint and stable model semantics can be combined to the purpose of clarifying the semantics of Datalog++ programs and supporting their efficient execution. Finally, we provide a more concrete implementation strategy on which basis the design of optimization techniques tailored for Datalog++ is addressed. Fosca Giannotti, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2000 | Logic-Based Knowledge Discovery in Databases
Fosca Giannotti, Mirco Nanni, Dino Pedreschi |
EJC | 3 |
| 1999 | Using Data Mining Techniques in Fiscal Fraud Detection
Francesco Bonchi, Fosca Giannotti, Gianni Mainetto, Dino Pedreschi |
DaWaK | 4 |
| 1999 | A Classification-Based Methodology for Planning Audit Strategies in Fraud DetectionabstractPlanning adequate audit strategies is a key success factor in a posterion' fraud detection, e.g., in the fiscal and insurance domains, where audits are intended to detect tax evasion and fraudulent claims.A case study is presented in this paper, which illustrates how techniques based on classification can be used to support the task of planning audit strategies.The proposed approach is sensible to some conflicting issues of audit planning, e.g., the trade-off between maximizing audit benefits vs. minimizing audit costs.A methodological scenario, common to a whole class of similar applications, is then abstracted away from the case study.The limitations of available systems to support the identified overall KDD process, bring us to point out the key aspects of a logic-based database language, integrated with mining mechanisms, which is used to provide a uniform, highly expressive environment for the various steps in the construction of the considered case-study. Francesco Bonchi, Fosca Giannotti, Gianni Mainetto, Dino Pedreschi |
KDD | 4 |
| 1998 | Query Answering in Nondeterministic, Nonmonotonic Logic Databases
Fosca Giannotti, Giuseppe Manco 0001, Mirco Nanni, Dino Pedreschi |
FQAS | 4 |
| 1998 | Weakest Preconditions for Pure Prolog Programs
Dino Pedreschi, Salvatore Ruggieri |
Inf. Process. Lett. | 1 |