Paolo Missier

dblp:28/1062 · DBLP profile ↗
← Back
33ranked-venue papers in the field
7as first author
9since 2021 · last 2026
0000-0002-0978-2446ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 23 (6 first)Big Data, Cloud & Distributed Data Systems · 4Information Retrieval & Web Search · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 2Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 C-SHIFT: Efficient Cluster-based Model Fairness Control under Data Drift
abstract
Machine learning models deployed in real-world scenarios must contend with data shifts that occur over time, resulting in degraded model performance and potentially exacerbating fairness concerns. Considerable research has focused separately on maintaining either model accuracy or algorithmic fairness under distribution shifts, but not both. Separately, previous results are also available for detecting the harmful regions of the training set, where data shifts have the highest impact, with the goal of reducing the cost of retraining. In this work, we propose C-SHIFT (Cluster-based Selective Harmful Shift Identification for Fairness-aware Training), a framework for efficiently managing model performance and fairness together in general data shift scenarios. Starting from an initial model with a satisfactory accuracy-fairness tradeoff, C-SHIFT activates on batches of serving data, using a novel cluster-based algorithm to identify harmful data regions that may appear in some of the clusters. C-SHIFT restores fairness and accuracy by either partially retraining or fine-tuning the original model, achieving efficiency by focusing only on the harmful data within a portion of the clusters. C-SHIFT works well with multiple fairness adjustment methods, and is not sensitive to the specific type of data shift. Because clusters are also agnostic to the data type (they only require a suitable distance metric), in principle, the approach applies to multiple data modes and is not restricted to tabular data. We evaluate the approach using three real-world datasets as well as synthetic datasets specifically designed to simulate harmful data shift scenarios. Our results indicate that C-SHIFT can restore accuracy-fairness balance with quality comparable to a baseline global retraining approach, but using a small fraction of the training/serving data, with similar results for Logistic Regression (LR) as well as nonlinear Neural Network (NN) models.
Paolo Missier
EDBT3
2025 PROLIT: Supporting the Transparency of Data Preparation Pipelines through Narratives over Data Provenance
Pasquale Leonardo Lazzaro, Marialaura Lazzaro, Paolo Missier, Riccardo Torlone
EDBT3
2024 Stacked Generalization for Overlapping Asymmetric Datasets
Matthew McTeer, Paolo Missier
MEDI2
2024 Fair and Private Data Preprocessing through Microaggregation
abstract
Privacy protection for personal data and fairness in automated decisions are fundamental requirements for responsible Machine Learning. Both may be enforced through data preprocessing and share a common target: data should remain useful for a task, while becoming uninformative of the sensitive information. The intrinsic connection between privacy and fairness implies that modifications performed to guarantee one of these goals, may have an effect on the other, e.g., hiding a sensitive attribute from a classification algorithm might prevent a biased decision rule having such attribute as a criterion. This work resides at the intersection of algorithmic fairness and privacy. We show how the two goals are compatible, and may be simultaneously achieved, with a small loss in predictive performance. Our results are competitive with both state-of-the-art fairness correcting algorithms and hybrid privacy-fairness methods. Experiments were performed on three widely used benchmark datasets: Adult Income , COMPAS, and German Credit .
Vladimiro González-Zelaya, Julián Salas, David Megías 0001, Paolo Missier
ACM Trans. Knowl. Discov. Data4
2024 Supporting Better Insights of Data Science Pipelines with Fine-grained Provenance
abstract
Successful data-driven science requires complex data engineering pipelines to clean, transform, and alter data in preparation for machine learning, and robust results can only be achieved when each step in the pipeline can be justified, and its effect on the data explained. In this framework, we aim at providing data scientists with facilities to gain an in-depth understanding of how each step in the pipeline affects the data, from the raw input to training sets ready to be used for learning. Starting from an extensible set of data preparation operators commonly used within a data science setting, in this work we present a provenance management infrastructure for generating, storing, and querying very granular accounts of data transformations, at the level of individual elements within datasets whenever possible. Then, from the formal definition of a core set of data science preprocessing operators, we derive a provenance semantics embodied by a collection of templates expressed in PROV, a standard model for data provenance. Using those templates as a reference, our provenance generation algorithm generalises to any operator with observable input/output pairs. We provide a prototype implementation of an application-level provenance capture library to produce, in a semi-automatic way, complete provenance documents that account for the entire pipeline. We report on the ability of that reference implementation to capture provenance in real ML benchmark pipelines and over TCP-DI synthetic data. We finally show how the collected provenance can be used to answer a suite of provenance benchmark queries that underpin some common pipeline inspection questions, as expressed on the Data Science Stack Exchange.
Adriane Chapman, Luca Lauro, Paolo Missier, Riccardo Torlone
ACM Trans. Database Syst.3
2023 Interpretable and robust hospital readmission predictions from Electronic Health Records
abstract
Rates of Hospital Readmission (HR), defined as unplanned readmission within 30 days of discharge, have been increasing over the years, and impose an economic burden on healthcare services worldwide. Despite recent research into predicting HR, few models provide sufficient discriminative ability. Three main drawbacks can be identified in the published literature: (i) imbalance in the target classes (readmitted or not), (ii) not including demographic and lifestyle predictors, and (iii) lack of interpretability of the models. In this work, we address these three points by evaluating class balancing techniques, performing a feature selection process including demographic and lifestyle features, and adding interpretability through a combination of SHapley Additive exPlanations (SHAP) and Accumulated Local Effects (ALE) post hoc methods. Our best classifier for this binary outcome achieves a UAC of 0.849 using a selection of 1296 features, extracted from patients’ Electronic Health Records (EHRs) and from their sociodemographics profiles. Using SHAP and ALE, we have established the importance of age, the number of long-term conditions, and the duration of the first admission as top predictors. In addition, we show through an ablation study that demographic and lifestyle features provide even better predictive capabilities than other features, suggesting their relevance toward HR.
Hugo Calero-Díaz, Rebeen Ali Hamad, Christian Atallah, John Casement, Dexter Canoy, Nick J. Reynolds, Michael Barnes, Paolo Missier
IEEE Big Data8
2022 Tracking trajectories of multiple long-term conditions using dynamic patient-cluster associations
abstract
Momentum has been growing into research to better understand the dynamics of multiple long-term conditions – multimorbidity (MLTC-M), defined as the co-occurrence of two or more long-term or chronic conditions within an individual. Several research efforts make use of Electronic Health Records (EHR), which represent patients’ medical histories. These range from discovering patterns of multimorbidity, namely by clustering diseases based on their co-occurrence in EHRs, to using EHRs to predict the next disease or other specific outcomes. One problem with the former approach is that it discards important temporal information on the co-occurrence, while the latter requires "big" data volumes that are not always available from routinely collected EHRs, limiting the robustness of the resulting models.In this paper we take an intermediate approach, where initially we use about 143,000 EHRs from UK Biobank to perform time-independent clustering using topic modelling, and Latent Dirichlet Allocation specifically. We then propose a metric to measure how strongly a patient is "attracted" into any given cluster at any point through their medical history. By tracking how such gravitational pull changes over time, we may then be able to narrow the scope for potential interventions and preventative measures to specific clusters, without having to resort to full-fledged predictive modelling.In this preliminary work we show exemplars of these dynamic associations, which suggest that further exploration may lead to actionable insights into patients’ medical trajectories.
Ron Kremer, Syed Mohib Raza, Fabiola Eto, John Casement, Christian Atallah, Sarah Finer, Dennis Lendrem, Michael Barnes, Nick J. Reynolds, Paolo Missier
IEEE Big Data10
2022 DPDS: Assisting Data Science with Data Provenance
abstract
Successful data-driven science requires a complex combination of data engineering pipelines and data modelling techniques. Robust and defensible results can only be achieved when each step in the pipeline that is designed to clean, transform and alter data in preparation for data modelling can be justified, and its effect on the data explained. The DPDS toolkit presented in this paper is designed to make such justification and explanation process an integral part of data science practice, adding value while remaining as un-intrusive as possible to the analyst. Catering to the broad community of python/pandas data engineers, DPDS implements an observer pattern that is able to capture the fine-grained provenance associated with each individual element of a dataframe, across multiple transformation steps. The resulting provenance graph is stored in Neo4j and queried through a UI, with the goal of helping engineers and analysts to justify and explain their choice of data operations, from raw data to model training, by highlighting the details of the changes through each transformation.
Adriane Chapman, Luca Lauro, Paolo Missier, Riccardo Torlone
Proc. VLDB Endow.3
2021 Optimising Fairness Through Parametrised Data Sampling
Vladimiro González-Zelaya, Julián Salas, Dennis Prangle, Paolo Missier
EDBT4
2020 Capturing and querying fine-grained provenance of preprocessing pipelines in data science
abstract
Data processing pipelines that are designed to clean, transform and alter data in preparation for learning predictive models, have an impact on those models' accuracy and performance, as well on other properties, such as model fairness. It is therefore important to provide developers with the means to gain an in-depth understanding of how the pipeline steps affect the data, from the raw input to training sets ready to be used for learning. While other efforts track creation and changes of pipelines of relational operators, in this work we analyze the typical operations of data preparation within a machine learning process, and provide infrastructure for generating very granular provenance records from it, at the level of individual elements within a dataset. Our contributions include: (i) the formal definition of a core set of preprocessing operators, and the definition of provenance patterns for each of them, and (ii) a prototype implementation of an application-level provenance capture library that works alongside Python. We report on provenance processing and storage overhead and scalability experiments, carried out over both real ML benchmark pipelines and over TCP-DI, and show how the resulting provenance can be used to answer a suite of provenance benchmark queries that underpin some of the developers' debugging questions, as expressed on the Data Science Stack Exchange.
Adriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo Torlone
Proc. VLDB Endow.2
2019 A Customisable Pipeline for Continuously Harvesting Socially-Minded Twitter Users
Flavio Primo, Paolo Missier, Alexander B. Romanovsky, Mickael Figueredo, Nélio Cacho
ICWE2
2018 Loom: Query-aware Partitioning of Online Graphs
Hugo Firth, Paolo Missier, Jack Aiston
EDBT2
2018 VazaDengue: An information system for preventing and combating mosquito-borne diseases with social networks
Leonardo da Silva Sousa, Rafael Maiani de Mello, Diego Cedrim, Alessandro F. Garcia 0001, Paolo Missier, Anderson G. Uchôa, Anderson Oliveira, Alexander B. Romanovsky
Inf. Syst.5
2017 Why-Diff: Explaining differences amongst similar workflow runs by exploiting scientific metadata
abstract
Majority of workflows executed nowadays need to process a massive amount of data. Re-execution of such dataintensive scientific workflows often results in different outputs. Scientific research progresses when discoveries are reproduced and verified. However, simply re-enacting a scientific computation, such as a workflow, does not guarantee the correctness of results because of unintentional changes that may have interfered with the re-enactment process. We investigate the hypothesis that the metadata of a workflow execution can be used to explain why the experimenter observes different results (cause analysis). Similarly, Scientific metadata can be used to determine the impact of intentional variations that the experimenter may have injected into a new version of the workflow. We explore these two complementary cases using a simple algorithm for traversing two metadata traces in lock-step mode, which we illustrate through two human genomics data analysis workflows.
Priyaa Thavasimani, Jacek Cala, Paolo Missier
IEEE BigData3
2017 Recruiting from the Network: Discovering Twitter Users Who Can Help Combat Zika Epidemics
Paolo Missier, Callum McClean, Jonathan Carlton, Diego Cedrim, Leonardo da Silva Sousa, Alessandro F. Garcia 0001, Alexandre Plastino 0001, Alexander B. Romanovsky
ICWE1
2017 TAPER: query-aware, partition-enhancement for large, heterogenous graphs
abstract
Graph partitioning has long been seen as a viable approach to addressing Graph DBMS scalability. A partitioning, however, may introduce extra query processing latency unless it is sensitive to a specific query workload, and optimised to minimise inter-partition traversals for that workload. Additionally, it should also be possible to incrementally adjust the partitioning in reaction to changes in the graph topology, the query workload, or both. Because of their complexity, current partitioning algorithms fall short of one or both of these requirements, as they are designed for offline use and as one-off operations. The TAPER system aims to address both requirements, whilst leveraging existing partitioning algorithms. TAPER takes any given initial partitioning as a starting point, and iteratively adjusts it by swapping chosen vertices across partitions, heuristically reducing the probability of inter-partition traversals for a given path queries workload. Iterations are inexpensive thanks to time and space optimisations in the underlying support data structures. We evaluate TAPER on two different large test graphs and over realistic query workloads. Our results indicate that, given a hash-based partitioning, TAPER reduces the number of inter-partition traversals by $$\sim $$ 80%; given an unweighted Metis partitioning, by $$\sim $$ 30%. These reductions are achieved within eight iterations and with the additional advantage of being workload-aware and usable online.
Hugo Firth, Paolo Missier
Distributed Parallel Databases2
2016 Facilitating reproducible research by investigating computational metadata
abstract
Computational workflows consist of a series of steps in which data is generated, manipulated, analysed and transformed. Researchers use tools and techniques to capture the provenance associated with the data to aid reproducibility. The metadata collected not only helps in reproducing the computation but also aids in comparing the original and reproduced computations. In this paper, we present an approach, “Why-Diff”, to analyse the difference between two related computations by changing the artifacts and how the existing tools “YesWorkflow” and “NoWorkflow” record the changed artifacts.
Priyaa Thavasimani, Paolo Missier
IEEE BigData2
2014 DistillFlow: removing redundancy in scientific workflows
abstract
Scientific workflows management systems are increasingly used by scientists to specify complex data processing pipelines. Workflows are represented using a graph structure, where nodes represent tasks and links represent the dataflow. However, the complexity of workflow structures is increasing over time, reducing the rate of scientific workflows reuse. Here, we introduce DistillFlow, a tool based on effective methods for workflow design, with a focus on the Taverna model. DistillFlow is able to detect "anti-patterns" in the structure of workflows (idiomatic forms that lead to over-complicated design) and replace them with different patterns to reduce the workflow's overall structural complexity. Rewriting workflows in this way is beneficial both in terms of user experience and workflow maintenance.
Jiuqiang Chen, Sarah Cohen Boulakia, Christine Froidevaux, Carole A. Goble, Paolo Missier, Alan R. Williams
SSDBM5
2013 Reference Architectures to Measure Data Completeness across Integrated Databases
Nurul A. Emran, Suzanne M. Embury, Paolo Missier, Norashikin Ahmad
ACIIDS (1)3
2013 Measuring Data Completeness for Microbial Genomics Database
Nurul A. Emran, Suzanne M. Embury, Paolo Missier, Mohd Noor Mat Isa, Azah Kamilah Muda
ACIIDS (1)3
2013 The W3C PROV family of specifications for modelling provenance metadata
abstract
Provenance, a form of structured metadata designed to record the origin or source of information, can be instrumental in deciding whether information is to be trusted, how it can be integrated with other diverse information sources, and how to establish attribution of information to authors throughout its history. The PROV set of specifications, produced by the World Wide Web Consortium (W3C), is designed to promote the publication of provenance information on the Web, and offers a basis for interoperability across diverse provenance management systems. The PROV provenance model is deliberately generic and domain-agnostic, but extension mechanisms are available and can be exploited for modelling specific domains. This tutorial provides an account of these specifications. Starting from intuitive and informal examples that present idiomatic provenance patterns, it progressively introduces the relational model of provenance along with the constraints model for validation of provenance documents, and concludes with example applications that show the extension points in use.
Paolo Missier, Khalid Belhajjame, James Cheney
EDBT1
2010 Fine-grained and efficient lineage querying of collection-based workflow provenance
abstract
The management and querying of workflow provenance data underpins a collection of activities, including the analysis of workflow results, and the debugging of workflows or services. Such activities require efficient evaluation of lineage queries over potentially complex and voluminous provenance logs. Näive implementations of lineage queries navigate provenance logs by joining tables that represent the flow of data between connected processors invoked from workflows. In this paper we provide an approach to provenance querying that: (i) avoids joins over provenance logs by using information about the workflow definition to inform the construction of queries that directly target relevant lineage results; (ii) provides fine grained provenance querying, even for workflows that create and consume collections; and (iii) scales effectively to address complex workflows, workflows with large intermediate data sets, and queries over multiple workflows.
Paolo Missier, Norman W. Paton, Khalid Belhajjame
EDBT1
2010 Taverna, Reloaded
Paolo Missier, Stian Soiland-Reyes, Stuart Owen, Wei Tan 0001, Aleksandra Nenadic, Ian Dunlop, Alan R. Williams, Thomas M. Oinn, Carole A. Goble
SSDBM1
2009 Time-completeness trade-offs in record linkage using adaptive query processing
abstract
Applications that involve data integration among multiple sources often require a preliminary step of data reconciliation in order to ensure that tuples match correctly across the sources. In dynamic settings such as data mashups, however, traditional offline data reconciliation techniques that require prior availability of the data may not be applicable. The alternative, performing similarity joins at query time, is computationally expensive, while ignoring the mismatch problem altogether leads to an incomplete integration. In this paper we make the assumption that, in some dynamic integration scenarios, users may agree to trade the completeness of a join result in return for a faster computation. We explore the consequences of this assumption by proposing a novel, hybrid join algorithm that involves a combination of exact and approximate join operators, managed using adaptive query processing techniques. The algorithm is optimistic: it can switch between physical join operators multiple times throughout query processing, but it only resorts to approximate join operators when there is statistical evidence that result completeness is compromised. Our experiments show that sensible savings in join execution time can be achieved in practice, at the expense of a modest reduction in result completeness.
Roald Lengu, Paolo Missier, Alvaro A. A. Fernandes, Giovanna Guerrini, Marco Mesiti
EDBT2
2007 Managing information quality in e-science: the qurator workbench
abstract
Data-intensive e-science applications often rely on third-party data found in public repositories, whose quality is largely unknown. Although scientists are aware that this uncertainty may lead to incorrect scientific conclusions, in the absence of a quantitative characterization of data quality properties they find it difficult to formulate precise data acceptability criteria. We present an Information Quality management workbench, called Qurator, that supports data experts in the specification of personal quality models, and lets them derive effective criteria for data acceptability. The demo of our working prototype will illustrate our approach on a real e-science workflow for a bioinformatics application.
Paolo Missier, Suzanne M. Embury, Robert Mark Greenwood, Alun D. Preece, Binling Jin
SIGMOD Conference1
2006 Managing Information Quality in e-Science Using Semantic Web Technology
Alun D. Preece, Binling Jin, Edoardo Pignotti, Paolo Missier, Suzanne M. Embury, David Stead, Al Brown
ESWC4
2006 Quality Views: Capturing and Exploiting the User Perspective on Data Quality
Paolo Missier, Suzanne M. Embury, Robert Mark Greenwood, Alun D. Preece, Binling Jin
VLDB1
2006 An overview of S-OGSA: A Reference Semantic Grid Architecture
Óscar Corcho, Pinar Alper, Ioannis Kotsiopoulos, Paolo Missier, Sean Bechhofer, Carole A. Goble
J. Web Semant.4
2005 Clustering Web pages based on their structure
Valter Crescenzi, Paolo Merialdo, Paolo Missier
Data Knowl. Eng.3
2004 Ontology-Based Question Answering in a Federation of University Sites: The MOSES Case Study
Paolo Atzeni, Roberto Basili 0001, Dorte Haltrup Hansen, Paolo Missier, Patrizia Paggio, Maria Teresa Pazienza, Fabio Massimo Zanzotto
NLDB4
2004 An Automatic Data Grabber for Large Web Sites
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo, Paolo Missier
VLDB4
2003 Improving Data Quality in Practice: A Case Study in the Italian Public Administration
Paolo Missier, Gail Lalk, Vassilios S. Verykios, F. Grillo, T. Lorusso, Paola Angeletti
Distributed Parallel Databases1
2000 Telcordia's Database Reconciliation and Data Quality Analysis Tool
Francesco Caruso, Munir Cochinwala, Uma Ganapathy, Gail Lalk, Paolo Missier
VLDB5