VLDB 2026 Research / reviewers in the wild / expert
Pedro Bizarro
dblp:b/PedroBizarro
· DBLP profile ↗
27ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-5281-1970ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 13 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SARSum: A Relevance and Comprehensiveness-Aware Abstractive Summarization Dataset for Suspicious Activity ReportsabstractExisting benchmarks that evaluate the ability of Large Language Models (LLMs) to summarize rely primarily on measuring a summary’s lexical similarity to a reference or on assessing whether its claims are factually consistent with the source document. These approaches fail to account for a summary’s comprehensiveness — the extent to which it captures important information, and relevance — the extent to which unessential elements are omitted. To bolster comprehensiveness and relevance evaluation in high-stakes domains, we propose SARSum, a dataset tailored to evaluate the summarization of notes taken by anti-money laundering (AML) analysts during the process of preparing a Suspicious Activity Report (SAR), a document filed by financial institutions to alert law enforcement about suspicious transactions or activities, where omission of key details can be extremely costly. To the best of our knowledge, SARSum is the first comprehensiveness and relevance-aware summarization dataset: each of the 2,000 sets of notes is accompanied by the key facts that must be retained in an ideal summary, along with 30 different summaries spanning six levels of information selection quality, created by either omitting key facts or introducing irrelevant information. These resources allow practitioners to evaluate not only a summary’s relevance and comprehensiveness, but also the ability of automatic metrics to assess them. These instances are generated using a variety of LLMs to rephrase templates approved by an AML expert, and we empirically verify that the resulting instances are highly abstractive and varied. While SARSum addresses a specific domain, the novel inclusion of key facts and a reference set with known levels of quality represents a crucial step with potential for broader application across high-stakes scenarios. These elements enable the use of techniques such as natural language inference and question-generation/question-answering to evaluate relevance and comprehensiveness. Jean V. Alves, Javier Liébana, Hugo M. Ferreira, Pedro Bizarro |
ECAI | 4 |
| 2025 | Evaluating Transfer Learning Methods on Real-World Data Streams: A Case Study in Financial Fraud Detection
Ricardo Ribeiro Pereira, Jacopo Bono, Hugo M. Ferreira, Pedro Ribeiro 0004, Carlos Soares, Pedro Bizarro |
ECML/PKDD (9) | 6 |
| 2024 | On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML MethodsabstractMost existing evaluations of explainable machine learning (ML) methods rely on simplifying assumptions or proxies that do not reflect real-world use cases; the handful of more robust evaluations on real-world settings have shortcomings in their design, generally leading to overestimation of methods' real-world utility. In this work, we seek to address this by conducting a study that evaluates post-hoc explainable ML methods in a setting consistent with the application context and provide a template for future evaluation studies. We modify and improve a prior study on e-commerce fraud detection by relaxing the original work's simplifying assumptions that departed from the deployment context. Our study finds no evidence for the utility of the tested explainable ML methods in the context, which is a drastically different conclusion from the earlier work. This highlights how seemingly trivial experimental design choices can yield misleading conclusions about method utility. In addition, our work carries lessons about the necessity of not only evaluating explainable ML methods using tasks, data, users, and metrics grounded in the intended application context but also developing methods tailored to specific applications, moving beyond general-purpose explainable ML methods. Kasun Amarasinghe, Kit T. Rodolfa, Sérgio M. Jesus, Valerie Chen, Vladimir Balayan, Pedro Saleiro, Pedro Bizarro, Ameet Talwalkar, Rayid Ghani |
AAAI | 7 |
| 2024 | Fair-OBNC: Correcting Label Noise for Fairer DatasetsabstractData used by automated decision-making systems, such as Machine Learning models, often reflects discriminatory behavior that occurred in the past. These biases in the training data are sometimes related to label noise, such as in COMPAS, where more African-American offenders are wrongly labeled as having a higher risk of recidivism when compared to their White counterparts. Models trained on such biased data may perpetuate or even aggravate the biases with respect to sensitive information, such as gender, race, or age. However, while multiple label noise correction approaches are available in the literature, these focus on model performance exclusively. In this work, we propose Fair-OBNC, a label noise correction method with fairness considerations, to produce training datasets with measurable demographic parity. The presented method adapts Ordering-Based Noise Correction, with an adjusted criterion of ordering, based both on the margin of error of an ensemble, and the potential increase in the observed demographic parity of the dataset. We evaluate Fair-OBNC against other different pre-processing techniques, under different scenarios of controlled label noise. Our results show that the proposed method is the overall better alternative within the pool of label correction methods, being capable of attaining better reconstructions of the original labels. Models trained in the corrected data have an increase, on average, of 150% in demographic parity, when compared to models trained in data with noisy labels, across the considered levels of label noise. Inês Oliveira e Silva, Sérgio M. Jesus, Hugo M. Ferreira, Pedro Saleiro, Inês Sousa, Pedro Bizarro, Carlos Soares |
ECAI | 6 |
| 2024 | Lessons Learned while Running ML Models in Harsh EnvironmentsabstractOnce a very large payment processor client told us: 'if we are down for 5 minutes, we open the evening news - so don't screw up'. Processing billions of dollars per day, many financial institutions, need to continuously fight organized crime in the form of transaction fraud, stolen cards, anti-money laundering, account opening fraud, impersonations scams, phishing, and many other exotic and ever changing attacks from organized crime groups worldwide. In fact, it is estimated that in 2023 the global losses in fraud scams and bank fraud reached 485.6 billion. However, in addition to having very good detection rates and very low false positive rates, financial institutions also need to maintain very high availability rates, very low latencies, very high throughputs, automatic fault tolerance, auto scale up and down, and more. In this talk we cover some lessons related to running ML models in harsh, mission critical environments. We describe data issues, scale issues, ethical issues, system issues, security issues, compliance issues, business and regulation issues, and some architectural tradeoffs and architectural evolutions. Pedro Bizarro |
KDD | 1 |
| 2024 | RIFF: Inducing Rules for Fraud Detection from Decision Trees
Lucas Martins, João Bravo, Ana Sofia Gomes, Carlos Soares, Pedro Bizarro |
RuleML+RR | 5 |
| 2024 | AutoVizuA11y: A Tool to Automate Screen Reader Accessibility in ChartsabstractAbstract Charts remain widely inaccessible on the web for users of assistive technologies like screen readers. This is, in part, due to data visualization experts still lacking the experience, knowledge, and time to consistently implement accessible charts. As a result, screen reader users are prevented from accessing information and are forced to resort to tabular alternatives (if available), limiting the insights that they can gather. We worked with both groups to develop AutoVizuA11y, a tool that automates the addition of accessible features to web‐based charts. It generates human‐like descriptions of the data using a large language model, calculates statistical insights from the data, and provides keyboard navigation between multiple charts and underlying elements. Fifteen screen reader users interacted with charts made accessible with AutoVizuA11y in a usability test, thirteen of which praised the tool for its intuitive design, short learning curve, and rich information. On average, they took 66 seconds to complete each of the eight analytical tasks presented and achieved a success rate of 89%. Through a SUS questionnaire, the participants gave AutoVizuA11y an “Excellent” score — 83.5/100 points. We also gathered feedback from two data visualization experts who used the tool. They praised the tool availability, ease of use and functionalities, and provided feedback to add AutoVizuA11y support for other technologies in the future. Diogo Duarte, Rita Costa, Pedro Bizarro, Carlos Duarte |
Comput. Graph. Forum | 3 |
| 2024 | Aequitas Flow: Streamlining Fair ML ExperimentationabstractAequitas Flow is an open-source framework and toolkit for end-to-end Fair Machine Learning (ML) experimentation, and benchmarking in Python. This package fills integration gaps that exist in other fair ML packages. In addition to the existing audit capabilities in Aequitas, the Aequitas Flow module provides a pipeline for fairness-aware model training, hyperparameter optimization, and evaluation, enabling easy-to-use and rapid experiments and analysis of results. Aimed at ML practitioners and researchers, the framework offers implementations of methods, datasets, metrics, and standard interfaces for these components to improve extensibility. By facilitating the development of fair ML practices, Aequitas Flow hopes to enhance the incorporation of fairness concepts in AI systems making AI systems more robust and fair. Sérgio M. Jesus, Pedro Saleiro, Inês Oliveira e Silva, Beatriz M. Jorge, Rita P. Ribeiro, João Gama 0001, Pedro Bizarro, Rayid Ghani |
J. Mach. Learn. Res. | 7 |
| 2023 | FairGBM: Gradient Boosting with Fairness Constraints
André F. Cruz, Catarina G. Belém, João Bravo, Pedro Saleiro, Pedro Bizarro |
ICLR | 5 |
| 2022 | Turning the Tables: Biased, Imbalanced, Dynamic Tabular Datasets for ML EvaluationabstractEvaluating new techniques on realistic datasets plays a crucial role in the development of ML research and its broader adoption by practitioners. In recent years, there has been a significant increase of publicly available unstructured data resources for computer vision and NLP tasks. However, tabular data — which is prevalent in many high-stakes domains — has been lagging behind. To bridge this gap, we present Bank Account Fraud (BAF), the first publicly available 1 privacy-preserving, large-scale, realistic suite of tabular datasets. The suite was generated by applying state-of-the-art tabular data generation techniques on an anonymized,real-world bank account opening fraud detection dataset. This setting carries a set of challenges that are commonplace in real-world applications, including temporal dynamics and significant class imbalance. Additionally, to allow practitioners to stress test both performance and fairness of ML methods, each dataset variant of BAF contains specific types of data bias. With this resource, we aim to provide the research community with a more realistic, complete, and robust test bed to evaluate novel and existing methods. Sérgio M. Jesus, José Pombal, Duarte M. Alves, André F. Cruz, Pedro Saleiro, Rita P. Ribeiro, João Gama 0001, Pedro Bizarro |
NeurIPS | 8 |
| 2021 | Promoting Fairness through Hyperparameter OptimizationabstractConsiderable research effort has been guided towards algorithmic fairness but real-world adoption of bias reduction techniques is still scarce. Existing methods are either metric-or model-specific, require access to sensitive attributes at inference time, or carry high development or deployment costs. This work explores the unfairness that emerges when optimizing ML models solely for predictive performance, and how to mitigate it with a simple and easily deployed intervention: fairness-aware hyperparameter optimization (HO). We propose and evaluate fairness-aware variants of three popular HO algorithms: Fair Random Search, Fair TPE, and Fairband. We validate our approach on a real-world bank account opening fraud case-study, as well as on three datasets from the fairness literature. Results show that, without extra training cost, it is feasible to find models with 111% mean fairness increase and just 6% decrease in performance when compared with fairness-blind HO.1 André F. Cruz, Pedro Saleiro, Catarina G. Belém, Carlos Soares, Pedro Bizarro |
ICDM | 5 |
| 2021 | TimeSHAP: Explaining Recurrent Models through Sequence PerturbationsabstractAlthough recurrent neural networks (RNNs) are state-of-the-art in numerous sequential decision-making tasks, there has been little research on explaining their predictions. In this work, we present TimeSHAP, a model-agnostic recurrent explainer that builds upon KernelSHAP and extends it to the sequential domain. TimeSHAP computes feature-, timestep-, and cell-level attributions. As sequences may be arbitrarily long, we further propose a pruning method that is shown to dramatically decrease both its computational cost and the variance of its attributions. We use TimeSHAP to explain the predictions of a real-world bank account takeover fraud detection RNN model, and draw key insights from its explanations: i) the model identifies important features and events aligned with what fraud analysts consider cues for account takeover; ii) positive predicted sequences can be pruned to only 10% of the original length, as older events have residual attribution values; iii) the most recent input event of positive predictions only contributes on average to 41% of the model's score; iv) notably high attribution to client's age, upheld on higher false positive rates for older clients. João Bento 0002, Pedro Saleiro, André F. Cruz, Mário A. T. Figueiredo, Pedro Bizarro |
KDD | 5 |
| 2021 | Railgun: managing large streaming windows under MAD requirementsabstractSome mission critical systems, e.g., fraud detection, require accurate, real-time metrics over long time sliding windows on applications that demand high throughput and low latencies. As these applications need to run "forever" and cope with large, spiky data loads, they further require to be run in a distributed setting. We are unaware of any streaming system that provides all those properties. Instead, existing systems take large simplifications, such as implementing sliding windows as a fixed set of overlapping windows, jeopardizing metric accuracy (violating regulatory rules) or latency (breaching service agreements). In this paper, we propose Railgun, a fault-tolerant, elastic, and distributed streaming system supporting real-time sliding windows for scenarios requiring high loads and millisecond-level latencies. We benchmarked an initial prototype of Railgun using real data, showing significant lower latency than Flink and low memory usage independent of window size. Further, we show that Railgun scales nearly linearly, respecting our msec-level latencies at high percentiles (<250ms @ 99.9%) even under a load of 1 million events per second. Ana Sofia Gomes, João Oliveirinha, Pedro Bizarro |
Proc. VLDB Endow. | 4 |
| 2020 | Interleaved Sequence RNNs for Fraud DetectionabstractPayment card fraud causes multibillion dollar losses for banks and merchants worldwide, often fueling complex criminal activities. To address this, many real-time fraud detection systems use tree-based models, demanding complex feature engineering systems to efficiently enrich transactions with historical data while complying with millisecond-level latencies. In this work, we do not require those expensive features by using recurrent neural networks and treating payments as an interleaved sequence, where the history of each card is an unbounded, irregular sub-sequence. We present a complete RNN framework to detect fraud in real-time, proposing an efficient ML pipeline from preprocessing to deployment. We show that these feature-free, multi-sequence RNNs outperform state-of-the-art models saving millions of dollars in fraud detection and using fewer computational resources. Bernardo Branco, Pedro Abreu, Ana Sofia Gomes, Mariana S. C. Almeida, João Tiago Ascensão, Pedro Bizarro |
KDD | 6 |
| 2017 | BreachRadar: Automatic Detection of Points-of-CompromiseabstractBank transaction fraud results in over $13B annual losses for banks, merchants, and card holders worldwide. Much of this fraud starts with a Point-of-Compromise (a data breach or a ”skimming” operation) where credit and debit card digital information is stolen, resold, and later used to perform fraud. We introduce this problem and present an automatic Points-of-Compromise (POC) detection procedure. BreachRadar is a distributed alternating algorithm that assigns a probability of being compromised to the different possible locations. We implement this method using Apache Spark and show its linear scalability in the number of machines and transactions. BreachRadar is applied to two datasets with billions of real transaction records and fraud labels where we provide multiple examples of real Points-of-Compromise we are able to detect. We further show the effectiveness of our method when injecting Points-of-Compromise in one of these datasets, simultaneously achieving over 90% precision and recall when only 10% of the cards have been victims of fraud. Miguel Araujo, Miguel Almeida, Jaime Ferreira, Luís Moura Silva, Pedro Bizarro |
SDM | 5 |
| 2013 | Towards a standard event processing benchmarkabstractThere has been an increasing interest both in academia and industry for systematic methods for evaluating the performance and scalability of event processing systems. A number of performance results have been disclosed over the last years, but there is still a lack of standardized benchmarks that allow an objective comparison of the different systems. In this paper, we present our work in progress: the BiCEP benchmark suite, a set of workloads, datasets and tools for evaluating different performance aspects of event processing platforms. In particular, we introduce 'Pairs', the first of the BiCEP benchmarks, aimed at assessing the ability of CEP engines in processing progressively larger volumes of events and simultaneous queries while providing quick answers. Marcelo R. N. Mendes, Pedro Bizarro |
ICPE | 2 |
| 2013 | Overcoming memory limitations in high-throughput event-based applicationsabstractThe last decade has witnessed the emergence of business critical applications processing streaming data for domains as diverse as credit card fraud detection, real-time recommendation systems, call-center monitoring, ad selection, network monitoring, and more. Most of those applications need to compute hundreds or thousands of metrics continuously while coping with very high event input rates. As a consequence, large amounts of state (i.e., moving windows) need to be maintained, very often exceeding the available memory resources. Nonetheless, current event processing platforms have little or no memory management capabilities, hanging or simply crashing when memory is exhausted. In this paper we report our experience in using secondary storage for solving the performance problems of memory-constrained event processing applications. For that, we propose SlideM, a novel buffer management algorithm that exploits the access pattern of sliding windows in order to efficiently handle memory shortages. The proposed algorithm was implemented in a real stream processing engine and validated through an extensive experimental performance evaluation. Results corroborate the efficacy of the approach: the system was able to sustain very high input rates (up to 300,000 events per second) for very large windows (about 30GB) while consuming small amounts of main memory (few kilobytes). Marcelo R. N. Mendes, Pedro Bizarro |
ICPE | 2 |
| 2013 | FINCoS: benchmark tools for event processing systemsabstractFINCoS is a set of benchmarking tools for load generation and performance measuring of event processing systems. It leverages the development of novel benchmarks by allowing researchers to create synthetic workloads, and enables users of the technology to evaluate candidate solutions using their own real datasets. An extensible set of adapters allows the framework to communicate with different CEP engines, and its architecture permits to distribute load generation across multiple nodes. In this paper we briefly review FINCoS, introducing its main characteristics and features, and discussing how it measures the performance of event processing platforms. Marcelo R. N. Mendes, Pedro Bizarro |
ICPE | 2 |
| 2011 | Deadline Queries: Leveraging the Cloud to Produce On-Time ResultsabstractMapReduce has become a widely used tool for computing complex tasks that process massive amounts of data in large clusters. Support for MapReduce tasks in cloud environments has been provided but it is left to users to make best guesses on the number of nodes needed for a task to complete within acceptable time. Moreover, the time a task will take to complete is often unknown beforehand. Previous research addressed this problem by establishing time constraints for query execution and, when needed, reduce the accuracy of queries using result approximation and/or sampling. However, in many situations reduced accuracy is not tolerable. In this paper we present Flood DQ, a MapReduce system that implements deadline queries -- queries that must finish before a deadline, never discarding data or reducing accuracy. Flood DQ produces timely, accurate results by adaptively increasing or decreasing computing power, at runtime, towards completing execution within the specified deadline. In Flood DQ, users only specify a deadline and the input data. The system monitors the progress of the task and extrapolates whether it will complete on time. If the task is deemed to complete after the specified time, the system requests more nodes from an IaaS Cloud provider, and adds them to the computation. On the other hand, if the task is deemed to complete before the specified time the system quiesces and releases surplus nodes, cutting costs to a minimum. This paper describes FloodDQ's architecture for supporting deadline queries and presents experimental results where the system always meets the deadline in spite of changes to the number of nodes, size of data or existence of perturbations. David Alves, Pedro Bizarro |
IEEE CLOUD | 2 |
| 2009 | A Training Process for Faculty Members in Collaborative Degree Programs: Design, Implementation and FeedbackabstractCollaborative degree programs in software engineering are becoming more common as universities try to expand their offering globally and leverage their knowledge and expertise.Faculty training program intended to help academics learn how to teach courses from collaborating institutions is a complicated undertaking considering the need to pass along course material, the dasiaspiritpsila of how the courses are taught and the quality standards to which they must adhere. Carnegie Mellon University developed a training process for teaching faculty members in its joint software engineering programs in India, Korea and Portugal. The process, its implementation and the feedback of using it with our overseas partners will be explored and described in detail. Gil Taran, Mário Zenha Rela, Pedro Bizarro |
CSEE&T | 4 |
| 2009 | Progressive Parametric Query OptimizationabstractCommercial applications usually rely on pre-compiled parameterized procedures to interact with a database. Unfortunately, executing a procedure with a set of parameters different from those used at compilation time may be arbitrarily sub-optimal. Parametric query optimization (PQO) attempts to solve this problem by exhaustively determining the optimal plans at each point of the parameter space at compile time. However, PQO is likely not cost-effective if the query is executed infrequently or if it is executed with values only within a subset of the parameter space. In this paper we propose instead to progressively explore the parameter space and build a parametric plan during several executions of the same query. We introduce algorithms that, as parametric plans are populated, are able to frequently bypass the optimizer but still execute optimal or near-optimal plans. Pedro Bizarro, Nicolas Bruno, David J. DeWitt |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Adaptive Query Processing in the Looking Glass
Shivnath Babu, Pedro Bizarro |
CIDR | 2 |
| 2005 | Proactive Re-optimizationabstractTraditional query optimizers rely on the accuracy of estimated statistics to choose good execution plans. This design often leads to suboptimal plan choices for complex queries, since errors in estimates for intermediate subexpressions grow exponentially in the presence of skewed and correlated data distributions. Reoptimization is a promising technique to cope with such mistakes. Current re-optimizers first use a traditional optimizer to pick a plan, and then react to estimation errors and resulting suboptimalities detected in the plan during execution. The effectiveness of this approach is limited because traditional optimizers choose plans unaware of issues affecting reoptimization. We address this problem using proactive reoptimization, a new approach that incorporates three techniques: i) the uncertainty in estimates of statistics is computed in the form of bounding boxes around these estimates, ii) these bounding boxes are used to pick plans that are robust to deviations of actual values from their estimates, and iii) accurate measurements of statistics are collected quickly and efficiently during query execution. We present an extensive evaluation of these techniques using a prototype proactive re-optimizer named Rio. In our experiments Rio outperforms current re-optimizers by up to a factor of three. Shivnath Babu, Pedro Bizarro, David J. DeWitt |
SIGMOD Conference | 2 |
| 2005 | Proactive re-optimization with RioabstractTraditional query optimizers rely on the accuracy of estimated statistics of intermediate subexpressions to choose good query execution plans. This design often leads to suboptimal plan choices for complex queries since errors in estimates grow exponentially in the presence of skewed and correlated data distributions. We propose to demonstrate the Rio prototype database system that uses proactive re-optimization to address the problems with traditional optimizers. Rio supports three new techniques:1. Intervals of uncertainty are considered around estimates of statistics during plan enumeration and costing2. These intervals are used to pick execution plans that are robust to deviations of actual values of statistics from estimated values, or to defer the choice of execution plan until the uncertainty in estimates can be resolved3. Statistics of intermediate subexpressions are collected quickly, accurately, and efficiently during query executionThese three features are fully functional in the current Rio prototype which is built using the Predator open-source DBMS [5]. In this proposal, we first describe the novel features of Rio, then we use an example query to illustrate the main aspects of our demonstration. Shivnath Babu, Pedro Bizarro, David J. DeWitt |
SIGMOD Conference | 2 |
| 2005 | Content-Based Routing: Different Plans for Different Data
Pedro Bizarro, Shivnath Babu, David J. DeWitt, Jennifer Widom |
VLDB | 1 |
| 2002 | Adding a Performance-Oriented Perspective to Data Warehouse Design
Pedro Bizarro, Henrique Madeira |
DaWaK | 1 |
| 1998 | JWarp: A Java Library for Parallel Discrete-Event SimulationsabstractJava is a very promising language for use in the simulation of physical models due to its object-oriented nature, portability, robustness and support for multithreading. This paper presents JWarp, a Java library for discrete-event parallel simulations. It is based on an optimistic model for synchronization of the simulation entities: the Time Warp mechanism. We introduce the main features of the library and discuss some of the implementation details. © 1998 John Wiley & Sons, Ltd. Pedro Bizarro, Luís Moura Silva, João Gabriel Silva |
Concurr. Pract. Exp. | 1 |