EDBT 2026 Demo / reviewers in the wild / expert
Paulo Orenstein
dblp:227/3490
· DBLP profile ↗
8ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-0907-704XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 39% Probabilistic and Bayesian machine learning · 18% Learning theory · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Environmental and earth informatics · 74% Medical and health informatics · 26% | |
| Theoretical computer science
1 paper |
Approximation and online algorithms · 33% Mathematical optimization · 33% Algorithmic game theory and mechanism design · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Environmental and earth informatics › weather forecasting
subseasonal forecasting |
1.0 | 2 | 2023 | SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and Benchmarking · NeurIPS 2023 Improving Subseasonal Forecasting in the Western U.S. with Machine Learning · KDD 2019 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
1.0 | 2 | 2024 | Split Conformal Prediction and Non-Exchangeable Data · J. Mach. Learn. Res. 2024 AmnioML: Amniotic Fluid Segmentation and Volume Prediction with Uncertainty Quantification · AAAI 2023 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.0 | 2 | 2024 | Split Conformal Prediction and Non-Exchangeable Data · J. Mach. Learn. Res. 2024 AmnioML: Amniotic Fluid Segmentation and Volume Prediction with Uncertainty Quantification · AAAI 2023 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.7 | 1 | 2023 | AmnioML: Amniotic Fluid Segmentation and Volume Prediction with Uncertainty Quantification · AAAI 2023 |
Environmental and earth informatics › climate science
climate informatics |
0.7 | 1 | 2023 | SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and Benchmarking · NeurIPS 2023 |
Information retrieval › evaluation
benchmark dataset |
0.7 | 1 | 2023 | SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and Benchmarking · NeurIPS 2023 |
Mathematical optimization › online optimization
online convex optimization |
0.5 | 1 | 2021 | Online Learning with Optimism and Delay · ICML 2021 |
Approximation and online algorithms
online learning |
0.5 | 1 | 2021 | Online Learning with Optimism and Delay · ICML 2021 |
Algorithmic game theory and mechanism design
regret minimization |
0.5 | 1 | 2021 | Online Learning with Optimism and Delay · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.4 | 1 | 2020 | Scalable Approximate MCMC Algorithms for the Horseshoe Prior · J. Mach. Learn. Res. 2020 |
Machine learning › Learning theory
high-dimensional statistics |
0.4 | 1 | 2020 | Scalable Approximate MCMC Algorithms for the Horseshoe Prior · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.4 | 1 | 2020 | Scalable Approximate MCMC Algorithms for the Horseshoe Prior · J. Mach. Learn. Res. 2020 |
Machine learning › Learning theory › high-dimensional statistics
sparse estimation |
0.4 | 1 | 2020 | Scalable Approximate MCMC Algorithms for the Horseshoe Prior · J. Mach. Learn. Res. 2020 |
Machine learning › Time series and sequential data
ensemble forecasting |
0.4 | 1 | 2019 | Improving Subseasonal Forecasting in the Western U.S. with Machine Learning · KDD 2019 |
Environmental and earth informatics
climate prediction |
0.1 | 1 | 2021 | Online Learning with Optimism and Delay · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
deep learning · 3.3meteorological baseline · 2.0dynamical model · 2.0conformal prediction · 1.3optimistic online learning · 1.0meta-learning · 1.0machine learning · 0.8decoupling · 0.8concentration inequalities · 0.8matrix product approximation · 0.4approximate MCMC · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BlockBoost: Scalable and Efficient Blocking through BoostingabstractAs datasets grow larger, matching and merging entries from different databases has become a costly task in modern data pipelines. To avoid expensive comparisons between entries, blocking similar items is a popular preprocessing step. In this paper, we introduce BlockBoost, a novel boosting-based method that generates compact binary hash codes for database entries, through which blocking can be performed efficiently. The algorithm is fast and scalable, resulting in computational costs that are orders of magnitude lower than current benchmarks. Unlike existing alternatives, BlockBoost comes with associated feature importance measures for interpretability, and possesses strong theoretical guarantees, including lower bounds on critical performance metrics like recall and reduction ratio. Finally, we show that BlockBoost delivers great empirical results, outperforming state-of-the-art blocking benchmarks in terms of both performance metrics and computational cost. Thiago Ramos, Rodrigo Schuller, Alex Akira Okuno, Lucas Nissenbaum, Roberto Oliveira 0001, Paulo Orenstein |
AISTATS | 6 |
| 2024 | Split Conformal Prediction and Non-Exchangeable DataabstractSplit conformal prediction (CP) is arguably the most popular CP method for uncertainty quantification, enjoying both academic interest and widespread deployment. However, the original theoretical analysis of split CP makes the crucial assumption of data exchangeability, which hinders many real-world applications. In this paper, we present a novel theoretical framework based on concentration inequalities and decoupling properties of the data, proving that split CP remains valid for many non-exchangeable processes by adding a small coverage penalty. Through experiments with both real and synthetic data, we show that our theoretical results translate to good empirical performance under non-exchangeability, e.g., for time series and spatiotemporal data. Compared to recent conformal algorithms designed to counter specific exchangeability violations, we show that split CP is competitive in terms of coverage and interval size, with the benefit of being extremely simple and orders of magnitude faster than alternatives. Roberto Oliveira 0001, Paulo Orenstein, Thiago Ramos, João Vitor Romano |
J. Mach. Learn. Res. | 2 |
| 2023 | AmnioML: Amniotic Fluid Segmentation and Volume Prediction with Uncertainty QuantificationabstractAccurately predicting the volume of amniotic fluid is fundamental to assessing pregnancy risks, though the task usually requires many hours of laborious work by medical experts. In this paper, we present AmnioML, a machine learning solution that leverages deep learning and conformal prediction to output fast and accurate volume estimates and segmentation masks from fetal MRIs with Dice coefficient over 0.9. Also, we make available a novel, curated dataset for fetal MRIs with 853 exams and benchmark the performance of many recent deep learning architectures. In addition, we introduce a conformal prediction tool that yields narrow predictive intervals with theoretically guaranteed coverage, thus aiding doctors in detecting pregnancy risks and saving lives. A successful case study of AmnioML deployed in a medical setting is also reported. Real-world clinical benefits include up to 20x segmentation time reduction, with most segmentations deemed by doctors as not needing any further manual refinement. Furthermore, AmnioML's volume predictions were found to be highly accurate in practice, with mean absolute error below 56mL and tight predictive intervals, showcasing its impact in reducing pregnancy complications. Daniel Csillag, Lucas Monteiro Paes, Thiago Ramos, João Vitor Romano, Rodrigo Schuller, Roberto de Beauclair Seixas, Roberto Oliveira 0001, Paulo Orenstein |
AAAI | 8 |
| 2023 | SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and BenchmarkingabstractSubseasonal forecasting of the weather two to six weeks in advance is critical for resource allocation and climate adaptation but poses many challenges for the forecasting community. At this forecast horizon, physics-based dynamical models have limited skill, and the targets for prediction depend in a complex manner on both local weather variables and global climate variables. Recently, machine learning methods have shown promise in advancing the state of the art but only at the cost of complex data curation, integrating expert knowledge with aggregation across multiple relevant data sources, file formats, and temporal and spatial resolutions.To streamline this process and accelerate future development, we introduce SubseasonalClimateUSA, a curated dataset for training and benchmarking subseasonal forecasting models in the United States. We use this dataset to benchmark a diverse suite of models, including operational dynamical models, classical meteorological baselines, and ten state-of-the-art machine learning and deep learning-based methods from the literature. Overall, our benchmarks suggest simple and effective ways to extend the accuracy of current operational models. SubseasonalClimateUSA is regularly updated and accessible via the https://github.com/microsoft/subseasonal_data/ Python package. Soukayna Mouatadid, Paulo Orenstein, Genevieve Flaspohler, Miruna Oprescu, Judah Cohen, Franklyn Wang, Sean Knight, Maria Geogdzhayeva, Sam Levang, Ernest Fraenkel, Lester Mackey |
NeurIPS | 2 |
| 2022 | ExactBoost: Directly Boosting the Margin in Combinatorial and Non-decomposable MetricsabstractMany classification algorithms require the use of surrogate losses when the intended loss function is combinatorial or non-decomposable. This paper introduces a fast and exact stagewise optimization algorithm, dubbed ExactBoost, that boosts stumps to the actual loss function. By developing a novel extension of margin theory to the non-decomposable setting, it is possible to provably bound the generalization error of ExactBoost for many important metrics with different levels of non-decomposability. Through extensive examples, it is shown that such theoretical guarantees translate to competitive empirical performance. In particular, when used as an ensembler, ExactBoost is able to significantly outperform other surrogate-based and exact algorithms available. Daniel Csillag, Carolina Piazza, Thiago Ramos, João Vitor Romano, Roberto Oliveira 0001, Paulo Orenstein |
AISTATS | 6 |
| 2021 | Online Learning with Optimism and DelayabstractInspired by the demands of real-time climate and weather forecasting, we develop optimistic online learning algorithms that require no parameter tuning and have optimal regret guarantees under delayed feedback. Our algorithms—DORM, DORM+, and AdaHedgeD—arise from a novel reduction of delayed online learning to optimistic online learning that reveals how optimistic hints can mitigate the regret penalty caused by delay. We pair this delay-as-optimism perspective with a new analysis of optimistic learning that exposes its robustness to hinting errors and a new meta-algorithm for learning effective hinting strategies in the presence of delay. We conclude by benchmarking our algorithms on four subseasonal climate forecasting tasks, demonstrating low regret relative to state-of-the-art forecasting models. Genevieve Flaspohler, Francesco Orabona, Judah Cohen, Soukayna Mouatadid, Miruna Oprescu, Paulo Orenstein, Lester Mackey |
ICML | 6 |
| 2020 | Scalable Approximate MCMC Algorithms for the Horseshoe PriorabstractThe horseshoe prior is frequently employed in Bayesian analysis of high-dimensional models, and has been shown to achieve minimax optimal risk properties when the truth is sparse. While optimization-based algorithms for the extremely popular Lasso and elastic net procedures can scale to dimension in the hundreds of thousands, algorithms for the horseshoe that use Markov chain Monte Carlo (MCMC) for computation are limited to problems an order of magnitude smaller. This is due to high computational cost per step and growth of the variance of time-averaging estimators as a function of dimension. We propose two new MCMC algorithms for computation in these models that have significantly improved performance compared to existing alternatives. One of the algorithms also approximates an expensive matrix product to give orders of magnitude speedup in high-dimensional applications. We prove guarantees for the accuracy of the approximate algorithm, and show that gradually decreasing the approximation error as the chain extends results in an exact algorithm. The scalability of the algorithm is illustrated in simulations with problem size as large as $N=5,000$ observations and $p=50,000$ predictors, and an application to a genome-wide association study with $N=2,267$ and $p=98,385$. The empirical results also show that the new algorithm yields estimates with lower mean squared error, intervals with better coverage, and elucidates features of the posterior that were often missed by previous algorithms in high dimensions, including bimodality of posterior marginals indicating uncertainty about which covariates belong in the model. James E. Johndrow, Paulo Orenstein, Anirban Bhattacharya |
J. Mach. Learn. Res. | 2 |
| 2019 | Improving Subseasonal Forecasting in the Western U.S. with Machine LearningabstractWater managers in the western United States (U.S.) rely on longterm forecasts of temperature and precipitation to prepare for droughts and other wet weather extremes. To improve the accuracy of these longterm forecasts, the U.S. Bureau of Reclamation and the National Oceanic and Atmospheric Administration (NOAA) launched the Subseasonal Climate Forecast Rodeo, a year-long real-time forecasting challenge in which participants aimed to skillfully predict temperature and precipitation in the western U.S. two to four weeks and four to six weeks in advance. Here we present and evaluate our machine learning approach to the Rodeo and release our SubseasonalRodeo dataset, collected to train and evaluate our forecasting system. Jessica Hwang, Paulo Orenstein, Judah Cohen, Karl Pfeiffer, Lester Mackey |
KDD | 2 |