VLDB 2026 Research / reviewers in the wild / expert
Piersilvio De Bartolomeis
dblp:313/2235
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 51% Reinforcement learning · 29% Trustworthy machine learning · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
1.7 | 2 | 2025 | Prediction-Powered Causal Inferences · NeurIPS 2025 Doubly robust identification of treatment effects from multiple environments · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning theory
convex reinforcement learning |
1.2 | 2 | 2023 | Convex Reinforcement Learning in Finite Trials · J. Mach. Learn. Res. 2023 Challenging Common Assumptions in Convex Reinforcement Learning · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
prediction-powered inference |
0.9 | 1 | 2025 | Prediction-Powered Causal Inferences · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning
statistical inference |
0.9 | 1 | 2025 | Efficient Randomized Experiments Using Foundation Models · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation |
0.9 | 1 | 2025 | Doubly robust identification of treatment effects from multiple environments · ICLR 2025 |
Computational social science and digital humanities
causal inference |
0.9 | 1 | 2025 | Prediction-Powered Causal Inferences · NeurIPS 2025 |
Computational social science and digital humanities › causal inference
treatment effect estimation |
0.9 | 1 | 2025 | Prediction-Powered Causal Inferences · NeurIPS 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.6 | 1 | 2022 | Challenging Common Assumptions in Convex Reinforcement Learning · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
foundation model |
0.3 | 1 | 2025 | Efficient Randomized Experiments Using Foundation Models · NeurIPS 2025 |
Machine learning › Learning theory
sample complexity |
0.2 | 1 | 2023 | Convex Reinforcement Learning in Finite Trials · J. Mach. Learn. Res. 2023 |
Mathematical optimization › continuous optimization
convex optimization |
0.2 | 1 | 2023 | Convex Reinforcement Learning in Finite Trials · J. Mach. Learn. Res. 2023 |
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
0.2 | 1 | 2022 | Challenging Common Assumptions in Convex Reinforcement Learning · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
foundational model fine-tuning · 1.7empirical risk minimization · 1.7conditional calibration · 1.7non-markovian policy · 1.3variance reduction · 0.9foundation model · 0.9asymptotic normality · 0.9convex analysis · 0.6approximation error analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Doubly robust identification of treatment effects from multiple environmentsabstractPractical and ethical constraints often require the use of observational data for causal inference, particularly in medicine and social sciences. Yet, observational datasets are prone to confounding, potentially compromising the validity of causal conclusions.
While it is possible to correct for biases if the underlying causal graph is known, this is rarely a feasible ask in practical scenarios. A common strategy is to adjust for all available covariates, yet this approach can yield biased treatment effect estimates, especially when post-treatment or unobserved variables are present.
We propose RAMEN, an algorithm that produces unbiased treatment effect estimates
by leveraging the heterogeneity of multiple data sources without the need to know or learn the underlying causal graph. Notably, RAMEN achieves *doubly robust identification*: it can identify the treatment effect whenever
the causal parents of the treatment or those of the outcome are observed, and the node whose parents are observed satisfies an invariance assumption. Empirical evaluations across synthetic, semi-synthetic, and real-world datasets show that our approach significantly outperforms existing methods. Piersilvio De Bartolomeis, Julia Kostin, Javier Abad, Fanny Yang |
ICLR | 1 |
| 2025 | Efficient Randomized Experiments Using Foundation ModelsabstractRandomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the other hand, in silico experiments leveraging foundation models offer a cost-effective alternative that can potentially attain higher statistical precision. However, the benefits of in silico experiments come with a significant risk: statistical inferences are not valid if the models fail to accurately predict experimental responses to interventions.
In this paper, we propose a novel approach that integrates the predictions from multiple foundation models with experimental data while preserving valid statistical inference. Our estimator is consistent and asymptotically normal, with asymptotic variance no larger than the standard estimator based on experimental data alone. Importantly, these statistical properties hold even when model predictions are arbitrarily biased. Empirical results across several randomized experiments show that our estimator offers substantial precision gains, equivalent to a reduction of up to 20\% in the sample size needed to match the same precision as the standard estimator based on experimental data alone. Piersilvio De Bartolomeis, Javier Abad, Guanbo Wang, Konstantin Donhauser, Raymond M. Duch, Fanny Yang, Issa J. Dahabreh |
NeurIPS | 1 |
| 2025 | Prediction-Powered Causal InferencesabstractIn many scientific experiments, the data annotating cost constraints the pace for testing novel hypotheses. Yet, modern machine learning pipelines offer a promising solution—provided their predictions yield correct conclusions. We focus on Prediction-Powered Causal Inferences (PPCI), i.e., estimating the treatment effect in an unlabeled target experiment, relying on training data with the same outcome annotated but potentially different treatment or effect modifiers. We first show that conditional calibration guarantees valid PPCI at population level. Then, we introduce a sufficient representation constraint transferring validity across experiments, which we propose to enforce in practice in Deconfounded Empirical Risk Minimization, our new model-agnostic training objective. We validate our method on synthetic and real-world scientific data, solving impossible problem instances for Empirical Risk Minimization even with standard invariance constraints. In particular, for the first time, we achieve valid causal inference on a scientific experiment with complex recording and no human annotations, fine-tuning a foundational model on our similar annotated experiment. Riccardo Cadei, Ilker Demirel, Piersilvio De Bartolomeis, Lukas Lindorfer, Sylvia Cremer, Cordelia Schmid, Francesco Locatello |
NeurIPS | 3 |
| 2024 | Hidden yet quantifiable: A lower bound for confounding strength using randomized trialsabstractIn the era of fast-paced precision medicine, observational studies play a major role in properly evaluating new treatments in clinical practice. Yet, unobserved confounding can significantly compromise causal conclusions drawn from non-randomized data. We propose a novel strategy that leverages randomized trials to quantify unobserved confounding. First, we design a statistical test to detect unobserved confounding above a certain strength. Then, we use the test to estimate an asymptotically valid lower bound on the unobserved confounding strength. We evaluate the power and validity of our statistical test on several synthetic and semi-synthetic datasets. Further, we show how our lower bound can correctly identify the absence and presence of unobserved confounding in a real-world example. Piersilvio De Bartolomeis, Javier Abad Martinez, Konstantin Donhauser, Fanny Yang |
AISTATS | 1 |
| 2024 | Detecting critical treatment effect bias in small subgroupsabstractRandomized trials are considered the gold standard for making informed decisions in medicine. However, they are often not representative of the patient population in clinical practice. Observational studies, on the other hand, cover a broader patient population but are prone to various biases. Thus, before using observational data for any downstream task, it is crucial to benchmark its treatment effect estimates against a randomized trial. We propose a novel strategy to benchmark observational studies on a subgroup level. First, we design a statistical test for the null hypothesis that the treatment effects – conditioned on a subset of relevant features – differ up to some tolerance value. Our test allows us to estimate an asymptotically valid lower bound on the maximum bias strength for any subgroup. We validate our lower bound in a real-world setting and show that it leads to conclusions that align with established medical knowledge. Piersilvio De Bartolomeis, Javier Abad, Konstantin Donhauser, Fanny Yang |
UAI | 1 |
| 2023 | Convex Reinforcement Learning in Finite TrialsabstractConvex Reinforcement Learning (RL) is a recently introduced framework that generalizes the standard RL objective to any convex (or concave) function of the state distribution induced by the agent's policy. This framework subsumes several applications of practical interest, such as pure exploration, imitation learning, and risk-averse RL, among others. However, the previous convex RL literature implicitly evaluates the agent's performance over infinite realizations (or trials), while most of the applications require excellent performance over a handful, or even just one, trials. To meet this practical demand, we formulate convex RL in finite trials, where the objective is any convex function of the empirical state distribution computed over a finite number of realizations. In this paper, we provide a comprehensive theoretical study of the setting, which includes an analysis of the importance of non-Markovian policies to achieve optimality, as well as a characterization of the computational and statistical complexity of the problem in various configurations. Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, Marcello Restelli |
J. Mach. Learn. Res. | 3 |
| 2022 | Challenging Common Assumptions in Convex Reinforcement LearningabstractThe classic Reinforcement Learning (RL) formulation concerns the maximization of a scalar reward function. More recently, convex RL has been introduced to extend the RL formulation to all the objectives that are convex functions of the state distribution induced by a policy. Notably, convex RL covers several relevant applications that do not fall into the scalar formulation, including imitation learning, risk-averse RL, and pure exploration. In classic RL, it is common to optimize an infinite trials objective, which accounts for the state distribution instead of the empirical state visitation frequencies, even though the actual number of trajectories is always finite in practice. This is theoretically sound since the infinite trials and finite trials objectives are equivalent and thus lead to the same optimal policy. In this paper, we show that this hidden assumption does not hold in convex RL. In particular, we prove that erroneously optimizing the infinite trials objective in place of the actual finite trials one, as it is usually done, can lead to a significant approximation error. Since the finite trials setting is the default in both simulated and real-world RL, we believe shedding light on this issue will lead to better approaches and methodologies for convex RL, impacting relevant research areas such as imitation learning, risk-averse RL, and pure exploration among others. Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, Marcello Restelli |
NeurIPS | 3 |