VLDB 2026 Research / reviewers in the wild / expert
Jeremiah Z. Liu
dblp:199/2301 · also Jeremiah Zhe Liu
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-7410-4502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 13 |
| 2024 | Uncertainty Estimation and Out-of-Distribution Detection for Deep Learning-Based Image Reconstruction Using the Local LipschitzabstractAccurate image reconstruction is at the heart of diagnostics in medical imaging. Supervised deep learning-based approaches have been investigated for solving inverse problems including image reconstruction. However, these trained models encounter unseen data distributions that are widely shifted from training data during deployment. Therefore, it is essential to assess whether a given input falls within the training data distribution. Current uncertainty estimation approaches focus on providing an uncertainty map to radiologists, rather than assessing the training distribution fit. In this work, we propose a method based on the local Lipschitz metric to distinguish out-of-distribution images from in-distribution with an area under the curve of 99.94% for True Positive Rate versus False Positive Rate. We demonstrate a very strong relationship between the local Lipschitz value and mean absolute error (MAE), supported by a Spearman's rank correlation coefficient of 0.8475, to determine an uncertainty estimation threshold for optimal performance. Through the identification of false positives, we demonstrate the local Lipschitz and MAE relationship can guide data augmentation and reduce uncertainty. Our study was validated using the AUTOMAP architecture for sensor-to-image Magnetic Resonance Imaging (MRI) reconstruction. We demonstrate our approach outperforms baseline techniques of Monte-Carlo dropout and deep ensembles as well as the state-of-the-art Mean Variance Estimation network approach. We expand our application scope to MRI denoising and Computed Tomography sparse-to-full view reconstructions using UNET architectures. We show our approach is applicable to various architectures and applications, especially in medical imaging, where preserving diagnostic accuracy of reconstructed images remains paramount. Danyal Bhutto, Bo Zhu 0011, Jeremiah Z. Liu, Neha Koonjoo, Hongwei Li 0004, Bruce R. Rosen, Matthew S. Rosen |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Using Domain Knowledge to Guide Dialog Structure Induction via Neural Probabilistic Soft LogicabstractConnor Pryor, Quan Yuan, Jeremiah Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Connor Pryor, Quan Yuan 0001, Jeremiah Z. Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor |
ACL (1) | 3 |
| 2023 | On Compositional Uncertainty Quantification for Seq2seq Graph Parsing
Zi Lin, Du Phan, Panupong Pasupat, Jeremiah Z. Liu, Jingbo Shang |
ICLR | 4 |
| 2023 | Pushing the Accuracy-Group Robustness Frontier with Introspective Self-play
Jeremiah Z. Liu, Krishnamurthy Dvijotham, Jihyeon Lee, Quan Yuan 0001, Balaji Lakshminarayanan, Deepak Ramachandran |
ICLR | 1 |
| 2023 | A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image ModelsabstractContrastively trained text-image models have the remarkable ability to perform zero-shot classification, that is, classifying previously unseen images into categories that the model has never been explicitly trained to identify. However, these zero-shot classifiers need prompt engineering to achieve high accuracy. Prompt engineering typically requires hand-crafting a set of prompts for individual downstream tasks. In this work, we aim to automate this prompt engineering and improve zero-shot accuracy through prompt ensembling. In particular, we ask “Given a large pool of prompts, can we automatically score the prompts and ensemble those that are most suitable for a particular downstream dataset, without needing access to labeled validation data?". We demonstrate that this is possible. In doing so, we identify several pathologies in a naive prompt scoring method where the score can be easily overconfident due to biases in pre-training and test data, and we propose a novel prompt scoring method that corrects for the biases. Using our proposed scoring method to create a weighted average prompt ensemble, our method overall outperforms equal average ensemble, as well as hand-crafted prompts, on ImageNet, 4 of its variants, and 11 fine-grained classification benchmarks. while being fully automatic, optimization-free, and not requiring access to labeled validation data. James Urquhart Allingham, Jie Ren 0006, Michael Dusenberry, Xiuye Gu, Yin Cui, Dustin Tran, Jeremiah Z. Liu, Balaji Lakshminarayanan |
ICML | 7 |
| 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-AwarenessabstractAccurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning. Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan |
J. Mach. Learn. Res. | 1 |
| 2022 | Bayesian Nonparametric Model Averaging Using Scalable Gaussian Process RepresentationsabstractBayesian nonparametric methods provide a convenient and well-founded framework for constructing spatio-temporally evolving ensemble models. They not only provide a flexible way to weight different underlying models in an ensemble according to the strengths of each model on different regions of the input space, but also allow for quantification of the ensemble’s uncertainty. However, computational costs can pose a challenge to kernel-based ensemble methods when spatio-temporal resolution is high, such as environmental models. We propose a Bayesian Nonparametric Ensemble (BNE) method for spatio-temporal ensemble learning based on the Gaussian process. Our ensemble relies on theoretically well-founded linearized approximations to the Gaussian process to adaptively weight the underlying models in a way that is scalable and amenable to stochastic learning methods. We investigate both the random Fourier feature and Nystrom approaches in this setting. We demonstrate the practicality and usefulness of the approximate model on the problem of air pollution prediction across the contiguous USA over a 6 year period. John W. Paisley, Sebastian Rowland, Jeremiah Z. Liu, Brent A. Coull, Marianthi-Anna Kioumourtzoglou |
IEEE Big Data | 3 |
| 2022 | Neural-Symbolic Inference for Robust Autoregressive Graph Parsing via Compositional Uncertainty QuantificationabstractPre-trained seq2seq models excel at graph semantic parsing with rich annotated data, but generalize worse to out-of-distribution (OOD) and long-tail examples.In comparison, symbolic parsers under-perform on populationlevel metrics, but exhibit unique strength in OOD and tail generalization.In this work, we study compositionality-aware approach to neural-symbolic inference informed by model confidence, performing fine-grained neuralsymbolic reasoning at subgraph level (i.e., nodes and edges) and precisely targeting subgraph components with high uncertainty in the neural parser.As a result, the method combines the distinct strength of the neural and symbolic approaches in capturing different aspects of the graph prediction, leading to well-rounded generalization performance both across domains and in the tail.We empirically investigate the approach in the English Resource Grammar (ERG) parsing problem on a diverse suite of standard in-domain and seven OOD corpora.Our approach leads to 35.26% and 35.60% error reduction in aggregated SMATCH score over neural and symbolic approaches respectively, and 14% absolute accuracy gain in key tail linguistic categories over the neural model, outperforming prior state-of-art methods that do not account for compositionality or uncertainty. Zi Lin, Jeremiah Z. Liu, Jingbo Shang |
EMNLP | 2 |
| 2022 | Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical GuaranteesabstractWe develop a simple and unified framework for nonlinear variable importance estimation that incorporates uncertainty in the prediction function and is compatible with a wide range of machine learning models (e.g., tree ensembles, kernel methods, neural networks, etc). In particular, for a learned nonlinear model $f(\mathbf{x})$, we consider quantifying the importance of an input variable $\mathbf{x}^j$ using the integrated partial derivative $\Psi_j = \Vert \frac{\partial}{\partial \mathbf{x}^j} f(\mathbf{x})\Vert^2_{P_\mathcal{X}}$. We then (1) provide a principled approach for quantifying uncertainty in variable importance by deriving its posterior distribution, and (2) show that the approach is generalizable even to non-differentiable models such as tree ensembles. Rigorous Bayesian nonparametric theorems are derived to guarantee the posterior consistency and asymptotic uncertainty of the proposed approach. Extensive simulations and experiments on healthcare benchmark datasets confirm that the proposed algorithm outperforms existing classical and recent variable selection methods. Wenying Deng, Beau Coker, Rajarshi Mukherjee, Jeremiah Z. Liu, Brent A. Coull |
NeurIPS | 4 |
| 2021 | Variable Selection with Rigorous Uncertainty Quantification using Deep Bayesian Neural Networks: Posterior Concentration and Bernstein-von Mises PhenomenonabstractThis work develops a theoretical basis for the deep Bayesian neural network (BNN)’s ability in performing high-dimensional variable selection with rigorous uncertainty quantification. We develop new Bayesian non-parametric theorems to show that a properly configured deep BNN (1) learns the variable importance effectively in high dimensions, and its learning rate can sometimes “break” the curse of dimensionality. (2) BNN’s uncertainty quantification for variable importance is rigorous, in the sense that its 95% credible intervals for variable importance indeed covers the truth 95% of the time (i.e. the Bernstein-von Mises (BvM) phenomenon). The theoretical results suggest a simple variable selection algorithm based on the BNN’s credible intervals. Extensive simulation confirms the theoretical findings and shows that the proposed algorithm outperforms existing classic and neural-network-based variable selection methods, particularly in high dimensions. Jeremiah Z. Liu |
AISTATS | 1 |
| 2021 | Training independent subnetworks for robust prediction
Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Z. Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew M. Dai, Dustin Tran |
ICLR | 4 |
| 2020 | Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessabstractBayesian neural networks (BNN) and deep ensembles are principled approaches to estimate the predictive uncertainty of a deep learning model. However their practicality in real-time, industrial-scale applications are limited due to their heavy memory and inference cost. This motivates us to study principled approaches to high-quality uncertainty estimation that require only a single deep neural network (DNN). By formalizing the uncertainty quantification as a minimax learning problem, we first identify input distance awareness, i.e., the model’s ability to quantify the distance of a testing example from the training data in the input space, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs, by adding a weight normalization step during training and replacing the output layer. On a suite of vision and language understanding tasks and on modern architectures (Wide-ResNet and BERT), SNGP is competitive with deep ensembles in prediction, calibration and out-of-domain detection, and outperforms the other single-model approaches. Jeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, Balaji Lakshminarayanan |
NeurIPS | 1 |
| 2019 | Personalized treatment for type 2 diabetes using weighted k-nearest neighbors
Wenyu Song, Linying Zhang, Emily Gill, Jeremiah Z. Liu, Adam Wright |
AMIA | 4 |
| 2019 | Accurate Uncertainty Estimation and Decomposition in Ensemble LearningabstractEnsemble learning is a standard approach to building machine learning systems that capture complex phenomena in real-world data. An important aspect of these systems is the complete and valid quantification of model uncertainty. We introduce a Bayesian nonparametric ensemble (BNE) approach that augments an existing ensemble model to account for different sources of model uncertainty. BNE augments a model’s prediction and distribution functions using Bayesian nonparametric machinery. It has a theoretical guarantee in that it robustly estimates the uncertainty patterns in the data distribution, and can decompose its overall predictive uncertainty into distinct components that are due to different sources of noise and error. We show that our method achieves accurate uncertainty estimates under complex observational noise, and illustrate its real-world utility in terms of uncertainty decomposition and model bias detection for an ensemble in predict air pollution exposures in Eastern Massachusetts, USA. Jeremiah Z. Liu, John W. Paisley, Marianthi-Anna Kioumourtzoglou, Brent A. Coull |
NeurIPS | 1 |
| 2017 | Robust Hypothesis Test for Nonlinear Effect with Gaussian ProcessesabstractThis work constructs a hypothesis test for detecting whether an data-generating function $h: \real^p \rightarrow \real$ belongs to a specific reproducing kernel Hilbert space $\mathcal{H}_0$, where the structure of $\mathcal{H}_0$ is only partially known. Utilizing the theory of reproducing kernels, we reduce this hypothesis to a simple one-sided score test for a scalar parameter, develop a testing procedure that is robust against the mis-specification of kernel functions, and also propose an ensemble-based estimator for the null model to guarantee test performance in small samples. To demonstrate the utility of the proposed method, we apply our test to the problem of detecting nonlinear interaction between groups of continuous features. We evaluate the finite-sample performance of our test under different data-generating functions and estimation strategies for the null model. Our results revealed interesting connection between notions in machine learning (model underfit/overfit) and those in statistical inference (i.e. Type I error/power of hypothesis test), and also highlighted unexpected consequences of common model estimating strategies (e.g. estimating kernel hyperparameters using maximum likelihood estimation) on model inference. Jeremiah Z. Liu, Brent A. Coull |
NIPS | 1 |