VLDB 2026 Research / reviewers in the wild / expert
Jerry Zhu
dblp:35/1775
· DBLP profile ↗
18ranked-venue papers
0as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perceptually Training Viewers against Misleading Data Visualizations with Informative feedback
Jihyun Rho, Shubham Kumar Bharti, Shiyun Cheng, Martina A. Rau, Jerry Zhu |
CogSci | 5 |
| 2025 | Collaborative Mean Estimation Among Heterogeneous Strategic Agents: Individual Rationality, Fairness, and Truthful ContributionabstractWe study a collaborative learning problem where $m$ agents aim to estimate a vector $\mu =(\mu_1,\ldots,\mu_d)\in \mathbb{R}^d$ by sampling from associated univariate normal distributions $(\mathcal{N}(\mu_k, \sigma^2))_{k\in[d]}$. Agent $i$ incurs a cost $c_{i,k}$ to sample from $\mathcal{N}(\mu_k, \sigma^2)$. Instead of working independently, agents can exchange data, collecting cheaper samples and sharing them in return for costly data, thereby reducing both costs and estimation error. We design a mechanism to facilitate such collaboration, while addressing two key challenges: ensuring individually rational (IR) and fair outcomes so all agents benefit, and preventing strategic behavior (e.g. non-collection, data fabrication) to avoid socially undesirable outcomes. We design a mechanism and an associated Nash equilibrium (NE) which minimizes the social penalty-sum of agents’ estimation errors and collection costs-while being IR for all agents. We achieve a $\mathcal{O}(\sqrt{m})$-approximation to the minimum social penalty in the worst case and an $\mathcal{O}(1)$-approximation under favorable conditions. Additionally, we establish three hardness results: no nontrivial mechanism guarantees (i) a dominant strategy equilibrium where agents report truthfully, (ii) is IR for every strategy profile of other agents, (iii) or avoids a worst-case $\Omega(\sqrt{m})$ price of stability in any NE. Finally, by integrating concepts from axiomatic bargaining, we demonstrate that our mechanism supports fairer outcomes than one which minimizes social penalty. Alex Clinton, Yiding Chen, Jerry Zhu, Kirthevasan Kandasamy |
ICML | 3 |
| 2024 | The Delusional Hedge Algorithm as a Model of Human Learning from Diverse Opinions
Yun-Shiuan Chuang, Jerry Zhu, Timothy T. Rogers |
CogSci | 2 |
| 2024 | Various Misleading Visual Features in Misleading Graphs: Do they truly deceive us?
Jihyun Rho, Martina A. Rau, Shubham Kumar Bharti, Rosanne Luu, Jeremy McMahan, Jerry Zhu |
CogSci | 7 |
| 2024 | Learning interactions to boost human creativity with bandits and GPT-4
Ara Vartanian, Xiaoxi Sun, Yun-Shiuan Chuang, Siddharth Suresh, Jerry Zhu, Timothy T. Rogers |
CogSci | 5 |
| 2024 | Minimally Modifying a Markov Game to Achieve Any Nash Equilibrium and ValueabstractWe study the game modification problem, where a benevolent game designer or a malevolent adversary modifies the reward function of a zero-sum Markov game so that a target deterministic or stochastic policy profile becomes the unique Markov perfect Nash equilibrium and has a value within a target range, in a way that minimizes the modification cost. We characterize the set of policy profiles that can be installed as the unique equilibrium of a game and establish sufficient and necessary conditions for successful installation. We propose an efficient algorithm that solves a convex optimization problem with linear constraints and then performs random perturbation to obtain a modification plan with a near-optimal cost. Young Wu, Jeremy McMahan, Yiding Chen, Yudong Chen 0001, Jerry Zhu, Qiaomin Xie |
ICML | 5 |
| 2023 | Non-parametric Outlier Synthesis
Leitian Tao, Xuefeng Du, Jerry Zhu, Yixuan Li 0001 |
ICLR | 3 |
| 2023 | Mechanism Design for Collaborative Normal Mean EstimationabstractWe study collaborative normal mean estimation, where $m$ strategic agents collect i.i.d samples from a normal distribution $\mathcal{N}(\mu, \sigma^2)$ at a cost. They all wish to estimate the mean $\mu$. By sharing data with each other, agents can obtain better estimates while keeping the cost of data collection small. To facilitate this collaboration, we wish to design mechanisms that encourage agents to collect a sufficient amount of data and share it truthfully, so that they are all better off than working alone. In naive mechanisms, such as simply pooling and sharing all the data, an individual agent might find it beneficial to under-collect and/or fabricate data, which can lead to poor social outcomes. We design a novel mechanism that overcomes these challenges via two key techniques: first, when sharing the others' data with an agent, the mechanism corrupts this dataset proportional to how much the data reported by the agent differs from the others; second, we design minimax optimal estimators for the corrupted dataset. Our mechanism, which is Nash incentive compatible and individually rational, achieves a social penalty (sum of all agents' estimation errors and data collection costs) that is at most a factor 2 of the global minimum. When applied to high dimensional (non-Gaussian) distributions with bounded variance, this mechanism retains these three properties, but with slightly weaker results. Finally, in two special cases where we restrict the strategy space of the agents, we design mechanisms that essentially achieve the global minimum. Yiding Chen, Jerry Zhu, Kirthevasan Kandasamy |
NeurIPS | 2 |
| 2023 | Dream the Impossible: Outlier Imagination with Diffusion ModelsabstractUtilizing auxiliary outlier datasets to regularize the machine learning model has demonstrated promise for out-of-distribution (OOD) detection and safe prediction. Due to the labor intensity in data collection and cleaning, automating outlier data generation has been a long-desired alternative. Despite the appeal, generating photo-realistic outliers in the high dimensional pixel space has been an open challenge for the field. To tackle the problem, this paper proposes a new framework Dream-OOD, which enables imagining photo-realistic outliers by way of diffusion models, provided with only the in-distribution (ID) data and classes. Specifically, Dream-OOD learns a text-conditioned latent space based on ID data, and then samples outliers in the low-likelihood region via the latent, which can be decoded into images by the diffusion model. Different from prior works [16, 95], Dream-OOD enables visualizing and understanding the imagined outliers, directly in the pixel space. We conduct comprehensive quantitative and qualitative studies to understand the efficacy of Dream-OOD, and show that training with the samples generated by Dream-OOD can significantly benefit OOD detection performance. Xuefeng Du, Yiyou Sun, Jerry Zhu, Yixuan Li 0001 |
NeurIPS | 3 |
| 2022 | Provable Defense against Backdoor Policies in Reinforcement LearningabstractWe propose a provable defense mechanism against backdoor policies in reinforcement learning under subspace trigger assumption. A backdoor policy is a security threat where an adversary publishes a seemingly well-behaved policy which in fact allows hidden triggers. During deployment, the adversary can modify observed states in a particular way to trigger unexpected actions and harm the agent. We assume the agent does not have the resources to re-train a good policy. Instead, our defense mechanism sanitizes the backdoor policy by projecting observed states to a `safe subspace', estimated from a small number of interactions with a clean (non-triggered) environment. Our sanitized policy achieves $\epsilon$ approximate optimality in the presence of triggers, provided the number of clean interactions is $O\left(\frac{D}{(1-\gamma)^4 \epsilon^2}\right)$ where $\gamma$ is the discounting factor and $D$ is the dimension of state space. Empirically, we show that our sanitization defense performs well on two Atari game environments. Shubham Kumar Bharti, Xuezhou Zhang, Adish Singla, Jerry Zhu |
NeurIPS | 4 |
| 2021 | Using Machine Teaching to Investigate Human Assumptions when Teaching Reinforcement Learners
Yun-Shiuan Chuang, Xuezhou Zhang, Yuzhe Ma, Mark K. Ho, Joseph L. Austerweil, Jerry Zhu |
CogSci | 6 |
| 2021 | Policy Gradient Bayesian Robust Optimization for Imitation LearningabstractThe difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the true reward function is. While most policy optimization approaches handle this uncertainty by optimizing for expected performance, many applications demand risk-averse behavior. We derive a novel policy gradient-style robust optimization approach, PG-BROIL, that optimizes a soft-robust objective that balances expected performance and risk. To the best of our knowledge, PG-BROIL is the first policy optimization algorithm robust to a distribution of reward hypotheses which can scale to continuous MDPs. Results suggest that PG-BROIL can produce a family of behaviors ranging from risk-neutral to risk-averse and outperforms state-of-the-art imitation learning algorithms when learning from ambiguous demonstrations by hedging against uncertainty, rather than seeking to uniquely identify the demonstrator’s reward function. Zaynah Javed, Daniel S. Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca D. Dragan, Kenneth Y. Goldberg |
ICML | 4 |
| 2020 | Enhancing generalization through an optimized sequential curriculum: Learning (to read) through machine teaching
Matthew Cooper Borkenhagen, Ayon Sen, Mark S. Seidenberg, Jerry Zhu, Christopher R. Cox |
CogSci | 4 |
| 2019 | A Unified Framework for Data Poisoning Attack to Graph-based Semi-supervised LearningabstractIn this paper, we proposed a general framework for data poisoning attacks to graph-based semi-supervised learning (G-SSL). In this framework, we first unify different tasks, goals and constraints into a single formula for data poisoning attack in G-SSL, then we propose two specialized algorithms to efficiently solve two important cases --- poisoning regression tasks under $\ell_2$-norm constraint and classification tasks under $\ell_0$-norm constraint. In the former case, we transform it into a non-convex trust region problem and show that our gradient-based algorithm with delicate initialization and update scheme finds the (globally) optimal perturbation. For the latter case, although it is an NP-hard integer programming problem, we propose a probabilistic solver that works much better than the classical greedy method. Lastly, we test our framework on real datasets and evaluate the robustness of G-SSL algorithms. For instance, on the MNIST binary classification problem (50000 training data with 50 labeled), flipping two labeled data is enough to make the model perform like random guess (around 50\% error). Xuanqing Liu, Si Si, Jerry Zhu, Yang Li 0058, Cho-Jui Hsieh |
NeurIPS | 3 |
| 2019 | Policy Poisoning in Batch Reinforcement Learning and ControlabstractWe study a security threat to batch reinforcement learning and control where the attacker aims to poison the learned policy. The victim is a reinforcement learner / controller which first estimates the dynamics and the rewards from a batch data set, and then solves for the optimal policy with respect to the estimates. The attacker can modify the data set slightly before learning happens, and wants to force the learner into learning a target policy chosen by the attacker. We present a unified framework for solving batch policy poisoning attacks, and instantiate the attack on two standard victims: tabular certainty equivalence learner in reinforcement learning and linear quadratic regulator in control. We show that both instantiation result in a convex optimization problem on which global optimality is guaranteed, and provide analysis on attack feasibility and attack cost. Experiments show the effectiveness of policy poisoning attacks. Yuzhe Ma, Xuezhou Zhang, Wen Sun 0002, Jerry Zhu |
NeurIPS | 4 |
| 2018 | For Teaching Perceptual Fluency, Machines Beat Human Experts
Ayon Sen, Purav Patel, Martina A. Rau, Blake Mason, Robert D. Nowak, Timothy T. Rogers, Jerry Zhu |
CogSci | 7 |
| 2014 | Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian ModelsabstractMuch research in Bayesian modeling has been done to elicit a prior distribution that incorporates domain knowledge. We present a novel and more direct approach by imposing First-Order Logic (FOL) rules on the posterior distribution. Our approach unifies FOL and Bayesian modeling under the regularized Bayesian framework. In addition, our approach automatically estimates the uncertainty of FOL rules when they are produced by humans, so that reliable rules are incorporated while unreliable ones are ignored. We apply our approach to latent topic modeling tasks and demonstrate that by combining FOL knowledge and Bayesian modeling, we both improve the task performance and discover more structured latent representations in unsupervised and supervised learning. Shike Mei, Jun Zhu 0001, Jerry Zhu |
ICML | 3 |
| 2006 | Implementation & evaluation of an IDS to safeguard OLSR integrity in MANETsabstractOLSR (Optimized Link State Routing), a table-driven proactive protocol, standardized (RFC3626) for Mobile Ad-hoc Networks (MANETs), only works properly if participating nodes cooperate in routing. Hence, if nodes inject invalid messages or withhold critical messages, routing protocol integrity may be compromised - the protocol will fail to consistently provide each node with the accurate MANET topology. Opportunities, therefore, exist to provide intrusion detection in OLSR based MANETs by detecting anomalies in OLSR semantics. In this paper we implement and analyze an Intrusion Detection System (IDS) in which each MANET node evaluates such non-conformances locally, and thereafter infers & detects possible attacks on the routing protocol. We quantify the effectiveness of the IDS in terms of false positive and false negative detection rates. Although our discussion and implementation is based on OLSR, the techniques can be applied to any link-state routing protocol. Danny Dhillon, Jerry Zhu, Tejinder S. Randhawa |
IWCMC | 2 |