VLDB 2026 Research / reviewers in the wild / expert
Chris Cundy
dblp:206/7233
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-4608-4110ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 33% Deep learning architectures and training · 22% Probabilistic and Bayesian machine learning · 19% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
imitation learning |
1.3 | 2 | 2024 | SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking · ICLR 2024 IQ-Learn: Inverse soft-Q Learning for Imitation · NeurIPS 2021 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Preference Learning with Lie Detectors can Induce Honesty or Evasion · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.9 | 1 | 2025 | Preference Learning with Lie Detectors can Induce Honesty or Evasion · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
autoregressive sequence generation |
0.8 | 1 | 2024 | SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking · ICLR 2024 |
Machine learning › Deep learning architectures and training
neural network expressivity |
0.7 | 1 | 2023 | Neural Networks and the Chomsky Hierarchy · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
bayesian causal discovery |
0.5 | 1 | 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal Discovery · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning |
0.5 | 1 | 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal Discovery · NeurIPS 2021 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.5 | 1 | 2021 | IQ-Learn: Inverse soft-Q Learning for Imitation · NeurIPS 2021 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.5 | 1 | 2021 | IQ-Learn: Inverse soft-Q Learning for Imitation · NeurIPS 2021 |
Machine learning › Reinforcement learning › value-based reinforcement learning
soft q-learning |
0.5 | 1 | 2021 | IQ-Learn: Inverse soft-Q Learning for Imitation · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.5 | 1 | 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal Discovery · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › recurrent neural network
linear recurrent neural network |
0.3 | 1 | 2018 | Parallelizing Linear Recurrent Neural Nets Over Sequence Length · ICLR (Poster) 2018 |
Machine learning › Efficient and distributed learning › distributed training › parallelization
parallel training |
0.3 | 1 | 2018 | Parallelizing Linear Recurrent Neural Nets Over Sequence Length · ICLR (Poster) 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Parallelizing Linear Recurrent Neural Nets Over Sequence Length · ICLR (Poster) 2018 |
Machine learning › Efficient and distributed learning › distributed training › model parallelism
sequence parallelism |
0.3 | 1 | 2018 | Parallelizing Linear Recurrent Neural Nets Over Sequence Length · ICLR (Poster) 2018 |
Machine learning › Reinforcement learning
preference learning |
0.3 | 1 | 2025 | Preference Learning with Lie Detectors can Induce Honesty or Evasion · NeurIPS 2025 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2024 | SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking · ICLR 2024 |
Automata and formal languages › formal grammars
chomsky hierarchy |
0.2 | 1 | 2023 | Neural Networks and the Chomsky Hierarchy · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
directed acyclic graph learning |
0.1 | 1 | 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal Discovery · NeurIPS 2021 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.1 | 1 | 2021 | IQ-Learn: Inverse soft-Q Learning for Imitation · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural equation models |
0.1 | 1 | 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal Discovery · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
lie detector · 0.9KL regularization · 0.9GRPO · 0.9DPO · 0.9imitation learning · 0.8divergence minimization · 0.8backtracking · 0.8variational inference · 0.5stochastic optimization · 0.5continuous relaxation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Preference Learning with Lie Detectors can Induce Honesty or EvasionabstractAs AI systems become more capable, deceptive behaviors can undermine evaluation and mislead users at deployment.
Recent work has shown that lie detectors can accurately classify deceptive behavior, but they are not typically used in the training pipeline due to concerns around contamination and objective hacking.
We examine these concerns by incorporating a lie detector into the labelling step of LLM post-training and evaluating whether the learned policy is genuinely more honest, or instead learns to fool the lie detector while remaining deceptive.
Using DolusChat, a novel 65k-example dataset with paired truthful/deceptive responses, we identify three key factors that determine the honesty of learned policies: amount of exploration during preference learning, lie detector accuracy, and KL regularization strength.
We find that preference learning with lie detectors and GRPO can lead to policies which evade lie detectors, with deception rates of over 85\%.
However, if the lie detector true positive rate (TPR) or KL regularization is sufficiently high, GRPO learns honest policies.
In contrast, off-policy algorithms (DPO) consistently lead to deception rates under 25\% for realistic TPRs.
Our results illustrate a more complex picture than previously assumed: depending on the context, lie-detector-enhanced training can be a powerful tool for scalable oversight, or a counterproductive method encouraging undetectable misalignment. Chris Cundy, Adam Gleave |
NeurIPS | 1 |
| 2024 | Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
Chris Cundy, Rishi Desai, Stefano Ermon |
AISTATS | 1 |
| 2024 | SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with BacktrackingabstractIn many domains, autoregressive models can attain high likelihood on the task of predicting the next observation. However, this maximum-likelihood (MLE) objective does not necessarily match a downstream use-case of autoregressively generating high-quality sequences. The MLE objective weights sequences proportionally to their frequency under the data distribution, with no guidance for the model's behaviour out of distribution (OOD): leading to compounding error during autoregressive generation. In order to address this compounding error problem, we formulate sequence generation as an imitation learning (IL) problem. This allows us to minimize a variety of divergences between the distribution of sequences generated by an autoregressive model and sequences from a dataset, including divergences with weight on OOD generated sequences. The IL framework also allows us to incorporate backtracking by introducing a backspace action into the generation process. This further mitigates the compounding error problem by allowing the model to revert a sampled token if it takes the sequence OOD. Our resulting method, SequenceMatch, can be implemented without adversarial training or major architectural changes. We identify the SequenceMatch-χ2 divergence as a more suitable training objective for autoregressive models which are used for generation. We show that empirically, SequenceMatch training leads to improvements over MLE on text generation with language models and arithmetic Chris Cundy, Stefano Ermon |
ICLR | 1 |
| 2023 | Neural Networks and the Chomsky Hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, Pedro A. Ortega |
ICLR | 7 |
| 2023 | Geo-knowledge-guided GPT models improve the extraction of location descriptions from disaster-related social media messagesabstractSocial media messages posted by people during natural disasters often contain important location descriptions, such as the locations of victims. Recent research has shown that many of these location descriptions go beyond simple place names, such as city names and street names, and are difficult to extract using typical named entity recognition (NER) tools. While advanced machine learning models could be trained, they require large labeled training datasets that can be time-consuming and labor-intensive to create. In this work, we propose a method that fuses geo-knowledge of location descriptions and a Generative Pre-trained Transformer (GPT) model, such as ChatGPT and GPT-4. The result is a geo-knowledge-guided GPT model that can accurately extract location descriptions from disaster-related social media messages. Also, only 22 training examples encoding geo-knowledge are used in our method. We conduct experiments to compare this method with nine alternative approaches on a dataset of tweets from Hurricane Harvey. Our method demonstrates an over 40% improvement over typically used NER approaches. The experiment results also show that geo-knowledge is indispensable for guiding the behavior of GPT models. The extracted location descriptions can help disaster responders reach victims more quickly and may even save lives. Yingjie Hu 0001, Gengchen Mai, Chris Cundy, Kristy Choi, Ni Lao, Gaurish Lakhanpal, Ryan Zhenqi Zhou, Kenneth Joseph |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | Towards a foundation model for geospatial artificial intelligence (vision paper)abstractLarge pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine tuning, few-shot, or even zero-shot learning. Despite their successes in language and vision tasks, we have yet to see an attempt to develop foundation models for geospatial artificial intelligence (GeoAI). In this work, we explore the promises and challenges for developing multimodal foundation models for GeoAI. We first show the advantages of this idea by testing the performance of existing Large pre-trained Language Models (LLMs) (e.g. GPT-2 and GPT-3) on two geospatial semantics tasks. Results indicate that these task-agnostic LLMs can outperform task-specific fully-supervised models on both tasks with 2--9% improvement in a few-shot learning setting. However, we also show the limitations of these existing foundation models given the multimodality nature of GeoAI, especially when dealing with geometries in conjunction with other modalities. So we discuss the possibility of a multimodal foundation model which can reason over various types of geospatial data through geospatial alignments. We conclude this paper by discussing the unique risks and challenges to develop such model for GeoAI. Gengchen Mai, Chris Cundy, Kristy Choi, Yingjie Hu 0001, Ni Lao, Stefano Ermon |
SIGSPATIAL/GIS | 2 |
| 2021 | BCD Nets: Scalable Variational Approaches for Bayesian Causal DiscoveryabstractA structural equation model (SEM) is an effective framework to reason over causal relationships represented via a directed acyclic graph (DAG).Recent advances have enabled effective maximum-likelihood point estimation of DAGs from observational data. However, a point estimate may not accurately capture the uncertainty in inferring the underlying graph in practical scenarios, wherein the true DAG is non-identifiable and/or the observed dataset is limited.We propose Bayesian Causal Discovery Nets (BCD Nets), a variational inference framework for estimating a distribution over DAGs characterizing a linear-Gaussian SEM.Developing a full Bayesian posterior over DAGs is challenging due to the the discrete and combinatorial nature of graphs.We analyse key design choices for scalable VI over DAGs, such as 1) the parametrization of DAGs via an expressive variational family, 2) a continuous relaxation that enables low-variance stochastic optimization, and 3) suitable priors over the latent variables.We provide a series of experiments on real and synthetic data showing that BCD Nets outperform maximum-likelihood methods on standard causal discovery metrics such as structural Hamming distance in low data regimes. Chris Cundy, Aditya Grover, Stefano Ermon |
NeurIPS | 1 |
| 2021 | IQ-Learn: Inverse soft-Q Learning for ImitationabstractIn many sequential decision-making problems (e.g., robotics control, game playing, sequential prediction), human or expert data is available containing useful information about the task. However, imitation learning (IL) from a small amount of expert data can be challenging in high-dimensional environments with complex dynamics. Behavioral cloning is a simple method that is widely used due to its simplicity of implementation and stable convergence but doesn't utilize any information involving the environment’s dynamics. Many existing methods that exploit dynamics information are difficult to train in practice due to an adversarial optimization process over reward and policy approximators or biased, high variance gradient estimators. We introduce a method for dynamics-aware IL which avoids adversarial training by learning a single Q-function, implicitly representing both reward and policy. On standard benchmarks, the implicitly learned rewards show a high positive correlation with the ground-truth rewards, illustrating our method can also be used for inverse reinforcement learning (IRL). Our method, Inverse soft-Q learning (IQ-Learn) obtains state-of-the-art results in offline and online imitation learning settings, significantly outperforming existing methods both in the number of required environment interactions and scalability in high-dimensional spaces, often by more than 3x. Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, Stefano Ermon |
NeurIPS | 3 |
| 2020 | Flexible Approximate Inference via Stratified Normalizing FlowsabstractA major obstacle to forming posterior distributions in machine learning is the difficulty of evaluating partition functions. Monte-Carlo approaches are unbiased, but can suffer from high variance. Variational methods are biased, but tend to have lower variance. We develop an approximate inference procedure that allows explicit control of the bias/variance tradeoff, interpolating between the sampling and the variational regime. We use a normalizing flow to map the integrand onto a uniform distribution. We then randomly sample regions from a partition of this uniform distribution and fit simpler, local variational approximations in the image of these regions through the flow. When a partition with only one region is used, we recover standard variational inference, and in the limit of an infinitely fine partition we recover Monte-Carlo sampling. We show experiments validating the effectiveness of our approach. Chris Cundy, Stefano Ermon |
UAI | 1 |
| 2018 | Parallelizing Linear Recurrent Neural Nets Over Sequence Length
Chris Cundy |
ICLR (Poster) | 2 |