Erik Jenner

dblp:295/8670 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 24% Trustworthy machine learning · 16% Language models and text generation · 16%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 50% Algorithms and data structures · 50%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Diffusion On Syntax Trees For Program Synthesis · ICLR 2025
Program synthesis and code generation
neural program synthesis
0.912025
Diffusion On Syntax Trees For Program Synthesis · ICLR 2025
Natural language and speech › Language models and text generation › alignment
deceptive alignment
0.812024
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.812024
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.812024
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network · NeurIPS 2024
Machine learning › Reinforcement learning
partial observability
0.812024
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback · NeurIPS 2024
Natural language and speech › Language models and text generation › alignment
reward hacking
0.812024
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback · NeurIPS 2024
Machine learning › Reinforcement learning
reward learning
0.812024
STARC: A General Framework For Quantifying Differences Between Reward Functions · ICLR 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network · NeurIPS 2024
Machine learning › Deep learning architectures and training
equivariant neural network
0.612022
Steerable Partial Differential Operators for Equivariant Neural Networks · ICLR 2022
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Computer vision › Segmentation and scene understanding › interactive segmentation
seeded segmentation
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Graph algorithms and graph theory
graph cut
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Algorithms and data structures › randomized algorithms
karger's algorithm
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Graph algorithms and graph theory
minimum cut
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Algorithms and data structures
randomized algorithms
0.512021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021
Image and video processing
image segmentation
0.112021
Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice · ICCV 2021

Methods — techniques the papers use, named apart from their topics

syntax tree editing · 1.7search · 1.7diffusion model · 1.7random walker · 1.0harmonic energy minimization · 1.0contraction algorithm · 1.0pseudometrics · 0.8linear probe · 0.8boltzmann rationality · 0.8attention head analysis · 0.8STARC metrics · 0.8RLHF · 0.8equivariance · 0.6
YearPublicationVenuePosition
2025 Diffusion On Syntax Trees For Program Synthesis
abstract
Large language models generate code one token at a time. Their autoregressive generation process lacks the feedback of observing the program's output. Training LLMs to suggest edits directly can be challenging due to the scarcity of rich edit data. To address these problems, we propose neural diffusion models that operate on syntax trees of any context-free grammar. Similar to image diffusion models, our method also inverts "noise" applied to syntax trees. Rather than generating code sequentially, we iteratively edit it while preserving syntactic validity, which makes it easy to combine this neural model with search. We apply our approach to inverse graphics tasks, where our model learns to convert images into programs that produce those images. Combined with search, our model is able to write graphics programs, see the execution result, and debug them to meet the required specifications. We additionally show how our system can write graphics programs for hand-drawn sketches. Video results can be found at https://tree-diffusion.github.io.
Shreyas Kapur, Erik Jenner, Stuart Russell 0001
ICLR2
2024 STARC: A General Framework For Quantifying Differences Between Reward Functions
abstract
In order to solve a task using reinforcement learning, it is necessary to first formalise the goal of that task as a *reward function*. However, for many real-world tasks, it is very difficult to manually specify a reward function that never incentivises undesirable behaviour. As a result, it is increasingly popular to use *reward learning algorithms*, which attempt to *learn* a reward function from data. However, the theoretical foundations of reward learning are not yet well-developed. In particular, it is typically not known when a given reward learning algorithm with high probability will learn a reward function that is safe to optimise. This means that reward learning algorithms generally must be evaluated empirically, which is expensive, and that their failure modes are difficult to anticipate in advance. One of the roadblocks to deriving better theoretical guarantees is the lack of good methods for *quantifying* the difference between reward functions. In this paper we provide a solution to this problem, in the form of a class of pseudometrics on the space of all reward functions that we call STARC (STAndardised Reward Comparison) metrics. We show that STARC metrics induce both an upper and a lower bound on worst-case regret, which implies that our metrics are tight, and that any metric with the same properties must be bilipschitz equivalent to ours. Moreover, we also identify a number of issues with reward metrics proposed by earlier works. Finally, we evaluate our metrics empirically, to demonstrate their practical efficacy. STARC metrics can be used to make both theoretical and empirical analysis of reward learning algorithms both easier and more principled.
Joar Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, Alessandro Abate
ICLR4
2024 Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
abstract
Do neural networks learn to implement algorithms such as look-ahead or search "in the wild"? Or do they rely purely on collections of simple heuristics? We present evidence of *learned look-ahead* in the policy and value network of Leela Chess Zero, the currently strongest deep neural chess engine. We find that Leela internally represents future optimal moves and that these representations are crucial for its final output in certain board states. Concretely, we exploit the fact that Leela is a transformer that treats every chessboard square like a token in language models, and give three lines of evidence: (1) activations on certain squares of future moves are unusually important causally; (2) we find attention heads that move important information "forward and backward in time," e.g., from squares of future moves to squares of earlier ones; and (3) we train a simple probe that can predict the optimal move 2 turns ahead with 92% accuracy (in board states where Leela finds a single best line). These findings are clear evidence of learned look-ahead in neural networks and might be a step towards a better understanding of their capabilities.
Erik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen, Scott Emmons, Stuart Russell 0001
NeurIPS1
2024 When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
abstract
Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on partial observations? We formally define two failure cases: deceptive inflation and overjustification. Modeling the human as Boltzmann-rational w.r.t. a belief over trajectories, we prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance, overjustify their behavior to make an impression, or both. Under the new assumption that the human's partial observability is known and accounted for, we then analyze how much information the feedback process provides about the return function. We show that sometimes, the human's feedback determines the return function uniquely up to an additive constant, but in other realistic cases, there is irreducible ambiguity. We propose exploratory research directions to help tackle these challenges and experimentally validate both the theoretical concerns and potential mitigations, and caution against blindly applying RLHF in partially observable settings.
Leon Lang, Davis Foote, Stuart Russell 0001, Anca D. Dragan, Erik Jenner, Scott Emmons
NeurIPS5
2022 Steerable Partial Differential Operators for Equivariant Neural Networks
Erik Jenner, Maurice Weiler
ICLR1
2021 Extensions of Karger's Algorithm: Why They Fail in Theory and How They Are Useful in Practice
abstract
The minimum graph cut and minimum s-t-cut problems are important primitives in the modeling of combinatorial problems in computer science, including in computer vision and machine learning. Some of the most efficient algorithms for finding global minimum cuts are randomized algorithms based on Karger’s groundbreaking contraction algorithm. Here, we study whether Karger’s algorithm can be successfully generalized to other cut problems. We first prove that a wide class of natural generalizations of Karger’s algorithm cannot efficiently solve the s-t-mincut or the normalized cut problem to optimality. However, we then present a simple new algorithm for seeded segmentation / graph-based semi-supervised learning that is closely based on Karger’s original algorithm, showing that for these problems, extensions of Karger’s algorithm can be useful. The new algorithm has linear asymptotic runtime and yields a potential that can be interpreted as the posterior probability of a sample belonging to a given seed / class. We clarify its relation to the random walker algorithm / harmonic energy minimization in terms of distributions over spanning forests. On classical problems from seeded image segmentation and graph-based semi-supervised learning on image data, the method performs at least as well as the random walker / harmonic energy minimization / Gaussian processes.
Erik Jenner, Enrique Fita Sanmartin, Fred A. Hamprecht
ICCV1