VLDB 2026 Research / reviewers in the wild / expert
Cem Anil
dblp:218/6350
· DBLP profile ↗
9ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 46% Trustworthy machine learning · 34% Deep learning architectures and training · 11% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.8 | 2 | 2019 | Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks · NeurIPS 2019 Sorting Out Lipschitz Function Approximation · ICML 2019 |
Natural language and speech › Language models and text generation
large language model safety |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
lipschitz-constrained networks |
0.8 | 2 | 2019 | Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks · NeurIPS 2019 Sorting Out Lipschitz Function Approximation · ICML 2019 |
Security and privacy of machine learning
adversarial attack |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
equilibrium models |
0.6 | 1 | 2022 | Path Independent Equilibrium Models Can Better Exploit Test-Time Computation · NeurIPS 2022 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | Exploring Length Generalization in Large Language Models · NeurIPS 2022 |
Natural language and speech › Language models and text generation › compositional generalization
length generalization |
0.6 | 1 | 2022 | Exploring Length Generalization in Large Language Models · NeurIPS 2022 |
Natural language and speech › Language models and text generation › mathematical reasoning
mathematical problem solving |
0.6 | 1 | 2022 | Solving Quantitative Reasoning Problems with Language Models · NeurIPS 2022 |
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning |
0.6 | 1 | 2022 | Solving Quantitative Reasoning Problems with Language Models · NeurIPS 2022 |
Natural language and speech › Language models and text generation › large language model inference
test-time compute |
0.6 | 1 | 2022 | Path Independent Equilibrium Models Can Better Exploit Test-Time Computation · NeurIPS 2022 |
Algorithmic game theory and mechanism design
social choice |
0.5 | 1 | 2021 | Learning to Elect · NeurIPS 2021 |
Algorithmic game theory and mechanism design › social choice › computational social choice
voting rules |
0.5 | 1 | 2021 | Learning to Elect · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
certified adversarial robustness |
0.4 | 1 | 2019 | Sorting Out Lipschitz Function Approximation · ICML 2019 |
Machine learning › Trustworthy machine learning › robustness
certified robustness |
0.4 | 1 | 2019 | Preventing Gradient Attenuation in Lipschitz Constrained Convolutional Networks · NeurIPS 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) Pipeline for Musical Timbre Transfer · ICLR (Poster) 2019 |
Security and privacy of machine learning
prompt injection |
0.2 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Machine learning › Learning theory
generalization |
0.2 | 1 | 2022 | Path Independent Equilibrium Models Can Better Exploit Test-Time Computation · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.2 | 1 | 2022 | Exploring Length Generalization in Large Language Models · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
pre-training |
0.2 | 1 | 2022 | Solving Quantitative Reasoning Problems with Language Models · NeurIPS 2022 |
Knowledge, reasoning and agents › Multi-agent systems › social choice
computational social choice |
0.1 | 1 | 2021 | Learning to Elect · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
long-context prompting · 1.5in-context learning · 1.5fine-tuning · 1.3chain-of-thought · 0.8scratchpad prompting · 0.6recurrent network · 0.6large language model pre-training · 0.6deep equilibrium model · 0.6chain-of-thought prompting · 0.6set transformer · 0.5graph neural network · 0.5deepsets · 0.5deep sets · 0.5wavenet · 0.4constant-q transform · 0.4CycleGAN · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Many-shot JailbreakingabstractWe investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This attack is newly feasible with the larger context windows recently deployed by language model providers like Google DeepMind, OpenAI and Anthropic. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs. Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomek Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger B. Grosse, David Duvenaud |
NeurIPS | 1 |
| 2024 | Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training DataabstractOne way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit information can remain scattered across various training documents. Could an LLM infer the censored knowledge by piecing together these implicit hints? As a step towards answering this question, we study inductive out-of-context reasoning (OOCR), a type of generalization in which LLMs infer latent information from evidence distributed across training documents and apply it to downstream tasks without in-context learning. Using a suite of five tasks, we demonstrate that frontier LLMs can perform inductive OOCR. In one experiment we finetune an LLM on a corpus consisting only of distances between an unknown city and other known cities. Remarkably, without in-context examples or Chain of Thought, the LLM can verbalize that the unknown city is Paris and use this fact to answer downstream questions. Further experiments show that LLMs trained only on individual coin flip outcomes can verbalize whether the coin is biased, and those trained only on pairs $(x,f(x))$ can articulate a definition of $f$ and compute inverses. While OOCR succeeds in a range of cases, we also show that it is unreliable, particularly for smaller LLMs learning complex structures. Overall, the ability of LLMs to "connect the dots" without explicit in-context learning poses a potential obstacle to monitoring and controlling the knowledge acquired by LLMs. Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger B. Grosse, Owain Evans |
NeurIPS | 5 |
| 2022 | Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationabstractDesigning networks capable of attaining better performance with an increased inference budget is important to facilitate generalization to harder problem instances. Recent efforts have shown promising results in this direction by making use of depth-wise recurrent networks. In this work, we reproduce the performance of the prior art using a broader class of architectures called equilibrium models, and find that stronger generalization performance on harder examples (which require more iterations of inference to get correct) strongly correlates with the path independence of the system—its ability to converge to the same attractor (or limit cycle) regardless of initialization, given enough computation. Experimental interventions made to promote path independence result in improved generalization on harder (and thus more compute-hungry) problem instances, while those that penalize it degrade this ability. Path independence analyses are also useful on a per-example basis: for equilibrium models that have good in-distribution performance, path independence on out-of-distribution samples strongly correlates with accuracy. Thus, considering equilibrium models and path independence jointly leads to a valuable new viewpoint under which we can study the generalization performance of these networks on hard problem instances. Cem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein, Yuhuai Wu, Shaojie Bai, J. Zico Kolter, Roger B. Grosse |
NeurIPS | 1 |
| 2022 | Exploring Length Generalization in Large Language ModelsabstractThe ability to extrapolate from short problem instances to longer ones is an important form of out-of-distribution generalization in reasoning tasks, and is crucial when learning from datasets where longer problem instances are rare. These include theorem proving, solving quantitative mathematics problems, and reading/summarizing novels. In this paper, we run careful empirical studies exploring the length generalization capabilities of transformer-based language models. We first establish that naively finetuning transformers on length generalization tasks shows significant generalization deficiencies independent of model scale. We then show that combining pretrained large language models' in-context learning abilities with scratchpad prompting (asking the model to output solution steps before producing an answer) results in a dramatic improvement in length generalization. We run careful failure analyses on each of the learning modalities and identify common sources of mistakes that highlight opportunities in equipping language models with the ability to generalize to longer problems. Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay V. Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, Behnam Neyshabur |
NeurIPS | 1 |
| 2022 | Solving Quantitative Reasoning Problems with Language ModelsabstractLanguage models have achieved remarkable performance on a wide range of tasks that require natural language understanding. Nevertheless, state-of-the-art models have generally struggled with tasks that require quantitative reasoning, such as solving mathematics, science, and engineering questions at the college level. To help close this gap, we introduce Minerva, a large language model pretrained on general natural language data and further trained on technical content. The model achieves strong performance in a variety of evaluations, including state-of-the-art performance on the MATH dataset. We also evaluate our model on over two hundred undergraduate-level problems in physics, biology, chemistry, economics, and other sciences that require quantitative reasoning, and find that the model can correctly answer nearly a quarter of them. Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, Vedant Misra |
NeurIPS | 8 |
| 2021 | Learning to ElectabstractVoting systems have a wide range of applications including recommender systems, web search, product design and elections. Limited by the lack of general-purpose analytical tools, it is difficult to hand-engineer desirable voting rules for each use case. For this reason, it is appealing to automatically discover voting rules geared towards each scenario. In this paper, we show that set-input neural network architectures such as Set Transformers, fully-connected graph networks and DeepSets are both theoretically and empirically well-suited for learning voting rules. In particular, we show that these network models can not only mimic a number of existing voting rules to compelling accuracy --- both position-based (such as Plurality and Borda) and comparison-based (such as Kemeny, Copeland and Maximin) --- but also discover near-optimal voting rules that maximize different social welfare functions. Furthermore, the learned voting rules generalize well to different voter utility distributions and election sizes unseen during training. Cem Anil, Xuchan Bao |
NeurIPS | 1 |
| 2019 | TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) Pipeline for Musical Timbre Transfer
Sicong Huang 0001, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, Roger B. Grosse |
ICLR (Poster) | 3 |
| 2019 | Sorting Out Lipschitz Function ApproximationabstractTraining neural networks under a strict Lipschitz constraint is useful for provable adversarial robustness, generalization bounds, interpretable gradients, and Wasserstein distance estimation. By the composition property of Lipschitz functions, it suffices to ensure that each individual affine transformation or nonlinear activation is 1-Lipschitz. The challenge is to do this while maintaining the expressive power. We identify a necessary property for such an architecture: each of the layers must preserve the gradient norm during backpropagation. Based on this, we propose to combine a gradient norm preserving activation function, GroupSort, with norm-constrained weight matrices. We show that norm-constrained GroupSort architectures are universal Lipschitz function approximators. Empirically, we show that norm-constrained GroupSort networks achieve tighter estimates of Wasserstein distance than their ReLU counterparts and can achieve provable adversarial robustness guarantees with little cost to accuracy. Cem Anil, James Lucas, Roger B. Grosse |
ICML | 1 |
| 2019 | Preventing Gradient Attenuation in Lipschitz Constrained Convolutional NetworksabstractLipschitz constraints under L2 norm on deep neural networks are useful for provable adversarial robustness bounds, stable training, and Wasserstein distance estimation. While heuristic approaches such as the gradient penalty have seen much practical success, it is challenging to achieve similar practical performance while provably enforcing a Lipschitz constraint. In principle, one can design Lipschitz constrained architectures using the composition property of Lipschitz functions, but Anil et al. recently identified a key obstacle to this approach: gradient norm attenuation. They showed how to circumvent this problem in the case of fully connected networks by designing each layer to be gradient norm preserving. We extend their approach to train scalable, expressive, provably Lipschitz convolutional networks. In particular, we present the Block Convolution Orthogonal Parameterization (BCOP), an expressive parameterization of orthogonal convolution operations. We show that even though the space of orthogonal convolutions is disconnected, the largest connected component of BCOP with 2n channels can represent arbitrary BCOP convolutions over n channels. Our BCOP parameterization allows us to train large convolutional networks with provable Lipschitz bounds. Empirically, we find that it is competitive with existing approaches to provable adversarial robustness and Wasserstein distance estimation. Qiyang Li, Saminul Haque, Cem Anil, James Lucas, Roger B. Grosse, Jörn-Henrik Jacobsen |
NeurIPS | 3 |