Baptiste Rozière

dblp:178/8263 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-9014-4379ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 44% Trustworthy machine learning · 14% Reinforcement learning · 13%
Software engineering, system software, and programming languages
7 papers
Program synthesis and code generation · 57% Software testing · 40% Program analysis · 3%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code translation
1.732023
Code Translation with Compiler Representations · ICLR 2023
Leveraging Automated Unit Tests for Unsupervised Code Translation · ICLR 2022
Unsupervised Translation of Programming Languages · NeurIPS 2020
Program synthesis and code generation
code generation with language models
1.632025
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark · ICLR 2025
DOBF: A Deobfuscation Pre-Training Objective for Programming Languages · NeurIPS 2021
Getting the most out of your tokenizer for pre-training and domain adaptation · ICML 2024
Software testing › test generation
unit test generation
1.422025
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark · ICLR 2025
Leveraging Automated Unit Tests for Unsupervised Code Translation · ICLR 2022
Software testing
test generation
0.912025
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark · ICLR 2025
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.812024
Getting the most out of your tokenizer for pre-training and domain adaptation · ICML 2024
Natural language and speech › Language models and text generation
large language model training
0.812024
Better & Faster Large Language Models via Multi-token Prediction · ICML 2024
Natural language and speech › Language models and text generation › language modeling
multi-token prediction
0.812024
Better & Faster Large Language Models via Multi-token Prediction · ICML 2024
Natural language and speech › Language models and text generation
tokenization
0.812024
Getting the most out of your tokenizer for pre-training and domain adaptation · ICML 2024
Machine learning › Deep learning architectures and training
training objective
0.812024
Better & Faster Large Language Models via Multi-token Prediction · ICML 2024
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial sample generation
adversarial image generation
0.512021
Inspirational Adversarial Image Generation · IEEE Trans. Image Process. 2021
Machine learning › Generative modeling
latent space optimization
0.512021
Inspirational Adversarial Image Generation · IEEE Trans. Image Process. 2021
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.412020
Adversarial Attacks on Linear Contextual Bandits · NeurIPS 2020
Machine learning › Reinforcement learning › bandit
contextual bandit
0.412020
Adversarial Attacks on Linear Contextual Bandits · NeurIPS 2020
Machine learning › Reinforcement learning › bandit › contextual bandit
linear contextual bandit
0.412020
Adversarial Attacks on Linear Contextual Bandits · NeurIPS 2020
Program analysis › program representation
source code representation
0.112021
DOBF: A Deobfuscation Pre-Training Objective for Programming Languages · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

tokenizer ablation · 1.5fine-tuning · 1.5chain-of-thought · 1.5byte-pair encoding · 1.5preference-based optimization · 1.0gradient-free optimization · 1.0gradient descent · 1.0large language model evaluation · 0.9next-token prediction · 0.8multi-token prediction · 0.8machine translation · 0.7compiler intermediate representation · 0.7unsupervised learning · 0.6unit tests · 0.6masked language modeling · 0.5adversarial perturbation · 0.4
YearPublicationVenuePosition
2025 LLM Compiler: Foundation Language Models for Compiler Optimization
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of software engineering and coding tasks. However, their application in the domain of code and compiler optimization remains underexplored. Training LLMs is resource-intensive, requiring substantial GPU hours and extensive data collection, which can be prohibitive. To address this gap, we introduce LLM Compiler, a suite of robust, openly available, pre-trained models specifically designed for compiler tasks. Built on the foundation of Code Llama, LLM Compiler enhances the understanding of compiler intermediate representations (IRs), assembly language, and optimization techniques. The models have been trained on a vast corpus of 546 billion tokens of LLVM-IR and assembly code and have undergone instruction fine-tuning to interpret compiler behavior. To demonstrate the utility of these research tools, we also present fine-tuned versions of the models with enhanced capabilities in optimizing code size and disassembling from x86_64 and ARM assembly back into LLVM-IR. These achieve 77% of the optimising potential of an autotuning search, and 45% disassembly round trip (14% exact match). LLM Compiler is released under a bespoke commercial license to allow wide reuse and is available in two sizes: 7 billion and 13 billion parameters. Our aim is to provide scalable, cost-effective foundational models for further research and development in compiler optimization by both academic researchers and industry practitioners. Since we released LLM Compiler the community has quantized, repackaged, and downloaded the models over 250k times.
Chris Cummins, Volker Seeker, Dejan Grubisic, Baptiste Rozière, Jonas Gehring, Gabriel Synnaeve, Hugh Leather
CC4
2025 TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
abstract
Code generation models can help improve many common software tasks ranging from code completion to defect prediction. Most of the existing benchmarks for code generation LLMs focus on code authoring or code completion. Surprisingly, there has been far less effort dedicated to benchmarking software testing, despite the strong correlation between well-tested software and effective bug detection. To address this gap, we create and release TestGenEval, a large-scale benchmark to measure test generation performance. Based on SWEBench, TestGenEval comprises 68,647 tests from 1,210 code and test file pairs across 11 well-maintained Python repositories. It covers initial tests authoring, test suite completion, and code coverage improvements. Test authoring simulates the process of a developer writing a test suite from scratch, while test completion mimics the scenario where a developer aims to improve the coverage of an existing test suite. We evaluate several popular models, with sizes ranging from 7B to 405B parameters. Our detailed analysis highlights TestGenEval's contribution to a comprehensive evaluation of test generation performance. In particular, models struggle to generate high-coverage test suites, with the best model, GPT-4o, achieving an average coverage of only 35.2\%. This is primarily due to models struggling to reason about execution, and their frequent assertion errors when addressing complex code paths.
Kush Jain, Gabriel Synnaeve, Baptiste Rozière
ICLR3
2024 Getting the most out of your tokenizer for pre-training and domain adaptation
abstract
Tokenization is an understudied and often neglected component of modern LLMs. Most published works use a single tokenizer for all experiments, often borrowed from another model, without performing ablations or analysis to optimize tokenization. Moreover, the tokenizer is generally kept unchanged when fine-tuning a base model. In this paper, we show that the size, pre-tokenization regular expression, and training data of a tokenizer can significantly impact the model’s generation speed, effective context size, memory usage, and downstream performance. We train specialized Byte-Pair Encoding code tokenizers, and conduct extensive ablations on the impact of tokenizer design on the performance of LLMs for code generation tasks such as HumanEval and MBPP, and provide recommendations for tokenizer hyper-parameters selection and switching the tokenizer in a pre-trained LLM. We perform our experiments on models trained from scratch and from pre-trained models, verifying their applicability to a wide range of use-cases. We find that when fine-tuning on more than 50 billion tokens, we can specialize the tokenizer of a pre-trained LLM to obtain large gains in generation speed and effective context size.
Gautier Dagan, Gabriel Synnaeve, Baptiste Rozière
ICML3
2024 Better & Faster Large Language Models via Multi-token Prediction
abstract
Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens at once results in higher sample efficiency. More specifically, at each position in the training corpus, we ask the model to predict the following $n$ tokens using $n$ independent output heads, operating on top of a shared model trunk. Considering multi-token prediction as an auxiliary training task, we measure improved downstream capabilities with no overhead in training time for both code and natural language models. The method is increasingly useful for larger model sizes, and keeps its appeal when training for multiple epochs. Gains are especially pronounced on generative benchmarks like coding, where our models consistently outperform strong baselines by several percentage points. Our 13B parameter models solves 12% more problems on Human Eval and 17% more on MBPP than comparable next-token models. Experiments on small algorithmic tasks demonstrate that multi-token prediction is favorable for the development of induction heads and algorithmic reasoning capabilities. As an additional benefit, models trained with 4-token prediction are up to $3\times$ faster at inference, even with large batch sizes.
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
ICML3
2024 CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
abstract
We present Code Reasoning, Understanding, and eXecution Evaluation, a benchmark consisting of 800 Python functions (3-13 lines). Each function comes with an input-output pair, leading to two natural tasks: input prediction and output prediction. First, we propose a general recipe for generating our execution benchmark by sampling from a model, which can be used for more challenging versions of the benchmark if needed. Second, we evaluate twenty code models on our benchmark and discover that many recent high-scoring models on HumanEval show no improvements on our benchmark. Third, we show that simple CoT and fine-tuning schemes can improve performance on our benchmark but remain far from solving it. The best setup, GPT-4 with chain of thought (CoT), achieves a pass@1 of 75% and 81% on input and output prediction, respectively. In contrast, Code Llama 34B achieves a pass@1 of 50% and 46% on input and output prediction. When it comes to reasoning about code, GPT-4 has a huge edge over other models but still fails consistently on some surprisingly simple Python programs.
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, Sida I. Wang
ICML2
2023 Code Translation with Compiler Representations
Marc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut, François Charton, Gabriel Synnaeve
ICLR2
2022 Leveraging Automated Unit Tests for Unsupervised Code Translation
Baptiste Rozière, Jie Zhang 0050, François Charton, Mark Harman, Gabriel Synnaeve, Guillaume Lample
ICLR1
2022 Black-Box Optimization Revisited: Improving Algorithm Selection Wizards Through Massive Benchmarking
abstract
Existing studies in black-box optimization suffer from low generalizability, caused by a typically selective choice of problem instances used for training and testing of different optimization algorithms. Among other issues, this practice promotes overfitting and poor-performing user guidelines. We address this shortcoming by introducing in this work a general-purpose algorithm selection wizard that was designed and tested on a previously unseen breadth of black-box optimization problems, ranging from academic benchmarks to real-world applications, from discrete over numerical to mixed-integer problems, from small to very large-scale problems, from noisy over dynamic to static problems, etc. Not only did we use the already very extensive benchmark environment available in Nevergrad, but we also extended it significantly by adding a number of additional benchmark suites, including Pyomo, Photonics, large-scale global optimization (LSGO), and MuJoCo. Our wizard achieves competitive performance on all benchmark suites. It significantly outperforms previous state-of-the-art algorithms on some of the suites, including YABBOB and LSGO. Its excellent performance is obtained without any task-specific parametrization. The algorithm selection wizard, all of its base solvers, as well as the benchmark suites are available for reproducible research in the open-source Nevergrad platform.
Laurent Meunier, Herilalaina Rakotoarison, Pak-Kan Wong, Baptiste Rozière, Jérémy Rapin, Olivier Teytaud, Antoine Moreau, Carola Doerr
IEEE Trans. Evol. Comput.4
2021 DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
abstract
Recent advances in self-supervised learning have dramatically improved the state of the art on a wide variety of tasks. However, research in language model pre-training has mostly focused on natural languages, and it is unclear whether models like BERT and its variants provide the best pre-training when applied to other modalities, such as source code. In this paper, we introduce a new pre-training objective, DOBF, that leverages the structural aspect of programming languages and pre-trains a model to recover the original version of obfuscated source code. We show that models pre-trained with DOBF significantly outperform existing approaches on multiple downstream tasks, providing relative improvements of up to 12.2% in unsupervised code translation, and 5.3% in natural language code search. Incidentally, we found that our pre-trained model is able to deobfuscate fully obfuscated source files, and to suggest descriptive variable names.
Marie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume Lample
NeurIPS2
2021 Inspirational Adversarial Image Generation
abstract
The task of image generation started receiving some attention from artists and designers, providing inspiration for new creations. However, exploiting the results of deep generative models such as Generative Adversarial Networks can be long and tedious given the lack of existing tools. In this work, we propose a simple strategy to inspire creators with new generations learned from a dataset of their choice, while providing some control over the output. We design a simple optimization method to find the optimal latent parameters corresponding to the closest generation to any input inspirational image. Specifically, we allow the generation given an inspirational image of the user's choosing by performing several optimization steps to recover optimal parameters from the model's latent space. We tested several exploration methods from classical gradient descents to gradient-free optimizers. Many gradient-free optimizers just need comparisons (better/worse than another image), so they can even be used without numerical criterion nor inspirational image, only with human preferences. Thus, by iterating on one's preferences we can make robust facial composite or fashion generation algorithms. Our results on four datasets of faces, fashion images, and textures show that satisfactory images are effectively retrieved in most cases.
Baptiste Rozière, Morgane Rivière, Olivier Teytaud, Jérémy Rapin, Yann LeCun, Camille Couprie
IEEE Trans. Image Process.1
2020 EvolGAN: Evolutionary Generative Adversarial Networks
Baptiste Rozière, Fabien Teytaud, Vlad Hosu, Hanhe Lin, Jérémy Rapin, Mariia Zameshina, Olivier Teytaud
ACCV (4)1
2020 Versatile black-box optimization
abstract
Choosing automatically the right algorithm using problem descriptors is a classical component of combinatorial optimization. It is also a good tool for making evolutionary algorithms fast, robust and versatile. We present Shiwa, an algorithm good at both discrete and continuous, noisy and noise-free, sequential and parallel, black-box optimization. Our algorithm is experimentally compared to competitors on YABBOB, a BBOB comparable testbed, and on some variants of it, and then validated on several real world testbeds.
Jialin Liu 0001, Antoine Moreau, Mike Preuss, Jérémy Rapin, Baptiste Rozière, Fabien Teytaud, Olivier Teytaud
GECCO5
2020 Tarsier: Evolving Noise Injection in Super-Resolution GANs
abstract
Super-resolution aims at increasing the resolution and level of detail within an image. The current state of the art in general single-image super-resolution is held by NESRGAN+, which injects a Gaussian noise after each residual layer at training time. In this paper, we harness evolutionary methods to improve NESRGAN+ by optimizing the noise injection at inference time. More precisely, we use Diagonal CMA to optimize the injected noise according to a novel criterion combining quality assessment and realism. Our results are validated by the PIRM perceptual score and a human study. Our method outperforms NESRGAN+ on several standard super-resolution datasets. More generally, our approach can be used to optimize any method based on noise injection.
Baptiste Rozière, Nathanaël Carraz Rakotonirina, Vlad Hosu, Andry Rasoanaivo, Hanhe Lin, Camille Couprie, Olivier Teytaud
ICPR1
2020 Adversarial Attacks on Linear Contextual Bandits
abstract
Contextual bandit algorithms are applied in a wide range of domains, from advertising to recommender systems, from clinical trials to education. In many of these domains, malicious agents may have incentives to force a bandit algorithm into a desired behavior For instance, an unscrupulous ad publisher may try to increase their own revenue at the expense of the advertisers; a seller may want to increase the exposure of their products, or thwart a competitor’s advertising campaign. In this paper, we study several attack scenarios and show that a malicious agent can force a linear contextual bandit algorithm to pull any desired arm T − o(T) times over a horizon of T steps, while applying adversarial modifications to either rewards or contexts with a cumulative cost that only grow logarithmically as O(log T). We also investigate the case when a malicious agent is interested in affecting the behavior of the bandit algorithm in a single context (e.g., a specific user). We first provide sufficient conditions for the feasibility of the attack and an efficient algorithm to perform an attack. We empirically validate the proposed approaches on synthetic and real-world datasets.
Evrard Garcelon, Baptiste Rozière, Laurent Meunier, Jean Tarbouriech, Olivier Teytaud, Alessandro Lazaric, Matteo Pirotta
NeurIPS2
2020 Unsupervised Translation of Programming Languages
abstract
A transcompiler, also known as source-to-source translator, is a system that converts source code from a high-level programming language (such as C++ or Python) to another. Transcompilers are primarily used for interoperability, and to port codebases written in an obsolete or deprecated language (e.g. COBOL, Python 2) to a modern one. They typically rely on handcrafted rewrite rules, applied to the source code abstract syntax tree. Unfortunately, the resulting translations often lack readability, fail to respect the target language conventions, and require manual modifications in order to work properly. The overall translation process is time-consuming and requires expertise in both the source and target languages, making code-translation projects expensive. Although neural models significantly outperform their rule-based counterparts in the context of natural language translation, their applications to transcompilation have been limited due to the scarcity of parallel data in this domain. In this paper, we propose to leverage recent approaches in unsupervised machine translation to train a fully unsupervised neural transcompiler. We train our model on source code from open source GitHub projects, and show that it can translate functions between C++, Java, and Python with high accuracy. Our method relies exclusively on monolingual source code, requires no expertise in the source or target languages, and can easily be generalized to other programming languages. We also build and release a test set composed of 852 parallel functions, along with unit tests to check the correctness of translations. We show that our model outperforms rule-based commercial baselines by a significant margin.
Baptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume Lample
NeurIPS1