EDBT 2026 Demo / reviewers in the wild / expert
François Charton
dblp:255/5318
· DBLP profile ↗
16ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-5912-3342ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pangenome-Informed Language Models for Synthetic Genome Sequence GenerationabstractLanguage Models (LM) have been extensively utilized for learning DNA sequence patterns and generating synthetic sequences. In this paper, we present a novel approach for the generation of synthetic DNA data using pangenomes in combination with LM. We introduce three innovative pangenome-based tokenization schemes that enhance DNA sequence generation. Our experimental results demonstrate the superiority of pangenome-based tokenization over classical methods in generating high-utility synthetic DNA sequences, highlighting significant improvements in training efficiency and sequence quality. Pengzhi Huang, François Charton, Jan-Niklas Schmelzle, Shelby S. Darnell, Pjotr Prins, Erik Garrison, G. Edward Suh |
BIBM | 2 |
| 2025 | Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationabstractLarge Language Models (LLM) are increasingly trained on data generated by other LLMs, either because generated text and images become part of the pre-training corpus, or because synthetized data is used as a replacement for expensive human-annotation. This raises concerns about *model collapse*, a drop in model performance when their training sets include generated data. Considering that it is easier for both humans and machines to tell between good and bad examples than to generate high-quality samples, we investigate the use of verification on synthesized data to prevent model collapse. We provide a theoretical characterization using Gaussian mixtures, linear classifiers, and linear verifiers to derive conditions with measurable proxies to assess whether the verifier can effectively select synthesized data that leads to optimal performance. We experiment with two practical tasks -- computing matrix eigenvalues with transformers and news summarization with LLMs -- which both exhibit model collapse when trained on generated data, and show that verifiers, even imperfect ones, can indeed be harnessed to prevent model collapse and that our proposed proxy measure strongly correlates with performance. Yunzhen Feng, Elvis Dohmatob, François Charton, Julia Kempe |
ICLR | 4 |
| 2025 | TAPAS: Datasets for Learning the Learning with Errors ProblemabstractAI-powered attacks on Learning with Errors (LWE)—an important hard math problem in post-quantum cryptography—rival or outperform "classical" attacks on LWE under certain parameter settings. Despite the promise of this approach, a dearth of accessible data limits AI practitioners' ability to study and improve these attacks. Creating LWE data for AI model training is time- and compute-intensive and requires significant domain expertise. To fill this gap and accelerate AI research on LWE attacks, we propose the TAPAS datasets, a ${\bf t}$oolkit for ${\bf a}$nalysis of ${\bf p}$ost-quantum cryptography using ${\bf A}$I ${\bf s}$ystems. These datasets cover several LWE settings and can be used off-the-shelf by AI practitioners to prototype new approaches to cracking LWE. This work documents TAPAS dataset creation, establishes attack performance baselines, and lays out directions for future work. Eshika Saxena, Alberto Alfarano, François Charton, Emily Wenger, Kristin E. Lauter |
NeurIPS | 3 |
| 2024 | Learning the greatest common divisor: explaining transformer predictionsabstractThe predictions of small transformers, trained to calculate the greatest common divisor (GCD) of two positive integers, can be fully characterized by looking at model inputs and outputs.
As training proceeds, the model learns a list $\mathcal D$ of integers, products of divisors of the base used to represent integers and small primes, and predicts the largest element of $\mathcal D$ that divides both inputs.
Training distributions impact performance. Models trained from uniform operands only learn a handful of GCD (up to $38$ GCD $\leq100$). Log-uniform operands boost performance to $73$ GCD $\leq 100$, and a log-uniform distribution of outcomes (i.e. GCD) to $91$. However, training from uniform (balanced) GCD breaks explainability. François Charton |
ICLR | 1 |
| 2024 | A Tale of Tails: Model Collapse as a Change of Scaling LawsabstractAs AI model size grows, neural scaling laws have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread use of popular models means that the ecosystem of online data and text will co-evolve to progressively contain increased amounts of synthesized data. In this paper we ask: How will the scaling laws change in the inevitable regime where synthetic data makes its way into the training corpus? Will future models, still improve, or be doomed to degenerate up to total (model) collapse? We develop a theoretical framework of model collapse through the lens of scaling laws. We discover a wide range of decay phenomena, analyzing loss of scaling, shifted scaling with number of generations, the ”un-learning" of skills, and grokking when mixing human and synthesized data. Our theory is validated by large-scale experiments with a transformer on an arithmetic task and text generation using the large language model Llama2. Elvis Dohmatob, Yunzhen Feng, François Charton, Julia Kempe |
ICML | 4 |
| 2024 | Global Lyapunov functions: a long-standing open problem in mathematics, with symbolic transformersabstractDespite their spectacular progress, language models still struggle on complex reasoning tasks, such as advanced mathematics.
We consider a long-standing open problem in mathematics: discovering a Lyapunov function that ensures the global stability of a dynamical system. This problem has no known general solution, and algorithmic solvers only exist for some small polynomial systems.
We propose a new method for generating synthetic training samples from random solutions, and show that sequence-to-sequence transformers trained on such datasets perform better than algorithmic solvers and humans on polynomial systems, and can discover new Lyapunov functions for non-polynomial systems. Alberto Alfarano, François Charton, Amaury Hayat |
NeurIPS | 2 |
| 2024 | Iteration Head: A Mechanistic Study of Chain-of-ThoughtabstractChain-of-Thought (CoT) reasoning is known to improve Large Language Models both empirically and in terms of theoretical approximation power.
However, our understanding of the inner workings and conditions of apparition of CoT capabilities remains limited.
This paper helps fill this gap by demonstrating how CoT reasoning emerges in transformers in a controlled and interpretable setting.
In particular, we observe the appearance of a specialized attention mechanism dedicated to iterative reasoning, which we coined "iteration heads".
We track both the emergence and the precise working of these iteration heads down to the attention level, and measure the transferability of the CoT skills to which they give rise between tasks. Vivien Cabannes, Charles Arnal, Wassim Bouaziz, François Charton, Julia Kempe |
NeurIPS | 5 |
| 2023 | SalsaPicante: A Machine Learning Attack on LWE with Binary SecretsabstractLearning with Errors (LWE) is a hard math problem underpinning many proposed post-quantum cryptographic (PQC) systems. The only PQC Key Exchange Mechanism (KEM) standardized by NIST [13] is based on module LWE [2], and current publicly available PQ Homomorphic Encryption (HE) libraries are based on ring LWE. The security of LWE-based PQ cryptosystems is critical, but certain implementation choices could weaken them. One such choice is sparse binary secrets, desirable for PQ HE schemes for efficiency reasons. Prior work SALSA[51] demonstrated a machine learning-based attack on LWE with sparse binary secrets in small dimensions (n ≤ = 128) and low Hamming weights (h ≤ = 4). However, this attack assumes access to millions of eavesdropped LWE samples and fails at higher Hamming weights or dimensions. Cathy Yuanchen Li, Jana Sotáková, Emily Wenger, Mohamed Malhou, Evrard Garcelon, François Charton, Kristin E. Lauter |
CCS | 6 |
| 2023 | Code Translation with Compiler Representations
Marc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut, François Charton, Gabriel Synnaeve |
ICLR | 5 |
| 2023 | SALSA VERDE: a machine learning attack on LWE with sparse small secretsabstractLearning with Errors (LWE) is a hard math problem used in post-quantum cryptography. Homomorphic Encryption (HE) schemes rely on the hardness of the LWE problem for their security, and two LWE-based cryptosystems were recently standardized by NIST for digital signatures and key exchange (KEM). Thus, it is critical to continue assessing the security of LWE and specific parameter choices. For example, HE uses secrets with small entries, and the HE community has considered standardizing small sparse secrets to improve efficiency and functionality. However, prior work, SALSA and PICANTE, showed that ML attacks can recover sparse binary secrets. Building on these, we propose VERDE, an improved ML attack that can recover sparse binary, ternary, and narrow Gaussian secrets. Using improved preprocessing and secret recovery techniques, VERDE can attack LWE with larger dimensions ($n=512$) and smaller moduli ($\log_2 q=12$ for $n=256$), using less time and power. We propose novel architectures for scaling. Finally, we develop a theory that explains the success of ML LWE attacks. Cathy Yuanchen Li, Emily Wenger, Zeyuan Allen Zhu, François Charton, Kristin E. Lauter |
NeurIPS | 4 |
| 2022 | Leveraging Automated Unit Tests for Unsupervised Code Translation
Baptiste Rozière, Jie Zhang 0050, François Charton, Mark Harman, Gabriel Synnaeve, Guillaume Lample |
ICLR | 3 |
| 2022 | Deep symbolic regression for recurrence predictionabstractSymbolic regression, i.e. predicting a function from the observation of its values, is well-known to be a challenging task. In this paper, we train Transformers to infer the function or recurrence relation underlying sequences of integers or floats, a typical task in human IQ tests which has hardly been tackled in the machine learning literature. We evaluate our integer model on a subset of OEIS sequences, and show that it outperforms built-in Mathematica functions for recurrence prediction. We also demonstrate that our float model is able to yield informative approximations of out-of-vocabulary functions and constants, e.g. $\operatorname{bessel0}(x)\approx \frac{\sin(x)+\cos(x)}{\sqrt{\pi x}}$ and $1.644934\approx \pi^2/6$. Stéphane d'Ascoli, Pierre-Alexandre Kamienny, Guillaume Lample, François Charton |
ICML | 4 |
| 2022 | End-to-end Symbolic Regression with TransformersabstractSymbolic regression, the task of predicting the mathematical expression of a function from the observation of its values, is a difficult task which usually involves a two-step procedure: predicting the "skeleton" of the expression up to the choice of numerical constants, then fitting the constants by optimizing a non-convex loss function. The dominant approach is genetic programming, which evolves candidates by iterating this subroutine a large number of times. Neural networks have recently been tasked to predict the correct skeleton in a single try, but remain much less powerful.In this paper, we challenge this two-step procedure, and task a Transformer to directly predict the full mathematical expression, constants included. One can subsequently refine the predicted constants by feeding them to the non-convex optimizer as an informed initialization. We present ablations to show that this end-to-end approach yields better results, sometimes even without the refinement step. We evaluate our model on problems from the SRBench benchmark and show that our model approaches the performance of state-of-the-art genetic programming with several orders of magnitude faster inference. Pierre-Alexandre Kamienny, Stéphane d'Ascoli, Guillaume Lample, François Charton |
NeurIPS | 4 |
| 2022 | SALSA: Attacking Lattice Cryptography with TransformersabstractCurrently deployed public-key cryptosystems will be vulnerable to attacks by full-scale quantum computers. Consequently, "quantum resistant" cryptosystems are in high demand, and lattice-based cryptosystems, based on a hard problem known as Learning With Errors (LWE), have emerged as strong contenders for standardization. In this work, we train transformers to perform modular arithmetic and mix half-trained models and statistical cryptanalysis techniques to propose SALSA: a machine learning attack on LWE-based cryptographic schemes. SALSA can fully recover secrets for small-to-mid size LWE instances with sparse binary secrets, and may scale to attack real world LWE-based cryptosystems. Emily Wenger, François Charton, Kristin E. Lauter |
NeurIPS | 3 |
| 2021 | Learning advanced mathematical computations from examples
François Charton, Amaury Hayat, Guillaume Lample |
ICLR | 1 |
| 2020 | Deep Learning For Symbolic Mathematics
Guillaume Lample, François Charton |
ICLR | 2 |