Lorenzo Rosset

dblp:339/7283 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0001-6122-5252ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 40% Probabilistic and Bayesian machine learning · 40% Efficient and distributed learning · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
energy-based model
1.722025
Fast and Functional Structured Data Generators Rooted in Out-of-Equilibrium Physics · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Fast training and sampling of Restricted Boltzmann Machines · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine
1.722025
Fast and Functional Structured Data Generators Rooted in Out-of-Equilibrium Physics · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Fast training and sampling of Restricted Boltzmann Machines · ICLR 2025

Methods — techniques the papers use, named apart from their topics

markov chain monte carlo · 2.6non-equilibrium sampling · 1.7parallel trajectory tempering · 0.9convex optimization · 0.9annealing · 0.9
YearPublicationVenuePosition
2025 Fast training and sampling of Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machines (RBMs) are powerful tools for modeling complex systems and extracting insights from data, but their training is hindered by the slow mixing of Markov Chain Monte Carlo (MCMC) processes, especially with highly structured datasets. In this study, we build on recent theoretical advances in RBM training and focus on the stepwise encoding of data patterns into singular vectors of the coupling matrix, significantly reducing the cost of generating new samples and evaluating the quality of the model, as well as the training cost in highly clustered datasets. The learning process is analogous to the thermodynamic continuous phase transitions observed in ferromagnetic models, where new modes in the probability measure emerge in a continuous manner. We leverage the continuous transitions in the training process to define a smooth annealing trajectory that enables reliable and computationally efficient log-likelihood estimates. This approach enables online assessment during training and introduces a novel sampling strategy called Parallel Trajectory Tempering (PTT) that outperforms previously optimized MCMC methods. To mitigate the critical slowdown effect in the early stages of training, we propose a pre-training phase. In this phase, the principal components are encoded into a low-rank RBM through a convex optimization process, facilitating efficient static Monte Carlo sampling and accurate computation of the partition function. Our results demonstrate that this pre-training strategy allows RBMs to efficiently handle highly structured datasets where conventional methods fail. Additionally, our log-likelihood estimation outperforms computationally intensive approaches in controlled scenarios, while the PTT algorithm significantly accelerates MCMC processes compared to conventional methods.
Nicolas Béreux, Aurélien Decelle, Cyril Furtlehner, Lorenzo Rosset, Beatriz Seoane
ICLR4
2025 Fast and Functional Structured Data Generators Rooted in Out-of-Equilibrium Physics
abstract
In this study, we address the challenge of using energy-based models to produce high-quality, label-specific data in complex structured datasets, such as population genetics, RNA or protein sequences data. Traditional training methods encounter difficulties due to inefficient Markov chain Monte Carlo mixing, which affects the diversity of synthetic data and increases generation times. To address these issues, we use a novel training algorithm that exploits non-equilibrium effects. This approach, applied to the Restricted Boltzmann Machine, improves the model's ability to correctly classify samples and generate high-quality synthetic data in only a few sampling steps. The effectiveness of this method is demonstrated by its successful application to five different types of data: handwritten digits, mutations of human genomes classified by continental origin, functionally characterized sequences of an enzyme protein family, homologous RNA sequences from specific taxonomies and real classical piano pieces classified by their composer.
Alessandra Carbone, Aurélien Decelle, Lorenzo Rosset, Beatriz Seoane
IEEE Trans. Pattern Anal. Mach. Intell.3