VLDB 2026 Research / reviewers in the wild / expert
Robin San-Roman
dblp:289/7209 · also Robin San Roman
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 100% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › audio security
audio watermarking |
0.8 | 1 | 2024 | Proactive Detection of Voice Cloning with Localized Watermarking · ICML 2024 |
Machine learning › Generative modeling
audio generation |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Audio and music processing
sound synthesis |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Audio and music processing › speech coding
vocoder |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
perceptual loss · 1.5generator-detector architecture · 1.5auditory masking · 1.5multi-band diffusion · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Watermarking of Audio Generative ModelsabstractThe advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to be useful, the watermark must resist typical modifications applied to the model or its outputs. The use case of an open-source model trained on proprietary data is challenging, as post-hoc watermarks can then be trivially removed. In response, we introduce a method that watermarks latent audio generative models by directly watermarking their training data. We show the method to be robust against a broad range of audio edits including filtering, compression or even to changing the model’s decoder, maintaining high detection rates with very few false positives. Interestingly, we show that even fine-tuning the model on another dataset can only significantly lower the detection rate at the cost of degrading the generation performance near the level of re-training the model without the protected training data. Robin San-Roman, Pierre Fernandez, Antoine Deleforge, Yossi Adi, Romain Serizel |
ICASSP | 1 |
| 2025 | MusicGen-Stem: Multi-stem music generation and edition through autoregressive modelingabstractWhile most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train one specialized compression algorithm per stem to tokenize the music into parallel streams of tokens. Then, we leverage recent improvements in the task of music source separation to train a multi-stream text-to-music language model on a large dataset. Finally, thanks to a particular conditioning method, our model is able to edit bass, drums or other stems on existing or generated songs as well as doing iterative composition (e.g. generating bass on top of existing drums). This gives more flexibility in music generation algorithms and it is to the best of our knowledge the first open-source multi-stem autoregressive music generation model that can perform good quality generation and coherent source editing. Code and model weights will be released and samples are available on simonrouard.github.io/musicgenstem. Simon Rouard, Robin San-Roman, Yossi Adi, Axel Röbel |
ICASSP | 2 |
| 2024 | Proactive Detection of Voice Cloning with Localized WatermarkingabstractIn the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning. We present AudioSeal, the first audio watermarking technique designed specifically for localized detection of AI-generated speech. AudioSeal employs a generator / detector architecture trained jointly with a localization loss to enable localized watermark detection up to the sample level, and a novel perceptual loss inspired by auditory masking, that enables AudioSeal to achieve better imperceptibility. AudioSeal achieves state-of-the-art performance in terms of robustness to real life audio manipulations and imperceptibility based on automatic and human evaluation metrics. Additionally, AudioSeal is designed with a fast, single-pass detector, that significantly surpasses existing models in speed, achieving detection up to two orders of magnitude faster, making it ideal for large-scale and real-time applications.Code is available at https://github.com/facebookresearch/audioseal Robin San-Roman, Pierre Fernandez, Hady ElSahar, Alexandre Défossez, Teddy Furon |
ICML | 1 |
| 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionabstractDeep generative models can generate high-fidelity audio conditioned on various
types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients
(MFCC)). Recently, such models have been used to synthesize audio
waveforms conditioned on highly compressed representations. Although such
methods produce impressive results, they are prone to generate audible artifacts
when the conditioning is flawed or imperfect. An alternative modeling approach is
to use diffusion models. However, these have mainly been used as speech vocoders
(i.e., conditioned on mel-spectrograms) or generating relatively low sampling
rate signals. In this work, we propose a high-fidelity multi-band diffusion-based
framework that generates any type of audio modality (e.g., speech, music, environmental
sounds) from low-bitrate discrete representations. At equal bit rate,
the proposed approach outperforms state-of-the-art generative techniques in terms
of perceptual quality. Training and evaluation code are available on the facebookresearch/
audiocraft github project. Samples are available on the following
link (https://ai.honu.io/papers/mbd/). Robin San-Roman, Yossi Adi, Antoine Deleforge, Romain Serizel, Gabriel Synnaeve, Alexandre Défossez |
NeurIPS | 1 |