Nathanaël Perraudin

dblp:139/7579 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-8285-1308ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Graph learning · 38% Trustworthy machine learning · 25% Generative modeling · 19%
Computer graphics and multimedia
3 papers
Audio and music processing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph generation
1.322024
Efficient and Scalable Graph Generation through Iterative Local Expansion · ICLR 2024
SPECTRE: Spectral Conditioning Helps to Overcome the Expressivity Limits of One-shot Graph Generators · ICML 2022
Machine learning › Graph learning
graph neural network
0.922021
Scalable Graph Networks for Particle Simulations · AAAI 2021
DeepSphere: a graph-based spherical CNN · ICLR 2020
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
audio inpainting
0.722019
A Context Encoder For Audio Inpainting · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Inpainting of Long Audio Segments With Similarity Graphs · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.612022
What You See is What You Classify: Black Box Attributions · NeurIPS 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
What You See is What You Classify: Black Box Attributions · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map
0.612022
What You See is What You Classify: Black Box Attributions · NeurIPS 2022
Computational science and engineering › numerical simulation › particle simulation
n-body simulation
0.512021
Scalable Graph Networks for Particle Simulations · AAAI 2021
Computational science and engineering › numerical simulation
particle simulation
0.512021
Scalable Graph Networks for Particle Simulations · AAAI 2021
Machine learning › Deep learning architectures and training › equivariant neural network
spherical CNN
0.412020
DeepSphere: a graph-based spherical CNN · ICLR 2020
Machine learning › Generative modeling
adversarial generative modeling
0.412019
Adversarial Generation of Time-Frequency Features with application in audio synthesis · ICML 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
A Context Encoder For Audio Inpainting · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Machine learning › Optimization for machine learning › convex optimization › duality
duality gap
0.412019
A Domain Agnostic Measure for Monitoring and Evaluating GANs · NeurIPS 2019
Machine learning › Generative modeling › generative model evaluation
generative adversarial network evaluation
0.412019
A Domain Agnostic Measure for Monitoring and Evaluating GANs · NeurIPS 2019
Machine learning › Generative modeling
generative model evaluation
0.412019
A Domain Agnostic Measure for Monitoring and Evaluating GANs · NeurIPS 2019
Machine learning › Graph learning
graph signal processing
0.412019
Large Scale Graph Learning From Smooth Signals · ICLR (Poster) 2019
Audio and music processing › sound synthesis
neural audio synthesis
0.412019
Adversarial Generation of Time-Frequency Features with application in audio synthesis · ICML 2019
Audio and music processing
acoustic signal processing
0.312018
Inpainting of Long Audio Segments With Similarity Graphs · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Machine learning › Generative modeling
generative adversarial network
0.212022
SPECTRE: Spectral Conditioning Helps to Overcome the Expressivity Limits of One-shot Graph Generators · ICML 2022
Graph algorithms and graph theory › graph processing
scalable graph algorithms
0.112019
Large Scale Graph Learning From Smooth Signals · ICLR (Poster) 2019
Audio and music processing
sound synthesis
0.112018
Inpainting of Long Audio Segments With Similarity Graphs · IEEE ACM Trans. Audio Speech Lang. Process. 2018

Methods — techniques the papers use, named apart from their topics

graph neural network · 1.0smooth signal models · 0.8iterative local expansion · 0.8generative adversarial network · 0.8denoising diffusion · 0.8spectral conditioning · 0.6perturbation · 0.6graph laplacian spectrum · 0.6deep network · 0.6GAN · 0.6short-time fourier transform · 0.4magnitude spectrogram prediction · 0.4complex-valued time-frequency representation · 0.4similarity graphs · 0.3adversarial generation of time-frequency features · 0.3
YearPublicationVenuePosition
2025 Coarse-to-Fine Text-to-Music Latent Diffusion
abstract
We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the Descript Audio Codec. We consolidate this coarse-to-fine design through an important observation that the audio latent representation can be split into a primary and secondary part, controlling music content and details accordingly. We validate the effectiveness of our approach and text-audio alignment through various objective metrics. Furthermore, we provide access to high-quality synthetic captions for the MTG-Jamendo and FMA datasets, as well as open-sourcing DiscoDiff’s codebase and model checkpoints.
Luca A. Lanzendörfer, Tongyu Lu, Nathanaël Perraudin, Dorien Herremans, Roger Wattenhofer
ICASSP3
2025 Bootstrapping Language-Audio Pre-training for Music Captioning
abstract
We introduce BLAP, a model capable of generating high-quality captions for music. BLAP leverages a fine-tuned CLAP audio encoder and a pre-trained Flan-T5 large language model. To achieve effective cross-modal alignment between music and language, BLAP utilizes a Querying Transformer, allowing us to obtain state-of-the-art performance using 6x less data compared to previous models. This is a critical consideration given the scarcity of descriptive music data and the subjective nature of music interpretation. We provide qualitative examples demonstrating BLAP’s ability to produce realistic captions for music, and perform a quantitative evaluation on three datasets. BLAP achieves a relative improvement on FENSE compared to previous models of 3.5%, 6.5%, and 7.5% on the MusicCaps, Song Describer, and YouTube8m-MTC datasets, respectively. The codebase is available at https://github.com/ETH-DISCO/blap.
Luca A. Lanzendörfer, Constantin Pinkl, Nathanaël Perraudin, Roger Wattenhofer
ICASSP3
2024 Efficient and Scalable Graph Generation through Iterative Local Expansion
abstract
In the realm of generative models for graphs, extensive research has been conducted. However, most existing methods struggle with large graphs due to the complexity of representing the entire joint distribution across all node pairs and capturing both global and local graph structures simultaneously. To overcome these issues, we introduce a method that generates a graph by progressively expanding a single node to a target graph. In each step, nodes and edges are added in a localized manner through denoising diffusion, building first the global structure, and then refining the local details. The local generation avoids modeling the entire joint distribution over all node pairs, achieving substantial computational savings with subquadratic runtime relative to node count while maintaining high expressivity through multiscale generation. Our experiments show that our model achieves state-of-the-art performance on well-established benchmark datasets while successfully scaling to graphs with at least 5000 nodes. Our method is also the first to successfully extrapolate to graphs outside of the training distribution, showcasing a much better generalization capability over existing methods.
Andreas Bergmeister, Karolis Martinkus, Nathanaël Perraudin, Roger Wattenhofer
ICLR3
2022 SPECTRE: Spectral Conditioning Helps to Overcome the Expressivity Limits of One-shot Graph Generators
abstract
We approach the graph generation problem from a spectral perspective by first generating the dominant parts of the graph Laplacian spectrum and then building a graph matching these eigenvalues and eigenvectors. Spectral conditioning allows for direct modeling of the global and local graph structure and helps to overcome the expressivity and mode collapse issues of one-shot graph generators. Our novel GAN, called SPECTRE, enables the one-shot generation of much larger graphs than previously possible with one-shot models. SPECTRE outperforms state-of-the-art deep autoregressive generators in terms of modeling fidelity, while also avoiding expensive sequential generation and dependence on node ordering. A case in point, in sizable synthetic and real-world graphs SPECTRE achieves a 4-to-170 fold improvement over the best competitor that does not overfit and is 23-to-30 times faster than autoregressive generators.
Karolis Martinkus, Andreas Loukas, Nathanaël Perraudin, Roger Wattenhofer
ICML3
2022 What You See is What You Classify: Black Box Attributions
abstract
An important step towards explaining deep image classifiers lies in the identification of image regions that contribute to individual class scores in the model's output. However, doing this accurately is a difficult task due to the black-box nature of such networks. Most existing approaches find such attributions either using activations and gradients or by repeatedly perturbing the input. We instead address this challenge by training a second deep network, the Explainer, to predict attributions for a pre-trained black-box classifier, the Explanandum. These attributions are provided in the form of masks that only show the classifier-relevant parts of an image, masking out the rest. Our approach produces sharper and more boundary-precise masks when compared to the saliency maps generated by other methods. Moreover, unlike most existing approaches, ours is capable of directly generating very distinct class-specific masks in a single forward pass. This makes the proposed method very efficient during inference. We show that our attributions are superior to established methods both visually and quantitatively with respect to the PASCAL VOC-2007 and Microsoft COCO-2014 datasets.
Steven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz, Michele Volpi
NeurIPS2
2021 Scalable Graph Networks for Particle Simulations
abstract
Learning system dynamics directly from observations is a promising direction in machine learning due to its potential to significantly enhance our ability to understand physical systems. However, the dynamics of many real-world systems are challenging to learn due to the presence of nonlinear potentials and a number of interactions that scales quadratically with the number of particles N, as in the case of the N-body problem. In this work we introduce an approach that transforms a fully-connected interaction graph into a hierarchical one which reduces the number of edges to O(N). This results in a linear time and space complexity while the pre-computation of the hierarchical graph requires O(N log (N)) time and O(N) space. Using our approach, we are able to train models on much larger particle counts, even on a single GPU. We evaluate how the phase space position accuracy and energy conservation depend on the number of simulated particles. Our approach retains high accuracy and efficiency even on large-scale gravitational N-body simulations which are impossible to run on a single machine if a fully-connected graph is used. Similar results are also observed when simulating Coulomb interactions. Furthermore, we make several important observations regarding the performance of this new hierarchical model, including: i) its accuracy tends to improve with the number of particles in the simulation and ii) its generalisation to unseen particle counts is also much better than for models that use all O(N^2) interactions.
Karolis Martinkus, Aurélien Lucchi, Nathanaël Perraudin
AAAI3
2020 DeepSphere: a graph-based spherical CNN
Michaël Defferrard, Martino Milani, Frédérick Gusset, Nathanaël Perraudin
ICLR4
2019 Large Scale Graph Learning From Smooth Signals
Vassilis Kalofolias, Nathanaël Perraudin
ICLR (Poster)2
2019 Adversarial Generation of Time-Frequency Features with application in audio synthesis
abstract
Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous attempts at unconditionally synthesizing audio from neurally generated invertible TF features still struggle to produce audio at satisfying quality. In this article, focusing on the short-time Fourier transform, we discuss the challenges that arise in audio synthesis based on generated invertible TF features and how to overcome them. We demonstrate the potential of deliberate generative TF modeling by training a generative adversarial network (GAN) on short-time Fourier features. We show that by applying our guidelines, our TF-based network was able to outperform a state-of-the-art GAN generating waveforms directly, despite the similar architecture in the two networks.
Andrés Marafioti, Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak
ICML2
2019 A Domain Agnostic Measure for Monitoring and Evaluating GANs
abstract
Generative Adversarial Networks (GANs) have shown remarkable results in modeling complex distributions, but their evaluation remains an unsettled issue. Evaluations are essential for: (i) relative assessment of different models and (ii) monitoring the progress of a single model throughout training. The latter cannot be determined by simply inspecting the generator and discriminator loss curves as they behave non-intuitively. We leverage the notion of duality gap from game theory to propose a measure that addresses both (i) and (ii) at a low computational cost. Extensive experiments show the effectiveness of this measure to rank different GAN models and capture the typical GAN failure scenarios, including mode collapse and non-convergent behaviours. This evaluation metric also provides meaningful monitoring on the progression of the loss during training. It highly correlates with FID on natural image datasets, and with domain specific scores for text, sound and cosmology data where FID is not directly suitable. In particular, our proposed metric requires no labels or a pretrained classifier, making it domain agnostic.
Paulina Grnarova, Kfir Y. Levy, Aurélien Lucchi, Nathanaël Perraudin, Ian J. Goodfellow, Thomas Hofmann 0001, Andreas Krause 0001
NeurIPS4
2019 A Context Encoder For Audio Inpainting
abstract
In this article, we study the ability of deep neural networks (DNNs) to restore missing audio content based on its context, i.e., inpaint audio gaps. We focus on a condition which has not received much attention yet: gaps in the range of tens of milliseconds. We propose a DNN structure that is provided with the signal surrounding the gap in the form of time-frequency (TF) coefficients. Two DNNs with either complex-valued TF coefficient output or magnitude TF coefficient output were studied by separately training them on inpainting two types of audio signals (music and musical instruments) having 64-ms long gaps. The magnitude DNN outperformed the complex-valued DNN in terms of signal-to-noise ratios and objective difference grades. Although, for instruments, a reference inpainting obtained through linear predictive coding performed better in both metrics, it performed worse than the magnitude DNN for music. This demonstrates the potential of the magnitude DNN, in particular for inpainting signals that are more complex than single instrument sounds.
Andrés Marafioti, Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Inpainting of Long Audio Segments With Similarity Graphs
abstract
Pre-trained models for:\n\nAdversarial Generation of Time-Frequency Features with application in audio synthesis\nA Marafioti, N Holighaus, N Perraudin, P Majdak - arXiv preprint arXiv:1902.04072, 2019\n\n
Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Towards stationary time-vertex signal processing
abstract
Graph-based methods for signal processing have shown promise for the analysis of data exhibiting irregular structure, such as those found in social, transportation, and sensor networks. Yet, though these systems are often dynamic, state-of-the-art methods for graph signal processing ignore the time dimension. To address this shortcoming, this paper considers the statistical analysis of time-varying graph signals. We introduce a novel definition of joint (time-vertex) stationarity, which generalizes the classical definition of time stationarity and the recent definition appropriate for graphs. This gives rise to a scalable Wiener optimization framework for denoising, semi-supervised learning, or more generally inverting a linear operator, that is provably optimal. Experimental results on real weather data demonstrate that taking into account graph and time dimensions jointly can yield significant accuracy improvements in the reconstruction effort.
Nathanaël Perraudin, Andreas Loukas, Francesco Grassi, Pierre Vandergheynst
ICASSP1
2016 PCA using graph total variation
abstract
Mining useful clusters from high dimensional data has received significant attention of the signal processing and machine learning community in the recent years. Linear and non-linear dimensionality reduction has played an important role to overcome the curse of dimensionality. However, often such methods are accompanied with problems such as high computational complexity (usually associated with the nuclear norm minimization), non-convexity (for matrix factorization methods) or susceptibility to gross corruptions in the data. In this paper we propose a convex, robust, scalable and efficient Principal Component Analysis (PCA) based method to approximate the low-rank representation of high dimensional datasets via a two-way graph regularization scheme. Compared to the exact recovery methods, our method is approximate, in that it enforces a piecewise constant assumption on the samples using a graph total variation and a piecewise smoothness assumption on the features using a graph Tikhonov regularization. Futhermore, it retrieves the low-rank representation in a time that is linear in the number of data samples. Clustering experiments on 3 benchmark datasets with different types of corruptions show that our proposed model outperforms 7 state-of-the-art dimensionality reduction models.
Nauman Shahid, Nathanaël Perraudin, Vassilis Kalofolias, Benjamin Ricaud, Pierre Vandergheynst
ICASSP2