Stefan Lattner

dblp:155/7947 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-3945-7580ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Adversarial Aesthetics: Using Spectrogram Steganography to Interrogate AI Ethics in Electronic Music
abstract
Audio steganography—the practice of concealing information within sound—has emerged as a creative tool in electronic music production, with artists embedding visual data in the time-frequency domain that is audible as texture but reveals itself clearly only when rendered as a spectrogram. Artists are increasingly adopting steganographic techniques, known as data poisoning, to prevent their works being used to train AI without their consent. By embedding imperceptible adversarial noise directly into their work, artists can corrupt the feature-extraction process of unauthorised AI model development, while the data remains perceptually unchanged to humans. This work, through the medium of two audio-visual artworks incorporating steganographic processes, places a critical lens on the socio-cultural impacts of technological development and its historically undulating relationship with electronic music. Through this, we hope to promote discussions on responsible technology development that recognises the value of human creativity, an imperative dialogue as data-driven AI continues its rapid expansion.
Alexander J. Williams, Yutian Hu, Stefan Lattner, Mathieu Barthet
Creativity & Cognition3
2025 Estimating Musical Surprisal in Audio
abstract
In modeling musical surprisal expectancy with computational methods, it has been proposed to use the information content (IC) of one-step predictions from an autoregressive model as a proxy for surprisal in symbolic music. With an appropriately chosen model, the IC of musical events has been shown to correlate with human perception of surprise and complexity aspects, including tonal and rhythmic complexity. This work investigates whether an analogous methodology can be applied to music audio. We train an autoregressive Transformer model to predict compressed latent audio representations of a pretrained autoencoder network. We verify learning effects by estimating the decrease in IC with repetitions. We investigate the mean IC of musical segment types (e.g., A or B) and find that segment types appearing later in a piece have a higher IC than earlier ones on average. We investigate the IC’s relation to audio and musical features and find it correlated with timbral variations and loudness and, to a lesser extent, dissonance, rhythmic complexity, and onset density related to audio and musical features. Finally, we investigate if the IC can predict EEG responses to songs and thus model humans’ surprisal in music. We provide code for our method on github.com/sonycslparis/audioic.
Mathias Rose Bjare, Giorgia Cantisani, Stefan Lattner, Gerhard Widmer
ICASSP3
2025 Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
abstract
Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoencoders, however, often struggle to achieve high compression ratios while preserving audio fidelity and facilitating efficient downstream applications. We introduce Music2Latent2, a novel audio autoencoder that addresses these limitations by leveraging consistency models and a novel approach to representation learning based on unordered latent embeddings, which we call summary embeddings. Unlike conventional methods that encode local audio features into ordered sequences, Music2Latent2 compresses audio signals into sets of summary embeddings, where each embedding can capture distinct global features of the input sample. This enables to achieve higher reconstruction quality at the same compression ratio. To handle arbitrary audio lengths, Music2Latent2 employs an autoregressive consistency model trained on two consecutive audio chunks with causal masking, ensuring coherent reconstruction across segment boundaries. Additionally, we propose a novel two-step decoding procedure that leverages the denoising capabilities of consistency models to further refine the generated audio at no additional cost. Our experiments demonstrate that Music2Latent2 outperforms existing continuous audio autoencoders regarding audio quality and performance on downstream tasks. Music2Latent2 paves the way for new possibilities in audio compression.
Marco Pasini, Stefan Lattner, György Fazekas
ICASSP2
2025 Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
abstract
In this paper, we tackle the task of musical stem retrieval. Given a musical mix, it consists in retrieving a stem that would fit with it, i.e., that would sound pleasant if played together. To do so, we introduce a new method based on Joint-Embedding Predictive Architectures, where an encoder and a predictor are jointly trained to produce latent representations of a context and predict latent representations of a target. In particular, we design our predictor to be conditioned on arbitrary instruments, enabling our model to perform zero-shot stem retrieval. In addition, we discover that pretraining the encoder using contrastive learning drastically improves the model’s performance.We validate the retrieval performances of our model using the MUSDB18 and MoisesDB datasets. We show that it significantly out-performs previous baselines on both datasets, showcasing its ability to support more or less precise (and possibly unseen) conditioning. We also evaluate the learned embeddings on a beat tracking task, demonstrating that they retain temporal structure and local information.
Alain Riou, Antonin Gagneré, Gaëtan Hadjeres, Stefan Lattner, Geoffroy Peeters
ICASSP4
2025 Hybrid Losses for Hierarchical Embedding Learning
abstract
In traditional supervised learning, the cross-entropy loss treats all incorrect predictions equally, ignoring the relevance or proximity of wrong labels to the correct answer. By leveraging a tree hierarchy for fine-grained labels, we investigate hybrid losses, such as generalised triplet and cross-entropy losses, to enforce similarity between labels within a multi-task learning framework. We propose metrics to evaluate the embedding space structure and assess the model’s ability to generalise to unseen classes, that is, to infer similar classes for data belonging to unseen categories. Our experiments on OrchideaSOL, a four-level hierarchical instrument sound dataset with nearly 200 detailed categories, demonstrate that the proposed hybrid losses outperform previous works in classification, retrieval, embedding space structure, and generalisation.
Haokun Tian, Stefan Lattner, Brian McFee, Charalampos Saitis
ICASSP2
2025 Temporal Considerations in DJ Mix Information Retrieval and Generation (Short Paper)
Alexander J. Williams, Gregor Meehan, Stefan Lattner, Johan Pauwels, Mathieu Barthet
TIME3
2024 Bass Accompaniment Generation Via Latent Diffusion
abstract
The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the core of our method are audio autoencoders that efficiently compress audio waveform samples into invertible latent representations, and a conditional latent diffusion model that takes as input the latent encoding of a mix and generates the latent encoding of a corresponding stem. To provide control over the timbre of generated samples, we introduce a technique to ground the latent space to a user-provided reference style during diffusion sampling. For further improving audio quality, we adapt classifier-free guidance to avoid distortions at high guidance strengths when generating an unbounded latent space. We train our model on a dataset of pairs of mixes and matching bass stems. Quantitative experiments demonstrate that, given an input mix, the proposed system can generate basslines with user-specified timbres. Our controllable conditional audio generation framework represents a significant step forward in creating generative AI tools to assist musicians in music production.
Marco Pasini, Maarten Grachten, Stefan Lattner
ICASSP3
2015 A Computational Approach to Modelling the Perception of Pitch and Tonality in Music
Kat Agres, Carlos Eduardo Cancino-Chacón, Maarten Grachten, Stefan Lattner
CogSci4
2015 Pseudo-Supervised Training Improves Unsupervised Melody Segmentation
Stefan Lattner, Carlos Eduardo Cancino-Chacón, Maarten Grachten
IJCAI1