Valentin Emiya

dblp:67/7645 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0001-7102-6943ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
3 papers
Information theory · 75% Mathematical optimization · 25%
Computer graphics and multimedia
4 papers
Audio and music processing · 64% Image and video processing · 23% Multimedia systems and quality of experience · 13%
Artificial intelligence
1 paper
Optimization for machine learning · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information theory › signal processing › compressed sensing
sparse recovery
1.122024
Straight-Through Meets Sparse Recovery: the Support Exploration Algorithm · ICML 2024
Recovery and Convergence Rate of the Frank-Wolfe Algorithm for the m-Exact-Sparse Problem · IEEE Trans. Inf. Theory 2019
Machine learning › Optimization for machine learning › gradient estimation
straight-through estimator
0.812024
Straight-Through Meets Sparse Recovery: the Support Exploration Algorithm · ICML 2024
Information theory › signal processing › compressed sensing
support recovery
0.812024
Straight-Through Meets Sparse Recovery: the Support Exploration Algorithm · ICML 2024
Mathematical optimization
frank-wolfe algorithm
0.412019
Recovery and Convergence Rate of the Frank-Wolfe Algorithm for the m-Exact-Sparse Problem · IEEE Trans. Inf. Theory 2019
Audio and music processing
music transcription
0.212016
Optimal spectral transportation with application to music transcription · NIPS 2016
Image and video processing › hyperspectral image analysis
spectral unmixing
0.212016
Optimal spectral transportation with application to music transcription · NIPS 2016
Mathematical optimization
optimal transport
0.212016
Optimal spectral transportation with application to music transcription · NIPS 2016
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
audio inpainting
0.112012
Audio Inpainting · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › acoustic signal processing › audio signal reconstruction
audio restoration
0.112012
Audio Inpainting · IEEE Trans. Speech Audio Process. 2012
Multimedia systems and quality of experience
objective quality assessment
0.112011
Subjective and Objective Quality Assessment of Audio Source Separation · IEEE Trans. Speech Audio Process. 2011
Audio and music processing
source separation
0.112011
Subjective and Objective Quality Assessment of Audio Source Separation · IEEE Trans. Speech Audio Process. 2011
Audio and music processing › music transcription
multipitch estimation
0.112010
Multipitch Estimation of Piano Sounds Using a New Probabilistic Spectral Smoothness Principle · IEEE Trans. Speech Audio Process. 2010
Image and video processing
sparse representation
0.012012
Audio Inpainting · IEEE Trans. Speech Audio Process. 2012
Multimedia systems and quality of experience › subjective quality assessment
subjective listening test
0.012011
Subjective and Objective Quality Assessment of Audio Source Separation · IEEE Trans. Speech Audio Process. 2011
Audio and music processing › music transcription
piano transcription
0.012010
Multipitch Estimation of Piano Sounds Using a New Probabilistic Spectral Smoothness Principle · IEEE Trans. Speech Audio Process. 2010

Methods — techniques the papers use, named apart from their topics

straight-through estimator · 1.5restricted isometry property · 1.5branch-and-bound · 1.5optimal transportation · 0.5nonnegative matrix factorization · 0.5convergence rate analysis · 0.4coherence analysis · 0.4sparse representation · 0.1orthogonal matching pursuit · 0.1matching pursuit · 0.1PEMO-Q perceptual salience measure · 0.1moving-average noise model · 0.1maximum likelihood · 0.1autoregressive spectral envelope model · 0.1
YearPublicationVenuePosition
2025 Learning Permutations in Monarch Factorization
abstract
In order to reduce the quadratic cost of matrix-vector multiplications in dense and attention layers, Monarch matrices have been recently introduced, achieving a sub-quadratic complexity. It consists in factorizing a matrix using fixed permutations and learned block diagonal matrices, at the price of a small performance drop. We propose a more general model where some permutations are learned. The optimization algorithm explores the space of permutations using a Straight-Through Estimator (STE) inspired by the support exploration algorithm designed for sparse support recovery. Our experimental results demonstrate performance improvement in the context of sparse matrix factorization and of end-to-end sparse learning.
Mimoun Mohamed, Valentin Emiya, Caroline Chaux
ICASSP2
2024 Straight-Through Meets Sparse Recovery: the Support Exploration Algorithm
abstract
The *straight-through estimator* (STE) is commonly used to optimize quantized neural networks, yet its contexts of effective performance are still unclear despite empirical successes. To make a step forward in this comprehension, we apply STE to a well-understood problem: *sparse support recovery*. We introduce the *Support Exploration Algorithm* (SEA), a novel algorithm promoting sparsity, and we analyze its performance in support recovery (a.k.a. model selection) problems. SEA explores more supports than the state-of-the-art, leading to superior performance in experiments, especially when the columns of $A$ are strongly coherent. The theoretical analysis considers recovery guarantees when the linear measurements matrix $A$ satisfies the *Restricted Isometry Property* (RIP). The sufficient conditions of recovery are comparable but more stringent than those of the state-of-the-art in sparse support recovery. Their significance lies mainly in their applicability to an instance of the STE.
Mimoun Mohamed, François Malgouyres, Valentin Emiya, Caroline Chaux
ICML3
2021 QuicK-means: accelerating inference for K-means by learning fast transforms
Luc Giffon, Valentin Emiya, Hachem Kadri, Liva Ralaivola
Mach. Learn.2
2020 Filtering Out Time-Frequency Areas Using Gabor Multipliers
abstract
We address the problem of filtering out localized time-frequency components in signals. The problem is formulated as a minimization of a suitable quadratic form, that involves a data fidelity term on the short-time Fourier transform outside the support of the undesired component, and an energy penalization term inside the support. The minimization yields a linear system whose solution can be expressed in closed form using Gabor multipliers. We provide an analysis of the solution and its approximation by truncated eigenvalue expansion, and illustrate its performance on synthetic mixtures of audio signals. The proposed approach outperforms approaches that are routinely used in applications.
A. Marina Krémé, Valentin Emiya, Caroline Chaux, Bruno Torrésani
ICASSP2
2019 Recovery and Convergence Rate of the Frank-Wolfe Algorithm for the m-Exact-Sparse Problem
abstract
We study the properties of the Frank-Wolfe algorithm to solve the m-EXACT-SPARSE reconstruction problem, where a signal y must be expressed as a sparse linear combination of a predefined set of atoms, called dictionary. We prove that when the signal is sparse enough with respect to the coherence of the dictionary, then the iterative process implemented by the Frank-Wolfe algorithm only recruits atoms from the support of the signal, is the smallest set of atoms from the dictionary that allows for a perfect reconstruction of y. We also prove that under this same condition, there exists an iteration beyond which the algorithm converges exponentially.
Farah Cherfaoui, Valentin Emiya, Liva Ralaivola, Sandrine Anthoine
IEEE Trans. Inf. Theory2
2018 Being Low-Rank in the Time-Frequency Plane
abstract
When using optimization methods with matrix variables in signal processing and machine learning, it is customary to assume some low-rank prior on the targeted solution. Nonnegative matrix factorization of spectrograms is a case in point in audio signal processing. However, this low-rank prior is not straightforwardly related to complex matrices obtained from a short-time Fourier - or discrete Gabor - transform (STFT), which is generally defined from and studied based on a modulation operator and a translation operator applied to a so-called window. This paper is a first study of the low-rankness property of time-frequency matrices. We characterize the set of signals with a rank- r (complex) STFT matrix in the case of a unit hop size and frequency step with few assumptions on the transform parameters. We discuss the scope of this result and its implications on low-rank approximations of STFT matrices.
Valentin Emiya, Ronan Hamon, Caroline Chaux
ICASSP1
2018 Sparse Non-Local Similarity Modeling for Audio Inpainting
abstract
Audio signals are highly structured from a low, signal level to high cognitive aspects. We investigate how to exploit the common sparse structure between similar audio frames in order to reconstruct missing data in audio signals. While joint sparse models and related algorithms have been widely studied, one important challenge is to locate such similar frames: the search must be adapted to the joint-sparse model and should be fast and one must deal with missing data in the frames. We propose, compare and discuss several similarity measures dedicated to this task. We then show how this strategy can lead to better reconstruction of missing data in audio signals.
Ichrak Toumi, Valentin Emiya
ICASSP2
2017 Assessment of musical noise using localization of isolated peaks in time-frequency domain
abstract
Musical noise is a recurrent issue that appears in spectral techniques for denoising or blind source separation. Due to localised errors of estimation, isolated peaks may appear in the processed spectrograms, resulting in annoying tonal sounds after synthesis known as “musical noise”. In this paper, we propose a method to assess the amount of musical noise in an audio signal, by characterising the impact of these artificial isolated peaks on the processed sound. It turns out that because of the constraints between STFT coefficients, the isolated peaks are described as time-frequency “spots” in the spectrogram of the processed audio signal. The quantification of these “spots”, achieved through the adaptation of a method for localisation of significant STFT regions, allows for an evaluation of the amount of musical noise. We believe that this will pave the way to an objective measure and a better understanding of this phenomenon.
Ronan Hamon, Valentin Emiya, Lucas Rencker, Wenwu Wang 0001, Mark D. Plumbley
ICASSP2
2016 Optimal spectral transportation with application to music transcription
abstract
Many spectral unmixing methods rely on the non-negative decomposition of spectral data onto a dictionary of spectral templates. In particular, state-of-the-art music transcription systems decompose the spectrogram of the input signal onto a dictionary of representative note spectra. The typical measures of fit used to quantify the adequacy of the decomposition compare the data and template entries frequency-wise. As such, small displacements of energy from a frequency bin to another as well as variations of timber can disproportionally harm the fit. We address these issues by means of optimal transportation and propose a new measure of fit that treats the frequency distributions of energy holistically as opposed to frequency-wise. Building on the harmonic nature of sound, the new measure is invariant to shifts of energy to harmonically-related frequencies, as well as to small and local displacements of energy. Equipped with this new measure of fit, the dictionary of note templates can be considerably simplified to a set of Dirac vectors located at the target fundamental frequencies (musical pitch values). This in turns gives ground to a very fast and simple decomposition algorithm that achieves state-of-the-art performance on real musical data.
Rémi Flamary, Cédric Févotte, Nicolas Courty, Valentin Emiya
NIPS4
2014 Compressed sensing with unknown sensor permutation
abstract
Compressed sensing is the ability to retrieve a sparse vector from a set of linear measurements. The task gets more difficult when the sensing process is not perfectly known. We address such a problem in the case where the sensors have been permuted, i.e., the order of the measurements is unknown. We propose a branch-and-bound algorithm that converges to the solution. The experimental study shows that our approach always retrieves the unknown permutation, while a simple convex relaxation strategy almost always fails. In terms of its time complexity, we show that the proposed algorithm converges quickly with respect to the combinatorial nature of the problem.
Valentin Emiya, Antoine Bonnefoy, Laurent Daudet, Rémi Gribonval
ICASSP1
2012 Sparse underwater acoustic imaging: A case study
abstract
Underwater acoustic imaging is traditionally performed with beamforming: beams are formed at emission to insonify limited angular regions; beams are (synthetically) formed at reception to form the image. We propose to exploit a natural sparsity prior to perform 3D underwater imaging using a newly built flexible-configuration sonar device. The computational challenges raised by the high-dimensionality of the problem are highlighted, and we describe a strategy to overcome them. As a proof of concept, the proposed approach is used on real data acquired with the new sonar to obtain an image of an underwater target. We discuss the merits of the obtained image in comparison with standard beamforming, as well as the main challenges lying ahead, and the bottlenecks that will need to be solved before sparse methods can be fully exploited in the context of underwater compressed 3D sonar imaging.
Nikolaos Stefanakis, Jacques Marchal, Valentin Emiya, Nancy Bertin, Rémi Gribonval, Pierre Cervenka
ICASSP3
2012 Audio Inpainting
abstract
We propose the audio inpainting framework that recovers portions of audio data distorted due to impairments such as impulsive noise, clipping, and packet loss. In this framework, the distorted data are treated as missing and their location is assumed to be known. The signal is decomposed into overlapping time-domain frames and the restoration problem is then formulated as an inverse problem per audio frame. Sparse representation modeling is employed per frame, and each inverse problem is solved using the Orthogonal Matching Pursuit algorithm together with a discrete cosine or a Gabor dictionary. The Signal-to-Noise Ratio performance of this algorithm is shown to be comparable or better than state-of-the-art methods when blocks of samples of variable durations are missing. We also demonstrate that the size of the block of missing samples, rather than the overall number of missing samples, is a crucial parameter for high quality signal restoration. We further introduce a constrained Matching Pursuit approach for the special case of audio declipping that exploits the sign pattern of clipped audio samples and their maximal absolute value, as well as allowing the user to specify the maximum amplitude of the signal. This approach is shown to outperform state-of-the-art and commercially available methods for audio declipping in terms of Signal-to-Noise Ratio.
Amir Adler, Valentin Emiya, Maria G. Jafari, Michael Elad, Rémi Gribonval, Mark D. Plumbley
IEEE Trans. Speech Audio Process.2
2011 A constrained matching pursuit approach to audio declipping
abstract
We present a novel sparse representation based approach for the restoration of clipped audio signals. In the proposed approach, the clipped signal is decomposed into overlapping frames and the declipping problem is formulated as an inverse problem, per audio frame. This problem is further solved by a constrained matching pursuit algorithm, that exploits the sign pattern of the clipped samples and their maximal absolute value. Performance evaluation with a collection of music and speech signals demonstrate superior results compared to existing algorithms, over a wide range of clipping levels.
Amir Adler, Valentin Emiya, Maria G. Jafari, Michael Elad, Rémi Gribonval, Mark D. Plumbley
ICASSP2
2011 Subjective and Objective Quality Assessment of Audio Source Separation
abstract
We aim to assess the perceived quality of estimated source signals in the context of audio source separation. These signals may involve one or more kinds of distortions, including distortion of the target source, interference from the other sources or musical noise artifacts. We propose a subjective test protocol to assess the perceived quality with respect to each kind of distortion and collect the scores of 20 subjects over 80 sounds. We then propose a family of objective measures aiming to predict these subjective scores based on the decomposition of the estimation error into several distortion components and on the use of the PEMO-Q perceptual salience measure to provide multiple features that are then combined. These measures increase correlation with subjective scores up to 0.5 compared to nonlinear mapping of individual state-of-the-art source separation measures. Finally, we released the data and code presented in this paper in a freely available toolkit called PEASS.
Valentin Emiya, Emmanuel Vincent 0001, Niklas Harlander, Volker Hohmann
IEEE Trans. Speech Audio Process.1
2010 Multipitch Estimation of Piano Sounds Using a New Probabilistic Spectral Smoothness Principle
abstract
A new method for the estimation of multiple concurrent pitches in piano recordings is presented. It addresses the issue of overlapping overtones by modeling the spectral envelope of the overtones of each note with a smooth autoregressive model. For the background noise, a moving-average model is used and the combination of both tends to eliminate harmonic and sub-harmonic erroneous pitch estimations. This leads to a complete generative spectral model for simultaneous piano notes, which also explicitly includes the typical deviation from exact harmonicity in a piano overtone series. The pitch set which maximizes an approximate likelihood is selected from among a restricted number of possible pitch combinations as the one. Tests have been conducted on a large homemade database called MAPS, composed of piano recordings from a real upright piano and from high-quality samples.
Valentin Emiya, Roland Badeau, Bertrand David 0002
IEEE Trans. Speech Audio Process.1
2009 Expectation-maximization algorithm for multi-pitch estimation and separation of overlapping harmonic spectra
abstract
This paper addresses the problem of multi-pitch estimation, which consists in estimating the fundamental frequencies of multiple harmonic sources, with possibly overlapping partials, from their mixture. The proposed approach is based on the expectation-maximization algorithm, which aims at maximizing the likelihood of the observed spectrum, by performing successive single-pitch and spectral envelope estimations. This algorithm is illustrated in the context of musical chord identification.
Roland Badeau, Valentin Emiya, Bertrand David 0002
ICASSP2
2007 A Parametric Method for Pitch Estimation of Piano Tones
abstract
The efficiency of most pitch estimation methods declines when the analyzed frame is shortened and/or when a wide fundamental frequency (Fo) range is targeted. The technique proposed herein jointly uses a periodicity analysis and a spectral matching process to improve the foestimation performance in such an adverse context: a 60 ms-long data frame together with the whole, 71/4-octaves, piano tessitura. The enhancements are obtained thanks to a parametric approach which, among other things, models the inharmonicity of piano tones. The performance of the algorithm is assessed, is compared to the results obtained from other estimators and is discussed in order to characterize their behavior and typical misestimations.
Valentin Emiya, Bertrand David 0002, Roland Badeau
ICASSP (1)1