Nancy Bertin

dblp:68/6673 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0002-7690-4378ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-authorArtificial intelligence and machine learning · 9 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 87% Multimedia systems and quality of experience · 13%
Artificial intelligence
1 paper
Motion planning and robot control · 67% Speech recognition and synthesis · 33%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › acoustic signal processing › audio signal reconstruction
audio restoration
0.512021
Sparsity-Based Audio Declipping Methods: Selected Overview, New Algorithms, and Large-Scale Evaluation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Robotics › Motion planning and robot control › robot control
feedback control
0.212016
First applications of sound-based control on a mobile robot equipped with two microphones · ICRA 2016
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition
0.212016
First applications of sound-based control on a mobile robot equipped with two microphones · ICRA 2016
Robotics › Motion planning and robot control
robot control
0.212016
First applications of sound-based control on a mobile robot equipped with two microphones · ICRA 2016
Multimedia systems and quality of experience › subjective quality assessment
subjective listening test
0.112021
Sparsity-Based Audio Declipping Methods: Selected Overview, New Algorithms, and Large-Scale Evaluation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › music transcription
multipitch estimation
0.112010
Adaptive Harmonic Spectral Decomposition for Multiple Pitch Estimation · IEEE Trans. Speech Audio Process. 2010
Audio and music processing
music information retrieval
0.112010
Enforcing Harmonicity and Smoothness in Bayesian Non-Negative Matrix Factorization Applied to Polyphonic Music Transcription · IEEE Trans. Speech Audio Process. 2010
Audio and music processing
music transcription
0.112010
Enforcing Harmonicity and Smoothness in Bayesian Non-Negative Matrix Factorization Applied to Polyphonic Music Transcription · IEEE Trans. Speech Audio Process. 2010
Audio and music processing › speech processing
pitch estimation
0.112010
Adaptive Harmonic Spectral Decomposition for Multiple Pitch Estimation · IEEE Trans. Speech Audio Process. 2010
Audio and music processing › music transcription
polyphonic music transcription
0.112010
Enforcing Harmonicity and Smoothness in Bayesian Non-Negative Matrix Factorization Applied to Polyphonic Music Transcription · IEEE Trans. Speech Audio Process. 2010

Methods — techniques the papers use, named apart from their topics

time-frequency sparsity · 0.5synthesis sparsity · 0.5structured sparsity · 0.5analysis sparsity · 0.5time difference of arrival · 0.2auditory servoing · 0.2space-alternating generalized expectation-maximization · 0.1nonnegative matrix factorization · 0.1multiplicative harmonic NMF · 0.1bayesian non-negative matrix factorization · 0.1adaptive basis estimation · 0.1
YearPublicationVenuePosition
2021 Sparsity-Based Audio Declipping Methods: Selected Overview, New Algorithms, and Large-Scale Evaluation
abstract
Recent advances in audio declipping have substantially improved the state of the art. Yet, practitioners need guidelines to choose a method, and while existing benchmarks have been instrumental in advancing the field, larger-scale experiments are needed to guide such choices. First, we show that the clipping levels in existing small-scale benchmarks are moderate and call for benchmarks with more perceptually significant clipping levels. We then propose a general algorithmic framework for declipping that covers existing and new combinations of variants of state-of-the-art techniques exploiting time-frequency sparsity: synthesisvs.analysis sparsity, with plain or structured sparsity. Finally, we systematically compare these combinations and a selection of state-of-the-art methods. Using a large-scale numerical benchmark and a smaller scale formal listening test, we provide guidelines for various clipping levels, both for speech and various musical genres. The code is made publicly available for the purpose of reproducible research and benchmarking.
Clément Gaultier, Srdan Kitic, Rémi Gribonval, Nancy Bertin
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Blaster: An Off-Grid Method for Blind and Regularized Acoustic Echoes Retrieval
abstract
Acoustic echoes retrieval is a research topic that is gaining importance in many speech and audio signal processing applications such as speech enhancement, source separation, dereverberation and room geometry estimation. This work proposes a novel approach to blindly retrieve the off-grid timing of early acoustic echoes from a stereophonic recording of an unknown sound source such as speech. It builds on the recent framework of continuous dictionaries. In contrast with existing methods, the proposed approach does not rely on parameter tuning nor peak picking techniques by working directly in the parameter space of interest. The accuracy and robustness of the method are assessed on challenging simulated setups with varying noise and reverberation levels and are compared to two state-of-the-art methods.
Diego Di Carlo, Clement Elvira, Antoine Deleforge, Nancy Bertin, Rémi Gribonval
ICASSP4
2019 Mirage: 2D Source Localization Using Microphone Pair Augmentation with Echoes
abstract
It is commonly observed that acoustic echoes hurt per mance of sound source localization (SSL) methods. We troduce the concept of microphone array augmentation echoes (MIRAGE) and show how estimation of early-e characteristics can in fact benefit SSL. We propose a learn based scheme for echo estimation combined with a phys based scheme for echo aggregation. In a simple scenario volving 2 microphones close to a reflective surface and source, we show using simulated data that the proposed proach performs similarly to a correlation-based metho azimuth estimation while retrieving elevation as well from 2 microphones only, an impossible task in anechoic settings.
Diego Di Carlo, Antoine Deleforge, Nancy Bertin
ICASSP3
2019 VoiceHome-2, an extended corpus for multichannel speech processing in real homes
Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent 0001, Sunit Sivasankaran, Irina Illina, Frédéric Bimbot
Speech Commun.1
2018 Cascade: Channel-Aware Structured Cosparse Audio Declipper
abstract
This work features a new algorithm, CASCADE, which leverages a structured cosparse prior across channels to address the multichannel audio declipping problem. CASCADE technique outperforms the state-of-the-art method A-SPADE applied on each channel separately in all tested settings, while retaining similar runtime.
Clément Gaultier, Nancy Bertin, Rémi Gribonval
ICASSP2
2018 Aural Servo: Sensor-Based Control From Robot Audition
abstract
This paper proposes a control framework based on auditory perception. Generally, in robot audition, the motion control of a robot from the sense of hearing relies on sound source localization. We propose in this paper an alternative approach, aural servo, which is derived from the sensor-based control framework. In this approach, robot motions are directly connected to the aural perception: The variation of low-level auditory features dictates the motions applied to the robot through a feedback loop. It has the advantage of being robust to spurious measurements and modeling approximations for a low computational cost. This paper presents the theoretical concept of the aural servo framework. Besides a theoretical analysis, the aural servo framework is validated through several experiments on different robotic platforms and under real-world conditions.
Aly Magassouba, Nancy Bertin, François Chaumette
IEEE Trans. Robotics2
2016 Joint estimation of sound source location and boundary impedance with physics-driven cosparse regularization
abstract
Indoor acoustic source localization can be efficiently performed by modeling the sound propagation in the room, and by solving the arising inverse problem by means of cosparse regularization and convex optimization techniques. However, previous methods relying on this approach used to assume the knowledge of a number of room characteristics: its geometry, the walls' absorption or reflexion properties, as well as the speed of sound. In this paper, we show that this model, and the corresponding algorithms, can be extended to the case where the specific acoustic impedance of the boundary is unknown. The proposed method allows to jointly estimate the boundary impedances and the sound pressure in the room, without any preliminary calibration phase, from the only knowledge of the room geometry. Validated on simulation, this new algorithm constitutes a important step towards practical applicability of sound field cosparse modeling.
Nancy Bertin, Srdan Kitic, Rémi Gribonval
ICASSP1
2016 Membrane shape and boundary conditions estimation using eigenmode decomposition
abstract
This paper investigates the problem of estimating the shape or the boundary impedance of a vibrating membrane from acoustic measurements in a limited sub-domain of the membrane. In acoustics, polygonal room shapes are usually estimated through room impulse response measurements. Impedance values of materials are, in turn, often calculated from the measurement of the acoustic reflection coefficients at the boundaries. In this work, we develop an alternative frequency-domain method to estimate the shape of a convex membrane with generalized Robin boundary conditions, from the measurement of its eigenmodes on a small portion of its surface. Reciprocally, we show that the same model allows to estimate the membrane borders' impedances when its shape is known.
Thibault Nowakowski, Nancy Bertin, Rémi Gribonval, Julien de Rosny, Laurent Daudet
ICASSP2
2016 First applications of sound-based control on a mobile robot equipped with two microphones
abstract
This paper validates experimentally a novel approach to robot audition, sound-based control, which consists in introducing auditory features directly as inputs of a closed-loop control scheme, that is, without any explicit localization process. The applications we present rely on the implicit bearings of the sound sources computed from the time difference of arrival (TDOA) between two microphones. By linking the motion of the robot to the aural perception of the environment, this approach has the benefit of being more robust to reverberation and noise. Therefore neither complex tracking method such as Kalman filtering nor TDOA enhancement with denoising or dereverberation methods are needed to track the correct TDOA measurements. The experiments conducted on a mobile robot instrumented with a pair of microphones show the validity of our approach. In a reverberating and noisy room, this approach is able to orient the robot to a mobile sound source in real time. A positioning task with respect to two sound sources is also performed while the robot perception is disturbed by altered and spurious TDOA measurements.
Aly Magassouba, Nancy Bertin, François Chaumette
ICRA2
2016 A French Corpus for Distant-Microphone Speech Processing in Real Homes
abstract
International audience
Nancy Bertin, Ewen Camberlein, Emmanuel Vincent 0001, Romain Lebarbenchon, Stéphane Peillon, Éric Lamande, Sunit Sivasankaran, Frédéric Bimbot, Irina Illina, Ariane Tom, Sylvain Fleury, Eric Jamet
INTERSPEECH1
2016 Audio-based robot control from interchannel level difference and absolute sound energy
abstract
This paper is a follow-up way to our previous works regarding audio-based control, that is an alternative method for auditory-based robot tasks. Conversely to classic methods oriented towards sound source localization, audio-based control is a sensor-based framework that does not localize the sound source. Instead, auditory features are used as inputs of a closed-loop control scheme. The audio-based control method presented in this paper relies on the sound signal energy measured by two microphones. By combining the interchannel level difference to the acoustic absolute energy level, the control scheme allows positioning the robot with respect to the sound source at a given distance and orientation. Moreover this method has the benefit of a low computation cost, since it only relies on the signal energy measurement. Experimental results conducted on a mobile robot validate the relevance and the robustness of this approach in dynamic and real world conditions.
Aly Magassouba, Nancy Bertin, François Chaumette
IROS2
2015 Sound-based control with two microphones
abstract
This paper presents a novel approach to robot audition by performing robotic tasks with auditory cues. Unlike many previous works, we propose a control scheme, that does not require any explicit sound source localization. This approach is capable of controlling all the three degrees of freedom in a plane from two microphones. Built upon the sensor-based control framework, this approach relies on implicit sound source direction obtained from the time difference of arrival (TDOA). We introduce an analytical modelling of auditory cues considering a robot equipped with a pair of microphones and multiple sound sources, from which a control scheme is designed. A stability analysis is provided as well. The results obtained in simulation show the feasibility and the suitability of this method even in reverberant area.
Aly Magassouba, Nancy Bertin, François Chaumette
IROS2
2014 Hearing behind walls: Localizing sources in the room next door with cosparsity
abstract
Acoustic source localization is traditionally performed using cues such as interchannel time of arrival and intensity differences to infer the geometric localization of emitting sources with respect to the receiving microphone array. However the presence of obstacles between the sources and the array makes it impossible to rely on the direct path, and more advanced techniques are needed. The huge body of work on sparse recovery suggests an approach where source localization is expressed as a linear inverse problem and the spatial sparsity of the sources is exploited. An inverse problem can be naturally expressed in the recently introduced cosparse framework, exploiting the fact that the acoustic pressure satisfies the homogeneous wave equation except in the few locations of the sources. The resulting optimization problem involves a discretized second derivative analysis operator, which is extremely sparse. In this paper, we demonstrate the performance of the cosparse approach on an extreme source localization problem, where the microphone array is installed in the room next door to the room where the emitting sources are located, somehow hearing behind a wall.
Srdan Kitic, Nancy Bertin, Rémi Gribonval
ICASSP2
2012 Sparse underwater acoustic imaging: A case study
abstract
Underwater acoustic imaging is traditionally performed with beamforming: beams are formed at emission to insonify limited angular regions; beams are (synthetically) formed at reception to form the image. We propose to exploit a natural sparsity prior to perform 3D underwater imaging using a newly built flexible-configuration sonar device. The computational challenges raised by the high-dimensionality of the problem are highlighted, and we describe a strategy to overcome them. As a proof of concept, the proposed approach is used on real data acquired with the new sonar to obtain an image of an underwater target. We discuss the merits of the obtained image in comparison with standard beamforming, as well as the main challenges lying ahead, and the bottlenecks that will need to be solved before sparse methods can be fully exploited in the context of underwater compressed 3D sonar imaging.
Nikolaos Stefanakis, Jacques Marchal, Valentin Emiya, Nancy Bertin, Rémi Gribonval, Pierre Cervenka
ICASSP4
2011 Stability analysis of multiplicative update algorithms for non-negative matrix factorization
abstract
Multiplicative update algorithms have encountered a great success to solve optimization problems with non-negativity constraints, such as the famous non-negative matrix factorization (NMF) and its many variants. However, despite several years of research on the topic, the understanding of their convergence properties is still to be improved. In this paper, we show that Lyapunov's stability theory provides a very enlightening viewpoint on the problem. We prove the stability of supervised NMF and study the more difficult case of unsupervised NMF. Numerical simulations illustrate those theoretical results, and the convergence speed of NMF multiplicative updates is analyzed.
Roland Badeau, Nancy Bertin, Emmanuel Vincent 0001
ICASSP2
2010 Enforcing Harmonicity and Smoothness in Bayesian Non-Negative Matrix Factorization Applied to Polyphonic Music Transcription
abstract
This paper presents theoretical and experimental results about constrained non-negative matrix factorization (NMF) in a Bayesian framework. A model of superimposed Gaussian components including harmonicity is proposed, while temporal continuity is enforced through an inverse-Gamma Markov chain prior. We then exhibit a space-alternating generalized expectation-maximization (SAGE) algorithm to estimate the parameters. Computational time is reduced by initializing the system with an original variant of multiplicative harmonic NMF, which is described as well. The algorithm is then applied to perform polyphonic piano music transcription. It is compared to other state-of-the-art algorithms, especially NMF-based. Convergence issues are also discussed on a theoretical and experimental point of view. Bayesian NMF with harmonicity and temporal continuity constraints is shown to outperform other standard NMF-based transcription systems, providing a meaningful mid-level representation of the data. However, temporal smoothness has its drawbacks, as far as transients are concerned in particular, and can be detrimental to transcription performance when it is the only constraint used. Possible improvements of the temporal prior are discussed.
Nancy Bertin, Roland Badeau, Emmanuel Vincent 0001
IEEE Trans. Speech Audio Process.1
2010 Adaptive Harmonic Spectral Decomposition for Multiple Pitch Estimation
abstract
Multiple pitch estimation consists of estimating the fundamental frequencies and saliences of pitched sounds over short time frames of an audio signal. This task forms the basis of several applications in the particular context of musical audio. One approach is to decompose the short-term magnitude spectrum of the signal into a sum of basis spectra representing individual pitches scaled by time-varying amplitudes, using algorithms such as nonnegative matrix factorization (NMF). Prior training of the basis spectra is often infeasible due to the wide range of possible musical instruments. Appropriate spectra must then be adaptively estimated from the data, which may result in limited performance due to overfitting issues. In this paper, we model each basis spectrum as a weighted sum of narrowband spectra representing a few adjacent harmonic partials, thus enforcing harmonicity and spectral smoothness while adapting the spectral envelope to each instrument. We derive a NMF-like algorithm to estimate the model parameters and evaluate it on a database of piano recordings, considering several choices for the narrowband spectra. The proposed algorithm performs similarly to supervised NMF using pre-trained piano spectra but improves pitch estimation performance by 6% to 10% compared to alternative unsupervised NMF algorithms.
Emmanuel Vincent 0001, Nancy Bertin, Roland Badeau
IEEE Trans. Speech Audio Process.2
2010 Stability Analysis of Multiplicative Update Algorithms and Application to Nonnegative Matrix Factorization
abstract
Multiplicative update algorithms have proved to be a great success in solving optimization problems with nonnegativity constraints, such as the famous nonnegative matrix factorization (NMF) and its many variants. However, despite several years of research on the topic, the understanding of their convergence properties is still to be improved. In this paper, we show that Lyapunov's stability theory provides a very enlightening viewpoint on the problem. We prove the exponential or asymptotic stability of the solutions to general optimization problems with nonnegative constraints, including the particular case of supervised NMF, and finally study the more difficult case of unsupervised NMF. The theoretical results presented in this paper are confirmed by numerical simulations involving both supervised and unsupervised NMF, and the convergence speed of NMF multiplicative updates is investigated.
Roland Badeau, Nancy Bertin, Emmanuel Vincent 0001
IEEE Trans. Neural Networks2
2009 A tempering approach for Itakura-Saito non-negative matrix factorization. With application to music transcription
abstract
In this paper we are interested in non-negative matrix factorization (NMF) with the Itakura-Saito (IS) divergence. Previous work has demonstrated the relevance of this cost function for the decomposition of audio power spectrograms. This is in particular due to its scale invariance, which makes it more robust to the wide dynamics of audio, a property which is not shared by other popular costs such as the Euclidean distance or the generalized Kulback-Leibler (KL) divergence. However, while the latter two cost functions are convex, the IS divergence is not, which makes it more prone to convergence to irrelevant local minima, as observed empirically. Thus, the aim of this paper is to propose a tempering scheme that favors convergence of IS-NMF to global minima. Our algorithm is based on NMF with the beta-divergence, where the shape parameter beta acts as a temperature parameter. Results on both synthetical and music data (in a transcription context) show the relevance of our approach.
Nancy Bertin, Cédric Févotte, Roland Badeau
ICASSP1
2009 Nonnegative Matrix Factorization with the Itakura-Saito Divergence: With Application to Music Analysis
abstract
This letter presents theoretical, algorithmic, and experimental results about nonnegative matrix factorization (NMF) with the Itakura-Saito (IS) divergence. We describe how IS-NMF is underlaid by a well-defined statistical model of superimposed gaussian components and is equivalent to maximum likelihood estimation of variance parameters. This setting can accommodate regularization constraints on the factors through Bayesian priors. In particular, inverse-gamma and gamma Markov chain priors are considered in this work. Estimation can be carried out using a space-alternating generalized expectation-maximization (SAGE) algorithm; this leads to a novel type of NMF algorithm, whose convergence to a stationary point of the IS cost function is guaranteed. We also discuss the links between the IS divergence and other cost functions used in NMF, in particular, the Euclidean distance and the generalized Kullback-Leibler (KL) divergence. As such, we describe how IS-NMF can also be performed using a gradient multiplicative algorithm (a standard algorithm structure in NMF) whose convergence is observed in practice, though not proven. Finally, we report a furnished experimental comparative study of Euclidean-NMF, KL-NMF, and IS-NMF algorithms applied to the power spectrogram of a short piano sequence recorded in real conditions, with various initializations and model orders. Then we show how IS-NMF can successfully be employed for denoising and upmix (mono to stereo conversion) of an original piece of early jazz music. These experiments indicate that IS-NMF correctly captures the semantics of audio and is better suited to the representation of music signals than NMF with the usual Euclidean and KL costs.
Cédric Févotte, Nancy Bertin, Jean-Louis Durrieu
Neural Comput.2
2008 Harmonic and inharmonic Nonnegative Matrix Factorization for Polyphonic Pitch transcription
abstract
Polyphonic pitch transcription consists of estimating the onset time, duration and pitch of each note in a music signal. This task is difficult in general, due to the wide range of possible instruments. This issue has been studied using adaptive models such as Nonnegative Matrix Factorization (NMF), which describe the signal as a weighted sum of basis spectra. However basis spectra representing multiple pitches result in inaccurate transcription. To avoid this, we propose a family of constrained NMF models, where each basis spectrum is expressed as a weighted sum of narrowband spectra consisting of a few adjacent partials at harmonic or inharmonic frequencies. The model parameters are adapted via combined multiplicative and Newton updates. The proposed method is shown to outperform standard NMF on a database of piano excerpts.
Emmanuel Vincent 0001, Nancy Bertin, Roland Badeau
ICASSP2
2007 Blind Signal Decompositions for Automatic Transcription of Polyphonic Music: NMF and K-SVD on the Benchmark
abstract
This paper investigates on the behavior of two blind signal decomposition algorithms, non negative matrix factorization (NMF) and non negative K-SVD (NKSVD), in a polyphonic music transcription task. State-of-the-art transcription systems are based on a frame-by-frame, low-level approach; blind systems could be an alternative to them. Two raw but effective audio-to-MIDI systems are proposed and evaluated. Performances are similar, but in favor of NMF, which is more robust to initialization, choice of the order and computationally less costly.
Nancy Bertin, Roland Badeau, Gaël Richard
ICASSP (1)1