Jean-Marc Valin

dblp:87/6271 · DBLP profile ↗
← Back
56ranked-venue papers
21as first author
16since 2021 · last 2024
0000-0002-9883-6927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 15 first-author · 16 since 2021Artificial intelligence and machine learning · 28 · 8 first-author · 4 since 2021Systems, architecture and hardware · 13 · 3 first-authorComputer networks · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 NOLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
abstract
Speech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexity and/or require delay. The recently proposed Linear Adaptive Coding Enhancer (LACE) addresses this problem by combining DNNs with classical long-term/short-term post-filtering resulting in a causal low-complexity model. A short-coming of the LACE model is, however, that quality quickly saturates when the model size is scaled up. To mitigate this problem, we propose a novel adatpive temporal shaping module that adds high temporal resolution to the LACE model resulting in the Non-Linear Adaptive Coding Enhancer (NoLACE). We adapt NoLACE to enhance the Opus codec and show that NoLACE significantly outperforms both the Opus baseline and an enlarged LACE model at 6, 9 and 12 kb/s. We also show that LACE and NoLACE are well-behaved when used with an ASR system.
Jan Büthe, Ahmed Mustafa, Jean-Marc Valin, Karim Helwani, Michael M. Goodwin
ICASSP3
2024 Noise-Robust DSP-Assisted Neural Pitch Estimation With Very Low Complexity
abstract
Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have been outperforming well-established DSP-based techniques. Unfortunately, these new estimators can be impractical to deploy in real-time systems, both because of their relatively high complexity, and the fact that some require significant lookahead. We show that a hybrid estimator using a small deep neural network (DNN) with traditional DSP-based features can match or exceed the performance of pure DNN-based models, with a complexity and algorithmic delay comparable to traditional DSP-based algorithms. We further demonstrate that this hybrid approach can provide benefits for a neural vocoding task.
Krishna Subramani, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP2
2024 Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
abstract
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures.
Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri, Umut Isik, Michael M. Goodwin
ICASSP2
2024 Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) With Pitch Prediction
abstract
Neural vocoders are now being used in a wide range of speech processing applications. In many of those applications, the vocoder can be the most complex component, so finding lower complexity algorithms can lead to significant practical benefits. In this work, we propose FARGAN, an autoregressive vocoder that takes advantage of long-term pitch prediction to synthesize high-quality speech in small subframes, without the need for teacher-forcing. Experimental results show that the proposed 600 MFLOPS FARGAN vocoder can achieve both higher quality and lower complexity than existing low-complexity vocoders. The quality even matches that of existing higher-complexity vocoders.
Jean-Marc Valin, Ahmed Mustafa, Jan Büthe
IEEE Signal Process. Lett.1
2023 Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity
abstract
GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN vocoders still challenging to run on normal CPUs without accelerators or parallel computers. In this work, we propose a new architecture for GAN vocoders that mainly depends on recurrent and fully-connected networks to directly generate the time domain signal in framewise manner. This results in considerable reduction of the computational cost and enables very fast generation on both GPUs and low-complexity CPUs. Experimental results show that our Framewise WaveGAN vocoder achieves significantly higher quality than auto-regressive maximum-likelihood vocoders such as LPCNet at a very low complexity of 1.2GFLOPS. This makes GAN vocoders more practical on edge and low-power devices.
Ahmed Mustafa, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP2
2023 Low-Bitrate Redundancy Coding of Speech Using A Rate-Distortion-Optimized Variational Autoencoder
abstract
Robustness to packet loss is one of the main ongoing challenges in real-time speech communication. Deep packet loss concealment (PLC) techniques have recently demonstrated improved quality compared to traditional PLC. Despite that, all PLC techniques hit fundamental limitations when too much acoustic information is lost. To reduce losses in the first place, data is commonly sent multiple times using various redundancy mechanisms. We propose a neural speech coder specifically optimized to transmit a large amount of overlapping redundancy at a very low bitrate, up to 50x redundancy using less than 32 kb/s. Results show that the proposed redundancy is more effective than the existing Opus codec redundancy, and that the two can be combined for even greater robustness.
Jean-Marc Valin, Jan Büthe, Ahmed Mustafa
ICASSP1
2023 A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
abstract
In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of the enhanced output and mitigate oversuppression, we experiment with re-weighting frames by the presence or absence of speech activity and applying augmentations to speaker embeddings. By training under a multi-task learning setting, we empirically show that the proposed unified model obtains promising results on both personalized and non-personalized speech enhancement benchmarks and reaches similar performance to models that are trained specialized for either task. The strong performance of the proposed method demonstrates that the unified model is a more economical alternative compared to keeping separate task-specific models during inference.
Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin, Michael M. Goodwin, Paris Smaragdis
ICASSP4
2022 Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet
abstract
Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the efficiency of LPCNet – targeting both algorithmic and computational improvements – to make it usable on a wide variety of devices. We demonstrate an improvement in synthesis quality while operating 2.5x faster. The resulting open-source1LPCNet algorithm can perform real-time neural synthesis on most existing phones and is even usable in some embedded devices.
Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy
ICASSP1
2022 Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing
abstract
Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing data and mildly improve model performance. We propose a novel data augmentation technique, chromagram-based pitch-aware remixing, where music segments with high pitch alignment are mixed. By performing controlled experiments in both supervised and semi-supervised settings, we demonstrate that training models with pitch-aware remixing significantly improves the test signal-to-distortion ratio (SDR).
Siyuan Yuan, Zhepei Wang, Umut Isik, Ritwik Giri, Jean-Marc Valin, Michael M. Goodwin, Arvindh Krishnaswamy
ICASSP5
2022 End-to-end LPCNet: A Neural Vocoder With Fully-Differentiable LPC Estimation
abstract
Neural vocoders have recently demonstrated high quality speech synthesis, but typically require a high computational complexity.LPCNet was proposed as a way to reduce the complexity of neural synthesis by using linear prediction (LP) to assist an autoregressive model.At inference time, LPCNet relies on the LP coefficients being explicitly computed from the input acoustic features.That makes the design of LPCNet-based systems more complicated, while adding the constraint that the input features must represent a clean speech spectrum.We propose an end-to-end version of LPCNet that lifts these limitations by learning to infer the LP coefficients from the input features in the frame rate network .Results show that the proposed end-toend approach equals or exceeds the quality of the original LPC-Net model, but without explicit LP analysis.Our open-source 1 end-to-end model still benefits from LPCNet's low complexity, while allowing for any type of conditioning features.
Krishna Subramani, Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy
INTERSPEECH2
2022 Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model
Jean-Marc Valin, Ahmed Mustafa, Christopher Montgomery, Timothy B. Terriberry, Michael Klingbeil, Paris Smaragdis, Arvindh Krishnaswamy
INTERSPEECH1
2021 Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders
abstract
Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Based on VQ-VAE autoencoders with WaveRNN decoders, we develop compressor-enhancer encoders and accompanying decoders, and show that they operate well in noisy conditions. We also observe that a compressor-enhancer model performs better on clean speech inputs than a compressor model trained only on clean speech.
Jonah Casebeer, Vinjai Vale, Umut Isik, Jean-Marc Valin, Ritwik Giri, Arvindh Krishnaswamy
ICASSP4
2021 Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet
abstract
Speech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging problem due to both real-world constraints like loudspeaker non-linearities, and to limited compute capabilities in some communication systems. In this work, we propose a system combining a traditional acoustic echo canceller, and a low-complexity joint residual echo and noise suppressor based on a hybrid signal processing/deep neural network (DSP/DNN) approach. We show that the proposed system outperforms both traditional and other neural approaches, while requiring only 5.5% CPU for real-time operation. We further show that the system can scale to even lower complexity levels.
Jean-Marc Valin, Srikanth V. Tenneti, Karim Helwani, Umut Isik, Arvindh Krishnaswamy
ICASSP1
2021 Semi-Supervised Singing Voice Separation With Noisy Self-Training
abstract
Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled data to improve the model’s performance. Following the noisy self-training framework, we first train a teacher network on the small labeled dataset and infer pseudo-labels from the large corpus of unlabeled mixtures. Then, a larger student network is trained on combined ground-truth and self-labeled datasets. Empirical results show that the proposed self-training scheme, along with data augmentation methods, effectively leverage the large unlabeled corpus and obtain superior performance compared to supervised methods.
Zhepei Wang, Ritwik Giri, Umut Isik, Jean-Marc Valin, Arvindh Krishnaswamy
ICASSP4
2021 Multi-Channel Opus Compression for Far-Field Automatic Speech Recognition with a Fixed Bitrate Budget
abstract
Automatic speech recognition (ASR) in the cloud allows the use of larger models and more powerful multi-channel signal processing front-ends compared to on-device processing. However, it also adds an inherent latency due to the transmission of the audio signal, especially when transmitting multiple channels of a microphone array. One way to reduce the network bandwidth requirements is client-side compression with a lossy codec such as Opus. However, this compression can have a detrimental effect especially on multi-channel ASR front-ends, due to the distortion and loss of spatial information introduced by the codec. In this publication, we propose an improved approach for the compression of microphone array signals based on Opus, using a modified joint channel coding approach and additionally introducing a multi-channel spatial decorrelating transform to reduce redundancy in the transmission. We illustrate the effect of the proposed approach on the spatial information retained in multi-channel signals after compression, and evaluate the performance on far-field ASR with a multi-channel beamforming front-end. We demonstrate that our approach can lead to a 37.5 % bitrate reduction or a 5.1 % relative word error rate reduction for a fixed bitrate budget in a seven channel setup.
Lukas Drude, Jahn Heymann, Andreas Schwarz, Jean-Marc Valin
Interspeech4
2021 Personalized PercepNet: Real-Time, Low-Complexity Target Voice Separation and Enhancement
abstract
The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized PercepNet, a real-time speech enhancement model that separates a target speaker from a noisy multi-talker mixture without compromising on complexity of the recently proposed PercepNet. To enable speaker-dependent speech enhancement, we first show how we can train a perceptually motivated speaker embedder network to produce a representative embedding vector for the given speaker. Personalized PercepNet uses the target speaker embedding as additional information to pick out and enhance only the target speaker while suppressing all other competing sounds. Our experiments show that the proposed model significantly outperforms PercepNet and other baselines, both in terms of objective speech enhancement metrics and human opinion scores.
Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik, Arvindh Krishnaswamy
Interspeech3
2020 PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss
abstract
Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is encountered in training data. We introduce several innovations that lead to better large neural networks for speech enhancement. The novel PoCoNet architecture is a convolutional neural network that, with the use of frequency-positional embeddings, is able to more efficiently build frequency-dependent features in the early layers. A semi-supervised method helps increase the amount of conversational training data by pre-enhancing noisy datasets, improving performance on real recordings. A new loss function biased towards preserving speech quality helps the optimization better match human perceptual opinions on speech quality. Ablation experiments and objective and human opinion metrics show the benefits of the proposed improvements.
Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, Arvindh Krishnaswamy
INTERSPEECH4
2020 Improving Opus Low Bit Rate Quality with Neural Speech Synthesis
abstract
The voice mode of the Opus audio coder can compress wideband speech at bit rates ranging from 6 kb/s to 40 kb/s. However, Opus is at its core a waveform matching coder, and as the rate drops below 10 kb/s, quality degrades quickly. As the rate reduces even further, parametric coders tend to perform better than waveform coders. In this paper we propose a backward-compatible way of improving low bit rate Opus quality by re-synthesizing speech from the decoded parameters. We compare two different neural generative models, WaveNet and LPCNet. WaveNet is a powerful, high-complexity, and high-latency architecture that is not feasible for a practical system, yet provides a best known achievable quality with generative models. LPCNet is a low-complexity, low-latency RNN-based generative model, and practically implementable on mobile phones. We apply these systems with parameters from Opus coded at 6 kb/s as conditioning features for the generative models. A listening test shows that for the same 6 kb/s Opus bit stream, synthesized speech using LPCNet clearly outperforms the output of the standard Opus decoder. This opens up ways to improve the decoding quality of existing speech and audio waveform coders without breaking compatibility.
Jan Skoglund, Jean-Marc Valin
INTERSPEECH2
2020 A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband Speech
abstract
Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time Fourier transform (STFT) domain, resulting in a high computational complexity. In this work, we propose PercepNet, an efficient approach that relies on human perception of speech by focusing on the spectral envelope and on the periodicity of the speech. We demonstrate high-quality, real-time enhancement of fullband (48 kHz) speech with less than 5% of a CPU core.
Jean-Marc Valin, Umut Isik, Neerad Phansalkar, Ritwik Giri, Karim Helwani, Arvindh Krishnaswamy
INTERSPEECH1
2019 LPCNET: Improving Neural Speech Synthesis through Linear Prediction
abstract
Neural speech synthesis models have recently demonstrated the ability to synthesize high quality speech for text-to-speech and compression applications. These new models often require powerful GPUs to achieve real-time operation, so being able to reduce their complexity would open the way for many new applications. We propose LPCNet, a WaveRNN variant that combines linear prediction with recurrent neural networks to significantly improve the efficiency of speech synthesis. We demonstrate that LPCNet can achieve significantly higher quality than WaveRNN for the same network size and that high quality LPCNet speech synthesis is achievable with a complexity under 3 GFLOPS. This makes it easier to deploy neural synthesis applications on lower-power devices, such as embedded systems and mobile phones.
Jean-Marc Valin, Jan Skoglund
ICASSP1
2019 A Real-Time Wideband Neural Vocoder at 1.6kb/s Using LPCNet
abstract
Neural speech synthesis algorithms are a promising new approach for coding speech at very low bitrate. They have so far demonstrated quality that far exceeds traditional vocoders, at the cost of very high complexity. In this work, we present a low-bitrate neural vocoder based on the LPCNet model. The use of linear prediction and sparse recurrent networks makes it possible to achieve real-time operation on general-purpose hardware. We demonstrate that LPCNet operating at 1.6 kb/s achieves significantly higher quality than MELP and that uncompressed LPCNet can exceed the quality of a waveform codec operating at low bitrate. This opens the way for new codec designs based on neural synthesis models.
Jean-Marc Valin, Jan Skoglund
INTERSPEECH1
2018 The Av1 Constrained Directional Enhancement Filter (Cdef)
abstract
This paper presents the constrained directional enhancement filter designed for the AV1 royalty-free video codec. The in-loop filter is based on a non-linear low-pass filter and is designed for vectorization efficiency. It takes into account the direction of edges and patterns being filtered. The filter works by identifying the direction of each block and then adaptively filtering with a high degree of control over the filter strength along the direction and across it. The proposed enhancement filter is shown to improve the quality of the Alliance for Open Media (AOM) AV1 and Thor video codecs in particular in low complexity configurations.
Steinar Midtskogen, Jean-Marc Valin
ICASSP2
2018 A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement
abstract
Despite noise suppression being a mature area in signal processing, it remains highly dependent on fine tuning of estimator algorithms and parameters. In this paper, we demonstrate a hybrid DSP/deep learning approach to noise suppression. We focus strongly on keeping the complexity as low as possible, while still achieving high-quality enhanced speech. A deep recurrent neural network with four hidden layers is used to estimate ideal critical band gains, while a more traditional pitch filter attenuates noise between pitch harmonics. The approach achieves significantly higher quality than a traditional minimum mean squared error spectral estimator, while keeping the complexity low enough for real-time operation at 48 kHz on a low-power CPU.
Jean-Marc Valin
MMSP1
2018 An Overview of Core Coding Tools in the AV1 Video Codec
abstract
AV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC.
Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz
PCS17
2016 Daala: A perceptually-driven still picture codec
abstract
Daala is a new royalty-free video codec based on perceptually-driven coding techniques. We explore using its keyframe format for still picture coding and show how it has improved over the past year. We believe the technology used in Daala could be the basis of an excellent, royalty-free image format.
Jean-Marc Valin, Nathan E. Egge, Thomas J. Daede, Timothy B. Terriberry, Christopher Montgomery
ICIP1
2016 Daala: Building a next-generation video codec from unconventional technology
abstract
Daala is a new royalty-free video codec that attempts to compete with state-of-the-art royalty-bearing codecs. To do so, it must achieve good compression while avoiding all of their patented techniques. We use technology that is as different as possible from traditional approaches to achieve this. This paper describes the technology behind Daala and discusses where it fits in the newly created AV1 codec from the Alliance for Open Media. We show that Daala is approaching the performance level of more mature, state-of-the art video codecs and can contribute to improving AV1.
Jean-Marc Valin, Timothy B. Terriberry, Nathan E. Egge, Thomas J. Daede, Yushin Cho, Christopher Montgomery, Michael Bebenita
MMSP1
2012 Integration of sound source localization and separation to improve Dialogue Management on a robot
abstract
To demonstrate the influence of an artificial audition system on speech recognition and dialogue management for a robot, this paper presents a case study involving soft coupling of ManyEars, a sound source localization, tracking and separation system, with the CSLU Dialogue Management system. Trials were conducted in a laboratory and a cafeteria. Results indicate that preprocessing of the audio signals by ManyEars improves speech recognition and dialogue management of the system, demonstrating the feasibility and the added flexibility provided by ManyEars for a robot to interact vocally with humans in a wide variety of contexts.
Maxime Fréchette, Dominic Létourneau, Jean-Marc Valin, François Michaud
IROS3
2010 A High-Quality Speech and Audio Codec With Less Than 10-ms Delay
abstract
With increasing quality requirements for multimedia communications, audio codecs must maintain both high quality and low delay. Typically, audio codecs offer either low delay or high quality, but rarely both. We propose a codec that simultaneously addresses both these requirements, with a delay of only 8.7 ms at 44.1 kHz. It uses gain-shape algebraic vector quantization in the frequency domain with time-domain pitch prediction. We demonstrate that the proposed codec operating at 48 kb/s and 64 kb/s out-performs both G.722.1C and MP3 and has quality comparable to AAC-LD, despite having less than one fourth of the algorithmic delay of these codecs.
Jean-Marc Valin, Timothy B. Terriberry, Christopher Montgomery, Gregory Maxwell
IEEE Trans. Speech Audio Process.1
2009 Priority Based Dynamic Rate Control for VoIP Traffic
abstract
This paper presents a novel mechanism for dynamic rate control of prioritised Voice Over IP (VoIP) traffic in real time. The system uses our proposed variable bit rate speech codec called Speex, which can dynamically adjust the encoding bit rate (and hence the voice quality) based on the feedback information about the network congestion, flow priority, and the instantaneous speech properties. Our extensive NS2 simulation results along with results from ITU-T standard of speech quality evaluation tool (PESQ) show that the proposed system indeed provides highest quality speech while maximising the bandwidth utilisation and reducing the network congestion.
Fariza Sabrina, Jean-Marc Valin
GLOBECOM2
2009 Reflected Simplex Codebooks for Limited Feedback MIMO Beamforming
abstract
This paper proposes reflected simplex codebooks for limited feedback beamforming in multiple-input multiple-output (MIMO) wireless systems. The codebooks are a geometric construction based on simplices and the Anlattice. We propose a fast codebook search and indexing algorithm. We show that such codebooks perform superior or comparable to other codebooks, with much lower implementation complexity.
Daniel J. Ryan, Iain B. Collings, Jean-Marc Valin
ICC3
2009 Evaluating real-time audio localization algorithms for artificial audition in robotics
abstract
Although research on localization of sound sources using microphone arrays has been carried out for years, providing such capabilities on robots is rather new. Artificial audition systems on robots currently exist, but no evaluation of the methods used to localize sound sources has yet been conducted. This paper presents an evaluation of various real-time audio localization algorithms using a medium-sized microphone array which is suitable for applications in robotics. The techniques studied here are implementations and enhancements of steered response power - phase transform beamformers, which represent the most popular methods for time difference of arrival audio localization. In addition, two different grid topologies for implementing source direction search are also compared. Results show that a direction refinement procedure can be used to improve localization accuracy and that more efficient and accurate direction searches can be performed using a uniform triangular element grid rather than the typical rectangular element grid.
Anthony P. Badali, Jean-Marc Valin, François Michaud, Parham Aarabi
IROS2
2008 Adaptive Rate Control for Aggregated VoIP Traffic
abstract
This paper presents a novel mechanism for dynamically adapting the quality of congestion controlled voice over IP (VoIP) applications on the Internet in real time. The system uses our proposed variable bit rate speech codec called Speex, which can dynamically adjust the encoding bit rate (and hence the speech quality) based on both the feedback information about the network congestion and the instantaneous speech properties. Our extensive NS2 simulation results prove that the proposed system indeed provides highest quality speech while maximising the bandwidth utilisation and reducing the network congestion.
Fariza Sabrina, Jean-Marc Valin
GLOBECOM2
2008 Embedded auditory system for small mobile robots
abstract
Auditory capabilities would allow small robots interacting with people to act according to vocal cues. In our recent work, we have demonstrated AUDIBLE, an auditory system capable of sound source localization, tracking and separation in real-time, using an array of eight microphones and running on a laptop computer. The system is able to localize and track up to four sources, while separating up to three sources in real-time in noisy environments. Signal processing techniques can be quite computer intensive, and the question of making it possible for this system to run on platforms that cannot carry a laptop computer onboard can be raised. This paper reports our investigation of the appropriate compromises to be made to AUDIBLE's implementation in order to port the system on an embedded DSP (Digital Signal Processor) platform. The DSP implementation is fully functional and performs well with minor limitations compared to the original system i.e., limitations on sound source duration and on the number of sources that can be processed simultaneously. Results demonstrate that it is feasible to port AUDIBLE on embedded platforms, opening up its use in field applications such as human-robot interaction in real life settings.
Simon Brière, Jean-Marc Valin, François Michaud, Dominic Létourneau
ICRA2
2007 Design and implementation of a robot audition system for automatic speech recognition of simultaneous speech
abstract
This paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ASRU5
2007 A New Robust Frequency Domain Echo Canceller with Closed-Loop Learning Rate Adaptation
abstract
One of the main difficulties in echo cancellation is the fact that the learning rate needs to vary according to conditions such as double-talk and echo path change. Several methods have been proposed to vary the learning. In this paper we propose a new closed-loop method where the learning rate is proportional to a misalignment parameter, which is in turn estimated based on a gradient adaptive approach. The method is presented in the context of a multidelay block frequency domain (MDF) echo canceller. We demonstrate that the proposed algorithm outperforms current popular double-talk detection techniques by up to 6 dB.
Jean-Marc Valin, Iain B. Collings
ICASSP (1)1
2007 Interference-Normalized Least Mean Square Algorithm
abstract
An interference-normalized least mean square (INLMS) algorithm for robust adaptive filtering is proposed. The INLMS algorithm extends the gradient-adaptive learning rate approach to the case where the signals are nonstationary. In particular, we show that the INLMS algorithm can work even for highly nonstationary interference signals, where previous gradient-adaptive learning rate algorithms fail.
Jean-Marc Valin, Iain B. Collings
IEEE Signal Process. Lett.1
2007 On Adjusting the Learning Rate in Frequency Domain Echo Cancellation With Double-Talk
abstract
One of the main difficulties in echo cancellation is the fact that the learning rate needs to vary according to conditions such as double-talk and echo path change. In this paper, we propose a new method of varying the learning rate of a frequency-domain echo canceller. This method is based on the derivation of the optimal learning rate of the normalized least mean square (NLMS) algorithm in the presence of noise. The method is evaluated in conjunction with the multidelay block frequency domain (MDF) adaptive filter. We demonstrate that it performs better than current double-talk detection techniques and is simple to implement
Jean-Marc Valin
IEEE Trans. Speech Audio Process.1
2007 Robust Recognition of Simultaneous Speech by a Mobile Robot
abstract
This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of geometric source separation (GSS) and a postfilter that gives a further reduction of interference from other sources. The postfllter is also used to estimate the reliability of spectral features and compute a missing feature mask. The mask is used in a missing feature theory-based speech recognition system to recognize the speech from simultaneous Japanese speakers in the context of a humanoid robot. Recognition rates are presented for three simultaneous speakers located at 2 m from the robot. The system was evaluated on a 200-word vocabulary at different azimuths between sources, ranging from 10deg to 90deg. Compared to the use of the microphone array source separation alone, we demonstrate an average reduction in relative recognition error rate of 24% with the postfllter and of 42% when the missing features approach is combined with the postfllter. We demonstrate the effectiveness of our multisource microphone array postfilter and the improvement it provides when used in conjunction with the missing features theory.
Jean-Marc Valin, Seiichi Yamamoto, Jean Rouat, François Michaud, Kazuhiro Nakadai, Hiroshi G. Okuno
IEEE Trans. Robotics1
2006 Robust 3D Localization and Tracking of Sound Sources Using Beamforming and Particle Filtering
abstract
In this paper we present a new robust sound source localization and tracking method using an array of eight microphones (US patent pending). The method uses a steered beamformer based on the reliability-weighted phase transform (RWPHAT) along with a particle filter-based tracking algorithm. The proposed system is able to estimate both the direction and the distance of the sources. In a videoconferencing context, the direction was estimated with an accuracy better than one degree while the distance was accurate within 10% RMS. Tracking of up to three simultaneous moving speakers is demonstrated in a noisy environment
Jean-Marc Valin, François Michaud, Jean Rouat
ICASSP (4)1
2006 Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE5
2006 Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real World
abstract
This paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2006 Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
PRICAI6
2005 A Brochette of Socially Interactive Robots
François Michaud, Dominic Létourneau, Pierre Lepage, Yan Morin, Frédéric Gagnon, Patrick Giguère, Eric Beaudry, Yannick Brosseau, Carle Côté, Audrey Duquette, Jean-François Laplante, Marc-Antoine Legault, Pierre Moisan, Arnaud Ponchon, Clément Raïevsky, Marc-André Roux, Tamie Salter, Jean-Marc Valin, Serge Caron, Patrice Masson, Froduald Kabanza, Michel Lauria
AAAI18
2005 Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative).
Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno
ICRA2
2005 Multiple moving speaker tracking by microphone array on mobile robot
abstract
Real-world applications often require tracking multiple moving speakers for improving human-robot interactions and/or sound source separation. This paper presents multiple moving speaker tracking using an 8ch microphone array system installed on a mobile robot. This problem is difficult because the system does not assume that sound sources and/or the microphone array are fixed. Our solutions consist of two key ideas – time delay of arrival estimation, and multiple Kalman filters. The former localizes multiple sound sources based on beamforming in real time. Non-linear movements are tracked by using a set of Kalman filters with different history lengths in order to reduce errors in tracking multiple moving speakers under noisy and echoic environments. For quantitative evaluation of the tracking, motion references of sound sources and a mobile robot, called SIG2, were measured accurately by ultrasonic 3D tag sensors. As a result, we showed that the system tracked three simultaneous sound sources even when SIG2 moved in a room with large reverberation due to glass walls. 1.
Masamitsu Murase, Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Kentaro Yamada, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH3
2005 Making a robot recognize three simultaneous sentences in real-time
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS3
2004 Microphone array post-filter for separation of simultaneous non-stationary sources
abstract
Microphone array post-filters have demonstrated their ability to greatly reduce noise at the output of a beamformer. However, current techniques only consider a single source of interest, most of the time assuming stationary background noise. We propose a microphone array post-filter that enhances the signals produced by the separation of simultaneous sources using common source separation algorithms. Our method is based on a loudness-domain optimal spectral estimator and on the assumption that the noise can be described as the sum of a stationary component and of a transient component that is due to leakage between the channels of the initial source separation algorithm. The system is evaluated in the context of mobile robotics and is shown to produce better results than current post-filtering techniques, greatly reducing interference while causing little distortion to the signal of interest, even at very low SNR.
Jean-Marc Valin, Jean Rouat, François Michaud
ICASSP (1)1
2004 Autonomous Initialization of Robot Formations
abstract
Real life deployment of robot formation cannot assume that robots are going to be correctly positioned to move in a particular configuration. To do so, we propose an approach that allows the group to determine autonomously the most appropriate assignment of positions in the formation. Our approach is distributed and uses directional visual perception to localize robots. Inter-robot communication allows them to share information on which robots are nearby, so that each can evaluate it ability to be the conductor of the group and assign formation positions to the other robots by minimizing repositioning. The assignment search is done using a distributed bounded depth-first with pruning search. The robot with the best score is selected as the conductor, and the other robots receive from the conductor their assignment in the formation. Validation of our work is done in simulation and with Pioneer 2 robots.
Mathieu Lemay, François Michaud, Dominic Létourneau, Jean-Marc Valin
ICRA4
2004 Localization of Simultaneous Moving Sound Sources for Mobile Robot Using a Frequency- Domain Steered Beamformer Approach
abstract
Mobile robots in real-life settings would benefit from being able to localize sound sources. Such a capability can nicely complement vision to help localize a person or an interesting event in the environment, and also to provide enhanced processing for other capabilities such as speech recognition. We present a robust sound source localization method in three-dimensional space using an array of 8 microphones. The method is based on a frequency-domain implementation of a steered beamformer along with a probabilistic post-processor. Results show that a mobile robot can localize in real time multiple moving sources of different types over a range of 5 meters with a response time of 200 ms.
Jean-Marc Valin, François Michaud, Brahim Hadjou, Jean Rouat
ICRA1
2004 Code reusability tools for programming mobile robots
abstract
This paper describes two initiatives aiming at improving code reusability for programming mobile robots: robotflow/flowdesigner, a data-flow programming environment; MARIE (mobile and autonomous robotics integration environment), a programming environment allowing multiple applications, programs and tools, to operate on one or multiple machines/OS and work together on a mobile robot implementation. Robotflow/flowdesigner's objective is to provide a modular, graphical programming environment that would help visualize and understand what is really happening in the robot's control loops, sensors, actuators, by using graphical probes. MARIE aims at avoiding making an exclusive choice on particular programming tools, making it possible to reuse code and applications.
Carle Côté, Dominic Létourneau, François Michaud, Jean-Marc Valin, Yannick Brosseau, Clément Raïevsky, Mathieu Lemay, Wctor Tran
IROS4
2004 Enhanced robot audition based on microphone array source separation with post-filter
abstract
We propose a system that gives a mobile robot the ability to separate simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of geometric source separation and a post-filter that gives us a further reduction of interferences from other sources. We present results and comparisons for separation of multiple non-stationary speech sources combined with noise sources. The main advantage of our approach for mobile robots resides in the fact that both the frequency domain geometric source separation algorithm and the post-filter are able to adapt rapidly to new sources and non-stationarity. Separation results are presented for three simultaneous interfering speakers in the presence of noise. A reduction of log spectral distortion (LSD) and increase of signal-to-noise ratio (SNR) of approximately 10 dB and 14 dB are observed.
Jean-Marc Valin, Jean Rouat, François Michaud
IROS1
2003 Textual message read by a mobile robot
abstract
Giving the ability to read characters and symbols is highly desirable for increased autonomy of mobile robots operating in the real world. The idea is fairly simple: give a robot the ability to acquire an image of a message to read, extract the symbols and recognize them. Image character recognition research has been going on for decades now, with good results. But compared to conventional character recognition systems, the challenge with a mobile robot is to find a textual message to capture in the world and to get a good view of the message, knowing that the viewpoint of the robot depends on its position in relation to the message, which cannot be pre-specified. In this paper we present our approach making it possible for an autonomous mobile robot to read messages. We outline the constraints under which the approach works, and present results obtained using a Pioneer 2 robot equipped with a Pentium 233 MHz and a pan-tilt-zoom camera.
Dominic Létourneau, François Michaud, Jean-Marc Valin, Catherine Proulx
IROS3
2003 Robust sound source localization using a microphone array on a mobile robot
abstract
The hearing sense on a mobile robot is important because it is omnidirectional and it does not require direct line-of-sight with the sound source. Such capabilities can nicely complement vision to help localize a person or an interesting event in the environment. To do so the robot auditory system must be able to work in noisy, unknown and diverse environmental conditions. In this paper, we present a robust sound source localization method in three-dimensional space using an array of 8 microphones. The method is based on time delay of arrival estimation. Results show that a mobile robot can localize in real time different types of sound sources over a range of 3 meters and with a precision of 3/spl deg/.
Jean-Marc Valin, François Michaud, Jean Rouat, Dominic Létourneau
IROS1
2003 Making a mobile robot read textual messages
abstract
With all the textual indications, messages and signs we find in urban settings to provide us with all types of information, it is only natural that we try to give to robots reading capabilities of interpreting such information. Equipped with optical character recognition algorithms, a mobile robot has to face the challenge of controlling its position in the world and its pan-tilt-zoom camera to find the textual message to capture, try to compensate for its viewpoint of the message, and use limited processing capabilities to decode the message. The robot also has to deal with non-uniform illumination and changing conditions in the world. In this work, we address the different aspects of the character recognition process to be incorporated into the higher level intelligence modules of a mobile robotic platform.
Dominic Létourneau, François Michaud, Jean-Marc Valin, Catherine Proulx
SMC3
2002 Dynamic robot formations using directional visual perception
abstract
Recent research projects have demonstrated that it is possible to make robots move in formation. The approaches differ by the various assumptions about what can be perceived and communicated by the robots, the strategies used to make the robots move in formation, the ability to deal with obstacles and to switch formations. After suggesting criteria to characterize problems associated with robot formations, this paper presents a distributed approach based on directional visual perception and inter-robot communication. Using a pan camera head, sonar readings and wireless communication, we demonstrate that robots are not only able to move in formation, avoid obstacles and switch formations, but also initialize and determine by themselves their positions in the formation. Validation of our work is done in simulation and with Pioneer 2 robots.
François Michaud, Dominic Létourneau, Matthieu Guilbert, Jean-Marc Valin
IROS4
1999 On the limits of speech recognition in noise
abstract
We consider the performance of speech recognition in noise and focus on its sensitivity to the acoustic feature set. In particular, we examine the perceived information reduction imposed on a speech signal using a feature extraction method commonly used for automatic speech recognition. We observe that the human recognition rates on noisy digit strings drop considerably as the speech signal undergoes the typical loss of phase and loss of frequency resolution. Steps are taken to ensure that human subjects are constrained in ways similar to that of an automatic recognizer. The high correlation between the performance of the human listeners and that of our connected digit recognizer leads us to some interesting conclusions, including that typical cepstral processing is insufficient to support speech information in noise.
Stephen D. Peters, Peter Stubley, Jean-Marc Valin
ICASSP3