Jean Rouat

dblp:26/5977 · DBLP profile ↗
← Back
51ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-9306-426XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 1 since 2021Systems, architecture and hardware · 9 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Adaptive Central Frequencies Locally Competitive Algorithm for Speech
abstract
Neuromorphic computing, inspired by nervous systems, revolutionizes information processing with its focus on efficiency and low power consumption. Using sparse coding, this paradigm enhances processing efficiency, which is crucial for edge devices with power constraints. The Locally Competitive Algorithm (LCA), adapted for audio with Gammatone and Gammachirp filter banks, provides an efficient sparse coding method for neuromorphic speech processing. Adaptive LCA (ALCA) further refines this method by dynamically adjusting modulation parameters, thereby improving reconstruction quality and sparsity. This paper introduces an enhanced ALCA version, the ALCA Central Frequency (ALCA-CF), which dynamically adapts both modulation parameters and central frequencies, optimizing the speech representation. Evaluations show that this approach improves reconstruction quality and sparsity while significantly reducing the power consumption of speech classification, without compromising classification accuracy, particularly on Intel’s Loihi 2 neuromorphic chip.
Soufiyan Bahadi, Eric Plourde, Jean Rouat
ICASSP3
2023 NAAQA: A Neural Architecture for Acoustic Question Answering
abstract
The goal of the Acoustic Question Answering (AQA) task is to answer a free-form text question about the content of an acoustic scene. It was inspired by the Visual Question Answering (VQA) task. In this paper, based on the previously introduced CLEAR dataset, we propose a new benchmark for AQA, namely CLEAR2, that emphasizes the specific challenges of acoustic inputs. These include handling of variable duration scenes, and scenes built with elementary sounds that differ between training and test set. We also introduce NAAQA, a neural architecture that leverages specific properties of acoustic inputs. The use of 1D convolutions in time and frequency to process 2D spectro-temporal representations of acoustic content shows promising results and enables reductions in model complexity. We show that time coordinate maps augment temporal localization capabilities which enhance performance of the network by ∼ 17 percentage points. On the other hand, frequency coordinate maps have little influence on this task. NAAQA achieves 79.5% of accuracy on the AQA task with ∼ four times fewer parameters than the previously explored VQA model. We evaluate the performance of NAAQA on an independent data set reconstructed from DAQA. We also test the addition of a MALiMo module in our model on both CLEAR2 and DAQA. We provide a detailed analysis of the results for the different question types. We release the code to produce CLEAR2 as well as NAAQA to foster research in this newly emerging machine learning task.
Jérôme Abdelnour, Jean Rouat, Giampiero Salvi
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Sonified Distance in Sensory Substitution Does Not Always Improve Localization: Comparison With a 2-D and 3-D Handheld Device
abstract
Early visual to auditory substitution devices encode 2-D monocular images into sounds while more recent devices use distance information from 3-D sensors. This study assesses whether the addition of sound-encoded distance in recent systems helps to convey the “where” information. This is important to the design of new sensory substitution devices. We conducted experiments for object localization and navigation tasks with a handheld visual to audio substitution system. It comprises 2-D and 3-D modes. Both encode in real-time the position of objects in images captured by a camera. The 3-D mode encodes in addition the distance between the system and the object. Experiments have been conducted with 16 blindfolded sighted participants. For the localization, participants were quicker to understand the scene with the 3-D mode that encodes distances. On the other hand, with the 2-D only mode, they were able to compensate for the lack of distance encoding after a small training. For the navigation, participants were as good with the 2-D only mode than with the 3-D mode encoding distance.
Louis Commère, Jean Rouat
IEEE Trans. Hum. Mach. Syst.2
2023 Evaluation of Short-Range Depth Sonifications for Visual-to-Auditory Sensory Substitution
abstract
Visual-to-auditory sensory substitution devices convert visual information into sound and can provide valuable assistance for blind people. Recent iterations of these devices rely on depth sensors. Rules for converting depth into sound (i.e., the sonifications) are often designed arbitrarily, with no strong evidence for choosing one over another. The purpose of this article is to compare and understand the effectiveness of five depth sonifications in order to assist the design process of future visual-to-auditory systems for blind people, which rely on depth sensors. The frequency, amplitude, and reverberation of the sound as well as the repetition rate of short high-pitched sounds and the signal-to-noise ratio of a mixture between pure sound and noise are studied. We conducted positioning experiments with 28 sighted blindfolded participants. Stage 1 incorporates learning phases followed by depth estimation tasks. Stage 2 adds the additional challenge of azimuth estimation to the first stage's protocol. Stage 3 tests learning retention by incorporating a 10-min break before retesting depth estimation. The best depth estimates in stage 1 were obtained with the sound frequency and the repetition rate of beeps. In stage 2, the beep repetition rate yielded the best depth estimation, and no significant difference was observed for the azimuth estimation. Results of stage 3 showed that the beep repetition rate was the easiest sonification to memorize. Based on the statistical analysis of the results, we discuss the effectiveness of each sonification and compare with other studies that encode depth into sounds. Finally, we provide recommendations for the design of depth encoding.
Louis Commère, Jean Rouat
IEEE Trans. Hum. Mach. Syst.2
2021 A proposal and evaluation of new timbre visualization methods for audio sample browsers
Etienne Richan, Jean Rouat
Pers. Ubiquitous Comput.2
2020 SECL-UMons Database for Sound Event Classification and Localization
abstract
We introduce the SECL-UMons dataset for sound event classification and localization in the context of office environments. The multichannel dataset is composed of 11 event classes recorded at several realistic positions in two different rooms. The dataset comprises two types of sequences according to the number of events in the sequence. 2662 unilabel sequences and 2724 multilabel sequences are recorded corresponding to a total of 5.24 hours. The database is publicly available to provide support for algorithm development and common ground for comparison of different techniques. The DCASE 2019 challenge baseline (SELDnet) employing a convolutional recurrent neural network is used to generate benchmark scores for the new dataset. We also slightly modify the model to introduce a benchmark score for real-time classification and localization for the new dataset.
Mathilde Brousmiche, Jean Rouat, Stéphane Dupont
ICASSP2
2017 Real-Time Speech Enhancement with GCC-NMF: Demonstration on the Raspberry Pi and NVIDIA Jetson
Sean U. N. Wood, Jean Rouat
INTERSPEECH2
2017 Real-Time Speech Enhancement with GCC-NMF
Sean U. N. Wood, Jean Rouat
INTERSPEECH2
2017 Blind Speech Separation and Enhancement With GCC-NMF
abstract
We present a blind source separation algorithm named GCC-NMF that combines unsupervised dictionary learning via non-negative matrix factorization (NMF) with spatial localization via the generalized cross correlation (GCC) method. Dictionary learning is performed on the mixture signal, with separation subsequently achieved by grouping dictionary atoms, at each point in time, according to their spatial origins. The resulting source separation algorithm is simple yet flexible, requiring no prior knowledge or information. Separation quality is evaluated for three tasks using stereo recordings from the publicly available SiSEC signal separation evaluation campaign: 3 and 4 concurrent speakers in reverberant environments, speech mixed with real-world background noise, and noisy recordings of a moving speaker. Performance is quantified using perceptually motivated and SNR-based measures with the PEASS and BSS Eval toolkits, respectively. We evaluate the effects of model parameters on separation quality, and compare our approach with other unsupervised and semi-supervised speech separation and enhancement approaches. We show that GCC-NMF is a flexible source separation algorithm, outperforming task-specific approaches in each of the three settings, including both blind as well as several informed approaches that require prior knowledge or information.
Sean U. N. Wood, Jean Rouat, Stéphane Dupont, Gueorgui Pironkov
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Audio Watermarking Using Spikegram and a Two-Dictionary Approach
abstract
This paper introduces a new audio watermarking technique based on a perceptual kernel representation of audio signals (spikegram). Spikegram is a recent method to represent audio signals. It is combined with a dictionary of gammatones to construct a robust representation of sounds. In traditional phase embedding methods, the phase of coefficients of a given signal in a specific domain (such as Fourier domain) is modified. In the encoder of the proposed method (two-dictionary approach), the signs and the phases of gammatones in the spikegram are chosen adaptively to maximize the strength of the decoder. Moreover, the watermark is embedded only into kernels with high amplitudes, where all masked gammatones have been already removed. The efficiency of the proposed spikegram watermarking is shown via several experimental results. First, robustness of the proposed method is shown against 32 kb/s MP3 with an embedding rate of 56.5 b/s. Second, we showed that the proposed method is robust against unified speech and audio codec (24-kb/s USAC, linear predictive, and Fourier domain modes) with an average payload of 5-15 b/s. Third, it is robust against simulated small real room attacks with a payload of roughly 1 b/s. Last, it is shown that the proposed method is robust against a variety of signal processing transforms while preserving quality.
Yousof Erfani, Ramin Pichevar, Jean Rouat
IEEE Trans. Inf. Forensics Secur.3
2016 Blind Speech Separation with GCC-NMF
Sean U. N. Wood, Jean Rouat
INTERSPEECH2
2016 A Flexible Bio-Inspired Hierarchical Model for Analyzing Musical Timbre
abstract
A flexible and multipurpose bio-inspired hierarchical model for analyzing musical timbre is presented in this paper. Inspired by findings in the fields of neuroscience, computational neuroscience, and psychoacoustics, not only does the model extract spectral and temporal characteristics of a signal, but it also analyzes amplitude modulations on different timescales. It uses a cochlear filter bank to resolve the spectral components of a sound, lateral inhibition to enhance spectral resolution, and a modulation filter bank to extract the global temporal envelope and roughness of the sound from amplitude modulations. The model was evaluated in three applications. First, it was used to simulate subjective data from two roughness experiments. Second, it was used for musical instrument classification using the k-NN algorithm and a Bayesian network. Third, it was applied to find the features that characterize sounds whose timbres were labeled in an audiovisual experiment. The successful application of the proposed model in these diverse tasks revealed its potential in capturing timbral information.
Mohammad Adeli, Jean Rouat, Sean U. N. Wood, Stephane Molotchnikoff, Eric Plourde
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Live demonstration: Efficient event-driven approach using synchrony processing for hardware spiking neural networks
abstract
Recent neuromorphic applications now use spiking neural networks (SNNs) because of their improved computational power compared to previous generations of neural networks. Efficient simulation is essential when using this type of neuron since many events have to be handled on a large number of neurons within the network. In this demonstration, a hardware simulator for SNNs that has applications in image recognition is presented. This SNN uses synchrony processing for efficient event-driven simulation (SPEEDS) which allows parallel computations of synchronized events. SPEEDS differs from common event-driven approaches that serialize every event and can improve significantly the computational efficiency of a SNN simulator. The hardware SNN is implemented on a Xilinx Virtex-6 XC6VLX240T field-programmable gate array (FPGA) and can contain 131 072 neurons. It can process approximately 70 million spikes per second on a 4-bank architecture clocked at 100 MHz. The presentation explains how such a system can be used for image processing tasks like image segmentation, feature extraction and pattern matching to realize a recognition system that can detect several objects in a given image.
Guillaume Seguin-Godin, Frédéric Mailhot 0001, Jean Rouat
ISCAS3
2015 Efficient event-driven approach using synchrony processing for hardware spiking neural networks
abstract
Current digital hardware implementations of spiking neural networks usually focus on a time-driven architecture to process the large number of events that occur during a typical simulation. While this type of implementation is practical for simulating biologically accurate neurons, most systems using a simpler neuron model can benefit from an event-driven architecture. In such cases, significant performance improvements are theoretically possible. In practice, however, such implementations do not maximize the available computational power because finding the next event often involves serializing computations. In this paper, a hardware architecture that offers the efficiency of an event-driven algorithm while allowing parallel computations is developed. The architecture uses multiple pipelined processing elements to compute spikes in parallel and a novel comparator tree structure to find the next event in a large network efficiently. The resulting system can implement up to 131 072 neurons on a single FPGA (Xilinx Virtex-6 XC6VLX240T) and processes approximately 70 million spikes per second when using a 4-bank architecture clocked at 100 MHz.
Guillaume Seguin-Godin, Frédéric Mailhot 0001, Jean Rouat
ISCAS3
2013 Event management for large scale event-driven digital hardware spiking neural networks
Louis-Charles Caron, Michiel D'Haene, Frédéric Mailhot 0001, Benjamin Schrauwen, Jean Rouat
Neural Networks5
2012 Regulation toward Self-organized Criticality in a Recurrent Spiking Neural Reservoir
Simon Brodeur, Jean Rouat
ICANN (1)2
2011 Variable frame rate hierarchical analysis for robust speech recognition
abstract
A new bio-inspired speech analysis system that extracts acoustical speech events is proposed and used in the design of a variable frame rate (VFR) speech recognizer. The same speech recognizer (Hidden Markov Model -HMM- and Mel Frequency Cepstrum Coefficients -MFCC-) has been used with the proposed VFR analysis and conventional fixed frame rate (FFR) approach. In comparison with other VFR recognizers, the hierarchical features in the proposed system have the potential to serve as classification parameters of a complete bio-inspired speech recognition system. Also, no voice activity detection is required and there are no hard decisions to be taken by the system. Events are used to label and identify the moments at which the acoustical properties of speech are stable or changing. These events are markers on which an analysis window can be positioned to perform the recognition. Inspired by our knowledge of the auditory and visual systems, hierarchical complex features like transients and energy orientation are used. Training has been done on clean speech and recognition on noisy (from 20dB to −10dB Signal to Noise Ratios -SNR) or reverberated speech by using the TI 46-word database corrupted with 4 noises taken from the Aurora 2 data. In comparison with a FFR recognizer, our VFR system yields more than 50% increase in recognition rates for a speaker independent isolated word recognition task when SNRs are between 0 and 20 dB.
Jean Rouat, Stéphane Loiselle, Stephane Molotchnikoff
IROS1
2011 FPGA implementation of a spiking neural network for pattern matching
abstract
A field programmable gate array (FPGA) implementation of a hardware spiking neural network is presented. The system is able to realize different signal processing tasks using the synchronization of oscillatory leaky integrate and fire neurons. The use of a bit slice architecture and short, local interconnections make it adaptable to projects of various scales. The system is also designed to efficiently process groups of synchronized neurons. A fully connected network of 648 neurons and 419904 synapses is implemented on a stand-alone Xilinx XC5VSX50T FPGA, processing up to 6M spikes/s. We describe the resource usage for the whole system as well as for each functional block, and illustrate the functioning of the circuit on a simple image recognition task.
Louis-Charles Caron, Frédéric Mailhot 0001, Jean Rouat
ISCAS3
2008 Non-negative sparse image coder via simulated annealing and pseudo-inversion
abstract
We propose a sparse non-negative image coding based on simulated annealing and matrix pseudo-inversion. We show that sparsity and non-negativity are both important to obtain part-based coding and we also show the impact of each of them on the coding. In contrast with other approaches in the literature, our method can constrain both weights and basis vectors to generate part-based bases suitable for image recognition and fiducial point extraction. We also propose a speed-up of the algorithm by implementing a hybrid system that mixes simulated annealing and pseudo-inverse computation of matrices.
Ramin Pichevar, Jean Rouat
ICASSP2
2007 Monophonic sound source separation with an unsupervised network of spiking neurones
Ramin Pichevar, Jean Rouat
Neurocomputing2
2007 Robust Recognition of Simultaneous Speech by a Mobile Robot
abstract
This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of geometric source separation (GSS) and a postfilter that gives a further reduction of interference from other sources. The postfllter is also used to estimate the reliability of spectral features and compute a missing feature mask. The mask is used in a missing feature theory-based speech recognition system to recognize the speech from simultaneous Japanese speakers in the context of a humanoid robot. Recognition rates are presented for three simultaneous speakers located at 2 m from the robot. The system was evaluated on a 200-word vocabulary at different azimuths between sources, ranging from 10deg to 90deg. Compared to the use of the microphone array source separation alone, we demonstrate an average reduction in relative recognition error rate of 24% with the postfllter and of 42% when the missing features approach is combined with the postfllter. We demonstrate the effectiveness of our multisource microphone array postfilter and the improvement it provides when used in conjunction with the missing features theory.
Jean-Marc Valin, Seiichi Yamamoto, Jean Rouat, François Michaud, Kazuhiro Nakadai, Hiroshi G. Okuno
IEEE Trans. Robotics3
2006 Cohesive Particle Filtering for Sound Source Localization
abstract
We present a novel resampling technique for particle filtering algorithms based on a model of attractive and repulsive forces. The new approach avoids the common degeneracy problem found in traditional resampling algorithms that prevents precise tracking of a source. The algorithm is applied to an 8 microphone source localization problem and is shown to accurately localize a single speaker.
Jonathan Fillion-Deneault, Jean Rouat
ICASSP (4)2
2006 Wavelet Based Independent Component Analysis for Multi-Channel Source Separation
abstract
We consider the problem of separating instantaneous mixtures of different sound sources in multi-channel audio signals. Several methods have been developed to solve this problem. Independent component analysis (ICA) is certainly the most known method and the most used. ICA exploits the non-Gaussianity of the sources in the mixtures. In this study, we propose an improved signal separation algorithm where simultaneously we increase the non-Gaussian nature of signals and we initiate the preliminary separation. For this, the observations are transformed into an adequate representation using the wavelet packets decomposition. In this study, we consider the instantaneous mixture of two sources using two sensors. We validate our approach by using synthetic and recorded audio signals. Preliminary results show a strong improvement when compared to conventional ICA (FastICA), with specific signals.
Rachid Moussaoui, Jean Rouat, Roch Lefebvre
ICASSP (5)2
2006 Robust 3D Localization and Tracking of Sound Sources Using Beamforming and Particle Filtering
abstract
In this paper we present a new robust sound source localization and tracking method using an array of eight microphones (US patent pending). The method uses a steered beamformer based on the reliability-weighted phase transform (RWPHAT) along with a particle filter-based tracking algorithm. The proposed system is able to estimate both the direction and the distance of the sources. In a videoconferencing context, the direction was estimated with an accuracy better than one degree while the distance was accurate within 10% RMS. Tracking of up to three simultaneous moving speakers is demonstrated in a noisy environment
Jean-Marc Valin, François Michaud, Jean Rouat
ICASSP (4)3
2006 The oscillatory dynamic link matcher for spiking-neuron-based pattern recognition
Ramin Pichevar, Jean Rouat, Le Tan Thanh Tai
Neurocomputing2
2006 Wavelet speech enhancement based on time-scale adaptation
Mohammed Bahoura, Jean Rouat
Speech Commun.2
2005 Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative).
Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno
ICRA4
2005 Exploration of rank order coding with spiking neural networks for speech recognition
abstract
Speech recognition is very difficult in the context of noisy and corrupted speech. Most conventional techniques need huge databases to estimate speech (or noise) density probabilities to perform recognition. We discuss the potential of perceptive speech analysis and processing in combination with biologically plausible neural network processors. We illustrate the potential of such non-linear processing of speech by means of a preliminary test with recognition of French spoken digits from a small speech database
Stéphane Loiselle, Jean Rouat, Daniel Pressnitzer, Simon J. Thorpe
IJCNN2
2005 Automatic music genre classification using second-order statistical measures for the prescriptive approach
Hassan Ezzaidi, Jean Rouat
INTERSPEECH2
2005 Making a robot recognize three simultaneous sentences in real-time
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2004 Microphone array post-filter for separation of simultaneous non-stationary sources
abstract
Microphone array post-filters have demonstrated their ability to greatly reduce noise at the output of a beamformer. However, current techniques only consider a single source of interest, most of the time assuming stationary background noise. We propose a microphone array post-filter that enhances the signals produced by the separation of simultaneous sources using common source separation algorithms. Our method is based on a loudness-domain optimal spectral estimator and on the assumption that the noise can be described as the sum of a stationary component and of a transient component that is due to leakage between the channels of the initial source separation algorithm. The system is evaluated in the context of mobile robotics and is shown to produce better results than current post-filtering techniques, greatly reducing interference while causing little distortion to the signal of interest, even at very low SNR.
Jean-Marc Valin, Jean Rouat, François Michaud
ICASSP (1)2
2004 Localization of Simultaneous Moving Sound Sources for Mobile Robot Using a Frequency- Domain Steered Beamformer Approach
abstract
Mobile robots in real-life settings would benefit from being able to localize sound sources. Such a capability can nicely complement vision to help localize a person or an interesting event in the environment, and also to provide enhanced processing for other capabilities such as speech recognition. We present a robust sound source localization method in three-dimensional space using an array of 8 microphones. The method is based on a frequency-domain implementation of a steered beamformer along with a probabilistic post-processor. Results show that a mobile robot can localize in real time multiple moving sources of different types over a range of 5 meters with a response time of 200 ms.
Jean-Marc Valin, François Michaud, Brahim Hadjou, Jean Rouat
ICRA4
2004 Enhanced robot audition based on microphone array source separation with post-filter
abstract
We propose a system that gives a mobile robot the ability to separate simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of geometric source separation and a post-filter that gives us a further reduction of interferences from other sources. We present results and comparisons for separation of multiple non-stationary speech sources combined with noise sources. The main advantage of our approach for mobile robots resides in the fact that both the frequency domain geometric source separation algorithm and the post-filter are able to adapt rapidly to new sources and non-stationarity. Separation results are presented for three simultaneous interfering speakers in the presence of noise. A reduction of log spectral distortion (LSD) and increase of signal-to-noise ratio (SNR) of approximately 10 dB and 14 dB are observed.
Jean-Marc Valin, Jean Rouat, François Michaud
IROS2
2003 Robust sound source localization using a microphone array on a mobile robot
abstract
The hearing sense on a mobile robot is important because it is omnidirectional and it does not require direct line-of-sight with the sound source. Such capabilities can nicely complement vision to help localize a person or an interesting event in the environment. To do so the robot auditory system must be able to work in noisy, unknown and diverse environmental conditions. In this paper, we present a robust sound source localization method in three-dimensional space using an array of 8 microphones. The method is based on time delay of arrival estimation. Results show that a mobile robot can localize in real time different types of sound sources over a range of 3 meters and with a precision of 3/spl deg/.
Jean-Marc Valin, François Michaud, Jean Rouat, Dominic Létourneau
IROS3
2002 Speech, music and songs discrimination in the context of handsets variability
abstract
The problem of speech, music and music with songs discrimination in telephony with handsets variability is addressed in this paper. Two systems are proposed. The first system uses three Gaussian Mixture Models (GMM) for speech, music and songs respectively. Each GMM comprises 8 Gaussians trained on very short sessions. Twenty six speakers (13 females, 13 males) have been randomly chosen from the SPIDRE corpus. The music were obtained from a large set of data and comprises various styles. For 138 minutes of testing time, a speech discrimination score of 97.9% is obtained when no channel normalization is used. These performance are obtained for a relatively short analysis frame (32ms sliding window, buffering of 100 ms). When using channel normalization, an important score reduction (on the order of 10 to 20%) is observed. The second system has been designed for applications requiring shorter processing times along with shorter training sessions. It is based on an empirical transformation of the #MFCC that enhances the dynamical evolution of tonality. It yields in average an acceptable discrimination rate of 90% (speech- /music) and 84% (speech, music and songs with music).
Hassan Ezzaidi, Jean Rouat
INTERSPEECH2
2001 A new approach for wavelet speech enhancement
abstract
We propose a new approach to improve the performance of speech enhancement techniques based on wavelet thresholding. First, space–adaptation of the threshold is obtained by extending the principle of the level–dependent threshold to the Wavelet Packet Transform (WPT). Next, the time–adaptation is introduced using the Teager Energy Operator (TEO) of the wavelets coefficients. Finally, the time–space adapted threshold is proposed. Comparisons with the Ephraim and Malah Filter are reported.
Mohammed Bahoura, Jean Rouat
INTERSPEECH2
2001 Towards combining pitch and MFCC for speaker recognition systems
Hassan Ezzaidi, Jean Rouat, Douglas D. O'Shaughnessy
INTERSPEECH2
2001 Wavelet speech enhancement based on the Teager energy operator
abstract
We propose a new speech enhancement method based on the time adaption of wavelet thresholds. The time dependence is introduced by approximating the Teager energy of the wavelets coefficients. This technique does not require an explicit estimation of the noise level or of the a priori knowledge of the SNR, which is usually needed in most of the popular enhancement methods. Performance of the proposed method is evaluated on speech recorded in real conditions and with artificial noise.
Mohammed Bahoura, Jean Rouat
IEEE Signal Process. Lett.2
2000 Comparison of MFCC and pitch synchronous AM, FM parameters for speaker identification
Hassan Ezzaidi, Jean Rouat
INTERSPEECH2
1997 A Novelty Detector Using a Network of Integrate and Fire Neurons
Hô Tuòng Vinh, Jean Rouat
ICANN2
1997 Spatio-Temporal Pattern Recognition with Neural Networks: Application to Speech
Jean Rouat
ICANN1
1997 A new algorithm for double talk detection and separation in the context of digital mobile radio telephone
abstract
The paper describes a new technique that enhances the voice activity detection (VAD) performance between the remote speaker (received signal) and the local speaker (located in the vehicle) in the context of mobile radio telephone environment. We use an auditory pitch and voiced/unvoiced detection (APD) algorithm in conjunction with an autoregressive (AR) analysis in order to remove the remote speaker's voice signal from the car hands-free microphone signal. The results are compared with a reference system that doesn't include the APD.
Hassan Ezzaidi, Ivan Bourmeyster, Jean Rouat
ICASSP3
1997 A pitch determination and voiced/unvoiced decision algorithm for noisy speech
Jean Rouat, Yong Chun Liu, Daniel Morissette
Speech Commun.1
1996 Modeling neurons in the anteroventral cochlear nucleus for amplitude modulation (AM) processing: application to speech sound
Jean Rouat
ICSLP2
1995 A pitch determination and voiced/unvoiced decision algorithm for noisy speech
abstract
The design of a pitch tracking system for noisy speech is a challenging and yet unsolved issue due to the association of "traditional" pitch determination problems with those of noise processing. We have developed a multi-channel pitch determination algorithm (PDA) that has been tested on three speech databases (0dB SNR telephone speech, speech recorded in a car and clean speech) involving fifty-eight speakers. Our system has been compared to a multi-channel PDA based on auditory modelling (AMPEX), to hand-labelled and to laryngograph pitch contours. Our PDA is comprised of an automatic channel selection module and a pitch extraction module that relies on a pseudo-periodic histogram (combination of normalised scalar products for the less corrupted channels) in order to find pitch. Our PDA excelled in performance over the reference system on 0dB telephone and car speech. The automatic selection of channels was effective on the very noisy telephone speech (0dB) but performed less significantly on car speech where the robustness of the system is mainly due to the pitch extraction module in comparison to AMPEX. This paper reports in details the voiced/unvoiced, unvoiced/voiced performance and pitch estimation errors for the proposed PDA and the reference system while utilising three speech databases.
Jean Rouat, Yong Chun Liu, Daniel Morissette
EUROSPEECH1
1992 Conception of speech filters based on a neural network
A. Ennaji, Jean Rouat
ICSLP2
1992 A spectro-temporal analysis of speech based on nonlinear operators
abstract
This paper proposes a spectro-temporal analysis based on a bank of cochlea filters in combinaison with a nonlinear operator for amplitude modulation enhancement in the medium and high frequency formants. The output of the spectrotemporal analysis is represented as a 3D image where it is possible to observe very short-term speech transitions and formant modulations. With such analysis, it is possible to obtain patterns characteristics of phonemes and transitions betwen phonemes, which can not be obtained by using other speech analysis ( FFT, LPC ) techniques. The paper presents 3D images of vowels, where the amplitude modulation of the formants is clearly visable. 1.INTRODUCTION Research in speech analysis is recognized to be an important field in the area of speech processing, with applications in speech coding, speech recognition, etc. Depending on the application, the speech analyzer has to extract the most appropriate parameters. This paper proposes an analysis to enhance the modulation properties in speech. The automatic of speech with non-linear operators based on perceptive knowledge is a problem which has not yet been fully addressed, and speech might assist the researcher in understanding speech and / or in the design of an efficient speech analysis. 2.MODULATED TONE PERCEPTION Since the auditory system does not resolve the high frequency components, the temporal features of vowel-like sounds are coded similarly as those of amplitude-modulated tones. Furthermore, research work on automatic demodulation of speech can be motivated by the hypothesis proposing that the human brain has neural cells which specialize in Amplitude Modulation (AM) and Frequency Modulation (FM) detection [3][18][19]. More recently, Schreiner and Langner [6][17] have studied the representation of amplitude modulation in the inferior colliculus of cats and have shown that the inferior colliculus of the cat contains a highly systematic topographic representation of amplitude modulation paremeters. 3.BASILAR MEMBRANE NONLINEARITIES Nonlinearity and perception of intermodulation distortion products (f1-f2, 2f1-f2, etc.) are a pressing issue in hearing research and it is not easy to understand exactly the origin of these nonlinearities. Recently Robles, Ruggero and Rich have observed distortion products on chincilla basilar membrane b y using a laser-velocimetry technique [16]. Their work suggests that the lived basilar membrane is a nonlinear system and thus, the perception of distortion products could be due to the basilar membrane response and not only to the neural postprocessing. 4.NONLINEAR SPEECH PROCESSING Non linear processing The proposed analysis attempts to consider the automatic demodulation of the signal, before it is transformed in neural pulses in the cochlea. In fact, we will show than nonlinear operations of the signal create distorsion products and can enhance the modulation properties of the signal. Two nonlinear operators will be included in the proposed analysis to enhance the Amplitude Modulation observed with vowels. Nonlinear filtering seems to be very attractive and much work has been done in that field, refer to [9] for examples. More recently, one can cite the work by P. Maragos et al [7] where it is shown that the nonlinear operator, called Teager energy operator [5], allows AM and FM demodulation. Furthermore, L. Atlas and J. Fang [1] have shown that quadratic detectors allow for a better representation of speech in the context of a noisy pitch tracker. The originality of the present work resides in the combination of a perceptive bank of filters with nonlinear operators to obtain a 3D representation of speech with the A M information enhanced. Nonlinear operators J.F. Kaiser [5] proposes the Teager energy operator as beeing able to extract the energy of a signal based on mechanical and physical considerations. It has been shown [7] that this operator is able to track either the amplitude of an A.M. signal or the frequency of an FM signal. Another nonlinear operator has been proposed [14] to take into consideration the changes in the instantaneous signal power in the cochlea. This operator, called Dyn, shows the ability to enhance the AM-FM modulation in speech. Generally speaking, nonlinear operators are simple tools with the ability to modify the signal spectrum by combining the spectrum information. This ability is particularly interesting for AM or FM demodulation and for spectrum shifting, which are not easy to perform with standart linear techniques. Figure 1 illustrates the output of the Teager Energy and Dyn operators for two tones. The first section of figure 1 is a 600Hz tone, the second section is a 1000Hz tone. Section 3 is the sum of the 600Hz and 1000Hz tones. Sections 4 and 5 are respectively the output of the Dyn and Teager energy operators for the signal from section 3. Let us consider the combination tone defined as : s(t) = A1cos (w1t) + A2cos (w2t). By using the analog version of the Teager energy operator [4], one can show that: Teager [s(t)] = (A1w1)2 + (A2w2)2 + (A1A2)( w12!+!w22 2 w1w2) . cos[(w1 + w2) t] + (A1A2)( w12!+!w22 2 + w1w2) . cos[(w1 w2) t] (1) The amplitude difference between the two tones from the Teager output is equal to 2A1A2w1w2. Therefore, the w1 w2 component will largely dominate in comparison with the w1 + w2 component, as it is observed in section 5 from figure 1 where w1 w2 = 2p(1000-600) rad/s. Similarly, by using the analog form of the Dyn operator [13], one can show that : Dyn[s(t)] = A12!w1 2 sin (2w1t) A22!w2 2 sin (2w2t) A1A2( w1!+!w2 2 ) sin [(w1 + w2)t] A1A2( w1!-!w2 2 ) sin [(w1 w2)t] (2) By comparing equation (2) with figure 1, we observe that the component w1 + w2 = 2p(1000+600) rad/s is predominant in the output of Dyn for the composite signal. In summary, nonlinear operators are simple and powerful tools to obtain distorsion components from a sum of pure tones and might be used in speech processing where, some of the distorsion components might be perceptively important. 5.THE ANALYSIS OF SPEECH In this section, we will describe how a perceptive filterbank has been used in conjonction with the Teager energy or Dyn operators to generate a 3D representation of speech where the amplitude modulation of formants has been enhanced. Filtering The actual version of the analyzer is comprised of a bank of twenty-four filters centred on 330Hz to 4700Hz. These filters partially simulate the frequency analysis performed by the cochlea. These are rounded exponential filters with the Equivalent Rectangular Bandwidths (ERB) proposed b y Patterson [11] and Moore and Glasberg [10]. The output of each filter is a bandpass signal with a narrow-band spectrum centred around fi where fi is the central frequency (C.F.) of channel i. According to communication theory [2] the output signal si(t) from channel i can be considered to have been modulated in amplitude and phase with a carrier frequency of fi. si(t) = Ai(t) cos [wit+fi(t)] (3) Ai(t) is the modulating amplitude and fi(t) is the modulating phase. It should be noticed that equation (3) is true only for a bandpass signal (bandwidth of Ai(t) and fi(t) small in
Jean Rouat, Sylvain Lemieux, Alain Migneault
ICSLP1
1989 Speaker and mother tongue independent analysis and recognition of some nasals, liquids and fricatives for integration in an automatic speech recognition system
Jean Rouat, Renato De Mori, Jean-Pierre Adoul
EUROSPEECH1
1989 Property extraction for automatic speech recognition
Régis Cardin, Renato De Mori, Jean Rouat
Pattern Recognit. Lett.3
1988 A network of actions for automatic speech recognition
Renato De Mori, Régis Cardin, Ettore Merlo, Mathew J. Palakal, Jean Rouat
Speech Commun.5
1987 Use of Procedural Knowledge for Automatic Speech Recognition
Renato De Mori, Ettore Merlo, Mathew J. Palakal, Jean Rouat
IJCAI4