Jean-Luc Schwartz

dblp:49/1045 · DBLP profile ↗
← Back
59ranked-venue papers
7as first author
6since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 39 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 Bayesian gates: a probabilistic modeling tool for temporal segmentation of sensory streams into sequences of perceptual accumulators
Mamady Nabé, Jean-Luc Schwartz, Julien Diard
CogSci2
2022 Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal Imitation
abstract
We propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward model predicting the sensory consequences of articulatory commands, and an internal inverse model based on a recurrent neural network recovering articulatory commands from the acoustic speech input. Both forward and inverse models are jointly trained in a self-supervised way from raw acoustic-only speech data from different speakers. The imitation simulations are evaluated objectively and subjectively and display quite encouraging performances.
Marc-Antoine Georges, Julien Diard, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber
ICASSP4
2022 Orofacial somatosensory inputs in speech perceptual training modulate speech production
abstract
International audience
Monica Ashokumar, Jean-Luc Schwartz, Takayuki Ito 0002
INTERSPEECH2
2022 Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE
abstract
International audience
Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber
INTERSPEECH2
2022 Isochronous is beautiful? Syllabic event detection in a neuro-inspired oscillatory model is facilitated by isochrony in speech
abstract
International audience
Mamady Nabé, Julien Diard, Jean-Luc Schwartz
INTERSPEECH3
2021 Learning Robust Speech Representation with an Articulatory-Regularized Variational Autoencoder
abstract
It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task.
Marc-Antoine Georges, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber
Interspeech3
2020 Evaluating the Potential Gain of Auditory and Audiovisual Speech-Predictive Coding Using Deep Learning
abstract
Sensory processing is increasingly conceived in a predictive framework in which neurons would constantly process the error signal resulting from the comparison of expected and observed stimuli. Surprisingly, few data exist on the accuracy of predictions that can be computed in real sensory scenes. Here, we focus on the sensory processing of auditory and audiovisual speech. We propose a set of computational models based on artificial neural networks (mixing deep feedforward and convolutional networks), which are trained to predict future audio observations from present and past audio or audiovisual observations (i.e., including lip movements). Those predictions exploit purely local phonetic regularities with no explicit call to higher linguistic levels. Experiments are conducted on the multispeaker LibriSpeech audio speech database (around 100 hours) and on the NTCD-TIMIT audiovisual speech database (around 7 hours). They appear to be efficient in a short temporal range (25-50 ms), predicting 50% to 75% of the variance of the incoming stimulus, which could result in potentially saving up to three-quarters of the processing power. Then they quickly decrease and almost vanish after 250 ms. Adding information on the lips slightly improves predictions, with a 5% to 10% increase in explained variance. Interestingly the visual gain vanishes more slowly, and the gain is maximum for a delay of 75 ms between image and predicted sound.
Thomas Hueber, Eric Tatulli, Laurent Girin, Jean-Luc Schwartz
Neural Comput.4
2020 The Fharvard corpus: A phonemically-balanced French sentence resource for audiology and intelligibility research
abstract
The current study describes the collection of a new phonemically-balanced sentence resource for French, known as the Fharvard corpus. The resource consists of 700 sentences inspired by the original English Harvard sentences, along with audio recordings from one female and one male native French talker. Each of the sentences contains five mono- or bisyllabic keywords and are grouped into 70 lists of 10 sentences using an automatic phoneme-balancing procedure. Twenty-three normal-hearing French listeners identified keywords in the Fharvard sentences in speech-shaped noise. Psychometric functions for the Fharvard sentences indicate mean speech reception thresholds of −4.48 and −3.87 dB and slopes of 10.55 and 12.52 percentage points per dB at the 50% keywords correct point for the female and male talkers respectively. The complete list of Fharvard sentences and the associated audio recordings are available online for speech perception testing.
Vincent Aubanel, C. Bayard, Antje Strauß, Jean-Luc Schwartz
Speech Commun.4
2018 COSMO SylPhon: A Bayesian Perceptuo-motor Model to Assess Phonological Learning
abstract
International audience
Marie-Lou Barnaud, Julien Diard, Pierre Bessière, Jean-Luc Schwartz
INTERSPEECH4
2018 Picture Naming or Word Reading: Does the Modality Affect Speech Motor Adaptation and Its Transfer?
abstract
International audience
Tiphaine Caudrelier, Pascal Perrier, Jean-Luc Schwartz, Amélie Rochet-Capellan
INTERSPEECH3
2018 What drives the perceptual change resulting from speech motor adaptation? Evaluation of hypotheses in a Bayesian modeling framework
abstract
Shifts in perceptual boundaries resulting from speech motor learning induced by perturbations of the auditory feedback were taken as evidence for the involvement of motor functions in auditory speech perception. Beyond this general statement, the precise mechanisms underlying this involvement are not yet fully understood. In this paper we propose a quantitative evaluation of some hypotheses concerning the motor and auditory updates that could result from motor learning, in the context of various assumptions about the roles of the auditory and somatosensory pathways in speech perception. This analysis was made possible thanks to the use of a Bayesian model that implements these hypotheses by expressing the relationships between speech production and speech perception in a joint probability distribution. The evaluation focuses on how the hypotheses can (1) predict the location of perceptual boundary shifts once the perturbation has been removed, (2) account for the magnitude of the compensation in presence of the perturbation, and (3) describe the correlation between these two behavioral characteristics. Experimental findings about changes in speech perception following adaptation to auditory feedback perturbations serve as reference. Simulations suggest that they are compatible with a framework in which motor adaptation updates both the auditory-motor internal model and the auditory characterization of the perturbed phoneme, and where perception involves both auditory and somatosensory pathways.
Jean-François Patri, Pascal Perrier, Jean-Luc Schwartz, Julien Diard
PLoS Comput. Biol.3
2016 Assessing Idiosyncrasies in a Bayesian Model of Speech Communication
abstract
International audience
Marie-Lou Barnaud, Julien Diard, Pierre Bessière, Jean-Luc Schwartz
INTERSPEECH4
2016 Does Auditory-Motor Learning of Speech Transfer from the CV Syllable to the CVCV Word?
abstract
International audience
Tiphaine Caudrelier, Pascal Perrier, Jean-Luc Schwartz, Amélie Rochet-Capellan
INTERSPEECH3
2016 Audiovisual Speech Scene Analysis in the Context of Competing Sources
abstract
International audience
Attigodu C. Ganesh, Frédéric Berthommier, Jean-Luc Schwartz
INTERSPEECH3
2014 No, There Is No 150 ms Lead of Visual Speech on Auditory Speech, but a Range of Audiovisual Asynchronies Varying from Small Audio Lead to Large Audio Lag
abstract
An increasing number of neuroscience papers capitalize on the assumption published in this journal that visual speech would be typically 150 ms ahead of auditory speech. It happens that the estimation of audiovisual asynchrony in the reference paper is valid only in very specific cases, for isolated consonant-vowel syllables or at the beginning of a speech utterance, in what we call "preparatory gestures". However, when syllables are chained in sequences, as they are typically in most parts of a natural speech utterance, asynchrony should be defined in a different way. This is what we call "comodulatory gestures" providing auditory and visual events more or less in synchrony. We provide audiovisual data on sequences of plosive-vowel syllables (pa, ta, ka, ba, da, ga, ma, na) showing that audiovisual synchrony is actually rather precise, varying between 20 ms audio lead and 70 ms audio lag. We show how more complex speech material should result in a range typically varying between 40 ms audio lead and 200 ms audio lag, and we discuss how this natural coordination is reflected in the so-called temporal integration window for audiovisual speech perception. Finally we present a toy model of auditory and audiovisual predictive coding, showing that visual lead is actually not necessary for visual prediction.
Jean-Luc Schwartz, Christophe Savariaux
PLoS Comput. Biol.1
2013 Effect of context, rebinding and noise, on audiovisual speech fusion
Ganesh Attigodu Chandrashekara, Frédéric Berthommier, Olha Nahorna, Jean-Luc Schwartz
INTERSPEECH4
2013 A computational model of perceptuo-motor processing in speech perception: learning to imitate and categorize synthetic CV syllables
abstract
This paper presents COSMO, a Bayesian computational model, which is expressive enough to carry out syllable production, perception and imitation tasks using motor, auditory or perceptuo-motor information.An imitation algorithm enables to learn the articulatory-to-acoustic mapping and the link between syllables and corresponding articulatory gestures, from acoustic inputs only: synthetic CV syllables generated with a human vocal tract model.We compare purely auditory, purely motor and perceptuo-motor syllable categorization under various noise levels.
Raphaël Laurent, Jean-Luc Schwartz, Pierre Bessière, Julien Diard
INTERSPEECH2
2010 Speech and face-to-face communication - An introduction
Marion Dohen, Jean-Luc Schwartz, Gérard Bailly
Speech Commun.2
2008 Invariance and variability in the production of the height feature in French vowels
Lucie Ménard, Jean-Luc Schwartz, Jérôme Aubin
Speech Commun.2
2007 Pointing to a target while naming it with /pata/ or /tapa/: the effect of consonants and stress position on jaw-finger coordination
abstract
International audience
Amélie Rochet-Capellan, Jean-Luc Schwartz, Rafael Laboissière, Arturo Galvàn
INTERSPEECH2
2006 An Analysis of Visual Speech Information Applied to Voice Activity Detection
abstract
We present a new approach to the voice activity detection (VAD) problem for speech signals embedded in non-stationary noise. The method is based on automatic lipreading: the objective is to detect voice activity or non-activity by exploiting the coherence between the speech acoustic signal and the speaker's lip movements. From a comprehensive analysis of lip shape parameters during speech and non-speech events, we show that a single appropriate visual parameter, defined to characterize the lip movements, can be used for the detection of sections of voice activity or more precisely, for the detection of silence sections. Detection scores obtained on spontaneous speech confirm the efficiency of the visual voice activity detector (VVAD)
David Sodoyer, Bertrand Rivet, Laurent Girin, Jean-Luc Schwartz, Christian Jutten
ICASSP (1)4
2005 The labial-coronal effect and CVCV stability during reiterant speech production: an acoustic analysis
abstract
International audience
Amélie Rochet-Capellan, Jean-Luc Schwartz
INTERSPEECH2
2005 The labial-coronal effect and CVCV stability during reiterant speech production: an articulatory analysis
abstract
In a companion paper [1], we showed that CVCV utterances with a labial consonant followed by a coronal one (LC sequences) are more stable than reverse CL sequences in speeded reiterant speech. We proposed that this could explain why human languages select LC sequences more often than CL ones (the “LC effect”). We provide here articulatory data explaining where the greater LC stability could come from, by investigating inter-articulator coordination during LC and CL utterances at an increasing rate. Rate increase leads variegated CVCV (e.g. /pata/) to be produced in a single jaw cycle but this is not the case for duplicated CVCV (e.g. /papa/. Furthermore, LC and CL sequences both evolve towards the same cycle with a progressive phasing of lips and tongue close together in the jaw cycle. Taken together, these results provide new elements to argue for motor control constraints shaping phonological patterns from economy principles.
Amélie Rochet-Capellan, Jean-Luc Schwartz
INTERSPEECH2
2005 Asymmetries in vowel perception, in the context of the Dispersion-Focalisation Theory
Jean-Luc Schwartz, Christian Abry, Louis-Jean Boë, Lucie Ménard, Nathalie Vallée
Speech Commun.1
2004 Modeling audio-visual speech perception: back on fusion architectures and fusion control
abstract
In a review paper about audio-visual (AV) fusion models in speech perception, we (Schwartz et al., 1998) proposed a taxonomy of models around two basic questions: architecture and control. Six years after, it appears that the proposals we made still seem rather convenient for discussing major questions about AV fusion. Moreover – and more importantly – recent experimental and theoretical progress seem to provide some elements of answer in both aspects. The aim of this paper is to review these elements, and to incorporate them into the general architecture-and-control framework. 1. FUSION ARCHITECTURES 1.1. The four architectures for audio-visual fusion In his well-known presentation of audio-visual models of speech perception, Summerfield (1987) introduced the concept of “metrics for audio-visual integration”, focusing on “the representations of the auditory and visual streams of information at their conflux”. The general literature on sensory interactions in cognitive psychology, and on sensor fusion in information processing, lead us conclude that there are four basic architectures (Schwartz et al., 1998). Their common point is that they should connect two separate inputs to one common output. The conception of a single output “loosing” in some sense the monosensorial nature of each input may be discussed (see the “convergence vs. association ” debate raised by Bernstein et al., in press). However, even in a conception of two separate routes in interaction from the input to the output, the questions addressed in this section remain valid, provided that they are rephrased in terms of: under what format are the A and V inputs represented in their sensory pathway when they interact in the route towards phonology or lexicon? In the Separate Identification (SI) model, the A and V representations are phonetic, that is mediated by the knowledge the subject has of his/her own language. In the Dominant Recoding (DR) model, the visual input is recoded into an equivalent sound, or into some of its spectro-temporal characteristics. In the Motor Recoding (MR) model, both the A and V inputs are in contact with a system analysing percepts in terms of the action able to have produced them. In the Direct Identification (DI) model, none of these process occur before phonetic identification which operates directly on a set of joined A and V parameters. In our view, the static vs. dynamic issue (or the shape vs. movement debate) is independent of the architecture. In consequence, the preference for static or dynamic parameters, if any, should not lead to select one or the other architecture. Notice that this position is itself controversial, and it has been often argued that movement and motor representation were linked topics (see e.g. Rosenblum & Saldana, 1998; Whalen et al., in press). However, we have several times advocated that recovering the vocal tract shape could be done without necessarily calling for dynamic features (Cathiard et al., 1996), and proposed a “shape from shading from movement” approach in line with recent neurophysiological computational models (Cathiard et al., 2003). 1.2. The “very early” route The four architectures share a common assumption of independence of the primitive monosensorial processing. That is, information would be first extracted separately in each sensorial channel before interaction and fusion. However, a number of recent studies on the detection of speech in noise have raised serious doubts about this assumption (since Grant and Seitz, 2000). Our own contribution was to determine if this gain in detection could contribute to a gain in identification. In a recent study (Schwartz et al., 2004), we showed, thanks to an original paradigm, that seeing the speaker’s lips does enable to better hear and hence better understand. The stimuli used in this set of experiments could not lead to lipreading per se since they corresponded to exactly the same lip gesture. However, intelligibility of these stimuli merged in noise was improved just because the acoustic cues were better extracted thanks to vision. The experimental trick consisted in dubbing the same lip gesture on a number of visually similar but auditorily different configurations, e.g. [y u ty tu ky ku dy du gy gu] in French. The visual stimulus did not enable to identify the syllable, but it provided a temporal cue improving the audio identification of these stimuli embedded in a large level of cocktail-party noise, and particularly the identification of plosive voicing. Replacing the visual speech cue (the lip rounding gesture) by a non-speech one with the same temporal pattern (a red bar on a black background, increasing and decreasing in synchrony with the lips) removed the benefit. Therefore, cross-modal interactions can occur early to enhance speech in noise and improve intelligibility. This indicates that there is, whatever the architecture, a preliminary set of interactions, that we called “very early” to make clear that they correspond to a contact point that should be distinguished from early interactions in the classical sense. This is likely to provide a number of interesting technological counterparts in terms of speech enhancement, source separation and audiovisual scene analysis (e.g. Girin et al., 2001; Sodoyer et al., 2002; Berthommier, 2003).
Jean-Luc Schwartz, Marie-Agnès Cathiard
INTERSPEECH1
2004 Using audiovisual speech processing to improve the robustness of the separation of convolutive speech mixtures
abstract
Looking at the speaker's face seems useful in hearing better a speech signal and extract it from the competing sources before identification. In this paper, we present a novel algorithm plugging audiovisual coherence of speech signals, estimated by statistical tools, on audio blind source separation (BSS) algorithms in the difficult case of convolutive mixtures. The algorithm mainly works in the frequency (transform) domain, where the convolutive mixture becomes an additive mixture for each frequency channel. Frequency by frequency separation is made by an audio BSS algorithm, and the audiovisual information is used to solve the standard source permutation problem at the output of the separation stage, for each frequency. The proposed method is shown to be efficient in the case of 2 /spl times/ 2 convolutive mixtures.
Bertrand Rivet, Laurent Girin, Christian Jutten, Jean-Luc Schwartz
MMSP4
2004 Visual perception of contrastive focus in reiterant French speech
Marion Dohen, Hélène Loevenbruck, Marie-Agnès Cathiard, Jean-Luc Schwartz
Speech Commun.4
2004 Editorial
Jean-Luc Schwartz, Frédéric Berthommier, Marie-Agnès Cathiard, Renato De Mori
Speech Commun.1
2004 Developing an audio-visual speech source separation algorithm
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz
Speech Commun.4
2003 Potential audiovisual correlates of contrastive focus in French
Marion Dohen, Hélène Loevenbruck, Marie-Agnès Cathiard, Jean-Luc Schwartz
INTERSPEECH4
2003 Extracting an AV speech source from a mixture of signals
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz
INTERSPEECH4
2002 Special session: issues in audiovisual spoken language processing (when, where, and how?)
Lynne E. Bernstein, Denis Burnham, Jean-Luc Schwartz
INTERSPEECH3
2002 Intrasyllabic articulatory control constraints in verbal working memory
Marc Sato, Jean-Luc Schwartz, Marie-Agnès Cathiard, Christian Abry, Hélène Loevenbruck
INTERSPEECH2
2002 Audio-visual scene analysis: evidence for a "very-early" integration process in audio-visual speech perception
Jean-Luc Schwartz, Frédéric Berthommier, Christophe Savariaux
INTERSPEECH1
2002 Motor specifications of a baby robot via the analysis of infants² vocalizations
Jihène Serkhane, Jean-Luc Schwartz, Louis-Jean Boë, Barbara L. Davis, Christine L. Matyear
INTERSPEECH2
2002 Audio-visual speech sources separation: a new approach exploiting the audio-visual coherence of speech stimuli
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz
INTERSPEECH4
2001 Perceptual identification and normalization of synthesized French vowels from birth to adulthood
Lucie Ménard, Jean-Luc Schwartz, Louis-Jean Boë, Sonia Kandel, Nathalie Vallée
INTERSPEECH2
2001 Speech signals separation: a new approach exploiting the coherence of audio and visual speech
abstract
We present a new approach to the source separation problem in the case of multiple speech signals. The method is based on the use of automatic lip reading: the objective is to extract an acoustic speech signal from other acoustic signals by exploiting its coherence with the speaker's lip movements. For this aim, a statistical model is used to quantify this coherence. The results, while very preliminary, are encouraging. They show that this method can achieve a good separation of a speech source in the case of simple 2/spl times/2 additive mixtures. Moreover, it presents some interesting complementarity with traditional pure audio techniques.
Laurent Girin, A. Allard, Jean-Luc Schwartz
MMSP3
1999 A reliability criterion for time-frequency labeling based on periodicity in an auditory scene
François Gaillard, Frédéric Berthommier, Gang Feng 0002, Jean-Luc Schwartz
EUROSPEECH4
1999 Comparing models for audiovisual fusion in a noisy-vowel recognition task
abstract
Audiovisual speech recognition involves fusion of the audio and video sensors for phonetic identification. There are three basic ways to fuse data streams for taking a decision such as phoneme identification: data-to-decision, decision-to-decision, and data-to-data. This leads to four possible models for audiovisual speech recognition, that is direct identification in the first case, separate identification in the second one, and two variants of the third early integration case, namely dominant recoding or motor recoding. However, no systematic comparison of these models is available in the literature. We propose an implementation of these four models, and submit them to a benchmark test. For this aim, we use a noisy-vowel corpus tested on two recognition paradigms in which the systems are tested at noise levels higher than those used for learning. In one of these paradigms, the signal-to-noise ratio (SNR) value is provided to the recognition systems, in the other it is not. We also introduce a new criterion for evaluating performances, based on transmitted information on individual phonetic features. In light of the compared performances of the four models with the two recognition paradigms, we discuss the advantages and drawbacks of these models, leading to proposals for data representation, fusion architecture, and control of the fusion process through sensor reliability.
Pascal Teissier, Jordi Robert-Ribes, Jean-Luc Schwartz, Anne Guérin-Dugué
IEEE Trans. Speech Audio Process.3
1998 Fusion of auditory and visual information for noisy speech enhancement: a preliminary study of vowel transitions
abstract
This paper deals with a noisy speech enhancement technique based on the fusion of auditory and visual information. We first present the global structure of the system, and then we focus on the tool we used to melt both sources of information. The whole noise reduction system is implemented in the context of vowel transitions corrupted with white noise. A complete evaluation of the system in this context is presented, including distance measures, Gaussian classification scores, and a perceptive test. The results are very promising.
Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz
ICASSP3
1998 A signal processing system for having the sound "pop-out" in noise thanks to the image of the speaker's lips: new advances using multi-layer perceptrons
Laurent Girin, Laurent Varin, Gang Feng 0002, Jean-Luc Schwartz
ICSLP4
1998 Audiovisual speech enhancement: new advances using multi-layer perceptrons
abstract
This paper deals with the improvement of a noisy speech enhancement system based on the fusion of auditory and visual information. The system was presented in previous papers and implemented with a simple stimuli corrupted with white noise. Its principle consists of an analysis-enhancement-synthesis process based on a linear prediction (LP) model of the signal: the LP filter is enhanced thanks to associative tools that estimate the LP cleaned parameters from both noisy audio and lip shape information. The structure of the system is reviewed and we focus on the improvement that concerns the associators: multi-layers perceptrons are used instead of linear regression. It is shown that in the context of VCV transitions corrupted with white noise, the performances of the system are improved in terms of the intelligibility gain, distance measures and classification tests.
Laurent Girin, Laurent Varin, Gang Feng 0002, Jean-Luc Schwartz
MMSP4
1997 A modified zero-crossing method for pitch detection in presence of interfering sources
François Gaillard, Frédéric Berthommier, Gang Feng 0002, Jean-Luc Schwartz
EUROSPEECH4
1997 Noisy speech enhancement by fusion of auditory and visual information: a study of vowel transitions
abstract
This paper deals with a noisy speech enhancement technique based on the fusion of auditory and visual information. We first present the global structure of the system, and then we focus on the tool we used to melt both sources of information. The whole noise reduction system is implemented in the context of vowel transitions corrupted with white noise. A complete evaluation of the system in this context is presented, including distance measures, gaussian classification scores, and a perceptive test. The results are very promising.
Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz
EUROSPEECH3
1997 Non-linear representations, sensor reliability estimation and context-dependent fusion in the audiovisual recognition of speech in noise
abstract
The paper involves the recognition of French audiovisual vowels at various signal-to-noise ratios (SNRs). It deals with a new non-linear preprocessing of the audio data which enables an estimation of the reliability of the audio sensor in relation to SNR, and a significant increase in the recognition performances at the output of the fusion process.
Pascal Teissier, Jean-Luc Schwartz, Anne Guérin-Dugué
EUROSPEECH2
1997 Models for audiovisual fusion in a noisy-vowel recognition task
abstract
This paper presents a comparison of four basic architectures dealing with audiovisual speech in a noisy-vowel recognition task. Provided contextual input (signal-to-noise ratio), three of the four architectures respect the "synergy" criterion which means that audiovisual (AV) recognition is better than audio-alone (A) or visual-alone (V) recognition, both in global terms and for each individual phonetic feature. Without contextual input, the performances collapse, but we propose for one model an original approach using an efficient non-linear data processing which leads to more simple algorithms and increases performances of the audiovisual fusion operator.
Pascal Teissier, Jean-Luc Schwartz, Anne Guérin-Dugué
MMSP2
1995 Noisy speech enhancement with filters estimated from the speaker's lips
Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz
EUROSPEECH3
1993 Integrating auditory and visual representations for audiovisual vowel recognition
Jordi Robert-Ribes, Tahar Lallouache, Pierre Escudier, Jean-Luc Schwartz
EUROSPEECH4
1993 An information theoretical investigation into the distribution of phonetic information across the auditory spectrogram
Andrew C. Morris, Jean-Luc Schwartz, Pierre Escudier
Comput. Speech Lang.2
1991 On and off units detect information bottle-necks for speech recognition
Andrew C. Morris, Pierre Escudier, Jean-Luc Schwartz
EUROSPEECH3
1990 Modeling spectral processing in the central auditory system
abstract
Several models of the central nervous system (CNS) which further process the spatio-temporal firing pattern of the auditory nerve are presented. All these models suppose that for the CNS, the auditory nerve fibers are organized as individual groups whose central frequencies vary continuously, each individual group being composed of localized neighboring fibers, and that what the CNS does is to measure the coherence and cooperation of activities of fibers in each individual group, i.e. local coherence. The local coherence is in turn defined as the sum of each individual coherence which is a measure of coherence between the activities of two neighboring fibers. An internal representation of a stimulus, given at the output of the model, is thus composed of a series of measures of local coherence and cooperation. Two derivatives of this general conceptual model are described in detail, i.e. when the individual coherence is based on either cross correlation analysis or covariance.>
Zong Liang Wu, Jean-Luc Schwartz, Pierre Escudier
ICASSP2
1989 Specialized physiology-based channels for the detection of articulatory-acoustic events. A preliminary scheme and its performance
abstract
Neural mechanisms for the detection of articulatory-acoustic events, which are based on a model of the peripheral auditory system, are studied. Several specialized neural networks, composed of 'on' neurons and large-scale spatial integration mechanisms, are proposed for further processing the firing pattern of the auditory nerve. It is shown that although they are very simple in structure and preliminary in nature, these channels are physiologically plausible and powerful enough for the detection of most of the articulatory-acoustic events considered.>
Zong Liang Wu, Pierre Escudier, Jean-Luc Schwartz
ICASSP3
1989 Auditory processing in a post-cochlear neural network vowel spectrum processing based on spike synchrony
Frédéric Berthommier, Jean-Luc Schwartz, Pierre Escudier
EUROSPEECH2
1989 Maximal vowel space
Louis-Jean Boë, Pascal Perrier, Bernard Guérin, Jean-Luc Schwartz
EUROSPEECH4
1989 Perceptual contract and stability in vowel systems: a 3-d simulation study
Jean-Luc Schwartz, Louis-Jean Boë, Pascal Perrier, Bernard Guérin, Pierre Escudier
EUROSPEECH1
1989 A theoretical study of neural mechanisms specialized in the detection of articulatory-acoustic events
Zong Liang Wu, Jean-Luc Schwartz, Pierre Escudier
EUROSPEECH2
1989 A strong evidence for the existence of a large-scale integrated spectral representation in vowel perception
Jean-Luc Schwartz, Pierre Escudier
Speech Commun.1
1985 Pulsation threshold patterns of synthetic vowels: Study of the second formant emergence and the "center of gravity" effects
Pierre Escudier, Jean-Luc Schwartz
Speech Commun.2