EDBT 2026 Demo / reviewers in the wild / expert
Bayya Yegnanarayana
dblp:43/1564 · also B. Yegnanarayana 0001
· DBLP profile ↗
200ranked-venue papers
37as first author
4since 2021 · last 2023
0000-0002-7080-1239ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 144 · 27 first-authorArtificial intelligence and machine learning · 115 · 18 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
21 papers |
Audio and music processing · 92% Image and video processing · 5% Multimedia analysis and retrieval · 1% | |
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 45% Face, body and person analysis · 19% Video understanding and tracking · 14% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speech processing |
1.5 | 13 | 2021 | Extraction and Utilization of Excitation Information of Speech: A Review · Proc. IEEE 2021 Extraction of Fundamental Frequency From Degraded Speech Using Temporal Envelopes at High SNR Frequencies · IEEE ACM Trans. Audio Speech Lang. Process. 2017 Single Frequency Filtering Approach for Discriminating Speech and Nonspeech · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing
speech analysis |
0.9 | 11 | 2021 | Extraction and Utilization of Excitation Information of Speech: A Review · Proc. IEEE 2021 Event-Based Instantaneous Fundamental Frequency Estimation From Speech Signals · IEEE Trans. Speech Audio Process. 2009 Epoch Extraction From Speech Signals · IEEE Trans. Speech Audio Process. 2008 |
Audio and music processing › speech analysis
fundamental frequency estimation |
0.5 | 3 | 2017 | Extraction of Fundamental Frequency From Degraded Speech Using Temporal Envelopes at High SNR Frequencies · IEEE ACM Trans. Audio Speech Lang. Process. 2017 Performance of an Event-Based Instantaneous Fundamental Frequency Estimator for Distant Speech Signals · IEEE Trans. Speech Audio Process. 2011 Event-Based Instantaneous Fundamental Frequency Estimation From Speech Signals · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing › speech processing
speech/nonspeech discrimination |
0.2 | 1 | 2015 | Single Frequency Filtering Approach for Discriminating Speech and Nonspeech · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing › speech processing
voice activity detection |
0.2 | 1 | 2015 | Single Frequency Filtering Approach for Discriminating Speech and Nonspeech · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing › sound source localization
time delay estimation |
0.1 | 2 | 2009 | Determining Mixing Parameters From Multispeaker Data Using Speech-Specific Information · IEEE Trans. Speech Audio Process. 2009 Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › speech analysis
glottal closure instant detection |
0.1 | 3 | 2008 | Epoch Extraction From Speech Signals · IEEE Trans. Speech Audio Process. 2008 Robustness of group-delay-based method for extraction of significant instants of excitation from speech signals · IEEE Trans. Speech Audio Process. 1999 Determination of instants of significant excitation in speech using group delay function · IEEE Trans. Speech Audio Process. 1995 |
Natural language and speech › Speech recognition and synthesis › voice conversion
spectral mapping |
0.1 | 1 | 2010 | Spectral Mapping Using Artificial Neural Networks for Voice Conversion · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis
voice conversion |
0.1 | 1 | 2010 | Spectral Mapping Using Artificial Neural Networks for Voice Conversion · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › source separation
speech separation |
0.1 | 1 | 2009 | Determining Mixing Parameters From Multispeaker Data Using Speech-Specific Information · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing › source separation › blind source separation
underdetermined source separation |
0.1 | 1 | 2009 | Determining Mixing Parameters From Multispeaker Data Using Speech-Specific Information · IEEE Trans. Speech Audio Process. 2009 |
Computer vision › Video understanding and tracking
activity recognition |
0.1 | 1 | 2008 | Activity Modeling Using Event Probability Sequences · IEEE Trans. Image Process. 2008 |
Machine learning › Time series and sequential data
anomaly detection |
0.1 | 1 | 2008 | Activity Modeling Using Event Probability Sequences · IEEE Trans. Image Process. 2008 |
Audio and music processing
speech enhancement |
0.1 | 2 | 2005 | Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 Enhancement of reverberant speech using LP residual signal · IEEE Trans. Speech Audio Process. 2000 |
Audio and music processing
speech synthesis |
0.1 | 2 | 2006 | Prosody modification using instants of significant excitation · IEEE Trans. Speech Audio Process. 2006 Source-system windowing for speech analysis and synthesis · IEEE Trans. Speech Audio Process. 1996 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2007 | Face Verification Using Template Matching · IEEE Trans. Inf. Forensics Secur. 2007 |
Computer vision › Face, body and person analysis › face recognition
face verification |
0.1 | 1 | 2007 | Face Verification Using Template Matching · IEEE Trans. Inf. Forensics Secur. 2007 |
Computer vision › Image recognition and object detection
template matching |
0.1 | 1 | 2007 | Face Verification Using Template Matching · IEEE Trans. Inf. Forensics Secur. 2007 |
Image and video processing › image fusion
feature fusion |
0.1 | 1 | 2005 | Combining evidence from source, suprasegmental and spectral features for a fixed-text speaker verification system · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing
microphone array processing |
0.1 | 1 | 2005 | Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › speaker recognition
speaker verification |
0.1 | 1 | 2005 | Combining evidence from source, suprasegmental and spectral features for a fixed-text speaker verification system · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › speaker recognition › speaker verification
text-dependent speaker verification |
0.1 | 1 | 2005 | Combining evidence from source, suprasegmental and spectral features for a fixed-text speaker verification system · IEEE Trans. Speech Audio Process. 2005 |
Wireless sensing and localization
acoustic source localization |
0.1 | 1 | 2005 | Speaker Localization Using Excitation Source Information in Speech · IEEE Trans. Speech Audio Process. 2005 |
Physical-layer communications › signal processing for communications › statistical signal processing › estimation theory
delay estimation |
0.1 | 1 | 2005 | Speaker Localization Using Excitation Source Information in Speech · IEEE Trans. Speech Audio Process. 2005 |
Image and video processing
feature extraction |
0.0 | 1 | 2004 | Finding axes of symmetry from potential fields · IEEE Trans. Image Process. 2004 |
Multimedia analysis and retrieval
image analysis |
0.0 | 1 | 2004 | Finding axes of symmetry from potential fields · IEEE Trans. Image Process. 2004 |
Geometric modeling and processing › shape analysis
symmetry detection |
0.0 | 1 | 2004 | Finding axes of symmetry from potential fields · IEEE Trans. Image Process. 2004 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 1 | 2011 | Performance of an Event-Based Instantaneous Fundamental Frequency Estimator for Distant Speech Signals · IEEE Trans. Speech Audio Process. 2011 |
Audio and music processing › speech synthesis
vocal tract modeling |
0.0 | 2 | 1998 | Extraction of vocal-tract system characteristics from speech signals · IEEE Trans. Speech Audio Process. 1998 Source-system windowing for speech analysis and synthesis · IEEE Trans. Speech Audio Process. 1996 |
Image and video processing
image segmentation |
0.0 | 2 | 1997 | Unsupervised texture classification using vector quantization and deterministic relaxation neural network · IEEE Trans. Image Process. 1997 Segmentation of Gabor-filtered textures using deterministic relaxation · IEEE Trans. Image Process. 1996 |
Methods — techniques the papers use, named apart from their topics
single frequency filtering · 0.5deep learning · 0.5autocorrelation · 0.3zero-crossing analysis · 0.2short-time spectrum · 0.2resonator cascade filtering · 0.2adaptive multi-rate VAD · 0.2excitation source feature extraction · 0.1gaussian mixture model · 0.1artificial neural network · 0.1zero-crossing detection · 0.1zero frequency resonator · 0.1hidden markov model · 0.1event probability sequence · 0.1autoassociative neural network · 0.11-d image processing · 0.1generalized cross correlation · 0.1excitation source features · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Analysis of Instantaneous Frequency Components of Speech Signals for Epoch ExtractionabstractThe major impulse-like excitation in the speech signal is due to abrupt closure of the vocal folds, which takes place at the glottal closure instant (GCI) or epoch in each cycle. GCIs are used in many areas of speech science and technology, such as in prosody modification, voice source analysis, formant extraction and speech synthesis. It is difficult to observe these discontinuities (corresponding to GCIs) in the speech signal because of the superimposed time-varying response of the vocal tract system. This paper examines the phase part of different frequency components of the speech signal to extract epochs. Three analysis methods to decompose the speech signal into different frequency components are considered. These methods are the short-time Fourier transform (STFT), narrow bandpass filtering (NBPF), and single frequency filtering (SFF). The locations of the discontinuities in the speech signal are obtained from the instantaneous frequency (IF) (i.e., the time derivative of the phase) of each of the frequency components. A method for automatic detection of epochs using the amplitude weighted IF is proposed. Performance of the proposed epoch detection method is compared with four state-of-the-art methods in clean and telephone quality speech. The performance of the proposed method is comparable with the performance of the existing epoch detection methods for clean speech but better for telephone quality speech. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Comput. Speech Lang. | 3 |
| 2021 | A neural network approach for speech activity detection for Apollo corpus
Vishala Pannala, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2021 | A study of vowel nasalization using instantaneous spectra
RaviShankar Prasad, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2021 | Extraction and Utilization of Excitation Information of Speech: A ReviewabstractSpeech production can be regarded as a process where a time-varying vocal tract system (filter) is excited by a time-varying excitation. In addition to its linguistic message, the speech signal also carries information about, for example, the gender and age of the speaker. Moreover, the speech signal includes acoustical cues about several speaker traits, such as the emotional state and the state of health of the speaker. In order to understand the production of these acoustical cues by the human speech production mechanism and utilize this information in speech technology, it is necessary to extract features describing both the excitation and the filter of the human speech production mechanism. While the methods to estimate and parameterize the vocal tract system are well established, the excitation appears less studied. This article provides a review of signal processing approaches used for the extraction of excitation information from speech. This article highlights the importance of excitation information in the analysis and classification of phonation type and vocal emotions, in the analysis of nonverbal laughter sounds, and in studying pathological voices. Furthermore, recent developments of deep learning techniques in the context of extraction and utilization of the excitation information are discussed. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Proc. IEEE | 3 |
| 2020 | Comparison of Glottal Closure Instants Detection Algorithms for Emotional SpeechabstractIn production of voiced speech, epochs or glottal closure instants (GCIs) refer to the instants of significant excitation of the vocal tract. Extraction of GCIs is used as a pre-processing stage in many areas of speech technology, such as in prosody modification, speech synthesis and voice source analysis. In the past decades, several GCI detection algorithms have been developed and most of them provide excellent results for speech signals produced using modal (normal) type of phonation. There are, however, no studies comparing multiple state-of-the-art GCI detection methods in emotional speech. In this paper, we compare six GCI detection algorithms using emotional speech and known evaluation metrics. We use the Berlin EMO-DB acted emotional speech database which contains seven emotions and simultaneous electroglottography (EGG) recordings as ground truth. The results show that all six GCI detection algorithms give best performance in processing speech of neutral emotion and that the performance degrade particularly in emotions of high arousal (anger and joy). To improve the performance of GCI detection in emotional speech, the study underlines the importance of local average pitch period estimates. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
ICASSP | 3 |
| 2020 | Instantaneous Time Delay Estimation of Broadband Signals
B. H. V. S. Narayanamurthy, J. V. Satyanarayana, Nivedita Chennupati, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2020 | Enhancing Formant Information in Spectrographic Display of Speech
Bayya Yegnanarayana, Joseph M. Anand, Vishala Pannala |
INTERSPEECH | 1 |
| 2020 | Determination of glottal closure instants from clean and telephone quality speech signals using single frequency filtering
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2020 | Analysis and classification of phonation types in speech and singing voice
Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2020 | Detection of glottal closure instant and glottal open region from speech signals using spectral flatness measure
Sudarsana Reddy Kadiri, RaviShankar Prasad, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2019 | Spectral and temporal manipulations of SFF envelopes for enhancement of speech intelligibility in noise
Nivedita Chennupati, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Comput. Speech Lang. | 3 |
| 2018 | Detection of Glottal Closure Instants in Degraded Speech Using Single Frequency Filtering Analysis
Gunnam Aneeja, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2018 | Breathy to Tense Voice Discrimination using Zero-Time Windowing Cepstral Coefficients (ZTWCCs)
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Analysis and Detection of Phonation Modes in Singing Voice using Excitation Source Features and Single Frequency Filtering Cepstral Coefficients (SFFCC)
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Estimation of Fundamental Frequency from Singing Voice Using Harmonics of Impulse-like Excitation Source
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Determining Speaker Location from Speech in a Practical Environment
B. H. V. S. Narayanamurthy, J. V. Satyanarayana, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2018 | Discriminating Nasals and Approximants in English Language Using Zero Time Windowing
RaviShankar Prasad, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2018 | Identification and Classification of Fricatives in Speech Using Zero Time Windowing Method
RaviShankar Prasad, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Significance of phase in single frequency filtering outputs of speech signals
Nivedita Chennupati, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2017 | Speech polarity detection using strength of impulse-like excitation extracted from speech epochsabstractIn this paper, we address the issue of speech polarity detection using strength of impulse-like excitation around epoch. The correct detection of speech polarity is a crucial step for many speech processing algorithms to extract suitable information. Occurrence of errors in the detection of speech polarity could have an impact on the performance of speech systems. Automatic detection of speech polarity has become an important preliminary step for many speech processing algorithms. We propose a method based on the knowledge of impulse-like excitation of speech production mechanism. The impulse-like excitation is reflected across all frequencies including the zero frequency (0 Hz). Using the slope around zero crossings of the zero frequency filtered signal, an automatic speech polarity detection method is proposed. Performance of the proposed method is demonstrated on 8 different speech corpora. The proposed method is compared with the three existing techniques such as gradient of the spurious glottal waveforms (GSGW), oscillating moments-based polarity detection (OMPD) and residual excitation skewness (RESKEW). From the experimental results, it is observed that the performance of the proposed method is comparable or better than the existing methods for the experiments considered. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
ICASSP | 2 |
| 2017 | A Signal Processing Approach for Speaker Separation Using SFF Analysis
Nivedita Chennupati, B. H. V. S. Narayanamurthy, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2017 | A Robust and Alternative Approach to Zero Frequency Filtering Method for Epoch Extraction
P. Gangamohan, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2017 | Locating Burst Onsets Using SFF Envelope and Phase Information
Bhanu Teja Nellore, RaviShankar Prasad, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 5 |
| 2017 | Epoch extraction from emotional speech using single frequency filtering approachabstractEpochs are instants of significant excitation of the vocal tract system during production of voiced speech. Existing methods for epoch extraction provide good results on neutral speech. But effectiveness of these methods has not been examined carefully for analysis of emotional speech, where the emotion characteristics are embedded mainly in the source component of the signal. Performance of the state-of-art epoch extraction methods on emotional speech data may be affected due to large variations in the pitch period. An approach, which exploits the nature of impulse-like excitation in the speech signal, instead of the pitch period information, is explored in this paper. The approach uses single frequency filtering (SFF) analysis of speech signals, which provides the high temporal resolution of some features of excitation source (such as impulse-like events) and high spectral resolution for some features of spectrum (such as harmonics and resonances). The Berlin emotional speech database (EMO-DB), which contains the simultaneous electroglottograph (EGG) recordings is used as the ground truth. For comparison, several epoch extraction methods are evaluated in terms of both reliability and accuracy measures for six different emotion categories and neutral speech. The results indicate that the performance of the proposed SFF-based methods for emotional speech is comparable to the results for neutral speech, and is better than the results from many of the standard methods. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2017 | Extraction of Fundamental Frequency From Degraded Speech Using Temporal Envelopes at High SNR FrequenciesabstractIn this paper we propose a method for extracting the fundamental frequency (fo) from degraded speech signals using single frequency filtering (SFF) approach. The SFF of frequency-shifted speech signal gives high signal-to-noise ratio (SNR) segments at some frequencies and hence the SFF approach can be exploited for foextraction using autocorrelation function of those segments. Since the fois computed from the envelope of a single frequency component of the signal, the vocal tract resonances do not affect the foextraction. The use of the high SNR frequency component in a given segment helps in overcoming the effects of degradations in the speech signal, without explicitly estimating the characteristics of noise. The proposed method of foextraction is shown to give better performance for several types of real and simulated degradations, in comparison with some of the methods reported recently in the literature. Gunnam Aneeja, Bayya Yegnanarayana |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Use of Vowels in Discriminating Speech-Laugh from Laughter and Neutral Speech
Sri Harsha Dumpala, P. Gangamohan, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2016 | Robust Vowel Landmark Detection Using Epoch-Based Features
Sri Harsha Dumpala, Bhanu Teja Nellore, Raghu Ram Nevali, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 5 |
| 2016 | Robust Estimation of Fundamental Frequency Using Single Frequency Filtering Approach
Vishala Pannala, Gunnam Aneeja, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2015 | Analysis of singing voice for epoch extraction using Zero Frequency Filtering methodabstractEpoch is the instant of significant excitation of the vocal tract system during the production of voiced speech. Estimation of epochs or Glottal closure instants (GCIs) is a well studied topic in the speech analysis. From the recent studies on GCI detection from singing voice with state-of-art methods proposed for speech, there exist a clear gap in accuracy between speech and singing voice. This is because of source-filter interaction in singing voice compared to speech. Performance of existing algorithms deteriorates as most of the techniques depends on the ability to model the vocal tract system in order to emphasize the excitation characteristics in the residual. The objective of this paper is to analyze the singing voice for the estimation of epochs by studying the characteristics of the source-filter interaction and the effect of wider range of pitch using the Zero Frequency Filtering (ZFF) method. It is observed that high source-filter interaction can be captured in the form of the impulse-like excitation by passing the signal through three ideal digital resonators having poles at zero frequency, and the effect of wider range of pitch can be controlled by processing short segment (0.4-0.5 sec) signal. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
ICASSP | 2 |
| 2015 | Robust features for sonorant segmentation in continuous speech
Sri Harsha Dumpala, Bhanu Teja Nellore, Raghu Ram Nevali, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 5 |
| 2015 | Analysis of excitation source features of speech for emotion recognition
Sudarsana Reddy Kadiri, P. Gangamohan, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2015 | Robust pitch estimation in noisy speech using ZTW and group delay function
RaviShankar Prasad, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2015 | Distinctive feature based representation of speech for query-by-example spoken term detection
Abhijeet Saxena, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2015 | Analysis of production characteristics of laughter
Vinay Kumar Mittal, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2015 | Single Frequency Filtering Approach for Discriminating Speech and NonspeechabstractIn this paper, a signal processing approach is proposed for speech/nonspeech discrimination. The approach is based on single frequency filtering (SFF), where the amplitude envelope of the signal is obtained at each frequency with high temporal and spectral resolution. This high resolution property helps to exploit the resulting high signal-to-noise ratio (SNR) regions in time and frequency. The variance of the spectral information across frequency is higher for speech and lower for many types of noises. The mean and variance of the noise-compensated weighted envelopes are computed across frequency at each time instant. Decision logic is applied to the feature derived from the mean and variance values on varieties of degradations, including NTIMIT, CTIMIT and distance speech, besides degradation due to standard noise types. In all cases, the proposed method gives significantly better performance than the standard Adaptive Multi-rate VAD2 (AMR2) method. AMR2 method is chosen for comparison, as the method adapts itself for different degradations, and is seen to give good performance over different SNR situations. The proposed method does not use training data to derive the characteristics of speech or noise, nor makes any assumption on the nonspeech beginning. The SFF method appears promising in other applications of speech processing, such as pitch extraction and speech enhancement. Gunnam Aneeja, Bayya Yegnanarayana |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Analysis of laughter and speech-laugh signals using excitation source informationabstractSpeech-laugh is a speech-synchronous form of laughter that often occurs in natural conversation. However, there are deviations in features of speech-laugh when compared with laughter and neutral speech individually. The objective of this study is to analyse the excitation source features to capture the deviations between laughter and speech-laughs in voiced regions. The features used in this analysis are based on instantaneous fundamental frequency and strength of excitation (β) at epochs. Modified zero frequency filtering (ZFF) method is used to extract the features. Kullback-Leibler (KL) distances obtained show that there are deviations in excitation source features which can be exploited to develop a method to discriminate speech-laughs from laughter. Experimental results show that features used are robust and speaker independent in discriminating speech-laughs from laughter. Results showing deviations of laughter and speech-laughs from neutral speech were also presented. Sri Harsha Dumpala, Karthik Venkat Sridaran, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
ICASSP | 4 |
| 2014 | Unsupervised query-by-example spoken term detection using segment-based Bag of Acoustic WordsabstractIn this work, we present an unsupervised framework to address the problem of spotting spoken terms in large speech databases. The segment-based Bag of Acoustic Words (BoAW) framework proposed is inspired from the Bag of Words (BoW) approach widely used in text retrieval systems. Since this model ignores the sequence information in speech samples for efficient indexing of the database, a Dynamic Time Warping (DTW) based temporal matching technique is used to re-rank the results and restore the time sequence information. The speech data is stored efficiently in an inverted index which makes the retrieval very fast, thus making this framework particularly useful for searching large databases. We address the issue of choosing the appropriate size of the segment of speech for reliable indexing. Comparison with other query-by-example spoken term detection systems shows that the proposed system outperforms the rest. Basil George, Bayya Yegnanarayana |
ICASSP | 2 |
| 2014 | Speech detection in transient noises
Gunnam Aneeja, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2014 | Excitation source features for discrimination of anger and happy emotionsabstractStudies on the emotion recognition task indicate that there is confusion in discrimination among higher activation states like ‘anger’ and ‘happy’. In this study, features related to excitation source of speech are examined for discriminating ‘anger’ and ‘happy’ emotions. The objective is to explore the features which are independent of lexical content, language, channel and speaker. The features like strength of excitation from zero frequency filtering method and spectral band magnitude energies from short-time spectral analysis are used. Experimental results show that these features can discriminate ‘anger’ and ‘happy’ emotion states to a good extent. Index Terms: Emotion recognition, zero frequency filtering method, KL distance measure. P. Gangamohan, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2014 | Unsupervised query-by-example spoken term detection using bag of acoustic words and non-segmental dynamic time warping
Basil George, Abhijeet Saxena, Gautam Varma Mantena, Kishore Prahallad, Bayya Yegnanarayana |
INTERSPEECH | 5 |
| 2014 | Significance of aperiodicity in the pitch perception of expressive voices
Vinay Kumar Mittal, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2014 | Study of changes in glottal vibration characteristics during laughter
Vinay Kumar Mittal, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2014 | Glottal source processing: From analysis to applications
Thomas Drugman, Paavo Alku, Abeer Alwan, Bayya Yegnanarayana |
Comput. Speech Lang. | 4 |
| 2014 | Extraction of formant bandwidths using properties of group delay functions
Anand Joseph Xavier Medabalimi, Guruprasad Seshadri, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2013 | Production features for detection of shouted speechabstractShouted speech or screaming signals have been studied mostly through spectral representation such as melcepstral coefficients. Intuitive evidence that the characteristics of the excitation source may vary in the case of shouted speech has drawn little attention yet. In this paper we examine how the characteristics of both components of speech production mechanism, especially the glottal excitation source, are modified during the production of shout signals. Shouted and normal speech signals are examined along with the corresponding Electro-glotto-graph (EGG) signals. Distinguishing features like the dominant frequency and the strength of excitation are explored, along with the instantaneous fundamental frequency. These features are computed using linear prediction analysis and zero frequency filtering of the speech signal. Efficacy of these features in discriminating between shouted and normal speech is tested in five different vowel contexts. Vinay Kumar Mittal, Bayya Yegnanarayana |
CCNC | 2 |
| 2013 | Syllable nuclei detection using perceptually significant features
Apoorv Reddy Arrabothu, Nivedita Chennupati, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2013 | Analysis of emotional speech at subsegmental level
P. Gangamohan, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2013 | Acoustic segmentation of speech using zero time liftering (ZTL)
RaviShankar Prasad, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2013 | Spectro-temporal analysis of speech signals using zero-time windowing and group delay function
Bayya Yegnanarayana, Dhananjaya Gowda |
Speech Commun. | 1 |
| 2012 | A Flexible Analysis Synthesis Tool (FAST) for studying the characteristic features of emotion in speechabstractThis paper aims to understand the components of speech that contribute to emotion characteristics in speech. Four components of speech (vocal tract, excitation, duration and intonation) are considered in this study. A Flexible Analysis Synthesis Tool (FAST) is developed to modify the features of an utterance from neutral to emotion or from emotion to neutral. The key ideas used in this work are the dynamic time warping algorithm for alignment of two utterances and a flexible prosody manipulation for incorporating the desired features. The tool is used for conversion of neutral to emotion speech. Subjective evaluation is performed based on listening tests. The tool has potential to convert neutral to emotion speech and vice-versa, which can lead to understanding the significance of various components contributing to emotional content in speech. P. Gangamohan, Vinay Kumar Mittal, Bayya Yegnanarayana |
CCNC | 3 |
| 2012 | Analysis of Mimicry SpeechabstractIn this paper, mimicry speech is analysed using features at suprasegmental, segmental and subsegmental levels. The possibility of the imitator getting close at each of these levels is examined here. The imitator cannot duplicate all features of the target, as imitation depends on the target speaker, utterance chosen, and his ability to imitate. To study the variation of features in the case of best and poor imitations, the source and system features are observed for different target speakers and for different utterances. Features such as pitch contour, duration, Itakura distance, strength of excitation and loudness measure are used for this analysis. Perceptual evaluation is performed to determine the closeness of imitation to the target. The closeness of features for best imitated and poorly imitated utterances is presented here. D. Gomathi, Sathya Adithya Thati, Karthik Venkat Sridaran, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2012 | Effect of Tongue Tip Trilling on the Glottal Excitation SourceabstractRecent studies have indicated changes in the glottal excitation source characteristics apart from vocal tract resonances due to tongue tip trilling. In this paper we study the significance of changing vocal tract system and the associated glottal excitation source characteristics due to trilling, from perception point of view. These studies are made by generating speech signal by either retaining the features of the vocal tract system or of the glottal excitation source of trill sounds. Experiments are conducted to understand the perceptual significance of the excitation source characteristics on production of different trill sounds. Speech sounds of sustained trill and approximant pair, and apical trills produced by four different places of articulation are considered. Features of the vocal tract system are extracted using linear prediction analysis, and those of the source by zero frequency filtering. Vinay Kumar Mittal, Dhananjaya Gowda, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2012 | Spotting glottal stop in Amharic in continuous speech
Hussien Seid Worku, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2011 | Acoustic-phonetic information from excitation source for refining manner hypotheses of a phone recognizerabstractReliable acoustic-phonetic (AP) information derived from the speech signal can be used to detect and correct errors in the output of a phone recognizer. In this paper, limited acoustic-phonetic information derived primarily by processing the excitation source information in the speech signal is used to improve the performance of detection of manner of articulation from a baseline phone recognition system. A context-independent HMM-based monophone system without any language information is used as the baseline system for this purpose. The performance of the phone recognizer in terms of its ability to detect the manners of articulation is studied. The errors in the hypothesis of the manner of articulation of phones are corrected using AP information such as voicing, voice bar and frication. It is shown that significant improvement can be achieved by using simple or limited AP information. Dhananjaya Gowda, Bayya Yegnanarayana, Suryakanth V. Gangashetty |
ICASSP | 2 |
| 2011 | Decomposition of speech signals for analysis of aperiodic components of excitationabstractThe motivation for this study is the need for careful analysis of aperiodicity of the excitation component in expressive voices. The paper proposes analysis methods which can preserve the excitation information corresponding to sequence of impulse-like excitation with variable strengths. To analyze the details of the excitation source characteristics, the epochs and the strength of the excitation at the epochs are obtained using the output of an ideal zero-frequency digital resonator. The vocal tract system characteristics are derived from the signal between two successive epochs using the numerator of the group delay function. The spectrogram of the zero-frequency filtered signal and the group delay spectrum correspond to characteristics of the excitation and the vocal tract system, respectively. Decomposition of the speech signal into these two components bring out the features of excitation and vocal tract system, which can be used to explain the perception of expressive voices in terms of features of aperiodicity, pitch, harmonics and sub-harmonics. The decomposition method is illustrated using examples from linguistically significant glottalized sounds (glottal stops and ejectives), singing voices and Noh voice. Bayya Yegnanarayana, Anand Joseph Xavier Medabalimi, Suryakanth V. Gangashetty, Dhananjaya Gowda |
ICASSP | 1 |
| 2011 | Study of robustness of zero frequency resonator method for extraction of fundamental frequencyabstractThe objective of this work is to develop and study the robustness of the zero frequency resonator (ZFR) based method for extraction of the fundamental frequency (F0) of speech signals. The proposed ZFR method for estimating F0consists of zero frequency filtering of the Hilbert envelope (HE) of the linear prediction (LP) residual of speech signal, followed by short-term spectrum analysis of the filtered output. The robustness of the proposed method is tested using speech signals collected in practical environments like distant, reverberant, telephone, mobile and multispeaker. Experimental results show that the proposed ZFR method estimates F0in majority of the cases. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Sunitha Guruprasad |
ICASSP | 1 |
| 2011 | Neutral to Target Emotion Conversion Using Source and Suprasegmental InformationabstractThis work uses instantaneous pitch and strength of excitation along with duration of syllable-like units as the parameters for emotion conversion. Instantaneous pitch and duration of the syllable-like units of the neutral speech are modified by the prosody modification of its linear prediction (LP) residual using the instants of significant excitation. The strength of excitation is modified by scaling the Hilbert envelope (HE) of the LP residual. The target emotion speech is then synthesized using the prosody and strength modified LP residual. The pitch, duration and strength modification factors for emotion conversion are derived using the syllable-like units of initial, middle and final regions from an emotion speech database having different speakers, texts and emotions. The effectiveness of the region wise modification of source and supra segmental features over the gross level modification is confirmed by the waveforms, spectrograms and subjective evaluations. Index Terms: Emotions, ZFF, strength of excitation, instantaneous pitch, duration D. Govind 0001, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2011 | Performance of an Event-Based Instantaneous Fundamental Frequency Estimator for Distant Speech SignalsabstractThis paper proposes a method for extracting the fundamental frequency of voiced speech from distant speech signals. The method is based on the impulse-like nature of excitation in voiced speech. The characteristics of impulse-like excitation are extracted by filtering the speech signal through a cascade of resonators located at zero frequency. The resulting filtered signal preserves information specific to the fundamental frequency, in the sequence of positive-to-negative zero crossings. Also, the filtered signal is free from the effects of resonances of the vocal tract. An estimate of the fundamental frequency is derived from the short-time spectrum of the filtered signal. This estimate is used to remove spurious zero crossings in the filtered signal. The proposed method depends only on the strengths of impulse-like excitations in the direct component of distant speech signals, and not on the similarity of speech signal in successive glottal cycles. Hence, the method is robust to the effects of reverberation and noise. Performance of the method is evaluated using a database of close-speaking and distant speech signals. Experiments show that the accuracy of the proposed method is significantly higher than that of existing methods based on time-domain and frequency-domain processing. Guruprasad Seshadri, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Analysis of instantaneous F0 contours from two speakers mixed signal using zero frequency filteringabstractInstantaneous fundamental frequency (F0) in voiced speech can be obtained from the sequence of epochs corresponding to the instants of significant excitation. The epoch sequence can be derived using the recently proposed epoch extraction method based on zero frequency filtering. The epoch extraction method is robust against additive noise degradation. But in a multispeaker mixed signal, the degradation is caused due to overlapping impulse-like excitations of two or more speakers. The feasibility of extracting the instantaneous F0contours from the two speaker mixed signal using zero frequency filtering is studied in this paper. The present study is based on deriving speaker-specific Hilbert Envelope (HE) signal which emphasizes peaks due to impulse-like excitation of one speaker and suppresses peaks due to other speaker. The epochs from this speaker-specific signal are obtained using the approach based on zero frequency filtering. The results of the proposed method is demonstrated for three different cases of mixed signals of two speakers data. Bayya Yegnanarayana, S. R. Mahadeva Prasanna |
ICASSP | 1 |
| 2010 | Exploring subsegmental and suprasegmental features for a text-dependent speaker verification in distant speech signals
B. Avinash, Sunitha Guruprasad, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2010 | Significance of pitch synchronous analysis for speaker recognition using AANN models
Sri Harish Reddy Mallidi, Kishore Prahallad, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2010 | Speaker-dependent mapping of source and system features for enhancement of throat microphone speechabstractA throat microphone (TM) produces speech which is perceptually poorer than that produced by a close speaking microphone (CSM) speech. Many attempts at improving the quality of TM speech have been made by mapping the features corresponding to the vocal tract system. These techniques are limited by the methods used to generate the excitation signal. In this paper a method to map the source (excitation) using multilayer feedforward neural networks is proposed for voiced segments. This method anchors the analysis windows at the regions around the instants of glottal closure, so that the non-linear characteristics in these region of TM and CSM microphone is emphasized in the mapping process. The features obtained from these regions for both TM and CSM speech are used to train a MLFFNN to capture the non-linear relation between them. An improved technique for mapping the system features is also proposed. Speech synthesized using the proposed techniques was evaluated through subjective tests and was found to be significantly better than TM speech. Anand Joseph Xavier Medabalimi, Sri Harish Reddy Mallidi, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2010 | Voiced/Nonvoiced Detection Based on Robustness of Voiced EpochsabstractIn this paper, a new method for voiced/nonvoiced detection based on epoch extraction is proposed. Zero-frequency filtered speech signal is used to extract the instants of significant excitation (or epochs). The robustness of the method to extract epochs in the voiced regions, even with small amount of additive white noise, is used to distinguish voiced epochs from random instants detected in nonvoiced regions. The main feature of the proposed method is that it uses the strength of glottal activity as against using the periodicity of the signal. Performance of the proposed algorithm is studied on TIMIT and CMU ARCTIC databases, for two different noise types, white and vehicle noise from the NOISEX database, at different signal-to-noise ratios (SNRs). The proposed method performs similar or better than the popular normalized crosscorrelation based voiced/nonvoiced detection used in the open source utilitywavesurfer, especially at lower SNRs. Dhananjaya Gowda, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 2 |
| 2010 | Spectral Mapping Using Artificial Neural Networks for Voice ConversionabstractIn this paper, we use artificial neural networks (ANNs) for voice conversion and exploit the mapping abilities of an ANN model to perform mapping of spectral features of a source speaker to that of a target speaker. A comparative study of voice conversion using an ANN model and the state-of-the-art Gaussian mixture model (GMM) is conducted. The results of voice conversion, evaluated using subjective and objective measures, confirm that an ANN-based VC system performs as good as that of a GMM-based VC system, and the quality of the transformed speech is intelligible and possesses the characteristics of a target speaker. In this paper, we also address the issue of dependency of voice conversion techniques on parallel data between the source and the target speakers. While there have been efforts to use nonparallel data and speaker adaptation techniques, it is important to investigate techniques which capture speaker-specific characteristics of a target speaker, and avoid any need for source speaker's data either for training or for adaptation. In this paper, we propose a voice conversion approach using an ANN model to capture speaker-specific characteristics of a target speaker and demonstrate that such a voice conversion approach can perform monolingual as well as cross-lingual voice conversion of an arbitrary source speaker. Srinivas Desai, Alan W. Black, Bayya Yegnanarayana, Kishore Prahallad |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Voice conversion using Artificial Neural NetworksabstractIn this paper, we propose to use artificial neural networks (ANN) for voice conversion. We have exploited the mapping abilities of ANN to perform mapping of spectral features of a source speaker to that of a target speaker. A comparative study of voice conversion using ANN and the state-of-the-art Gaussian mixture model (GMM) is conducted. The results of voice conversion evaluated using subjective and objective measures confirm that ANNs perform better transformation than GMMs and the quality of the transformed speech is intelligible and has the characteristics of the target speaker. Srinivas Desai, E. Veera Raghavendra, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
ICASSP | 3 |
| 2009 | Speaker dependent mapping for low bit rate coding of throat microphone speechabstractThroat microphones (TM) which are robust to background noise can be used in environments with high levels of background noise. Speech collected using TM is perceptually less natural. The objective of this paper is to map the spectral features (represented in the form of cepstral features) of TM and close speaking microphone (CSM) speech to improve the former’s perceptual quality, and to represent it in an efficient manner for coding. The spectral mapping of TM and CSM speech is done using a multilayer feed-forward neural network, which is trained from features derived from TM and CSM speech. The sequence of estimated CSM spectral features is quantized and coded as a sequence of codebook indices using vector quantization. The sequence of codebook indices, the pitch contour and the energy contour derived from the TM signal are used to store/transmit the TM speech information efficiently. At the receiver, the allpole system corresponding to the estimated CSM spectral vectors is excited by a synthetic residual to generate the speech signal. Joseph M. Anand, Bayya Yegnanarayana, M. R. Kesheorey |
INTERSPEECH | 2 |
| 2009 | Analysis of Lombard speech using excitation source informationabstractThis paper examines the Lombard effect on the excitation fea-tures in speech production. These features correspond mostly to the acoustic features at subsegmental (< pitch period) level. The instantaneous fundamental frequency F0 (i.e., pitch), the strength of excitation at the instants of significant excitation and a loudness measure reflecting the sharpness of the impulse-like excitation around epochs are used to represent the excitation features at the subsegmental level. The Lombard effect influ-ences the pitch and the loudness. The extent of Lombard effect on speech depends on the nature and level (or intensity) of the external feedback that causes the Lombard effect. Index Terms: Lombard effect, excitation source, loudness 1. G. Bapineedu, B. Avinash, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2009 | Analysis of laugh signals for detecting in continuous speechabstractLaughter is a nonverbal vocalization that occurs often in speech communication. Since laughter is produced by the speech production mechanism, spectral analysis methods are used mostly for the study of laughter acoustics. In this paper the significance of excitation features for discriminating laughter and speech is discussed. New features describing the excitation characteristics are used to analyze the laugh signals. The features are based on instantaneous pitch and strength of excitation at epochs. An algorithm is developed based on these features to detect laughter regions in continuous speech. The results are illustrated by detecting laughter regions in a TV broadcast program. Index Terms: Laughter detection, epoch, strength of excitation K. Sudheer Kumar, Sri Harish Reddy Mallidi, K. Sri Rama Murty, Bayya Yegnanarayana |
INTERSPEECH | 4 |
| 2009 | Acoustic characteristics of ejectives in amharicabstractIn this paper, a preliminary investigation of the acoustic characteristics of Amharic ejectives in comparison with their unvoiced conjugates is presented. The normalized error from linear prediction residual and a zero frequency resonator output are used to locate the instant of release of the oral closure and the instant of the start of voicing, respectively. Amharic ejectives are found to have longer closure duration and smaller VOT than their unvoiced conjugates. Cross-linguistic comparisons reveal that no ejectives of two languages behave acoustically in a similar manner despite similarity in their articulation. Index Terms: Amharic, ejectives, glottalized stops Hussien Seid Worku, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2009 | Intonation modeling for Indian languages
K. Sreenivasa Rao, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2009 | Duration modification using glottal closure instants and vowel onset points
K. Sreenivasa Rao, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2009 | Characterization of Glottal Activity From Speech SignalsabstractThe objective of this work is to characterize certain important features of excitation of speech, namely, detecting the regions of glottal activity and estimating the strength of excitation in each glottal cycle. The proposed method is based on the assumption that the excitation to the vocal-tract system can be approximated by a sequence of impulses of varying strengths. The effect due to an impulse in the time-domain is spread uniformly across the frequency-domain including at zero-frequency. We propose the use of a zero-frequency resonator to extract the characteristics of excitation source from speech signals by filtering out most of the time-varying vocal-tract information. The regions of glottal activity and the strengths of excitation estimated from the speech signal are in close agreement with those observed from the simultaneously recorded electro-glotto-graph signals. The performance of the proposed glottal activity detection is evaluated under different noisy environments at varying levels of degradation. K. Sri Rama Murty, Bayya Yegnanarayana, Joseph M. Anand |
IEEE Signal Process. Lett. | 2 |
| 2009 | Event-Based Instantaneous Fundamental Frequency Estimation From Speech SignalsabstractExploiting the impulse-like nature of excitation in the sequence of glottal cycles, a method is proposed to derive the instantaneous fundamental frequency from speech signals. The method involves passing the speech signal through two ideal resonators located at zero frequency. A filtered signal is derived from the output of the resonators by subtracting the local mean computed over an interval corresponding to the average pitch period. The positive zero crossings in the filtered signal correspond to the locations of the strong impulses in each glottal cycle. Then the instantaneous fundamental frequency is obtained by taking the reciprocal of the interval between successive positive zero crossings. Due to filtering by zero-frequency resonator, the effects of noise and vocal-tract variations are practically eliminated. For the same reason, the method is also robust to degradation in speech due to additive noise. The accuracy of the fundamental frequency estimation by the proposed method is comparable or even better than many existing methods. Moreover, the proposed method is also robust against rapid variation of the pitch period or vocal-tract changes. The method works well even when the glottal cycles are not periodic or when the speech signals are not correlated in successive glottal cycles. Bayya Yegnanarayana, K. Sri Rama Murty |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Determining Mixing Parameters From Multispeaker Data Using Speech-Specific InformationabstractIn this paper, we propose an approach for processing multispeaker speech signals collected simultaneously using a pair of spatially separated microphones in a real room environment. Spatial separation of microphones results in a fixed time-delay of arrival of speech signals from a given speaker at the pair of microphones. These time-delays are estimated by exploiting the impulse-like characteristic of excitation during speech production. The differences in the time-delays for different speakers are used to determine the number of speakers from the mixed multispeaker speech signals. There is difference in the signal levels due to differences in the distances between the speaker and each of the microphones. The differences in the signal levels dictate the values of the mixing parameters. Knowledge of speech production, especially the excitation source characteristics, is used to derive an approximate weight function for locating the regions specific to a given speaker. The scatter plots of the weighted and delay-compensated mixed speech signals are used to estimate the mixing parameters. The proposed method is applied on the data collected in actual laboratory environment for an underdetermined case, where the number of speakers is more than the number of microphones. Enhancement of speech due to a speaker is also examined using the information of the time-delays and the mixing parameters, and is evaluated using objective measures proposed in the literature. Bayya Yegnanarayana, R. Kumaraswamy 0001, K. Sri Rama Murty |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Exploiting contextual information for improved phoneme recognitionabstractIn this paper, we investigate the significance of contextual information in a phoneme recognition system using the hidden Markov model - artificial neural network paradigm. Contextual information is probed at the feature level as well as at the output of the multilayered perceptron. At the feature level, we analyze and compare different methods to model sub-phonemic classes. To exploit the contextual information at the output of the multilayered perceptron, we propose the hierarchical estimation of phoneme posterior probabilities. The best phoneme (excluding silence) recognition accuracy of 73.4% on the TIMIT database is comparable to that of the state-of- the-art systems, but more emphasis is on analysis of the contextual information. Joel Pinto, Bayya Yegnanarayana, Hynek Hermansky, Mathew Magimai-Doss |
ICASSP | 2 |
| 2008 | Video Shot Segmentation Using Late Fusion TechniqueabstractIn this paper, a new method for detecting shot boundaries in video sequences using a late fusion technique is proposed. The method uses color histogram as the feature, and processes each bin separately for detecting shot boundaries. The decisions from individual bins are combined later for hypothesizing the presence of shot boundaries. The method provides a certain degree of robustness against illumination and camera/object motion, as it ignores small changes in the bins. While the early fusion techniques rely on the extent of change in color information, the proposed technique relies on the number of significant changes. Experimental results successfully validate the new method and show that it can effectively detect both abrupt and gradual transitions. C. Krishna Mohan, Dhananjaya Gowda, Bayya Yegnanarayana |
ICMLA | 3 |
| 2008 | AANN-HMM models for speaker verification and speech recognitionabstractPattern classification is an important task in speech recognition and speaker verification. Given the feature vectors of an input the goal is to capture the characteristics of these features unique to each class. This paper deals with exploring Auto Associative Neural Network (AANN) models for the task of speaker verification and speech recognition. We show that AANN models produce comparable performance with that of GMM based speaker verification and speech recognition. Sachin Joshi, Kishore Prahallad, Bayya Yegnanarayana |
IJCNN | 3 |
| 2008 | Features for automatic detection of voice bars in continuous speechabstractIn this paper we propose features for automatic detection of voice bar, which is an essential component of voiced stop consonants, in continuous speech. The acoustic-phonetic and production based knowledge such as, the presence of voicing, low strength of excitation compared to other voiced phones and a predominant low-frequency spectral energy, are mapped onto a set of acoustic features that can be automatically extracted from the signal. The usefulness of the proposed features in the detection of voice bars is studied using a knowledge-based as well as a neural network based approach. The performance of the proposed features and approaches is studied on phones from databases of two languages, namely English and Hindi. Dhananjaya Gowda, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2008 | Efficient representation of throat microphone speechabstractThe objective of this work is to represent the information in the speech signal picked up by a throat microphone (TM) in an efficient manner in terms of number of bits required.Since the TM signal is unaffected by ambient noise, it is possible to extract the required information effectively under different environmental conditions.A spectral mapping technique is proposed from the TM speech to normal microphone (NM) speech to improve the perceptual quality.The mapping is done using vector quantization of pairwise spectral feature vectors derived from each frame of TM and the corresponding NM speech signals.Once the codebook is formed, the spectral features from a TM signal are represented as a sequence of codebook indices.The sequence of codebook indices, the pitch contour and the energy contour derived from the TM signal are used to store/transmit the TM speech information efficiently.From the received sequence of codebook indices, the NM spectral vectors are retrieved due to pairwise vector quantization of the feature vectors.A synthetic residual signal is generated at the receiver from prestored residual templates by incorporating the pitch and the energy.The synthetic residual signal is used to excite the system corresponding to the NM spectral vectors to generate the speech signal. K. Sri Rama Murty, Saurav Khurana, Yogendra Umesh Itankar, M. R. Kesheorey, Bayya Yegnanarayana |
INTERSPEECH | 5 |
| 2008 | Building sleek synthesizers for multi-lingual screen readerabstractIn this paper, we are investigating the unit size: syllable, half-phone and quarter-phone to be used for speech synthesis in multi-lingual screen reader in phonetic languages such as Telugu and non-phonetic language English. Perceptual studies show that syllable-level unit performs better for Telugu and half-phone units perform better for English. While syllable based synthesizers produce better sounding speech, the coverage of all syllables is a non-trivial issue. We address the issue of coverage of syllables through approximate matching of syllable and show that such approximation produces intelligible and better quality speech than diphone units. In this paper, we also propose a hybrid synthesizer within the framework of unit selection and also show that the hybrid synthesizer built from pruned database performs as well as hybrid synthesizer built from unpruned database. Index Terms: speech synthesis, unit selection, unit size, database pruning, and hybrid speech synthesis. E. Veera Raghavendra, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
INTERSPEECH | 2 |
| 2008 | Analysis of glottal stops in speech signalsabstractDuring production of glottal stops the glottal vibration has un-equal cycles and is caused by laryngealization. While one can perceive the features of laryngealization in the speech, it is dif-ficult to analyse the signal to detect these source features from the standard spectrum-based analysis methods. In this paper we propose methods to extract the voice source vibration charac-teristics, and show that in the region of glottal stop, the pitch periods will be irregular, and the crosscorrelation coefficient of the signal in successive pitch periods will be low. This analy-sis enable us to locate the regions of glottal stops in continuous speech, and also help us to study the characteristics of creaky voice. Index Terms: glottal stop, laryngealization, creaky voice, glot-tal vibration, voice source, pitch period. Bayya Yegnanarayana, Hussien Seid Worku, Dhananjaya Gowda |
INTERSPEECH | 1 |
| 2008 | Global syllable set for building speech synthesis in Indian languagesabstractIndian languages are syllabic in nature where many syllables are found common across its languages. This motivates us to build a global syllable set by combining multiple language syllables to build a synthesizer which can borrow units from a different language when the required syllable is not found. Such synthesizer make use of speech database in different languages spoken by different speakers, whose output is likely to pick units from multiple languages and hence the synthesized utterance contains units spoken by multiple speakers which would annoy the user. We intend to use a cross lingual Voice Conversion framework using Artificial Neural Networks (ANN) to transform such an utterance to a single target speaker. E. Veera Raghavendra, Srinivas Desai, Bayya Yegnanarayana, Alan W. Black, Kishore Prahallad |
SLT | 3 |
| 2008 | Speech synthesis using approximate matching of syllablesabstractIn this paper we propose a technique for a syllable based speech synthesis system. While syllable based synthesizers produce better sounding speech than diphone and phone, the coverage of all syllables is a non-trivial issue. We address the issue of coverage of syllables through approximating the syllable when the required syllable is not found. To verify our hypothesis, we conducted perceptual studies on manually modified sentences and found that our assumption is valid. Similar approaches have been used in speech synthesis and it shows that such approximation produces intelligible and better quality speech than diphone units. E. Veera Raghavendra, Bayya Yegnanarayana, Kishore Prahallad |
SLT | 2 |
| 2008 | Multimodal person authentication using speech, face and visual speech
S. Palanivel, Bayya Yegnanarayana |
Comput. Vis. Image Underst. | 2 |
| 2008 | Speaker change detection in casual conversations using excitation source features
Dhananjaya Gowda, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2008 | Extraction and representation of prosodic features for language and speaker recognition
Leena Mary, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2008 | Epoch Extraction From Speech SignalsabstractEpoch is the instant of significant excitation of the vocal-tract system during production of speech. For most voiced speech, the most significant excitation takes place around the instant of glottal closure. Extraction of epochs from speech is a challenging task due to time-varying characteristics of the source and the system. Most epoch extraction methods attempt to remove the characteristics of the vocal-tract system, in order to emphasize the excitation characteristics in the residual. The performance of such methods depends critically on our ability to model the system. In this paper, we propose a method for epoch extraction which does not depend critically on characteristics of the time-varying vocal-tract system. The method exploits the nature of impulse-like excitation. The proposed zero resonance frequency filter output brings out the epoch locations with high accuracy and reliability. The performance of the method is demonstrated using CMU-Arctic database using the epoch information from the electroglottograph as reference. The proposed method performs significantly better than the other methods currently available for epoch extraction. The interesting part of the results is that the epoch extraction by the proposed method seems to be robust against degradations like white noise, babble, high-frequency channel, and vehicle noise. K. Sri Rama Murty, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Activity Modeling Using Event Probability SequencesabstractChanges in motion properties of trajectories provide useful cues for modeling and recognizing human activities. We associate an event with significant changes that are localized in time and space, and represent activities as a sequence of such events. The localized nature of events allows for detection of subtle changes or anomalies in activities. In this paper, we present a probabilistic approach for representing events using the hidden Markov model (HMM) framework. Using trained HMMs for activities, an event probability sequence is computed for every motion trajectory in the training set. It reflects the probability of an event occurring at every time instant. Though the parameters of the trained HMMs depend on viewing direction, the event probability sequences are robust to changes in viewing direction. We describe sufficient conditions for the existence of view invariance. The usefulness of the proposed event representation is illustrated using activity recognition and anomaly detection. Experiments using the indoor University of Central Florida human action dataset, the Carnegie Mellon University Credo Intelligence, Inc., Motion Capture dataset, and the outdoor Transportation Security Administration airport tarmac surveillance dataset show encouraging results. Naresh P. Cuntoor, Bayya Yegnanarayana, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2007 | Detection of instants of glottal closure using characteristics of excitation sourceabstractIn this paper, we propose a method for detection of glottal clo-sure instants (GCI) in the voiced regions of speech signals. The method is based on periodicity of significant excitations of the vocal tract system. The key idea is the computation of coherent covariance sequence, which overcomes the effect of dynamic range of the excitation source signal, while preserv-ing the locations of significant excitations. The Hilbert enve-lope of linear prediction residual is used as an estimate of the source of excitation of the vocal tract system. Performance of the proposed method is evaluated in terms of the deviation be-tween true GCIs and hypothesized GCIs, using clean speech and degraded speech signals. The signal-to-noise ratio (SNR) of speech signals in the vicinity of GCIs has significant bear-ing on the performance of the proposed method. The proposed method is accurate and robust for detection of GCIs, even in the presence of degradations. Index Terms: glottal closure instants, excitation source, peri-odicity, coherent covariance sequence Sunitha Guruprasad, Bayya Yegnanarayana, K. Sri Rama Murty |
INTERSPEECH | 2 |
| 2007 | Voice activity detection in degraded speech using excitation source informationabstractThis paper proposes a method for detection of voiced regions from speech signals collected in noisy environment.The proposed method is based on the characteristics of excitation source of speech production.The degraded speech signal is processed by linear prediction analysis for deriving the linear prediction residual.Hilbert envelope of the linear prediction residual is processed using covariance analysis to obtain coherentlyadded covariance signal.The periodicity property of the coherently added covariance signal is exploited to detect the voiced regions using autocorrelation analysis.The performance of the proposed voice activity detection algorithm is evaluated under different noise environments and at different levels of degradation. K. Sri Rama Murty, Bayya Yegnanarayana, Sunitha Guruprasad |
INTERSPEECH | 2 |
| 2007 | Modeling durations of syllables using neural networks
K. Sreenivasa Rao, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2007 | Determination of Instants of Significant Excitation in Speech Using Hilbert Envelope and Group Delay FunctionabstractThis letter proposes a time-effective method for determining the instants of significant excitation in speech signals. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. The proposed method consists of two phases: the first phase determines the approximate epoch locations using the Hilbert envelope of the linear prediction residual of the speech signal. The second phase determines the accurate locations of the instants of significant excitation by computing the group delay around the approximate epoch locations derived from the first phase. The accuracy in determining the instants of significant excitation and the time complexity of the proposed method is compared with the group delay based approach. K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 3 |
| 2007 | Determining Number of Speakers From Multispeaker Speech Signals Using Excitation Source InformationabstractIn this letter, we address the issue of determining the number of speakers from multispeaker speech signals collected simultaneously using a pair of spatially separated microphones. The spatial separation of the microphones results in time delay of arrival of speech signals from a given speaker. The differences in the time delays for different speakers are exploited to determine the number of speakers from the multispeaker signals. The key idea is that for a given speaker, the relative spacings of the instants of significant excitation of the vocal tract system remain unchanged in the direct components of the speech signals at the two microphones. The time delays can be estimated from the cross-correlation of the Hilbert envelopes of the linear prediction residuals of the multispeaker signals collected at the two microphones. R. Kumaraswamy 0001, K. Sri Rama Murty, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 3 |
| 2007 | Face Verification Using Template MatchingabstractHuman faces are similar in structure with minor differences from person to person. These minor differences may average out while trying to synthesize the face image of a given person, or while building a model of face image in automatic face recognition. In this paper, we propose a template-matching approach for face verification, which neither synthesizes the face image nor builds a model of the face image. Template matching is performed using an edginess-based representation of the face image. The edginess-based representation of face images is computed using 1-D processing of images. An approach is proposed based on autoassociative neural network models to verify the identity of a person. The issues of pose and illumination in face verification are addressed. Anil Kumar Sao, Bayya Yegnanarayana |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | Extracting formants from short segments of speech using group delay functionsabstractSpeech is a non-stationary signal, with the shape of the vocal tract changing over several pitch periods, and also within the open and closed glottis phases. The effect of these changes is reflected in the locations of the formants which correspond to the resonant frequencies of the vocal tract. To observe these changes, the analysis window should be small enough (relative to a pitch period), and appropriately anchored. A non-model based method is proposed in this paper to accurately determine formants from short segments (less than a pitch period) of speech signals. It makes use of high resolution properties of group delay function to estimate formants from segments of duration less than a pitch period. The main advantage of this method is its lack of dependence on the parameters of a model. Analysis segments are synchronised with instants of glottal closure, to increase the robustness of formant extraction. Since continuity or additional acoustic-phonetic knowledge are not used, this method is fairly reliable and robust. Joseph M. Anand, Sunitha Guruprasad, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2006 | Prosodic features for speaker verificationabstractIn this paper we study the effectiveness of prosodic features for speaker verification. We hypothesize that prosody is linked to linguistic units such as syllables and prosodic features can be better represented with reference to the syllabic sequence. For extracting prosodic features, speech is segmented into syllablelike regions using the knowledge of vowel onset points (VOP). We use a technique based on excitation source information to detect VOPs automatically. The location of VOPs serve as reference for extracting prosodic features directly from speech signal. Various parameters are used to represent the pitch and energy dynamics of the region between two consecutive VOPs. The effectiveness of the derived prosodic features for speaker verification is demonstrated on NIST SRE 2003 extended data. The complementary nature of prosodic features and spectral features help to improve the accuracy of the combined speaker verification system. Index Terms: prosody, speaker verification, syllable, vowel onset point, F0 contour. Leena Mary, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2006 | Mapping neural networks for bandwidth extension of narrowband speechabstractThis paper exploits the nonlinear mapping property of feedforward neural networks for estimation of high frequency components (4-8kHz) of the speech signals from the band-limited (04kHz) signals. Cepstral coefficients are used to represent the feature vectors of each frame of data. This paper also proposes an approach that uses the autocorrelation method to derive the Linear Prediction (LP) coefficients from the estimated cepstral coefficients that are obtained from the mapping network. This method guarantees the stability of the LP synthesis filter. Informal listenings indicate the effectiveness of the proposed method for estimation of wideband frequency components of speech. The enhanced speech sounds similar to the original wideband speech. Also, it does not contain any distortion that may arise due to spectral discontinuities between adjacent frames. Index Terms: speech enhancement, bandwidth extension, narrowband speech, mapping neural networks. A. Shahina, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2006 | Extraction of speaker-specific excitation information from linear prediction residual of speech
S. R. Mahadeva Prasanna, Cheedella S. Gupta, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2006 | Combining evidence from residual phase and MFCC features for speaker recognitionabstractThe objective of this letter is to demonstrate the complementary nature of speaker-specific information present in the residual phase in comparison with the information present in the conventional mel-frequency cepstral coefficients (MFCCs). The residual phase is derived from speech signal by linear prediction analysis. Speaker recognition studies are conducted on the NIST-2003 database using the proposed residual phase and the existing MFCC features. The speaker recognition system based on the residual phase gives an equal error rate (EER) of 22%, and the system using the MFCC features gives an EER of 14%. By combining the evidence from both the residual phase and the MFCC features, an EER of 10.5% is obtained, indicating that speaker-specific excitation information is present in the residual phase. This information is useful since it is complementary to that of MFCCs. K. Sri Rama Murty, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 2 |
| 2006 | Prosody modification using instants of significant excitationabstractProsody modification involves changing the pitch and duration of speech without affecting the message and naturalness. This paper proposes a method for prosody (pitch and duration) modification using the instants of significant excitation of the vocal tract system during the production of speech. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. Instants of significant excitation are computed from the linear prediction (LP) residual of speech signals by using the property of average group-delay of minimum phase signals. The modification of pitch and duration is achieved by manipulating the LP residual with the help of the knowledge of the instants of significant excitation. The modified residual is used to excite the time-varying filter, whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is good and is without any significant distortion. The proposed method is evaluated using waveforms, spectrograms, and listening tests. The performance of the method is compared with linear prediction pitch synchronous overlap and add (LP-PSOLA) method, which is another method for prosody manipulation based on the modification of the LP residual. The original and the synthesized speech signals obtained by the proposed method and by the LP-PSOLA method are available for listening at http://speech.cs.iitm.ernet.in/Main/result/prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Interpretation of State Sequences in HMM for Activity RepresentationabstractWe propose a method for activity representation based on semantic events, using the HMM framework. For every time instant, the probability of event occurrence is computed by exploring a subset of state sequences. The idea is that while activity trajectories may have large variations at the data or the state levels, they may exhibit similarities at the event level. Our experiments show the application of these events to activity recognition in an office environment and to anomalous trajectory detection using surveillance video data. Naresh P. Cuntoor, Bayya Yegnanarayana, Rama Chellappa |
ICASSP (2) | 2 |
| 2005 | Detection of vowel onset point events using excitation informationabstractThis paper proposes a method for the detection of Vowel Onset Point (VOP) events in speech using excitation information. VOP event is defined as the instant at which the onset of vowel takes place. For syllable-like units such as Consonant Vowel (CV) type, VOP event is the instant at which the consonant ends and the vowel begins. The speech signal is processed by the Linear Prediction (LP) analysis to extract the LP residual. The LP residual mostly contains the excitation information. The Hilbert envelope of the LP residual is derived using the analytic signal concept. A method is developed for detecting the VOP events using the Hilbert envelope of the LP residual and a modulated Gaussian window function. The performance of the proposed method is evaluated using reference VOP markings. The performance of the proposed method is also compared with the existing methods based on the vocal tract system features. The comparison shows that the excitation source also contains significant information about the VOP events. S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2005 | Speaker Localization Using Excitation Source Information in SpeechabstractThis paper presents the results of simulation and real room studies for localization of a moving speaker using information about the excitation source of speech production. The first step in localization is the estimation of time-delay from speech collected by a pair of microphones. Methods for time-delay estimation generally use spectral features that correspond mostly to the shape of vocal tract during speech production. Spectral features are affected by degradations due to noise and reverberation. This paper proposes a method for localizing a speaker using features that arise from the excitation source during speech production. Experiments were conducted by simulating different noise and reverberation conditions to compare the performance of the time-delay estimation and source localization using the proposed method with the results obtained using the spectrum-based generalized cross correlation (GCC) methods. The results show that the proposed method shows lower number of discrepancies in the estimated time-delays. The bias, variance and the root mean square error (RMSE) of the proposed method is consistently equal or less than the GCC methods. The location of a moving speaker estimated using the time-delays obtained by the proposed method are closer to the actual values, than those obtained by the GCC method. Vikas C. Raykar, Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Processing of reverberant speech for time-delay estimationabstractIn this paper, we present a method of extracting the time-delay between speech signals collected at two microphone locations. Time-delay estimation from microphone outputs is the first step for many sound localization algorithms, and also for enhancement of speech. For time-delay estimation, speech signals are normally processed using short-time spectral information (either magnitude or phase or both). The spectral features are affected by degradations in speech caused by noise and reverberation. Features corresponding to the excitation source of the speech production mechanism are robust to such degradations. We show that these source features can be extracted reliably from the speech signal. The time-delay estimate can be obtained using the features extracted even from short segments (50-100 ms) of speech from a pair of microphones. The proposed method for time-delay estimation is found to perform better than the generalized cross-correlation (GCC) approach. A method for enhancement of speech is also proposed using the knowledge of the time-delay and the information of the excitation source. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami, Dmitry N. Zotkin |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Combining evidence from source, suprasegmental and spectral features for a fixed-text speaker verification systemabstractThis paper proposes a text-dependent (fixed-text) speaker verification system which uses different types of information for making a decision regarding the identity claim of a speaker. The baseline system uses the dynamic time warping (DTW) technique for matching. Detection of the end-points of an utterance is crucial for the performance of the DTW-based template matching. A method based on the vowel onset point (VOP) is proposed for locating the end-points of an utterance. The proposed method for speaker verification uses the suprasegmental and source features, besides spectral features. The suprasegmental features such as pitch and duration are extracted using the warping path information in the DTW algorithm. Features of the excitation source, extracted using the neural network models, are also used in the text-dependent speaker verification system. Although the suprasegmental and source features individually may not yield good performance, combining the evidence from these features seem to improve the performance of the system significantly. Neural network models are used to combine the evidence from multiple sources of information. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Jinu Mariam Zachariah, Cheedella S. Gupta |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Extraction of pitch in adverse conditionsabstractThe paper proposes a method for the extraction of pitch in adverse conditions. The real environment, in which degradation is due to several unpredictable sources, like additive noise, reverberation and channel noise, is treated as an adverse condition. The proposed method is based on knowledge of glottal closure (GC) events. A GC event is the instant at which closure of vocal folds takes place within a pitch period. The Hilbert envelope of the linear prediction (LP) residual gives information about the location of GC events. Autocorrelation analysis is performed on the Hilbert envelope of the LP residual. The properties of the Hilbert envelope of the LP residual are exploited for the extraction of pitch from the autocorrelation sequence. The results of the proposed method are compared with the simple inverse filtering technique (SIFT) algorithm. The performance of the proposed algorithm is found to be superior, even in adverse conditions. S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICASSP (1) | 2 |
| 2004 | Modeling syllable duration in Indian languages using neural networksabstractWe propose a neural network model for predicting the syllable duration in Indian languages. A four layer feedforward neural network trained with a backpropagation algorithm is used for modeling the syllable duration. Analysis is performed on broadcast news data in Hindi, Telugu and Tamil in order to predict the duration of syllables in these languages using a neural network model. The input to the neural network consists of a set of phonological, positional and contextual features extracted from the text. About 88% of the syllable durations are predicted within 25% of the actual duration. The relative importance of the positional and contextual features are examined separately. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICASSP (5) | 2 |
| 2004 | Speaker Segmentation Based on Subsegmental Features and Neural Network Models
Dhananjaya Gowda, Sunitha Guruprasad, Bayya Yegnanarayana |
ICONIP | 3 |
| 2004 | Two-Stage Duration Model for Indian Languages Using Neural Networks
K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICONIP | 3 |
| 2004 | Content-Based Video Classification Using Support Vector Machines
Vakkalanka Suresh, C. Krishna Mohan, R. Kumaraswamy 0001, Bayya Yegnanarayana |
ICONIP | 4 |
| 2004 | Acoustic model combination for recognition of speech in multiple languages using support vector machinesabstractWe study the performance of support vector machine based classifiers in acoustic model combination for recognition of context dependent sub word units of speech in multiple languages. In acoustic model combination, the data for similar sub word units across languages are shared to train acoustic models for multilingual speech. Sharing of data across languages leads to an increase in the number of training examples for a subword unit common to the languages. It may also lead to increase in the variability of the data for a subword unit. In This work, we study the effect of data sharing on the classification accuracy and complexity of acoustic models built using support vector machines. We compare the performance of multilingual acoustic models with that of monolingual acoustic models in the recognition of a large number of consonant-vowel units in the broadcast news corpus of three Indian languages. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
IJCNN | 3 |
| 2004 | Enhancement of reverberant speech using excitation source informationabstractThis paper proposes a method for the enhancement of re-verberant speech using the knowledge of the excitation source of speech production. The degradation level in the reverberant speech is measured in terms of Speech-to-Reverberation component Ratio (SRR). From percep-tion and processing point of view high SRR regions are important. Hence the proposed method identifies and en-hances the speech in high SRR regions. The high SRR re-gions are identified using the Hilbert envelope of the Lin-ear Prediction (LP) residual, which contains information about the excitation source of speech production. The Hilbert envelope of the LP residual derived from the re-verberant speech is processed by the covariance analysis to derive the weight function. The LP residual of the re-verberant speech is multiplied with the weight function to enhance the excitations of speech in the high SRR re-gions. The speech signal synthesized from the modified LP residual is found to be less reverberant. 1. M. Chaitanya, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2004 | Detection of vowel on set points in continuous speech using autoassociative neural network modelsabstractDetection of vowel onset points (VOPs) is important for spotting subword units in continuous speech. For consonant-vowel (CV) utterances, VOP is the instant at which the consonant part ends and the vowel part begins. Accurate detection of VOPs is important for recognition of CV units in continuous speech. In this paper, we propose an approach for detection of VOPs using autoassociative neural network (AANN) models. A pair of AANN models are trained for each CV class to capture the characteristics of speech signal in the consonant and vowel regions of that class. The trained AANN models are then used to detect VOPs in continuous speech. The results of studies show that the proposed approach leads to significantly less number of spurious hypotheses. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2004 | Speech enhanced multi-Span language modelabstractTo capture local and global constraints in a language, sta-tistical -grams are used in combination with multi-span language models for improved language modelling. Use of latent semantic analysis (LSA) to capture the global semantic constraints and bigram models to capture lo-cal constraints, is shown to reduce the perplexity of the model. In this paper we propose a method in which the multi-span LSA language model can be developed based on the speech signal. Reference pattern vectors are de-rived from the speech signal for each word in the vocabu-lary. Based on the normalised distance between the refer-ence word pattern vector and the pattern vector of a word in the training data, the LSA model is developed. We show that this model in combination with a standard bi-gram model performs better than the conventional bigram + LSA model. The results are demonstrated for a limited vocabulary on a database for the Indian language, Tamil. 1. A. Nayeemulla Khan, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2004 | Latent semantic analysis for speaker recognitionabstractThere exists certain traits specifi ct oa speaker that help in easy identification of the speaker among a familiar set of speakers. These include certain dis-fluencies, and mannerisms like stress for certain words, frequent usage of certain phrases, manner of pronunciation and back channels. The focus of this paper is identification of a speaker using such idiolectic traits in conversational speech. Every normal conversation by a speaker contains his idiolectic signature. A model is developed in the latent semantic analysis framework to capture this signature. The similarity of the idiolectic signature in the test utterance to that captured by the model is used to hypothesise the target speaker. The technique is demonstrated for the NIST 2003 extended data task. A. Nayeemulla Khan, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2004 | Intonation modeling for indian languagesabstractIn this paper we propose models for predicting the intonation for the sequence of syllables present in the utterance.The term intonation refers to the temporal changes of the fundamental frequency ðF 0 Þ.Neural networks are used to capture the implicit intonation knowledge in the sequence of syllables of an utterance.We focus on the development of intonation models for predicting the sequence of fundamental frequency values for a given sequence of syllables.Labeled broadcast news data in the languages Hindi, Telugu and Tamil is used to develop neural network models in order to predict the F 0 of syllables in these languages.The input to the neural network consists of a feature vector representing the positional, contextual and phonological constraints.The interaction between duration and intonation constraints can be exploited for improving the accuracy further.From the studies we find that 88% of the F 0 values (pitch) of the syllables could be predicted from the models within 15% of the actual F 0 .The performance of the intonation models is evaluated using objective measures such as average prediction error ðlÞ, standard deviation ðrÞ and correlation coefficient ðcÞ.The prediction accuracy of the intonation models is further evaluated using listening tests.The prediction performance of the proposed intonation models using neural networks is compared with Classification and Regression Tree (CART) models. K. Sreenivasa Rao, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2004 | Throat microphone signal for speaker recognitionabstractSpeaker recognition systems perform better when clean speech signals are used for the task. In the presence of high levels of background noise, speech recorded from a close speaking microphone will be degraded and hence the performance of the speaker recognition system. Use of a transducer held at the throat results in a signal that is clean even in a noisy environment. This paper discusses the prospect of using such signals for speaker recognition. A study of a text-independent speaker recognition system based on features extracted from speech simultaneously recorded using a throat microphone and a closespeaking microphone in clean and simulated noisy conditions is conducted. Autoassociative neural networks are used to model the speaker characteristics based on the vocal tract system and excitation source features represented by weighted linear prediction cepstral coefficients and linear prediction residual, respectively. The results of experimental studies show that the speech collected from the throat microphone can be used for tasks like speaker recognition, especially in noisy conditions. Bayya Yegnanarayana, A. Shahina, M. R. Kesheorey |
INTERSPEECH | 1 |
| 2004 | Finding axes of symmetry from potential fieldsabstractThis paper addresses the problem of detecting axes of bilateral symmetry in images. In order to achieve robustness to variation in illumination, only edge-gradient information is used. To overcome the problem of edge breaks, a potential field is developed from the edge map which spreads the information in the image plane. Pairs of points in the image plane are made to vote for their axes of symmetry with some confidence values. To make the method robust to overlapping objects, only local features in the form of Taylor coefficients are used for quantifying symmetry. We define an axis of symmetry histogram, which is used to accumulate the weighted votes for all possible axes of symmetry. To reduce the computational complexity of voting, a hashing scheme is proposed, wherein pairs of points, whose potential fields are too asymmetric, are pruned by not being counted for the vote. Experimental results indicate that the proposed method is fairly robust to edge breaks and is able to detect symmetries even when only 0.05% of the possible pairs are used for voting. V. Shiv Naga Prasad, Bayya Yegnanarayana |
IEEE Trans. Image Process. | 2 |
| 2003 | Constraint satisfaction model for enhancement of evidence in recognition of consonant-vowel utterancesabstractWe address the issues in recognition of a large number of subword units of speech with high confusability among several units. Evidence available from the classification models trained with a limited number of training examples may not be strong to correctly recognize the subword units. We present a constraint satisfaction neural network model that can be used to enhance the evidence for a particular unit with the supporting evidence available for a subset of units confusable with that unit. We demonstrate the enhancement of evidence by the proposed model in recognition of utterances of 145 consonant-vowel units. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
ICASSP (2) | 3 |
| 2003 | Real time face recognition system using autoassociative neural network modelsabstractThis paper proposes a novel method for video-based real time face recognition. The proposed method uses motion information to detect the face region, and the region is processed in YC/sub r/C/sub b/ color space to determine the location of the eyes. The system extracts only the gray level features relative to the location of the eyes. autoassociative neural network (AANN) model is used to capture the distribution of the extracted gray level features. Experimental results show that the proposed system gives an average recognition rate of 99% in real time for 25 subjects. The performance of the proposed method is invariant to size, tilt of the face and is also not sensitive to natural lighting conditions. S. Palanivel, B. S. Venkatesh, Bayya Yegnanarayana |
ICASSP (2) | 3 |
| 2003 | Prosodic manipulation using instants of significant excitationabstractThe paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICASSP (1) | 2 |
| 2003 | Constraint satisfaction model for enhancement of evidence in recognition of consonant-vowel utterancesabstractIn this paper, we address the issues in recognition of a large number of subword units of speech with high confusability among several units. Evidence available from the classification models trained with a limited number of training examples may not be strong to correctly recognize the subword units. We present a constraint satisfaction neural network model that can be used to enhance the evidence for a particular unit with the supporting evidence available for a subset of units confusable with the unit. We demonstrate the enhancement of evidence by the proposed model in recognition of utterances of 145 consonant-vowel units. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
ICME | 3 |
| 2003 | Real time face authentication system using autoassociative neural network modelsabstractThis paper proposes a novel method for video-based real time face authentication. The proposed method uses motion information to detect the face region, and the face region is processed in YC/sub r/C/sub b/ color space to determine the location of the eyes. The system extracts only the gray level features relative to the location of the eyes. Autoassociative neural network (AANN) model is used to capture the distribution of the extracted gray level features. Experimental results show that the proposed system gives an equal error rate of less than 1% in real time for 25 subjects. The performance of the proposed method is invariant to size and tilt of the face, and is also insensitive to variations in natural lighting conditions. S. Palanivel, B. S. Venkatesh, Bayya Yegnanarayana |
ICME | 3 |
| 2003 | Prosodic manipulation using instants of significant excitationabstractThis paper proposes a technique for prosodic (pitch and duration) manipulation using instants of significant excitation. Instants of significant excitation correspond to the instants of glottal closure (epochs) in voiced speech and to some random excitations like burst onset in the case of nonvoiced speech. Instants of significant excitation are computed from the average group delay of minimum phase signals. The manipulation of pitch and duration is achieved by modifying the linear prediction (LP) residual with the help of instants of significant excitation as pitch markers. The modified residual is used to excite the time-varying filter whose parameters are derived from the original speech signal. Perceptual quality of the synthesized speech is found to be natural, and is without any distortion. The original and corresponding synthesized speech signals from the proposed approach are available for listening at http://speech.cs.iitm.ernet.in/Main/Results/Prosody.html. K. Sreenivasa Rao, Bayya Yegnanarayana |
ICME | 2 |
| 2003 | Combining evidence from multiple modular networks for recognition of consonant-vowel units of speechabstractIn this paper, we present a method to combine evidence from multiple classifiers to recognize a large number of subword units of speech using small size training data sets. Grouping criteria based on phonetic description are considered, to build multiple modular networks for recognition of the large number of units. Nonlinear compression of feature vectors is carried out to obtain reduced dimensional patterns, and multiple classifiers are trained separately using the uncompressed feature vectors and compressed feature vectors. Evidence from multiple classifiers at different stages in the recognition system is combined using the sum rule. Effectiveness of the proposed method is demonstrated for recognition of isolated utterances of 145 consonant-vowel units of speech. Suryakanth V. Gangashetty, K. Sreenivasa Rao, A. Nayeemulla Khan, Chellu Chandra Sekhar, Bayya Yegnanarayana |
IJCNN | 5 |
| 2003 | AANN models for speaker recognition based on difference cepstralsabstractThis paper presents a novel method for representing speaker characteristics present in the speech signal, by the way of deemphasizing the linguistic content of the signal. Cepstral coefficients that are widely employed as features for automatic speaker recognition task, contain considerable speech information in addition to the speaker information, and hence do not highlight the latter. The proposed method is based on using the difference between all-pole spectra due to higher order and lower order of linear prediction analysis. Distribution of the feature vectors in the multi-dimensional feature space is captured by employing autoassociative neural network models. A speaker recognition system is developed using the proposed method of feature extraction, whose performance is evaluated against that of the system based on cepstral coefficients. The complementary nature of evidence due to the proposed feature is also examined, so as to improve the overall system performance. Sunitha Guruprasad, Dhananjaya Gowda, Bayya Yegnanarayana |
IJCNN | 3 |
| 2003 | A novel vector quantizer for pattern classification tasksabstractWe present a novel vector quantization method for pattern classification tasks. The input space is quantized into volume regions by code-vectors formed by weights of neurons. During training, the volume regions are merged and split, depending upon the ambiguity in classification, measured using Kullback-Leibler divergence. The heuristic followed is to split ambiguous regions, and merge two volume regions if they contain predominant populations of the same class. The neural network forms a generalized Delaunay graph, whose topology changes dynamically with the merging and splitting. The simulation results indicate the utility of the proposed method. V. Shiv Naga Prasad, Bayya Yegnanarayana, Sunitha Guruprasad |
IJCNN | 2 |
| 2003 | Tracking a moving speaker using excitation source informationabstractMicrophone arrays are widely used to detect, locate, and track a stationary or moving speaker. The first step is to estimate the time delay, between the speech signals received by a pair of microphones. Conventional methods like generalized crosscorrelation are based on the spectral content of the vocal tract system in the speech signal. The spectral content of the speech signal is affected due to degradations in the speech signal caused by noise and reverberation. However, features corresponding to the excitation source of speech are less affected by such degradations. This paper proposes a novel method to estimate the time delays using the excitation source information in speech. The estimated delays are used to get the position of the moving speaker. The proposed method is compared with the spectrumbased approach using real data from a microphone array setup. 1. Vikas C. Raykar, Ramani Duraiswami, Bayya Yegnanarayana, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2003 | Enhancement of speech in multispeaker environmentabstractIn this paper a method based on the excitation source information is proposed for enhancement of speech, degraded by speech from other speakers. Speech from multiple speakers is simultaneously collected over two spatially distributed microphones. Time-delay of each speaker with respect to the two microphones is estimated using the excitation source information. A weight function is derived for each speaker using the knowledge of the timedelay and the excitation source information. Linear prediction (LP) residuals of the microphone signals are processed separately using the weight functions. Speech signals are synthesized from the modified residuals. One speech signal per speaker is derived from each microphone signal. The synthesized speech signals of each speaker are combined to produce enhanced speech. Significant enhancement of the speech of one speaker relative to other was observed from the combined signal. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Mathew Magimai-Doss |
INTERSPEECH | 1 |
| 2003 | Speaker-specific mapping for text-independent speaker recognition
Hemant Misra, Shajith Ikbal, Bayya Yegnanarayana |
Speech Commun. | 3 |
| 2002 | Linear and nonlinear compression of feature vectors for speech recognitionabstractIn this paper, we consider approaches for linear and nonlinear compression of feature vectors for recognition of utterances of syllable-like units in Indian languages. The distribution capturing ability of an autoassociative neural network model is exploited to derive the components for compressing the feature vectors. The nonlinear compression is accomplished by a five layer autoassociative neural network model. Linear compression is realized by principal component analysis. Both linear and nonlinear compressions are performed on each subgroup of the sound units separately. The results show that it is indeed possible to compress the feature vectors from 50 to 19 dimension without affecting the performance of the classifier. Suryakanth V. Gangashetty, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICASSP | 3 |
| 2002 | Speech enhancement using excitation source informationabstractThis paper proposes an approach for processing speech from multiple microphones to enhance speech degraded by noise and reverberation. The approach is based on exploiting the features of the excitation source in speech production. In particular, the characteristics of voiced speech can be used to derive a coherently added signal from the linear prediction (LP) residuals of the degraded speech data from different microphones. A weight function is derived from the coherently added signal. For coherent addition the time-delay between a pair of microphones is estimated using the knowledge of the source information present in the LP residual. The enhanced speech is generated by exciting the time varying all-pole filter with the weighted LP residual. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, K. Sreenivasa Rao |
ICASSP | 1 |
| 2002 | AANN: an alternative to GMM for pattern recognition
Bayya Yegnanarayana, Kishore Prahallad |
Neural Networks | 1 |
| 2002 | A constraint satisfaction model for recognition of stop consonant-vowel (SCV) utterancesabstractWe propose a model for recognition of utterances of consonant-vowel (CV) units. The acoustic-phonetic knowledge of the CV classes is incorporated in the form of constraints of a constraint satisfaction model. The model combines evidence from multiple classifiers. The significant feature of this model is that discrimination of the CV units could be enhanced by a combination of even weak evidence derived from the features. The evidence is obtained from multilayer feedforward neural networks trained for subgroups of CV classes. The evidence is enhanced using a set of feedback subnetworks in the constraint satisfaction model. The weights for the connections in the feedback subnetworks are derived using acoustic-phonetic knowledge and the performance statistics of the trained networks. The performance of the proposed model is demonstrated for recognition of utterances of a large number (80) of stop consonant-vowel units for the Indian language Hindi. Chellu Chandra Sekhar, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | A study of two dimensional linear discriminants for ASRabstractWe study the information in the joint time-frequency domain using 1515 dimensional-15 spectral energies and temporal span of 1s-block of spectrogram as features. In this feature space, we first derive 20 joint linear discriminants (JLDs) using linear discriminant analysis (LDA). Using principal component analysis (PCA), we conclude that information in this block of the spectrogram can be analyzed independently across the time and frequency domains. Under this assumption, we propose a sequential design of two dimensional discriminants (CLDs), i.e., spectral discriminants followed by temporal discriminants. We show that these CLDs are similar to first few JLDs and the discriminant features derived from the CLDs outperform those obtained from JLDs in the continuous-digit recognition task. Sachin S. Kajarekar, Bayya Yegnanarayana, Hynek Hermansky |
ICASSP | 2 |
| 2001 | Source and system features for speaker recognition using AANN modelsabstractWe study the effectiveness of the features extracted from the source and system components of the speech production process for the purpose of speaker recognition. The source and system components are derived using linear prediction (LP) analysis of short segments of speech. The source component is the LP residual derived from the signal, and the system component is a set of weighted linear prediction cepstral coefficients. The features are captured implicitly by a feedforward autoassociative neural network (AANN). Two separate speaker models are derived by training two AANN models using feature vectors corresponding to source and system components. A speaker recognition system for 20 speakers is built and tested using both the models to evaluate the performance of source and system features. The study demonstrates the complementary nature of the two components. Bayya Yegnanarayana, K. Sharat Reddy, Kishore Prahallad |
ICASSP | 1 |
| 2000 | Speaker verification: minimizing the channel effects using autoassociative neural network modelsabstractThe characteristics of the telephone channel and handset have a significant effect on the performance of speaker verification systems. The channel/handset mismatch between the training and testing data degrades the performance of speaker verification systems. In this paper, we show that the autoassociative neural network (AANN) models can be used to minimize the effects of channel characteristics on the performance of a text-independent speaker verification system. This paper also compares two approaches to represent the background model for an AANN based speaker verification system. Kishore Prahallad, Bayya Yegnanarayana |
ICASSP | 2 |
| 2000 | Enhancement of reverberant speech using LP residual signalabstractWe propose a new method of processing speech degraded by reverberation. The method is based on analysis of short (2 ms) segments of data to enhance the regions in the speech signal having a high signal-to-reverberant component ratio (SRR). The short segment analysis shows that SRR is different in different segments of speech. The processing method involves identifying and manipulating the linear prediction residual signal in three different regions of the speech signal, namely, high SRR region, low SRR region, and only reverberation component region. A weight function is derived to modify the linear prediction residual signal. The weighted residual signal samples are used to excite a time-varying all-pole filter to obtain perceptually enhanced speech. The method is robust to noise present in the recorded speech signal. The performance is illustrated through spectrograms, subjective and objective evaluations. Bayya Yegnanarayana, P. Satyanarayana Murthy |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | Analysis of autoassociative mapping neural networksabstractIn this paper we analyse the mapping behavior of an autoassociative neural network (AANN). The mapping in an AANN is achieved by using a dimension reduction followed by a dimension expansion. One of the major results of the analysis is that, the network performs better autoassociation as the size increases. This is because, a network of a given size can deal with only a certain level of nonlinearity. Performance of autoassociative mapping is illustrated with 2D examples. We have shown the utility of the mapping feature of an AANN for speaker verification. Shajith Ikbal, Hemant Misra, Bayya Yegnanarayana |
IJCNN | 3 |
| 1999 | Noise-invariant representation for speech signalsabstractBased on the concept of multiple-stream prior evolution and posterior pooling, we propose a new incremental adaptive Bayesian learning framework for e cient on-line adaptation of the continuous density hidden Markov model (CDHMM) parameters. As a rst step, we apply the a ne transformations to the mean vectors of CDHMMs to control the evolution of their prior distribution. This new stream of prior distribution can be combined with another stream of prior distribution evolved without any constraints applied. In a series of comparative experiments on the task of continuous Mandarin speech recognition, we show that the new adaptation algorithm achieves a similar fast-adaptation performance as that of incremental MLLR (maximum likelihood linear regression) in the case of small amount of adaptation data, while maintains the good asymptotic convergence property as that of our previously proposed quasi-Bayes adaptation algorithms. Aruna Bayya, Bayya Yegnanarayana |
EUROSPEECH | 2 |
| 1999 | A neural network-based text-dependent speaker verification system using suprasegmental featuresabstractWe describe the results of the ISAEUS project (TIDE DE 3004) under development. Its objective is to develop a prototype for training deaf people in three languages: French, German, and Spanish. M. Mathew, Bayya Yegnanarayana, R. Sundar |
EUROSPEECH | 2 |
| 1999 | Speech enhancement using linear prediction residual
Bayya Yegnanarayana, Carlos Avendaño, Hynek Hermansky, P. Satyanarayana Murthy |
Speech Commun. | 1 |
| 1999 | Robustness of group-delay-based method for extraction of significant instants of excitation from speech signalsabstractWe study the robustness of a group-delay-based method for determining the instants of significant excitation in speech signals. These instants correspond to the instants of glottal closure for voiced speech. The method uses the properties of the global phase characteristics of minimum phase signals. Robustness of the method against noise and distortion is due to the fact that the average phase characteristics of a signal is determined mainly by the strength of the excitation impulse. The strength of excitation is determined by the energy of the residual error signal around the instant of excitation. We propose a measure for the strength of the excitation based on Frobenius norm of the differenced signal. The robustness of the group-delay-based method is illustrated for speech under different types of degradations and for speech from different speakers. P. Satyanarayana Murthy, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | Enhancement of reverberant speech using LP residualabstractIn this paper we propose a new method of processing speech degraded by reverberation. The method is based on analysis of short (2 ms) segments of data to enhance the regions in the speech signal having high signal to reverberant component ratio (SRR). The short segment analysis shows that SRR is different in different segments of speech. The processing method involves identifying and manipulating the linear prediction residual in three different regions of the speech signal, namely, the high SRR region, the low SRR region and only the reverberation component region. A weighting function is derived to modify the LP residual. The weighted residual samples are used to excite the time-varying LP all-pole filter to obtain perceptually enhanced speech. Bayya Yegnanarayana, P. Satyanarayana Murthy, Carlos Avendaño, Hynek Hermansky |
ICASSP | 1 |
| 1998 | Robust features for speech recognition systemsabstractIn this paper we propose a set of features based on group delay spectrum for speech recognition systems. These features appear to be more robust to channel variations and environmental changes compared to features based on Melspectral coefficients. The main idea is to derive cepstrumlike features from group delay spectrum instead of deriving them from power spectrum. The group delay spectrum is computed from modified auto-correlation-like function. The effectiveness of the new feature set is demonstrated by the results of both speaker-independent (SI) and speaker-dependent (SD) recognition tasks. Preliminary results indicate that using the new features, we can obtain results comparable to Mel cepstra and PLP cepstra in most of the cases and a slight improvement in noisy cases. More optimization of the parameters is needed to fully exploit the nature of the new features. Aruna Bayya, Bayya Yegnanarayana |
ICSLP | 2 |
| 1998 | A review on merging some recent techniques with artificial neural networksabstractIn the last few years, there has been a large upswing in research activities aimed at synthesizing artificial neural networks with other well-established paradigms, like evolutionary computation, fuzzy logic, rough sets and chaos. In this paper, we briefly discuss the merits of these paradigms from artificial neural networks point of view, and we illustrate how these paradigms can be fused with the existing artificial neural network models to make the later one more efficient. Manish Sarkar, Bayya Yegnanarayana |
SMC | 2 |
| 1998 | Fuzzy-rough membership functionsabstractThis paper generalizes the concepts of rough membership functions in pattern classification tasks to fuzzy-rough membership functions. Unlike the rough membership value of a pattern, which is sensitive only towards the rough uncertainty associated with the pattern, the fuzzy-rough membership value of the pattern signifies the rough uncertainty as well as the fuzzy uncertainty associated it. In absence of fuzziness, the fuzzy-rough membership functions reduce to the existing rough membership functions. Moreover, under certain conditions the fuzzy-rough membership functions are equivalent to fuzzy membership functions or characteristic functions. In this paper, various set theoretic properties of the fuzzy-rough membership functions are exploited to characterize the concept of fuzzy-rough sets. Some measures of the fuzzy-rough ambiguity associated with a given output class are also discussed. Manish Sarkar, Bayya Yegnanarayana |
SMC | 2 |
| 1998 | Fuzzy-rough neural networks for vowel classificationabstractIn many real life applications two patterns from the same cluster belong to different classes, and hence, classification based on mere similarity property is inadequate. This problem arises because the available features are not sufficient to discriminate the classes. It implies that the fuzzy clusters generated by the input features have rough uncertainty. This paper proposes a fuzzy-rough set based network which exploits fuzzy-rough membership functions to reduce this problem. The proposed network is theoretically a powerful classifier as it is equivalent to a universal approximator. Moreover, its activity is transparent as it can easily be mapped to a Takagi-Sugeno type fuzzy rule base system. The efficacy of the proposed method is studied on a vowel recognition problem. Manish Sarkar, Bayya Yegnanarayana |
SMC | 2 |
| 1998 | Backpropagation learning algorithms for classification with fuzzy mean square error
Manish Sarkar, Bayya Yegnanarayana, Deepak Khemani |
Pattern Recognit. Lett. | 2 |
| 1998 | Extraction of vocal-tract system characteristics from speech signalsabstractWe propose methods to track natural variations in the characteristics of the vocal-tract system from speech signals. We are especially interested in the cases where these characteristics vary over time, as happens in dynamic sounds such as consonant-vowel transitions. We show that the selection of appropriate analysis segments is crucial in these methods, and we propose a selection based on estimated instants of significant excitation. These instants are obtained by a method based on the average group-delay property of minimum-phase signals. In voiced speech, they correspond to the instants of glottal closure. The vocal-tract system is characterized by its formant parameters, which are extracted from the analysis segments. Because the segments are always at the same relative position in each pitch period, in voiced speech the extracted formants are consistent across successive pitch periods. We demonstrate the results of the analysis for several difficult cases of speech signals. Bayya Yegnanarayana, Raymond N. J. Veldhuis |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | An iterative algorithm for decomposition of speech signals into periodic and aperiodic componentsabstractThe speech signal may be considered as the output of a time-varying vocal tract system excited with quasiperiodic and/or random sequences of pulses. The quasiperiodic part may be considered as the deterministic or periodic component and the random part as the stochastic or aperiodic component of the excitation. We discuss issues involved in identifying and separating the periodic and aperiodic components of the source. The decomposition is performed on an approximation to the excitation signal, instead of decomposing the speech signal directly. The linear prediction residual signal is used as an approximation to the excitation signal of the vocal tract system. Speech is first analyzed to determine the voiced and unvoiced parts of the signal. Decomposition of the voiced part into periodic and aperiodic components is then accomplished by first identifying the frequency regions of harmonic and noise components in the spectral domain. The signal corresponding to the noise regions is used as a first approximation to the aperiodic component. An iterative algorithm is proposed which reconstructs the aperiodic component in the harmonic regions. The periodic component is obtained by subtracting the reconstructed aperiodic component signal from the residual signal. The individual components of the residual are then used to excite the derived all-pole model of the vocal tract system to obtain the corresponding components of the speech signal. Experiments were conducted using synthetic speech. They demonstrated the ability of the algorithm for decomposition of a synthetic speech signal made of a mixture of periodic and aperiodic components. Application to natural speech is also discussed. Bayya Yegnanarayana, Christophe d'Alessandro, Vassilios Darsinos |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Effectiveness of a periodic and aperiodic decomposition method for analysis of voice sourcesabstractDecomposition of speech into periodic and aperiodic components is useful in analyzing and describing the characteristics of voice sources. Such a decomposition is also useful in controlling the excitation source for synthesis. This paper addresses the issue of decomposition of speech into periodic and aperiodic components in the context of speech production. The effectiveness of a recently proposed algorithm for decomposing speech into these components is examined for analysis of voice sources. Synthetic signals are generated using formant synthesis. Different sources of aperiodicity encountered in normal speech production are considered, using a set of parameters to control the synthetic signals. The sources of aperiodicity studied are: (1) additive pulsed or continuous random noise, and (2) modulation aperiodicities due to variation in the fundamental frequency, jitter, and shimmer. Three types of measures are used to characterize these voices: ratio of energies in the periodic and aperiodic components, perceptual spectral distance, and spectrograms. The results demonstrate the effectiveness of the periodic-aperiodic decomposition algorithm for analyzing aperiodicities for a wide variety of voices, and point out the limitations of the algorithm. Christophe d'Alessandro, Vassilios Darsinos, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 3 |
| 1998 | Supervised texture classification using a probabilistic neural network and constraint satisfaction modelabstractIn this paper, the texture classification problem is projected as a constraint satisfaction problem. The focus is on the use of a probabilistic neural network (PNN) for representing the distribution of feature vectors of each texture class in order to generate a feature-label interaction constraint. This distribution of features for each class is assumed as a Gaussian mixture model. The feature-label interactions and a set of label-label interactions are represented on a constraint satisfaction neural network. A stochastic relaxation strategy is used to obtain an optimal classification of textures in an image. The advantage of this approach is that all classes in an image are determined simultaneously, similar to human perception of textures in an image. P. P. Raghu, Bayya Yegnanarayana |
IEEE Trans. Neural Networks | 2 |
| 1997 | Processing linear prediction residual for speech enhancementabstractIn this paper we propose a method for enhancement of speech in the presence of additive noise. The objective is to selectively enhance the high SNR regions in the noisy speech in the temporal and spectral domains, without causing significant distortion in the resulting enhanced speech. This is proposed to be done at three different levels: (a) At the gross level, by identifying the regions of speech and noise in the temporal domain, (b) At the finer level, by identifying the regions of high and low SNR portions in the noisy speech, and (c) At the short--time spectrum level, by enhancing the spectral peaks over spectral valleys. Processing of noisy speech for enhancement involves mostly weighting the LP residual samples. The weighted residual samples are used to excite the time-- varying LP filter to produce enhanced speech. 1. INTRODUCTION Speech signal collected under normal environmental conditions is usually degraded due to noise and distortions. Performance of speech systems depe... Bayya Yegnanarayana, Carlos Avendaño, Hynek Hermansky, P. Satyanarayana Murthy |
EUROSPEECH | 1 |
| 1997 | Multispectral Image Classification Using Gabor Filters and Stochastic Relaxation Neural Network
P. P. Raghu, Bayya Yegnanarayana |
Neural Networks | 2 |
| 1997 | A clustering algorithm using an evolutionary programming-based approach
Manish Sarkar, Bayya Yegnanarayana, Deepak Khemani |
Pattern Recognit. Lett. | 2 |
| 1997 | Unsupervised texture classification using vector quantization and deterministic relaxation neural networkabstractThis paper describes the use of a neural network architecture for classifying textured images in an unsupervised manner using image-specific constraints. The texture features are extracted by using two-dimensional (2-D) Gabor filters arranged as a set of wavelet bases. The classification model comprises feature quantization, partition, and competition processes. The feature quantization process uses a vector quantizer to quantize the features into codevectors, where the probability of grouping the vectors is modeled as Gibbs distribution. A set of label constraints for each pixel in the image are provided by the partition and competition processes. An energy function corresponding to the a posteriori probability is derived from these processes, and a neural network is used to represent this energy function. The state of the network and the codevectors of the vector quantizer are iteratively adjusted using a deterministic relaxation procedure until a stable state is reached. The final equilibrium state of the vector quantizer gives a classification of the textured image. A cluster validity measure based on modified Hubert index is used to determine the optimal number of texture classes in the image. P. P. Raghu, R. Poongodi, Bayya Yegnanarayana |
IEEE Trans. Image Process. | 3 |
| 1996 | Word boundary hypothesization for continuous speech in Hindi based on F0 patterns
Bayya Yegnanarayana |
Speech Commun. | 2 |
| 1996 | Source-system windowing for speech analysis and synthesisabstractA new method of analysis of speech is proposed that will bring out variations in vocal tract system characteristics in short (2-4 ms) segments. In this method, the source and system components of the speech signal are suitably windowed to reduce the effects of truncation of conventional waveform windowing. Bayya Yegnanarayana, P. Satyanarayana Murthy |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | Segmentation of Gabor-filtered textures using deterministic relaxationabstractA supervised texture segmentation scheme is proposed in this article. The texture features are extracted by filtering the given image using a filter bank consisting of a number of Gabor filters with different frequencies, resolutions, and orientations. The segmentation model consists of feature formation, partition, and competition processes. In the feature formation process, the texture features from the Gabor filter bank are modeled as a Gaussian distribution. The image partition is represented as a noncausal Markov random field (MRF) by means of the partition process. The competition process constrains the overall system to have a single label for each pixel. Using these three random processes, the a posteriori probability of each pixel label is expressed as a Gibbs distribution. The corresponding Gibbs energy function is implemented as a set of constraints on each pixel by using a neural network model based on Hopfield network. A deterministic relaxation strategy is used to evolve the minimum energy state of the network, corresponding to a maximum a posteriori (MAP) probability. This results in an optimal segmentation of the textured image. The performance of the scheme is demonstrated on a variety of images including images from remote sensing. P. P. Raghu, Bayya Yegnanarayana |
IEEE Trans. Image Process. | 2 |
| 1995 | Decomposition of speech signals into deterministic and stochastic componentsabstractThis paper presents a new method for decomposition of the speech signal into a deterministic and a stochastic component. The method is based on iterative signal reconstruction. The method involves: (1) separation of speech into an approximate excitation and filter components using linear predictive (LP) analysis; (2) identification of frequency regions of noise and deterministic components of excitation using cepstrum; (3) reconstruction of the two excitation components of the residual using an iterative algorithm; (4) and finally, the deterministic and stochastic components of the excitation are then obtained by combining the reconstructed frames of data using an overlap-add procedure. The deterministic and stochastic components are then passed through the time varying all-pole filter to obtain the components of the speech signal. The algorithm is able to decompose varying mixtures of stochastic and deterministic signals, like the noise bursts produced at the glottal closure and the deterministic glottal pulses. This new algorithm is a powerful tool for analysis of relevant features of the source component of speech signals. Christophe d'Alessandro, Bayya Yegnanarayana, Vassilios Darsinos |
ICASSP | 2 |
| 1995 | A robust method for determining instants of major excitations in voiced speechabstractWe propose a method for determining the instants of significant excitation in speech signals using the negative derivative of the unwrapped phase (group delay) function of the short time Fourier transform. Here significant excitation refers primarily to the instants of glottal closure in voiced speech. The method computes the average slope of the unwrapped phase spectrum as a function of time. The instants where the phase slope function makes a positive zero-crossing correspond to the major excitations in the signal. For an analysis window size in the range of one to two pitch periods, these instants coincide with the instants of glottal closure in each pitch period. The method is robust, as it depends only on the average phase slope value, and further, it depends only on the positive zero-crossing instants of the average phase slope function. Bayya Yegnanarayana, R. L. H. M. Smits |
ICASSP | 1 |
| 1995 | Stereo-correspondence using Gabor logons and neural networksabstractStereo-correspondence is the most important issue in stereopsis. Feature extraction and matching are the basic steps involved in the solution of the stereo-correspondence problem. The article examines the effectiveness of Gabor logons as a pre-processing technique compared to the intensity image. The matching is performed using a Hopfield network and simulated annealing. The performance of these matching techniques with respect to their accuracy and execution speed is analysed. The effect of weightages to constraints and network parameters is also analysed. Simulated annealing is found to give much faster convergence compared to the Hopfield network. Babu Thomas, Bayya Yegnanarayana |
ICIP | 2 |
| 1995 | Evaluation of a periodic/aperiodic speech decomposition algorithm
Vassilios Darsinos, Christophe d'Alessandro, Bayya Yegnanarayana |
EUROSPEECH | 3 |
| 1995 | A combined neural network approach for texture classification
P. P. Raghu, R. Poongodi, Bayya Yegnanarayana |
Neural Networks | 3 |
| 1995 | Studies on object recognition from degraded images using neural networks
A. Ravichandran, Bayya Yegnanarayana |
Neural Networks | 2 |
| 1995 | Transformation of formants for voice conversion using artificial neural networks
M. Narendranath, Hema A. Murthy, Bayya Yegnanarayana |
Speech Commun. | 4 |
| 1995 | Determination of instants of significant excitation in speech using group delay functionabstractA new method for determining the instants of significant excitation in speech signals is proposed. In the paper, significant excitation refers primarily to the instant of glottal closure within a pitch period in voiced speech. The method is based on the global phase characteristics of minimum phase signals. The average slope of the unwrapped phase of the short-time Fourier transform of linear prediction residual is calculated as a function of time. Instants where the phase slope function makes a positive zero-crossing are identified as significant excitations. The method is discussed in a source-filter context of speech production. The method is not sensitive to the characteristics of the filter. The influence of the type, length, and position of the analysis window is discussed. The method works well for all types of voiced speech in male as well as female speech but, in all cases, under noise-free conditions only.> R. Smits, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 2 |
| 1994 | A speaker verification system using prosodic features
Bayya Yegnanarayana, S. P. Wagh |
ICSLP | 1 |
| 1993 | Intonation component of a text-to-speech system for Hindi
A. S. Madhukumar, Bayya Yegnanarayana |
Comput. Speech Lang. | 3 |
| 1991 | A two-stage neural network for translation, rotation and size-invariant visual pattern recognitionabstractA two-stage neural network is described for transformation-invariant visual pattern recognition. In the first stage, features are extracted after normalizing the image. It is shown how parameters of spatial transformation can be estimated even in the presence of noise by using knowledge about rigid objects. Circular arcs in the normalized image are used as generalized features to describe the input pattern. Each image pixel contributes to the features which it can constitute. Contributions from noisy pixels are distributed over the feature space, whereas meaningful parts contribute to clusters that correspond to features of the image. In the second stage, the image is classified on the basis of these features by a multilayer perceptron network trained using a backpropagation algorithm.> A. Ravichandran, Bayya Yegnanarayana |
ICASSP | 2 |
| 1991 | Processing of noisy speech using modified group delay functionsabstractA novel method of processing noisy speech is presented. The method exploits the properties of the negative derivative of the Fourier transform phase spectrum (group delay function) to derive the features of the vocal tract system and the excitation from the speech signal. The key idea used is that the properties of group delay functions for noise and a stable all-pole filter are distinct. Estimation of the spectrum of the vocal tract system and fundamental frequency are treated as problems of spectrum estimation from noisy data. Results of these studies show that intelligible speech can be synthesized from parameters derived from noisy data with an overall signal-to-noise ratio (SNR) as low as 3 dB.> Bayya Yegnanarayana, Hema A. Murthy, V. R. Ramachandran |
ICASSP | 1 |
| 1991 | Synthesizing intonation for speech in hindi
A. S. Madhukumar, Chellu Chandra Sekhar, Bayya Yegnanarayana |
EUROSPEECH | 4 |
| 1991 | Speech processing using group delay functions
Hema A. Murthy, Bayya Yegnanarayana |
Signal Process. | 2 |
| 1991 | Formant extraction from group delay function
Hema A. Murthy, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 1990 | An algorithm for thinning noisy imagesabstractA two-step algorithm is presented for thinning noisy images. The first step consists of removing noisy spurs from an image using a combination of morphological operations. In the second step, thinning is performed by using the standard morphological skeleton transform (MST). This approach is useful in mode-based vision applications where the features of the image are required for object recognition. The use of this algorithm in sensor array imaging situations is demonstrated.> Bayya Yegnanarayana, R. Ramaseshan, A. Ravichandran |
ICASSP | 1 |
| 1990 | Speech enhancement using group delay functions
Bayya Yegnanarayana, Hema A. Murthy, V. R. Ramachandran |
ICSLP | 1 |
| 1990 | A maximum entropy approach to interpolation
S. Tanveer Fathima, Bayya Yegnanarayana |
Signal Process. | 2 |
| 1989 | A nonparametric method of formant estimation using group delay spectraabstractA novel minimum-phase group delay technique is discussed that is a nonparametric spectral analysis method possessing the ability to demerge closely coupled formants and detect weak formants. It therefore requires no assumptions concerning the underlying nature of the signal, other than that it originates from an LTI filter system. The technique provides a level of performance in formant detection normally associated only with larynx-synchronous techniques. Moreover, it is extremely easy to implement, requiring only three FFT operations per analysis frame.> G. Duncan, Bayya Yegnanarayana, Hema A. Murthy |
ICASSP | 2 |
| 1989 | Formant extraction from Fourier transform phaseabstractA method of extracting formant information from the short-time Fourier transform phase spectrum of speech is proposed. Fourier transform phase has not been used for formant extraction because it appears to be noisy and difficult to interpret. The effects of wrapping of phase (due to zeros close to the unit circle and the linear phase component) make it difficult to derive useful information. The authors develop algorithms to reduce the effects of wrapping.> Hema A. Murthy, K. V. Madhu Murthy, Bayya Yegnanarayana |
ICASSP | 3 |
| 1989 | Parsing spoken utterances in an inflectional language
M. Prakash 0002, Venkata Ramana Rao Gadde, Chellu Chandra Sekhar, Bayya Yegnanarayana |
EUROSPEECH | 4 |
| 1989 | Word boundary hypothesisation in hindi speech
Venkata Ramana Rao Gadde, M. Prakash 0002, Bayya Yegnanarayana |
EUROSPEECH | 3 |
| 1989 | Analysis of short time speech segments based on linear prediction
Bayya Yegnanarayana, K. V. Madhu Murthy |
EUROSPEECH | 1 |
| 1989 | Voice conversion
Donald G. Childers, D. M. Hicks, Bayya Yegnanarayana |
Speech Commun. | 4 |
| 1987 | Reconstruction from Fourier transform phase with applications to speech analysisabstractThis paper addresses the problem of signal reconstruction from Fourier transform phase. In particular, we examine two aspects of this problem. First, we discuss signal reconstruction from the phase spectrum of the short-time Fourier transform(STFT). Next, we examine the problem of signal recovery from partial phase information. We present the results of our studies on reconstruction from partial phase and discuss the application of these results in speech analysis and coding. Bayya Yegnanarayana, S. Tanveer Fathima, Hema A. Murthy |
ICASSP | 1 |
| 1986 | An algorithm for bandlimited signal interpolationabstractIn this paper we present an algorithm for band-limited signal interpolation assuming the samples to be known at some randomly distributed instants. The algorithm is a modification of the Papoulis-Gerchberg algorithm. The key idea in our algorithm is to use some nonzero values at the missing points. The values are obtained using an interpolation scheme based on relaxation method for constraint propagation. The algorithm is illustrated with several examples. Bayya Yegnanarayana, S. Tanveer Fathima |
ICASSP | 1 |
| 1985 | Measuring source-tract interaction from speechabstractThe objective of our study is to define and measure source-tract interaction using the speech signal. Numerous researchers have conjectured that the subglottal, glottal, and supraglottal portions of our own speech production mechanism may interact, affecting the quality of our voice. Several methods for incorporating the effects of source-tract interaction into a synthetic speech model have been suggested. Speech synthesized with source-tract interaction sounds more natural than speech generated without such interaction. Can source-tract interaction be parameterized using the speech signal and incorporated into vocoders and speech synthesizers? Can this measurement be accomplished on a pitch period by pitch period basis? A. S. Ananth, Donald G. Childers, Bayya Yegnanarayana |
ICASSP | 3 |
| 1985 | Voice conversion: Factors responsible for qualityabstractA flexible analysis-synthesis system with signal dependent features is described and used to realize some desired voice characteristics in synthesized speech. The intelligibility of synthetic speech appears to depend on the ability to reproduce dynamic sounds such as stops, whereas the quality of voice is mainly determined by the true reproduction of voiced segments. We describe our work in converting the speech of one speaker to sound like that of another. A number of factors are important for maintaining the quality of the voice during this conversion process. These factors are derived from both the speech and electroglottograph signals. Donald G. Childers, Bayya Yegnanarayana |
ICASSP | 2 |
| 1985 | Processing of noisy speech using group delay functionsabstractA new noniterative technique for all pole modelling of noisy speech is proposed. We exploit the additive and high resolution properties of a group delay function to identify the high signal to noise ratio (SNR) regions of the short time spectrum and to separate out these regions from the low SNR regions. A modified spectrum is derived from the group delay function in these regions. The new spectrum is used to derive the all pole model for the speech segment. The main advantage of this method is that it is not necessary to have prior knowledge of the noise or speech characteristics as in the other methods of processing noisy speech. Joy A. Thomas, Bayya Yegnanarayana, Raghuram Karinthi, V. Venkateswar |
ICASSP | 2 |
| 1984 | Voice Simulation: Factors Affecting Quality And NaturalnessabstractIn this paper we describe a flexible analysis-synthesis system which can be used for a number of studies in speech research. The main objective is to have a synthesis system whose characteristics can be controlled through a set of parameters to realize any desired voice characteristics. The basic synthesis scheme consists of two steps: Generation of an excitation signal from pitch and gain contours and excitation of the linear system model described by linear prediction coefficients. We show that a number of basic studies such as time expansion/compression, pitch modifications and spectral expansion/compression can be made to study the effect of these parameters on the quality of synthetic speech. A systematic study is made to determine factors responsible for unnaturalness in synthetic speech. It is found that the shape of the glottal pulse determines the quality to a large extent. We have also made some studies to determine factors responsible for loss of intelligibility in some segments of speech. A signal dependent analysis-synthesis scheme is proposed to improve the intelligibility of dynamic sounds such as stops. A simple implementation of the signal dependent analysis is proposed. Bayya Yegnanarayana, Jayant M. Naik, Donald G. Childers |
COLING | 1 |
| 1984 | Performance of isolated word recognition system for confusable vocabularyabstractIn this paper we discuss some of the limitations of the existing isolated word speech recognition system (IWSR) when applied to confusable vocabulary. For our study we have chosen a subset of Hindi stop consonants as the confusable word set. The members of this set differ among themselves primarily in the short leading consonant part and at the interface of the consonant and the following dominant vowel part. We adopt a signal-dependent approach for parameter extraction and matching strategy. This approach gives better performance compared with the conventional approach, but the performance still falls far short of the desired goal of 100% recognition. Refined signal processing suitable for appropriate segments of speech appear to be the way out of this problem. We discuss our studies in this direction. We use a new measure, called Performance Index, to evaluate the changes in performance due to innovations carried out on small data sets. S. Raman 0001, Bayya Yegnanarayana |
ICASSP | 2 |
| 1984 | Performance of isolated word recognition system for degraded speechabstractThe performance of an isolated word speech recognition (IWSR) system is known to drop rapidly with increase in the degradation of the input speech. In this paper we propose a recognition scheme which adapts itself to mild degradations in speech. The scheme does not need apriori information regarding the nature and extent of noise. We suggest techniques which adaptively discriminate between noisy and noise-free parameters by using a selective weighting procedure in the final distance calculations. A suitable index is used to study the performance of the recognition system for small data sets. Our scheme lends itself to greater flexibility in handling degradations in speech input than do the existing recognition schemes. We illustrate our scheme by simulating an adaptive differential pulse code modulated (ADPCM) speech, where the main distortion is contributed by the quatization noise. Bayya Yegnanarayana, Sarat Chandran |
ICASSP | 1 |
| 1983 | Noniterative techniques for minimum phase signal reconstruction from phase or magnitudeabstractNew noniterative techniques are proposed for reconstruction of signal from samples of magnitude or phase of the Fourier transform of the signal. The only condition for reconstruction is that the signal is a minimum phase one. The basis for these new techniques is the relation between the magnitude and phase functions through cepstral coefficients The techniques are illustrated through several examples. In all the cases we find that phase from magnitude can be obtained exactly and magnitude from phase can be obtained to within a scale factor. Effects of truncation of minimum phase signals and aliasing due to sampling in the frequency domain are discussed. These studies show that effective noniterative techniques can be evolved for signal reconstruction instead of cumbersome iterative procedures suggested in literature recently. Bayya Yegnanarayana, A. Dhayalan |
ICASSP | 1 |
| 1981 | A pole-zero model for cepstrally smoothed speech spectraabstractIn this paper a new method for representation of cepstrally smoothed speech spectra by a pole-zero model is presented. In this method the cepstrally smoothed log spectrum is split into two parts, one corresponding to the response of the numerator polynomial of the model transfer function and the other part to the response of the denominator polynomial of the model transfer function. The decomposition is achieved by using the properties of the derivative of phase spectra of minimum phase signals. The inverse of each of these responses is approximated by a small number of auto-regressive coefficients. The method is illustrated with several examples of speech spectra. The residual from the inverse pole-zero model system can be used to obtain information about the excitation signal. The technique proposed in this paper can be used to represent any arbitrary smoothed log spectrum by a pole-zero model of appropriate order. Bayya Yegnanarayana |
ICASSP | 1 |
| 1980 | Pole-zero decomposition: A new technique for design of digital filtersabstractA new technique for design of digital filters is presented in this paper. The technique consists of splitting the given log magnitude response into two parts, one corresponding to the response of the numerator polynomial of the filter transfer function and the other part to the response of the denominator polynomial. The inverse of each of these polynomials is considered as an all-pole filter and the response of the all-pole filter is approximated by a small number of autoregressive coefficients. The autoregressive coefficients obtained for the numerator polynomial represent the zero part of the final filter and the coefficients obtained for the denominator polynomial represent the pole part of the final filter. With equal number of poles and zeros, the overall filter response can be made nearly equiripple in the passband and stopband. The amplitude of the ripple can be traded with the width of the transition band. The ripple characterstics can be controlled by appropriately choosing the number of poles and zeros of the filter. Bayya Yegnanarayana |
ICASSP | 1 |
| 1979 | A distance measure based on the derivative of linear prediction phase spectrumabstractA new distance measure based on the derivative of linear prediction (LP) phase spectrum is proposed for comparison of speech spectra. Relationships among several distance measures based on the linear prediction coefficients (LPCs) are discussed. The advantages of the new measure and an efficient method of computing it are also discussed. Bayya Yegnanarayana, Raj Reddy |
ICASSP | 1 |
| 1978 | Epoch extraction from linear prediction residualabstractAn interpretation of linear prediction (LP) residual is presented by considering the effect of following factors: shape of glottal pulse, phase angles of formants at the instant of excitation, inaccurate estimation of formants and bandwidths, zeroes in vocal tract system transfer function. Effect of improper phase cancellation on the accuracy of estimated epoch position is also discussed. A method for unambiguous identification of epochs from LP residual is presented. T. V. Ananthapadmanabha, Bayya Yegnanarayana |
ICASSP | 2 |
| 1978 | Nearest neighbour decision rule for vowel and digit recognitionabstractMinimum distance to mean is usually used as a classification rule in speech and speaker recognition studies. In this paper it is shown that the nearest neighbour decision rule gives significant improvement in classification score for vowel and digit recognition schemes. Autocorrelation coefficients of lags two to five sampling instants are used to form the feature vector. Pour samples per class have been used. Minimum squared Euclidean distance of the test vector from the nearest reference is chosen as the classification rule. For sustained vowels the recognition score is cent percent. for the same feature the minimum distance to mean gives 70 % recognition score. When the reference samples of a given speaker is tested over the vowels spoken by different speaker(up to 10), this scheme gives the recognition score of about 95 %. for digits without any time warping the recognition score of about 86 % to 92 % is obtained. T. K. Raja, Bayya Yegnanarayana |
ICASSP | 2 |
| 1976 | Cascade realization of digital inverse filter for extracting speaker dependent featuresabstractQuest for new speaker dependent features is a constant problem in the design of automatic speaker recognition systems. In speech, information about the speaker usually arises along with the semantic information which makes its independent use difficult. In this paper, a method based on linear prediction (LP) analysis is described which yields features that are more speaker dependent than the usual linear predictor coefficients (LPC). In this method the LPC contours are obtained through cascade realization of digital inverse filtering (DIF) for speech signals. A low order (2-4) DIF removes the gross spectral characteristics such as the large dynamic range and some significant peaks which tend to mask the weaker formants. Visual comparison of the contours and a preliminary statistical analysis indicate that the LPC contours obtained by processing the output signal of the first stage contain better features for speaker dependency than the direct LPC contours. V. V. S. Sarma, Bayya Yegnanarayana |
ICASSP | 2 |
| 1976 | Effect of noise and distortion in speech on parametric extractionabstractParameter or feature extraction from speech signal forms the basis for systems designed for speech recognition, speaker verification, speech bandwidth compression etc. The parameters in general are critically dependent upon the short-time spectrum of speech. The input speech waveform is however, subjected to several types of noises and distortions due to background noise sources, reverberation, close speaking into a microphone, telephone system imperfections etc. These factors modify the spectrum of the speech signal and hence the parameters extracted. Characteristics of common sources of noise and distortion are described in this paper and their effect in shaping the spectrum of speech is discussed. Steps to reduce the influence of some noises while producing speech input to a system are suggested. Methods of normalization of spectral distortions due to noise and the effect of such normalization on parametric extraction are also discussed. Bayya Yegnanarayana |
ICASSP | 1 |