Saeed Vaseghi

dblp:34/6509 · DBLP profile ↗
← Back
70ranked-venue papers
13as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 12 first-authorArtificial intelligence and machine learning · 39 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Audio and music processing · 55% Image and video processing · 45%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › video frame interpolation › interpolation › image interpolation
edge-directed interpolation
0.212016
Edge-Guided Image Gap Interpolation Using Multi-Scale Transformation · IEEE Trans. Image Process. 2016
Image and video processing
image restoration
0.212016
Edge-Guided Image Gap Interpolation Using Multi-Scale Transformation · IEEE Trans. Image Process. 2016
Audio and music processing
speech processing
0.222008
Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing · IEEE Trans. Multim. 2008
Analysis and Synthesis of Formant Spaces of British, Australian, and American Accents · IEEE Trans. Speech Audio Process. 2007
Audio and music processing › speech coding
digital speech interpolation
0.112008
Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing · IEEE Trans. Multim. 2008
Audio and music processing › speech coding
packet loss concealment
0.112008
Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing · IEEE Trans. Multim. 2008
Audio and music processing › speech analysis
formant tracking
0.112007
Analysis and Synthesis of Formant Spaces of British, Australian, and American Accents · IEEE Trans. Speech Audio Process. 2007
Audio and music processing › sound synthesis
harmonic plus noise model
0.112007
Noisy Speech Enhancement Using Harmonic-Noise Model and Codebook-Based Post-Processing · IEEE Trans. Speech Audio Process. 2007
Audio and music processing
speech enhancement
0.112007
Noisy Speech Enhancement Using Harmonic-Noise Model and Codebook-Based Post-Processing · IEEE Trans. Speech Audio Process. 2007
Audio and music processing
linear prediction
0.012008
Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing · IEEE Trans. Multim. 2008
Audio and music processing
speech coding
0.012008
Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing · IEEE Trans. Multim. 2008
Audio and music processing
speech synthesis
0.012007
Analysis and Synthesis of Formant Spaces of British, Australian, and American Accents · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011997
Noise compensation methods for hidden Markov model speech recognition in adverse environments · IEEE Trans. Speech Audio Process. 1997
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
noise-robust speech recognition
0.011997
Noise compensation methods for hidden Markov model speech recognition in adverse environments · IEEE Trans. Speech Audio Process. 1997

Methods — techniques the papers use, named apart from their topics

multi-scale pyramid transform · 0.2discrete wavelet transform · 0.2discrete cosine transform · 0.2canny edge detection · 0.2linear prediction · 0.1harmonic noise model · 0.1codebook post-processing · 0.1autoregressive interpolation · 0.1probability distribution modeling · 0.1linear prediction analysis · 0.1wiener filter · 0.0spectral subtraction · 0.0noise-adaptive HMM · 0.0
YearPublicationVenuePosition
2016 Edge-Guided Image Gap Interpolation Using Multi-Scale Transformation
abstract
This paper presents improvements in image gap restoration through the incorporation of edge-based directional interpolation within multi-scale pyramid transforms. Two types of image edges are reconstructed: 1) the local edges or textures, inferred from the gradients of the neighboring pixels and 2) the global edges between image objects or segments, inferred using a Canny detector. Through a process of pyramid transformation and downsampling, the image is progressively transformed into a series of reduced size layers until at the pyramid apex the gap size is one sample. At each layer, an edge skeleton image is extracted for edge-guided interpolation. The process is then reversed; from the apex, at each layer, the missing samples are estimated (an iterative method is used in the last stage of upsampling), up-sampled, and combined with the available samples of the next layer. Discrete cosine transform and a family of discrete wavelet transforms are utilized as alternatives for pyramid construction. Evaluations over a range of images, in regular and random loss pattern, at loss rates of up to 40%, demonstrate that the proposed method improves peak-signal-to-noise-ratio by 1-5 dB compared with a range of best published works.
Bahareh Langari, Saeed Vaseghi, Ales Procházka, Babak Vaziri, Farzad Tahmasebi Aria
IEEE Trans. Image Process.2
2014 Audio packet loss concealment using spectral motion
abstract
This paper presents a packet loss concealment (PLC) method with applications to VoIP, audio broadcast and streaming. The problem of modeling of time-varying frequency spectrum in the context of PLC is addressed and a novel solution is proposed for tracking and using the temporal motion of spectral flow. The proposed PLC utilizes a time-frequency motion (TFM) matrix representation of the audio signal where each frequency is tagged with a motion vector estimate. The spectral motion vectors are estimated by cross-correlating the movement of spectral energy within subbands across time frames. The missing packets are estimated in TFM domain and inverse transformed to the time-domain. The proposed method is compared with conventional approaches using objective performance evaluation of speech quality (PESQ), and subjective mean opinion scores (MOS) in a range of packet loss from 5% to 20%. The results demonstrate that the proposed algorithm improves performance.
Seyed Kamran Pedram, Saeed Vaseghi, Bahareh Langari
ICASSP2
2011 Fundamental Frequency Estimation Using Modified Higher Order Moments and Multiple Windows
abstract
This paper proposes a set of higher-order modified moments for estimation of the fundamental frequency of speech and explores the impact of the speech window length on pitch estimation error. The pitch extraction methods are evaluated in a range of noise types and SNRs. For calculation of errors, pitch reference values are calculated from manually-corrected estimates of the periods obtained from laryngograph signals. The results obtained for the 3rd and 4th order modified moment compare well with methods based on correlation and magnitude difference criteria and the YIN method; with improved pitch accuracy and less occurrence of large errors
Alipah Pawi, Saeed Vaseghi, Ben P. Milner, Seyed Ghorshi
INTERSPEECH2
2010 Pitch extraction using modified higher order moments
abstract
This paper proposes a set of higher-order modified moments as alternative objective criteria for pitch extraction and explores the impact of the speech window length on pitch estimation error. To obtain the Kthorder modified moment, each speech frame is split into a positive-valued signal and a negative-valued signal. The magnitudes of the Kthorder moments for the positive and the negative valued signals are obtained and combined. The proposed objective criteria form a relatively sharp peak around the true pitch value compared to the correlation function. For calculation of errors, pitch reference (`ground truth') values are calculated from manually-corrected estimates of the periods obtained from laryngograph signals. The results obtained for the third order modified moment are compared with the results for correlation and magnitude difference criteria and the YIN method. The modified moments provide improved pitch accuracy with less occurrence of large errors (e.g. half or double pitch estimation errors).
Alipah Pawi, Saeed Vaseghi, Ben P. Milner
ICASSP2
2010 Modeling and synthesis of English regional accents with pitch and duration correlates
Qin Yan, Saeed Vaseghi
Comput. Speech Lang.2
2008 Applying noise compensation methods to robustly predict acoustic speech features from MFCC vectors in noise
abstract
This paper examines the effect of applying noise compensation to improve acoustic speech feature prediction from noise contaminated MFCC vectors, as may be encountered in distributed speech recognition (DSR). A brief review of maximum a posteriori prediction of acoustic speech features (voicing, fundamental and formant frequencies) from MFCC vectors is made. Two noise compensation methods are then applied; spectral subtraction and model adaptation. Spectral subtraction is used to filter noise from the received MFCC vectors, while model adaptation is applied to adapt the joint models of acoustic features and MFCCs to account for noise contamination. Experiments examine acoustic feature prediction accuracy in noise and results show that the two noise compensation methods significantly improve prediction accuracy in noise. The technique of model adaptation was found to be better than spectral subtraction and could restore performance close to that achieved in matched training and testing.
Ben P. Milner, Jonathan Darch, Saeed Vaseghi
ICASSP3
2008 Kalman tracking of linear predictor and harmonic noise models for noisy speech enhancement
Qin Yan, Saeed Vaseghi, Esfandiar Zavarehei, Ben P. Milner, Jonathan Darch, Paul R. White, Ioannis Andrianakis
Comput. Speech Lang.2
2008 Cross-entropic comparison of formants of British, Australian and American English accents
Seyed Ghorshi, Saeed Vaseghi, Qin Yan
Speech Commun.2
2008 Interpolation of Lost Speech Segments Using LP-HNM Model With Codebook Post-Processing
abstract
This paper presents a method for interpolation of lost speech segments. The interpolation method can be used for packet loss concealment in voice communication over mobile phones, for voice over IP or for restoration of lost segments in speech recordings. The interpolation method employs a combination of a linear prediction (LP) model of the spectral envelope and a harmonic noise model (HNM) of the excitation of speech. The speech interpolation problem is transformed to the modeling and interpolation of the trajectories of LP parameters and the amplitude, phase and harmonicity of HNM tracks of speech excitation. In particular, the interpolation of harmonicity results in a smooth transition from voiced to unvoiced speech and vice versa. Crucially, the proposed interpolation method does not suffer from the consequences of zero-excitation of conventional autoregressive (AR) interpolation. Different combinations of linear and autoregressive interpolation methods are evaluated for the estimation of the time-varying parameters of LP-HNM tracks. Furthermore, a post-processing codebook mapping, employed to enhance the interpolation of the spectral envelope of speech, results in improved output quality for longer length speech gaps. For different packet loss rates and patterns of distributions of missing speech gaps, the proposed interpolation methods are evaluated and compared with popular AR-based interpolation methods and the speech packet recovery method specified in the ITU G.711 standard, as a reference. The evaluation results show that the proposed methods substantially improve the restoration of formants and harmonic tracks and consistently results in significant performance gain and improved perceptual quality of speech.
Esfandiar Zavarehei, Saeed Vaseghi
IEEE Trans. Multim.2
2007 Interpolation of lost speech segments using LP-HNM model with codebook-mapping post-processing
abstract
This paper presents a method for interpolation of lost speech segments. The short-time spectral amplitude (STSA) of speech is modeled using a linear prediction (LP) model of the spectral envelop and a harmonic plus noise model (HNM) of the excitation. The restoration algorithm is based on interpolation of the parameters of LP-HNM models of speech from both side of the gap. A codebook mapping (CBM) technique is used to fit the interpolated parameters to a pre-trained speech model. Experiments show that the CBM module mitigates the artifacts that may result from interpolation of relatively long speech gaps. Evaluations demonstrate that the proposed interpolation method results in a superior quality in comparison to alternative restoration methods.
Esfandiar Zavarehei, Saeed Vaseghi
ASRU2
2007 Visually-Derived Wiener Filters for Speech Enhancement
abstract
This work begins by examining the correlation between audio and visual speech features and reveals higher correlation to exist within individual phoneme sounds rather than globally across all speech. Utilising this correlation, a visually-derived Wiener filter is proposed in which clean power spectrum estimates are obtained from visual speech features. Two methods of extracting clean power spectrum estimates are made; first from a global estimate using a single Gaussian mixture model (GMM), and second from phoneme-specific estimates using a hidden Markov model (HMM)-GMM structure. Measurement of estimation accuracy reveals that the phoneme-specific (HMM-GMM) system leads to lower estimation errors than the global (GMM) system. Finally, the effectiveness of visually-derived Wiener filtering is examined.
Ibrahim Almajai, Ben P. Milner, Jonathan Darch, Saeed Vaseghi
ICASSP (4)4
2007 An Investigation into the Correlation and Prediction of Acoustic Speech Features from MFCC Vectors
abstract
This work develops a statistical framework to predict acoustic features (fundamental frequency, formant frequencies and voicing) from MFCC vectors. An analysis of correlation between acoustic features and MFCCs is made both globally across all speech and within phoneme classes, and also from speaker-independent and speaker-dependent speech. This leads to the development of both a global prediction method, using a Gaussian mixture model (GMM) to model the joint density of acoustic features and MFCCs, and a phoneme-specific prediction method using a combined hidden Markov model (HMM)-GMM. Prediction accuracy measurements show the phoneme-dependent HMM-GMM system to be more accurate which agrees with the correlation analysis. Results also show prediction to be more accurate from speaker-dependent speech which also agrees with the correlation analysis.
Jonathan Darch, Ben P. Milner, Ibrahim Almajai, Saeed Vaseghi
ICASSP (4)4
2007 Formant tracking linear prediction model using HMMs and Kalman filters for noisy speech processing
Qin Yan, Saeed Vaseghi, Esfandiar Zavarehei, Ben P. Milner, Jonathan Darch, Paul R. White, Ioannis Andrianakis
Comput. Speech Lang.2
2007 Analysis and Synthesis of Formant Spaces of British, Australian, and American Accents
abstract
In this paper, the probability distribution functions (pdf's) of the formant spaces of three major accents of the English language, namely, British Received Pronunciation (RP), General American, and Broad Australian, are modeled and compared. The statistical differences across the formant spaces of these accents are employed for accent conversion. An improved formant tracking method, based on linear prediction (LP) feature analysis and a two-dimensional hidden Markov model (2-D-HMM) of format trajectories, is used for estimation of the formant trajectories of vowels and diphthongs of each accent. Comparative analysis of the formant spaces of the three accents indicates that these accents are partly conveyed by the differences of the formants of vowels. The estimates of the probability distributions of the formants for each accent are used in a speech synthesis system for accent conversion. Accent synthesis, through modification of the acoustic parameters of speech, provides a means of assessing the perceptual contribution of each formant parameter on conveying an accent. The results of perceptual evaluations of accent conversion illustrate that formants play an important role in conveying accents
Qin Yan, Saeed Vaseghi, Dimitrios Rentzos, Ching-Hsiang Ho
IEEE Trans. Speech Audio Process.2
2007 Noisy Speech Enhancement Using Harmonic-Noise Model and Codebook-Based Post-Processing
abstract
This paper presents a post-processing speech restoration module for enhancing the performance of conventional speech enhancement methods. The restoration module aims to retrieve parts of speech spectrum that may be lost to noise or suppressed when using conventional speech enhancement methods. The proposed restoration method utilizes a harmonic plus noise model (HNM) of speech to retrieve damaged speech structure. A modified HNM of speech is proposed where, instead of the conventional binary labeling of the signal in each subband as voiced or unvoiced, the concept of harmonicity is introduced which is more adaptable to the codebook mapping method used in the later stage of enhancement. To restore the lost or suppressed information, an HNM codebook mapping technique is proposed. The HNM codebook is trained on speaker-independent speech data. To reduce the sensitivity of the HNM codebook to speaker variability, a spectral energy normalization process is introduced. The proposed post-processing method is tested as an add-on module with several popular noise reduction methods. Evaluations of the performance gain obtained from the proposed post-processing are presented and compared to standard speech enhancement systems which show substantial improvement gains in perceptual quality
Esfandiar Zavarehei, Saeed Vaseghi, Qin Yan
IEEE Trans. Speech Audio Process.2
2006 Speech Bandwidth Extension: Extrapolations of Spectral Envelop and Harmonicity Quality of Excitation
abstract
This paper presents a method for restoration of the missing bandwidth of narrowband speech signals. Speech is decomposed into a linear prediction (LP) model of the spectral envelop and a harmonic plus noise model (HNM) of speech excitation. The LP spectral envelope and HNM excitation parameters of the narrowband speech are extrapolated using codebooks trained on narrowband and wideband speech. A novel contribution of this paper is the introduction of a parametric measure of the harmonicity of excitation in harmonically-spaced sub-bands. The wideband LSF parameters and the degree of harmonicity of missing excitation are estimated from those of the narrowband speech via codebook mapping. The method is successful in restoring the harmonicity of speech and converts telephone quality speech to perceptually high quality wideband speech.
Saeed Vaseghi, Esfandiar Zavarehei, Qin Yan
ICASSP (3)1
2006 Temporal Modelling and Kalman Filtering of DFT Trajectories for Enhancement of Noisy Speech
abstract
This paper presents a time-frequency estimator for enhancement of noisy speech in the DFT domain. The time-varying trajectories of the DFT of speech and noise in each channel are modeled by low order autoregressive processes incorporated in the state equation of Kalman filters. The parameters of the Kalman filters are estimated recursively from the signal and noise in DFT channels. The issue of convergence of the Kalman filters to noise statistics during the noise-dominated periods is addressed and a method is incorporated for restarting of Kalman filters after long periods of noise-dominated activity in each DFT channel. The performance of the proposed method is compared with cases where the noise trajectories are not explicitly modeled. Evaluations show that the proposed method results in substantial improvement in perceived quality of speech.
Esfandiar Zavarehei, Saeed Vaseghi, Qin Yan
ICASSP (1)2
2006 Comparative analysis of formants of British, american and australian accents
abstract
This paper compares and quantifies the differences between formants of speech across accents. The cross entropy information measure is used to compare the differences between the formants of the vowels of three major English accents namely British, American and Australian. An improved formant estimation method, based on a linear prediction (LP) model feature analysis and a hidden Markov model (HMM) of formants, is employed for estimation of formant trajectories of vowels and diphthongs. Comparative analysis of the formant space of the three accents indicates that these accents are mostly conveyed by the first two formants. The third and fourth formants exhibit some significant differences across accents for only a few phonemes most notably the variants of vowel ‘r’ in the American (rhotic) accent compared to British (non-rhotic accent). The issue of speaker variability versus accent variability is examined by comparing the cross-entropies of speech models trained on different groups of speakers within and across the accents.
Seyed Ghorshi, Saeed Vaseghi, Qin Yan
INTERSPEECH2
2006 Weighted codebook mapping for noisy speech enhancement using harmonic-noise model
abstract
Most noisy speech enhancement methods result in partial suppression and distortion of speech spectrum. At instances when the local signal-to-noise ratio at a frequency band is very low speech partials are often obliterated. In this paper a method for enhancement and restoration of noisy speech based on a harmonic-noise model (HNM) is introduced. A HNM imposes a temporal-spectral structure that may reduce processing artifacts. The restoration process is enhanced through incorporation of a prior HNM of clean speech stored in a pre-trained codebook. The restored speech is a SNRdependent combination of the de-noised observation and the speech obtained from weighted codebook mapping. The additional improvements of speech quality resulting from the proposed method in comparison to conventional and modern speech enhancement systems are evaluated. The results show that the proposed method improves the quality of noisy speech and restores much of the information lost to noise.
Esfandiar Zavarehei, Saeed Vaseghi, Qin Yan
INTERSPEECH2
2006 MAP prediction of formant frequencies and voicing class from MFCC vectors in noise
Jonathan Darch, Ben P. Milner, Saeed Vaseghi
Speech Commun.3
2006 Inter-frame modeling of DFT trajectories of speech and noise for speech enhancement using Kalman filters
Esfandiar Zavarehei, Saeed Vaseghi, Qin Yan
Speech Commun.2
2005 Predicting Formant Frequencies from MFCC Vectors
abstract
This work proposes a novel method of predicting formant frequencies from a stream of mel-frequency cepstral coefficients (MFCC) feature vectors. Prediction is based on modelling the joint density of MFCCs and formant frequencies using a Gaussian mixture model (GMM). Using this GMM and an input MFCC vector, two maximum a posteriori (MAP) prediction methods are developed. The first method predicts formants from the closest, in some sense, cluster to the input MFCC vector, while the second method takes a weighted contribution of formants predicted from all clusters. Experimental results are presented using the ETSI Aurora connected digit database and show that predicted formant frequencies are within 3.2% of reference formant frequencies.
Jonathan Darch, Ben P. Milner, Xu Shao, Saeed Vaseghi, Qin Yan
ICASSP (1)4
2005 Speaker Identification in Unknown Noisy Conditions - A Universal Compensation Approach
abstract
We consider speaker identification involving background noise, assuming no knowledge about the noise characteristics. A new method, namely universal compensation (UC), is studied as a solution to the problem. The UC method is an extension of the missing-feature method, i.e. recognition based only on reliable data but robust to any corruption type, including full corruption that affects all time-frequency components of the speech representation. The UC technique achieves robustness to unknown, full noise corruption through a novel combination of the multi-condition training method and the missing-feature method. The combination of these two strategies makes the new method potentially capable of dealing with arbitrary additive noise - with arbitrary temporal-spectral characteristics - based only on clean speech training data and simulated noise data, without requiring knowledge about the actual noise. The SPIDRE database is used for the evaluation, assuming various corruptions from real-world noise data. The results obtained are encouraging.
Ji Ming, Darryl Stewart, Saeed Vaseghi
ICASSP (1)3
2005 Formant frequency prediction from MFCC vectors in noisy environments
abstract
This paper proposes a method of predicting the formant frequencies of a frame of speech from its mel-frequency cepstral coefficient (MFCC) representation. Prediction is achieved through the creation of a Gaussian mixture model (GMM) which models the joint density of formant frequencies and MFCCs. Using this GMM and an input MFCC vector, a maximum a posteriori (MAP) prediction of the formant frequencies is generated. Formant prediction accuracy is evaluated on both a constrained vocabulary connected digits database and on a 5000 word large vocabulary database. Experiments first examine the accuracy of formant frequency prediction as the number of clusters in the GMM is varied with a best formant frequency prediction error of 3.72% being obtained. Secondly the effect of noise on formant prediction accuracy is examined. A fall in accuracy is observed with reducing signal-to-noise ratios, but by using a GMM matched to the noise conditions formant prediction accuracy is significantly improved.
Jonathan Darch, Ben P. Milner, Saeed Vaseghi
INTERSPEECH3
2005 Formant-tracking linear prediction models for speech processing in noisy environments
abstract
This paper presents a formant-tracking method for estimation of the time-varying trajectories of a linear prediction (LP) model of speech in noise. The main focus of this work is on the modelling of the non-stationary temporal trajectories of the formants of speech for improved LP model estimation in noise. The proposed approach provides a systematic framework for modelling the inter-frame correlation of speech parameters across successive frames, the intra-frame correlations are modelled by LP parameters. The formant-tracking LP model estimation is composed of two stages: (a) a pre-cleaning intra-frame spectral amplitude estimation stage where an initial estimate of the magnitude frequency response of the LP model of clean speech is obtained and (b) an inter-frame signal processing stage where formant classification and Kalman filters are combined to estimate the trajectory of formants. The effects of car and train noise on the observations and estimation of formants tracks are investigated. The average formant tracking errors at different signal to noise ratios (SNRs) are computed. The evaluation results demonstrate that after noise reduction and Kalman filtering the formant tracking errors are significantly reduced. 1.
Qin Yan, Saeed Vaseghi, Esfandiar Zavarehei, Ben P. Milner
INTERSPEECH2
2005 Speech enhancement in temporal DFT trajectories using Kalman filters
Esfandiar Zavarehei, Saeed Vaseghi
INTERSPEECH2
2004 Voice conversion through transformation of spectral and intonation features
abstract
This paper presents a voice conversion method based on transformation of the characteristic features of a source speaker towards a target. Voice characteristic features are grouped into two main categories: (a) the spectral features at formants and (b) the pitch and intonation patterns. Signal modelling and transformation methods for each group of voice features are outlined. The spectral features at formants are modelled using a set of two-dimensional phoneme-dependent HMM. Subband frequency warping is used for spectrum transformation with the subbands centred on the estimates of the formant trajectories. The F0 contour is used for modelling the pitch and intonation patterns of speech. A PSOLA based method is employed for transformation of pitch, intonation patterns and speaking rate. The experiments present illustrations and perceptual evaluations of the results of transformations of the various voice features.
Dimitrios Rentzos, Saeed Vaseghi, Qin Yan, Ching-Hsiang Ho
ICASSP (1)2
2004 Analysis by synthesis of acoustic correlates of British, Australian and American accents
abstract
This paper presents analysis through synthesis of the acoustic correlates of British, Australian and American accents by transforming the correlates individually across the accents. The acoustic correlates of accents are grouped into three main categories: (a) the spectral features at formants, (b) the pitch intonation pattern and (c) duration. The modeling and transformation methods for each group of voice features are outlined. The spectral features at formants are modeled using two-dimensional (2D) phoneme-dependent HMM. Subband frequency warping is used for spectrum transformation where the subbands are centred on estimates of the formant trajectories. The F0 contour is used for modeling the pitch and intonation patterns of speech. A method based on the time domain pitch synchronous overlap and add algorithm (TD-PSOLA) is used for transformation of pitch intonation and duration pattern. Perceptual tests based on mean opinion score (MOS) are conducted to rank the main features of accents. Formants are regarded as the most important features of accents, followed by intonation pattern and duration.
Qin Yan, Saeed Vaseghi, Dimitrios Rentzos, Ching-Hsiang Ho
ICASSP (1)2
2004 Modelling and ranking of differences across formants of british, australian and american accents
abstract
The differences between formants of British, Australian and American English accents are analysed and ranked. An improved formant model based on linear prediction (LP) feature analysis and a two-dimensional(2D) hidden Markov model (HMM) of formants is employed for estimation of the formant frequencies and bandwidths of vowels and diphthongs. Comparative analysis of the formant trajectories, the formant target points and the bandwidth of the spectral resonance at formants of British, Australian and American accents are presented. British vowels and diphthongs have smaller formant bandwidth than Australian. A method for ranking the contribution of different formants in conveying an accent is proposed whereby formants are ranked according to the normalized distances between the formants across accents. The first two formants are considered more sensitive to accents than other formants.
Qin Yan, Saeed Vaseghi, Dimitrios Rentzos, Ching-Hsiang Ho
INTERSPEECH2
2004 A formant tracking LP model for speech processing
Qin Yan, Esfandiar Zavarehei, Saeed Vaseghi, Dimitrios Rentzos
INTERSPEECH3
2003 Evaluation of methods for parameteric formant transformation in voice conversion
abstract
This paper explores methods of estimation and mapping of parametric formant-based models for voice transformation. The main focus is the transformation of the parameters of a model of the vocal tract of a source speaker to a target speaker. The vocal tract parameters are represented with the linear prediction (LP) model coefficients and the associated formant frequencies, bandwidths, intensities and their temporal trajectories. Two methods are explored for vocal tract (formant) mapping. The first method is based on nonuniform frequency warping and the second is based on pole rotation. Both methods transform all parameters of the formants (frequency, bandwidth and intensity). In addition, the factors that affect the selection of the warping ratios for the mapping functions are presented. Experimental evaluation of voice morphing based on parametric models are presented.
Emir Turajlic, Dimitrios Rentzos, Saeed Vaseghi, Ching-Hsiang Ho
ICASSP (1)3
2003 Analysis, modelling and synthesis of formants of British, American and Australian accents
abstract
The formant space of three major English accents namely British, American and Australian are modelled and used for accent conversion. Accent synthesis, through modification of the acoustic parameters of speech, provides a means for assessing the perceptual contribution of each parameter on conveying an accent. An improved method based on a linear prediction (LP) model feature analysis and a 2D hidden Markov model (HMM) is employed for estimation of formant trajectories of vowels and diphthongs. Comparative analysis of the formant space of the three accents indicates that these accents are partly conveyed by the fronting and backing of vowels. It is found that the first formants of the vowels of British and American English accents are higher than those in Australian accent while Australians have higher second formants in vowels compared to Americans and British. The estimates of the distributions of formants for each accent are used in a speech synthesis system for accent conversion. Perceptual evaluations of accent conversion results illustrate that formants, in particular the second formant, play an important role in conveying accents.
Qin Yan, Saeed Vaseghi
ICASSP (1)2
2003 Robust speaker identification using posterior union models
Ji Ming, Darryl Stewart, Philip Hanna 0001, Pat Corr, Francis Jack Smith, Saeed Vaseghi
INTERSPEECH6
2003 Probability models of formant parameters for voice conversion
abstract
This paper explores the estimation and mapping of probability models of formant parameter vectors for voice conversion. The formant parameter vectors consist of the frequency, bandwidth and intensity of resonance at formants. Formant parameters are derived from the coefficients of a linear prediction (LP) model of speech. The formant distributions are modelled with phonemedependent two-dimensional hidden Markov models with state Gaussian mixture densities. The HMMs are subsequently used for re-estimation of the formant trajectories of speech. Two alternative methods are explored for voice morphing. The first is a non-uniform frequency warping method and the second is based on spectral mapping via rotation of the formant vectors of the source towards those of the target. Both methods transform all formant parameters (Frequency, Bandwidth and Intensity). In addition, the factors that affect the selection of the warping ratios for the mapping function are presented. Experimental evaluation of voice morphing examples is presented.
Dimitrios Rentzos, Saeed Vaseghi, Qin Yan, Ching-Hsiang Ho, Emir Turajlic
INTERSPEECH2
2003 Comparative analysis and synthesis of formant trajectories of british and broad australian accents
abstract
Abstract The differences between the formant trajectories of British and broad Australian English accents are analysed and used for accent conversion. An improved formant model based on linear prediction (LP) feature analysis and a 2-D hidden Markov model (HMM) of formants is employed for estimation of the formant trajectories of vowels and diphthongs. Comparative analysis of the formant values, the formant trajectories and the formant target points of British and broad Australian accents are presented. A method for ranking the contribution of formants to accent identity is proposed whereby formants are ranked according to the normalised distances between formants across accents. The first two formants are considered more sensitive to accents than other formants. Finally a set of experiments on accent conversion is presented to transform the broad Australian accent of a speaker to British Received Pronunciation (RP) accent by formant mapping and prosody modification. Perceptual evaluations of accent conversion results illustrate that besides prosodic correlates such as pitch and duration, formants also play an important role in conveying accents.
Qin Yan, Saeed Vaseghi, Ching-Hsiang Ho, Dimitrios Rentzos, Emir Turajlic
INTERSPEECH2
2002 Synthesis of unseen context and spectral and pitch contour smoothing in concatenated text to speech synthesis
abstract
The availability and perceptual clarity of speech units, and how these units are put together during synthesis have always been the cornerstones of any high quality concatenative text-to-speech synthesis (TTS) system. The speech units are usually obtained from different sentences and contexts in a speaker-dependent speech database. One of the problems with speech units obtained this way is the occurrence of unseen contexts. Here, unseen contexts denote phonological sequences that are not acoustically represented in the selection pool during synthesis. Unseen units are expected in any concatenative TTS system because it is difficult to obtain an: acoustic representation of all possible existing contexts that could occur in speech. This paper proposes a pitch synchronous, overlap and merge method to synthesise the acoustic representation of unseen contexts from existing similar units found in the inventory. It also gives a brief description of spectral and pitch contour smoothing across concatenated units.
Phuay Hui Low, Saeed Vaseghi
ICASSP2
2002 A comparative analysis of UK and US English accents in recognition and synthesis
abstract
In this paper, we present a comparative study of the acoustic speech features of two major English accents: British English and American English. Experiments examined the deterioration in speech recognition resulting from the mismatch between English accents of the input speech and the speech models. Mismatch in accents can increase the error rates by more than 100%. Hence a detailed study of the acoustic correlates of accent using intonation pattern and pitch characteristics was performed. Accents differences are acoustic manifestations of differences in duration, pitch and intonation pattern and of course the differences in phonetic transcriptions. Particularly, British speakers possess much steeper pitch rise and fall pattern and lower average pitch in most of vowels. Finally a possible means to convert English accents is suggested based on above analysis.
Qin Yan, Saeed Vaseghi
ICASSP2
2002 Formant model estimation and transformation for voice morphing
Ching-Hsiang Ho, Dimitrios Rentzos, Saeed Vaseghi
INTERSPEECH3
2002 Application of microprosody models in text to speech synthesis
Phuay Hui Low, Saeed Vaseghi
INTERSPEECH2
2000 State based sub-band LP Wiener filters for speech enhancement in car environments
abstract
The performance of Wiener filters in restoring the quality and intelligibility of noisy speech depends on: (i) the accuracy of the estimates of the power spectra or the correlation values of the noise and the speech processes, and (ii) on the Wiener filter structure. In this paper a Bayesian method is proposed where model combination and model decomposition are employed for the estimation of parameters required to implement subband LP Wiener filters. The use of subband LP Wiener filters provides advantages in terms of improved parameter estimates and also in restoring the temporal-spectral composition of speech. The method is evaluated, and compared with the parallel model combination, using the TIMIT continuous speech database with BMW and VOLVO car noise databases.
Aimin Chen, Saeed Vaseghi, Paul M. McCourt
ICASSP2
2000 Full covariance modelling and adaptation in sub-bands
abstract
With regard to the current interest in sub-band based modelling in the ASR community, this paper explores the gains in recognition performance and complexity reduction achieved by sub-band based full covariance modelling and speaker adaptation. With sub-band features, instead of a single large covariance matrix, it is now possible to have a set of smaller matrices making it practical to use Gaussian distributions employing full covariance matrices. This benefit is further demonstrated to give a significant complexity reduction in the implementation of speaker adaptation by maximum likelihood linear regression. The use of sub-band cepstra moreover presents the opportunity of capturing localised discriminative cues which contribute to increased recognition. In light of these gains, this paper explores the advantages of sub-band full covariance modelling and presents experimental evaluation on the WSJCAMO continuous speech database.
Bernard Doherty, Saeed Vaseghi, Paul M. McCourt
ICASSP2
2000 State based sub-band Wiener filters for speech enhancement in car environments
Aimin Chen, Saeed Vaseghi
INTERSPEECH2
2000 Multi-resolution sub-band features and models for HMM-based phonetic modelling
Paul M. McCourt, Saeed Vaseghi, Bernard Doherty
Comput. Speech Lang.2
1999 Discriminative spectral-temporal multiresolution features for speech recognition
abstract
Multi-resolution features, which are based on the premise that there may be more cues for phonetic discrimination in a given sub-band than in another, have been shown to outperform the standard MFCC feature set for both classification and recognition tasks on the TIMIT database. This paper presents an investigation into possible strategies to extend these ideas from the spectral domain into both the spectral and temporal domains. Experimental work on the integration of segmental models, which are better at capturing the longer term phonetic correlation of a phonetic unit, into the discriminative multi-resolution framework is presented. Results are presented which show that including this supplementary temporal information offers an improvement performance for the phoneme classification task over the standard multi-resolution MFCC feature set with time derivatives appended. Possible strategies for the extension of theses techniques into the area of continuous speech recognition are discussed.
Philip McMahon, Naomi Harte, Saeed Vaseghi, Paul M. McCourt
ICASSP3
1999 Decision tree micro-prosody structures for text to speech synthesis
Aimin Chen, Shu Lian Wong, Saeed Vaseghi, Charles Ho
EUROSPEECH3
1999 Linear transformations in sub-band groups for speech recognition
abstract
Linear transforms have been demonstrated to successfully achieve on-line speaker and environmental adaptation for robust recognition. This paper explores the gains in computational speed, speaker adaptation convergence rate and recognition performance obtained through the use of multi-resolution sub-band linear transforms in speech recognition. A useful feature of multiresolution processing is that significant savings can be attained as regards transform calculation. In this paper we appraise the relative merits of multiband processing over that of full-band and present evaluation results on the WSJCAM0 continuous speech database.
Bernard Doherty, Saeed Vaseghi, Paul M. McCourt
EUROSPEECH2
1999 Voice conversion between UK and US accented English
abstract
The performance of voice dialling systems often degrades rapidly as the intensity of the background noise increases. In this paper, we describe a neural network based speech enhancement technique for improving the speech recognition performance of a voice dialling system in very noisy real world type conditions. The speech samples were recorded in laboratory conditions and afterwards corrupted by adding car noise or babble noise recorded in a cafe. These noise corrupted speech samples were enhanced in cepstral domain by a context dependent multilayer perceptron (MLP) network before performing the recognition using a hidden Markov model (HMM) based speech recognition system. The accuracy of the test set increased 58%, 55% and 46% in the car noise environments having -5 dB, 0 dB and 5 dB SNRs, respectively. The accuracy of the test set increased 44%, 48% and 39% in the babble noise environments having SNR 5 dB, 10 dB and 15 dB, respectively. The accuracy remained approximately same for both car and babble noise environments when having SNR of 20 dB.
Ching-Hsiang Ho, Saeed Vaseghi, Aimin Chen
EUROSPEECH2
1999 Combined temporal and spectral multi-resolution phonetic modelling
Paul M. McCourt, Naomi Harte, Saeed Vaseghi
EUROSPEECH3
1998 MIMIC : a voice-adaptive phonetic-tree speech synthesiser
Aimin Chen, Saeed Vaseghi, Charles Ho
ICSLP2
1998 Context dependent tree based transforms for phonetic speech recognition
Bernard Doherty, Saeed Vaseghi, Paul M. McCourt
ICSLP2
1998 Joint recognition and segmentation using phonetically derived features and a hybrid phoneme model
Naomi Harte, Saeed Vaseghi, Ben P. Milner
ICSLP2
1998 Bayesian constrained frequency warping HMMS for speaker normalisation
Ching-Hsiang Ho, Saeed Vaseghi, Aimin Chen
ICSLP2
1998 Discriminative weighting of multi-resolution sub-band cepstral features for speech recognition
abstract
This paper explores possible strategies for the recombination of independent multi-resolution sub-band based recognisers. The multi-resolution approach is based on the premise that additional cues for phonetic discrimination may exist in the spectral correlates of a particular sub-band, but not in another. Weights are derived via discriminative training using the ‘Minimum Classification Error’ (MCE) criterion on loglikelihood scores. Using this criterion the weights for correct and competing classes are adjusted in opposite directions, thus conveying the sense of enforcing separation of confusable classes. Discriminative re-combination is shown to provide significant increases for both phone classification and continuous recognition tasks on the TIMIT database. Weighted recombination of independent multi-resolution subband models is also shown to provide robustness improvements in broadband noise.
Philip McMahon, Paul M. McCourt, Saeed Vaseghi
ICSLP3
1998 Capturing discriminative information using multiple modeling techniques
Ji Ming, Philip Hanna 0001, Darryl Stewart, Saeed Vaseghi, Francis Jack Smith
ICSLP4
1998 Multi-phone strings as subword units for speech recognition
Philip O'Neill, Saeed Vaseghi, Bernard Doherty, Wooi-Haw Tan, Paul M. McCourt
ICSLP2
1997 Multi-resolution phonetic/segmental features and models for HMM-based speech recognition
abstract
This paper explores the modelling of phonetic segments of speech with multi-resolution spectral/time correlates. For spectral representation a set of multi-resolution cepstral features are proposed. Cepstral features obtained from a DCT of the log energy-spectrum over the full voice-bandwidth (100-4000 Hz) are combined with higher resolution features obtained from the DCT of upper subband (say 100-2100) and lower subband (2100-4000) halves. This approach can be extended to several levels of different resolutions. For representation of the temporal structure of speech segments or phonetic units, the conventional cepstral and dynamic cepstral features representing speech at the sub-phonetic levels, are supplemented by a set of phonetic features that describe the trajectory of speech over the duration of a phonetic unit. A conditional probability model for phonetic and sub-phonetic features is considered. Experiments demonstrate that the inclusion of the segmental features result in about 10% decrease in error rates.
Saeed Vaseghi, Naomi Harte, Ben P. Milner
ICASSP1
1997 Noise compensation methods for hidden Markov model speech recognition in adverse environments
abstract
Several noise compensation schemes for speech recognition in impulsive and nonimpulsive noise are considered. The noise compensation schemes are spectral subtraction, HMM-based Wiener (1949) filters, noise-adaptive HMMs, and a front-end impulsive noise removal. The use of the cepstral-time matrix as an improved speech feature set is explored, and the noise compensation methods are extended for use with cepstral-time features. Experimental evaluations, on a spoken digit database, in the presence of ear noise, helicopter noise, and impulsive noise, demonstrate that the noise compensation methods achieve substantial improvement in recognition across a wide range of signal-to-noise ratios. The results also show that the cepstral-time matrix is more robust than a vector of identical size, which is composed of a combination of cepstral and differential cepstral features.
Saeed Vaseghi, Ben P. Milner
IEEE Trans. Speech Audio Process.1
1996 Dynamic features for segmental speech recognition
Naomi Harte, Saeed Vaseghi, Ben P. Milner
ICSLP2
1996 A comparitive analysis of channel-robust features and channel equalization methods for speech recognition
Saeed Vaseghi, Ben P. Milner
ICSLP1
1995 Speech recognition in impulsive noise
abstract
This paper presents experimental results on the use of noise compensation schemes with hidden Markov model (HMM) speech recognition systems operating in the presence of impulsive noise. A measure of signal to impulsive noise ratio is introduced, and the effects of varying the percentage of impulsive noise contamination, and the power of impulsive noise, on speech recognition are investigated. For the modelling of an impulsive noise process, an amplitude-modulated binary sequence model and a binary-state HMM are considered. For impulsive noise compensation a front-end method and a noise-adaptive method are evaluated. Experiments demonstrate that the noise compensation methods achieve a substantial improvement in speech recognition accuracy across a wide range of signal to impulsive noise ratios.
Saeed Vaseghi, Ben P. Milner
ICASSP1
1995 An analysis of cepstral-time matrices for noise and channel robust speech recognition
abstract
This paper presents an analysis of the cepstral-time matrix. The coefficients of the cepstral-time matrix are found to be similar to the standard cepstral vector with differential features augmented on. It is also shown that the cepstral-time matrix is inherently robust to convolutional channel distortion. Spectral-subtraction, Wiener filtering and model combination are extended into two-dimensions where improved noise robustness is achieved. Experimental results using the NOISEX database with noise and channel distorted speech are presented.,
Ben P. Milner, Saeed Vaseghi
EUROSPEECH2
1994 Speech modelling using cepstral-time feature matrices and hidden Markov models
abstract
Conventional HMMs assume that speech spectral vectors are uncorrelated. The use of information on the temporal evolution of spectral features, within each state, can improve recognition accuracy and produce a more robust recognition system. The authors present experimental results on improvements in speech recognition using cepstral-time matrix units. Experimental evaluation using a spoken digit data base and a spoken alphabet data base, indicates that the use of cepstral-time matrix features in noisy conditions can provide an improvement in recognition of as much as 20% in comparison to a conventional spectral vector comprising of cepstral, delta cepstral and delta-delta cepstral features.>
Ben P. Milner, Saeed Vaseghi
ICASSP (1)2
1994 Noisy speech recognition using cepstral-time features and spectral-time filters
abstract
This paper explores the advantages of using cepstral-time feature matrices, and spectral-time filters, for noisy speech recognition within a hidden Markov model framework. The use of cepstral-time features with spectral subtraction and state-based time-varying Wiener filters is investigated. Experimental results indicate that cepstral-time features, and spectral-time noise processing, provide an effective framework for robust speech recognition in noisy environments.>
Saeed Vaseghi, Ben P. Milner, Jason J. Humphries
ICASSP (2)1
1993 Noisy speech recognition based on HMMs, Wiener filters and re-evaluation of most likely candidates
Saeed Vaseghi, Ben P. Milner
ICASSP (2)1
1993 Speech modelling using cepstral-time feature matrices
Saeed Vaseghi, P. N. Conner, Ben P. Milner
EUROSPEECH1
1993 Noise-adaptive hidden Markov models based on wiener filters
Saeed Vaseghi, Ben P. Milner
EUROSPEECH1
1992 On increasing structural complexity of finite state speech models
abstract
Some methods of modeling speech spectral features and duration within the framework of finite-state models are discussed. On observation modeling, the use of cepstral-time matrices, instead of cepstral vectors, as the observation unit is investigated. On duration modeling, a new HMM is introduced in which state transition and duration probabilities are combined to form duration-dependent transition probabilities. The duration dependent transitions are derived from the cumulative density function (CDF) of state duration.>
Saeed Vaseghi, P. N. Conner
ICASSP1
1992 Speech recognition in noisy environments
Saeed Vaseghi, Ben P. Milner
ICSLP1
1990 Finite state CELP for variable rate speech coding
abstract
A finite-state code excited linear prediction (CELP) system is proposed for variable-rate speech coding. The encoding system consists of a number of CELP coders with different linear predictive coding parameter quantization patterns, code book sizes, and population densities. The selection of the encoding state for each input vector depends on the input signal characteristics, the desired bit rate/signal to quantized noise ratio, and the current state of the encoder. In CELP coders the greater part of the bit resources, about 70%, is used for encoding of the excitation signal. However, the excitation accuracy needed to encode a speech segment with a desired level of fidelity strongly depends on its short-term spectral characteristics. The use of a finite-state system involves some implicit clustering of speech and allows variable-rate coding of the excitation. The gain in compression that can be obtained from variable-rate coding of the excitation signal is investigated. Experiments with a four-state variable-rate CELP coder produce good-quality encoded speech at 5 kbit/s.>
Saeed Vaseghi
ICASSP1
1989 The effects of non-stationary signal characteristics on the performance of adaptive audio restoration systems
abstract
The degrading effects of nonstationary and abruptly changing signal characteristics on the performance of a particular class of restoration systems, namely those which use two sources of information, are discussed. The system considered is the adaptive noise suppression system. In this system abrupt changes in signal statistics result in transient degradations in the form of excessive suppression of some signal frequency components and or inadequate suppression of noise frequency components. Abrupt changes are detected using an appropriate distance between two autoregressive models of the signal. After detection of an abrupt change the adaptation gain is increased and/or the filter parameters are reinitialized in order to minimize the degradation.>
Saeed Vaseghi, Peter J. W. Rayner
ICASSP1