Wai-Yip Chan

dblp:45/4855 · DBLP profile ↗
← Back
75ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-5322-2449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 22 · 5 since 2021Computer networks · 8 · 2 first-authorSystems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 TSIP-Net: No-reference speech intelligibility prediction in the presence of competing speech
Haolan Wang, Wai-Yip Chan, Jesper Jensen 0001
Speech Commun.2
2025 Intelligibility Prediction for Time-Modified Speech Signals Using Spectro-Temporal Modulation Features
abstract
Reference-based speech intelligibility prediction algorithms (RB-SIPAs) are limited to speech degradations that maintain the time alignment between the clean and degraded speech signals. We address this limitation by augmenting the existing RB-SIPAs framework with time alignment. To achieve robust time alignment at low signal-to-noise ratios, we propose using spectro-temporal modulations (STM) of the speech signals for dynamic time warping (DTW). Our experiments demonstrate that in DTW, the use of specific STM components as features for clean and time-modified degraded speech achieves the baseline time alignment obtained using clean speech and noise-free time-modified speech. Moreover, we propose two methods to incorporate time alignment into existing RB-SIPAs. Using these methods, we compare the output scores of the RB-SIPA with listening test scores and show better correlation results using the STM features as compared to MFCCs.
Aymen Bashir, Haolan Wang, Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001
INTERSPEECH4
2025 Analysis and Extension of a Near-End Listening Enhancement Method Based on Long-Term Fractile Noise Statistics
abstract
This paper addresses the problem of near-end listening enhancement (NELE), where a clean speech signal is modified prior to playback and under an energy constraint to improve intelligibility in noise. We analyze a recently proposed NELE method, optimized using a Speech Intelligibility Index that has been modified to incorporate temporal aspects of the noise via long-term fractile noise statistics. Specifically, we explain the energy allocation strategy adopted by the algorithm, and show that, in contrast to many existing methods, the spectral energy distribution of the modified speech is a function of that of the background noise, but not that of the input speech. Our simulation experiments show that this simple method outperforms well-established spectral shaping NELE methods. In addition, we extend the algorithm by appending an off-the-shelf dynamic range compressor, and show that it performs generally better than state-of-the-art methods for NELE.
Filippo Villani, Wai-Yip Chan, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001
INTERSPEECH2
2024 Speaker Adaptation For Enhancement Of Bone-Conducted Speech
abstract
Deep neural network (DNN)-based speech enhancement models often face challenges in maintaining their performance for speakers not encountered during training. This challenge is exacerbated in applications such as enhancement and bandwidth extension of bone-conducted speech, where the distortion exhibits a close correlation with speaker-specific characteristics. We address this issue by introducing a bottleneck module aimed at disentangling speaker-specific characteristics from speech content in speech enhancement DNNs. A DNN model is trained for enhancement of bone-conducted speech and modified with the proposed bottleneck module. We evaluate the DNN’s adaptability to unseen speakers through fine-tuning the network with a limited amount of adaptation data. The results show that the proposed bottleneck module can enhance adaptation performance to new unseen speakers, especially when limited amount of speaker-specific adaptation data is available.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
ICASSP2
2024 No-Reference Speech Intelligibility Prediction Leveraging a Noisy-Speech ASR Pre-Trained Model
abstract
Recent advances in deep learning have improved the capabilities of data-driven speech intelligibility prediction (SIP) algorithms. Nevertheless, the scarcity of speech intelligibility datasets limits the development of data-driven algorithms. This study introduces a set of no-reference SIP algorithms leveraging a pre-trained wav2vec 2.0 backbone. We adapt wav2vec 2.0 for automatic speech recognition under additive noise conditions with a parameter-efficient methodology, low-rank adaptation. We demonstrate no-reference SIP algorithms designed with this approach using a moderate amount of training data. The best designs perform on par or even better than a state-of-the-art reference-based SIP algorithm across a variety of datasets comprising different degradation types.
Haolan Wang, Amin Edraki, Wai-Yip Chan, Iván López-Espejo, Jesper Jensen 0001
INTERSPEECH3
2023 On the deficiency of intelligibility metrics as proxies for subjective intelligibility
abstract
A recent trend in deep neural network (DNN)-based speech enhancement consists of using intelligibility and quality metrics as loss functions for model training with the aim of achieving high subjective speech intelligibility and perceptual quality in real-life conditions. In this study, we analyze a variety of loss functions, including some based on state-of-the-art intelligibility and quality metrics, to train an end-to-end speech enhancement system based on a fully convolutional neural network. The loss functions include perceptual metric for speech quality evaluation (PMSQE), scale-invariant signal-to-distortion ratio (SI-SDR), SI-SDR integrating speech pre-emphasis, short-time objective intelligibility (STOI), extended STOI (ESTOI), spectro-temporal glimpsing index (STGI), and a composite loss function combining STGI and SI-SDR. While DNNs trained with these loss functions produce notable speech intelligibility (and quality) gains according to pertinent objective metrics, we conduct a subjective intelligibility test that contradicts this result, showing no intelligibility improvement. From the results of this study, our conclusion is twofold: (1) subjective intelligibility evaluation is currently not replaceable by objective intelligibility evaluation, and (2) both the development of meaningful intelligibility metrics and DNN-based speech enhancement systems that can consistently improve the intelligibility of noisy speech for human listening remain open problems.
Iván López-Espejo, Amin Edraki, Wai-Yip Chan, Zheng-Hua Tan, Jesper Jensen 0001
Speech Commun.3
2021 A Spectro-Temporal Glimpsing Index (STGI) for Speech Intelligibility Prediction
abstract
We propose a monaural intrusive speech intelligibility prediction (SIP) algorithm called STGI based on detecting glimpses in short-time segments in a spectro-temporal modulation decomposition of the input speech signals. Unlike existing glimpse-based SIP methods, the application of STGI is not limited to additive uncorrelated noise; STGI can be employed in a broad range of degradation conditions. Our results show that STGI performs consistently well across 15 datasets covering degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, checkerboard noise, and gated noise.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
Interspeech2
2021 Speech Intelligibility Prediction Using Spectro-Temporal Modulation Analysis
abstract
Spectro-temporal modulations are believed to mediate the analysis of speech sounds in the human primary auditory cortex. Inspired by humans' robustness in comprehending speech in challenging acoustic environments, we propose an intrusive speech intelligibility prediction (SIP) algorithm, wSTMI, for normal-hearing listeners based on spectro-temporal modulation analysis (STMA) of the clean and degraded speech signals. In the STMA, each of 55 modulation frequency channels contributes an intermediate intelligibility measure. A sparse linear model with parameters optimized using Lasso regression results in combining the intermediate measures of 8 of the most salient channels for SIP. In comparison with a suite of 10 SIP algorithms, wSTMI performs consistently well across 13 datasets, which together cover degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, and speech interruption. We show that the optimized parameters of wSTMI may be interpreted in terms of modulation transfer functions of the human auditory system. Thus, the proposed approach offers evidence affirming previous studies of the perceptual characteristics underlying speech signal intelligibility.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Improvement and Assessment of Spectro-Temporal Modulation Analysis for Speech Intelligibility Estimation
abstract
Several recent high-performing intelligibility estimators of acoustically degraded speech signals employ temporal modulation analysis. In this paper, we investigate the utility of using both spectro- and temporal-modulation for estimating speech intelligibility. We modified a pre-existing speech intelligibility estimation scheme (STMI) that was inspired by human auditory spectro-temporal modulation analysis. We produced several variants of the modified STMI and assessed their intelligibility prediction accuracy, in comparison with several high-performing estimators. Among the estimators tested, one of the STMI variants and eSTOI performed consistently well on both noisy and reverberated speech. These results suggest that spectro-temporal modulation analysis is useful for certain degradation conditions such as modulated noise and reverberation.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
INTERSPEECH2
2016 Source-Interference Recovery Over Broadcast Channels: Asymptotic Bounds and Analog Codes
abstract
We consider the problem of joint recovery of a bivariate Gaussian source and of interference over the two-user Gaussian degraded broadcast channel in the presence of common interference. The interference, that is available non-causally at the encoder, is assumed to be Gaussian and correlated to the sources. The tradeoff between the distortion of the sources and the interference estimation error is studied; information-theoretic outer and inner bounds based on ideas from rate-distortion theory and hybrid coding are derived, respectively. More precisely, the outer bound is found by assuming additional knowledge at each user; the inner bound, however, is obtained by analyzing the distortion of a layered hybrid scheme based on proper power splitting, Costa and Wyner-Ziv coding. Low delay and complexity coding schemes based on analog mapping are next proposed. More specifically, parametric mappings based on linear and sawtooth curves are studied and optimized by minimizing an upper bound on the system's distortion; nonparametric mappings based on joint optimization between the encoder and the decoder using an iterative algorithm are designed. Numerical results show that for the special cases that are previously considered by Abou Saleh et al. (with no fading), the derived outer bound is tighter and the proposed hybrid scheme has a lower complex structure with no loss in performance. In addition, the proposed low delay nonlinear schemes outperform the linear scheme and perform relatively close to the inner bound under certain system settings.
Ahmad Abou Saleh, Fady Alajaji, Wai-Yip Chan
IEEE Trans. Commun.3
2015 Compressed Sensing with Non-Gaussian Noise and Partial Support Information
abstract
We study the problem of recovering sparse and compressible signals using a weighted${\ell _p}$minimization with$0 < p \leq 1$from noisy compressed sensing measurements when part of the support is known a priori. To better model different types of non-Gaussian (bounded) noise, the minimization program is subject to a data-fidelity constraint expressed as the${\ell _q}(2 \leq q < \infty)$norm of the residual error. We show theoretically that the reconstruction error of this optimization is bounded (stable) if the sensing matrix satisfies an extended restricted isometry property. Numerical results show that the proposed method, which extends the range of$p$and$q$comparing with previous works, outperforms other noise-aware basis pursuit programs. For$p < 1$, since the optimization is not convex, we use a variant of an iterative reweighted${\ell _2}$algorithm for computing a local minimum.
Ahmad Abou Saleh, Fady Alajaji, Wai-Yip Chan
IEEE Signal Process. Lett.3
2014 Source-Channel Coding for Fading Channels With Correlated Interference
abstract
We consider the problem of sending a Gaussian source over a fading channel with Gaussian interference known to the transmitter. We study joint source-channel coding schemes for the case of unequal bandwidth between the source and the channel and when the source and the interference are correlated. An outer bound on the system's distortion is first derived by assuming additional information at the decoder side. We then propose layered coding schemes based on proper combination of power splitting, bandwidth splitting, Wyner-Ziv and hybrid coding. More precisely, a hybrid layer, that uses the source and the interference, is concatenated (superimposed) with a purely digital layer to achieve bandwidth expansion (reduction). The achievable (square error) distortion region of these schemes under matched and mismatched noise levels is then analyzed. Numerical results show that the proposed schemes perform close to the best derived bound and to be resilient to channel noise mismatch. As an application of the proposed schemes, we derive both inner and outer bounds on the source-channel-state distortion region for the fading channel with correlated interference; the receiver, in this case, aims to jointly estimate both the source signal as well as the channel-state (interference).
Ahmad Abou Saleh, Wai-Yip Chan, Fady Alajaji
IEEE Trans. Commun.2
2013 Late reverberation suppression using MMSE modulation spectral estimation
Chenxi Zheng, Wai-Yip Chan
INTERSPEECH2
2013 Hybrid digital-analog coding for interference broadcast channels
abstract
We consider the transmission of bivariate Gaussian sources (V1, V2) over the two-user Gaussian broadcast channel in the presence of interference that is correlated to the source and known to the transmitter. Each user i is interested in estimating Vi. We study hybrid digital-analog (HDA) schemes and analyze the achievable (square-error) distortion region under matched and expansion bandwidth regimes. These schemes require proper combinations of power splitting, bandwidth splitting, rate splitting, Wyner-Ziv and HDA Costa coding. An outer bound on the distortion region is also derived by assuming knowledge of V1at the second user and full/partial knowledge of the interference at both users. Numerical results show that the HDA schemes outperform tandem and linear schemes and perform close to the derived bound for certain system settings.
Ahmad Abou Saleh, Fady Alajaji, Wai-Yip Chan
ISIT3
2012 Multi-scalable video multicast for heterogeneous playback requirements using a perceptual utility measure
abstract
A novel best-effort utility maximization scheme for video multicast with heterogeneous clients is proposed. Clients are assumed to have heterogeneous media playback requirements and channels with different capacities. The multicast employs H.264/SVC coded video which permits combined temporal, spatial, and amplitude scalability. The perceptual effects of the scalable video on clients with different terminal capabilities are modeled, and utilized in the optimization. For various scenarios with clients of different media adaptation capabilities, the proposed optimization of multi-scalable video multicasting reveals notable utility gains over single-scalable video multicasting. With a complexity independent of the number of clients, the proposed optimization scheme is suitable for large-scale multimedia multicast.
Ali Bakhshali, Wai-Yip Chan, Steven D. Blostein
MMSP2
2012 Characterization of atypical vocal source excitation, temporal dynamics and prosody for objective measurement of dysarthric word intelligibility
Tiago H. Falk, Wai-Yip Chan, Fraser Shein
Speech Commun.2
2012 Compressed Sensing With Nonlinear Analog Mapping in a Noisy Environment
abstract
We propose a low delay and low complexity sensor system based on the combination of Shannon–Kotel'nikov mapping and compressed sensing (CS). The proposed system uses nonlinear analog mappings on the CS measurements to increase their immunity against channel noise. Numerical results show that the proposed purely-analog system outperforms the state-of-the-art purely CS systems in terms of signal-to-distortion ratio. In addition to sparsity knowledge, we use a statistical characterization of the observed signal to further improve system performance.
Ahmad Abou Saleh, Wai-Yip Chan, Fady Alajaji
IEEE Signal Process. Lett.2
2011 Quantifying perturbations in temporal dynamics for automated assessment of spastic dysarthric speech intelligibility
abstract
Spastic dysarthric speech is often associated with imprecise placement of articulators which, in turn, cause perturbations in speech temporal dynamics, such as unclear distinctions between adjacent phonemes. While these perturbations can lead to a significant reduction in intelligibility, measures to objectively assess their detrimental effect on intelligibility are lacking. In this paper, short- and long-term temporal dynamics measures are proposed and evaluated as correlates of subjective intelligibility. The former is based on log-energy temporal dynamics information, whereas the latter is based on an auditory-inspired modulation spectral signal representation. A composite measure is also developed based on linearly combining the proposed measures with a tone-unit duration parameter. Experiments with the publicly-available 'Universal Access' database of spastic dysarthric speech show that the proposed composite measure can achieve rank correlations with subjective ratings as high as 0.87, thus providing a tool to automatically diagnose speech disorder severity and to evaluate dysarthria treatment outcomes.
Tiago H. Falk, Richard Hummel, Wai-Yip Chan
ICASSP3
2011 MPtracker: A new multi-pitch detection and separation algorithm for mixed speech signals
abstract
We present MPtracker, a new algorithm for tracking and separating the pitch frequencies of two speakers from their mixture. The pitch frequencies are detected by introducing a novel spectral distortion optimization which takes into account the sinusoidal modeling of the speech signal. The detected pitch frequencies are grouped, separated, and finally an interpolation method is applied to estimate missing pitch frequencies. We evaluated the performance of the proposed technique on 196 mixtures including 48 male-male, 48 female-female, and 96 male-female mixtures with target-to-interference ratios (TIR) ranging from 0 dB to +18 dB. The results show our simple but effective and fast technique significantly outperforms two widely-used approaches.
Mohammad H. Radfar, Richard M. Dansereau, Wai-Yip Chan, Willy Wong
ICASSP3
2011 Spectral Features for Automatic Blind Intelligibility Estimation of Spastic Dysarthric Speech
abstract
In this paper, we explore the use of the standard ITU-T P.563 speech quality estimation algorithm for automatic assessment of dysarthric speech intelligibility. A linear mapping consisting of three salient P.563 internal fea-tures is proposed and shown to accurately estimate spas-tic dysarthric speech intelligibility. Delta-energy features are further proposed in order to characterize the atypi-cal spectral dynamics and limited vowel space observed with spastic dysarthria. Experiments using the publicly-available Universal Access database (10 speaker patients) show that when salient delta-energy and internal P.563 features are used, correlations with subjective intelligi-bility ratings as high as 0.98 can be attained. Index Terms: Dysarthria, intelligibility, P.563, subjec-tive quality, delta-energy
Richard Hummel, Wai-Yip Chan, Tiago H. Falk
INTERSPEECH2
2011 An Assessment of the Improvement Potential of Time-Frequency Masking for Speech Dereverberation
abstract
The effect of ideal time-frequency masking (ITFM) on the intel-ligibility of reverberated speech is tested using objective mea-surement, namely STI and PESQ scores. The best choice of ITFM threshold is determined for a range of reverberation times (RTs). Four existing dereverberation algorithms are also as-sessed. Objective test results and informal subjective listen-ing show that IFTM provides great intelligibility improvement for all RTs and outperforms the existing dereverberation algo-rithms, one of which assumes perfect knowledge of the room impulse response. While ITFM provides only a best possible performance bound, our results demonstrate the potential im-provement that could be obtained using time-frequency mask-ing for speech dereverberation. Index Terms: Dereverberation, speech intelligibility, speech transmission index, speech quality, time-frequency masking
Chenxi Zheng, Tiago H. Falk, Wai-Yip Chan
INTERSPEECH3
2011 Optimization of rateless-coded asynchronous multimedia multicast
abstract
We investigate the optimization of asynchronous multimedia multicast to heterogenous users. A priority encoding transmission (PET) based packetization scheme in combination with rateless coding has recently been proposed for multicasting. In previous work, a suboptimal greedy search algorithm is proposed to find the allocation of source symbols over different layers. However, the complexity of the algorithm is high for large numbers of user classes and the optimality of the algorithm is not proven. In this paper, we first show that the problem can be transformed to an equivalent convex optimization problem under the relaxation of integer constraints. In addition, a solution is found analytically for the optimal allocation. Numerical results demonstrate that the optimal analytical solution matches the results using the greedy search algorithm as well as results obtained using convex optimization software.
Steven D. Blostein, Wai-Yip Chan
PIMRC3
2011 Automatic speech emotion recognition using modulation spectral features
Siqing Wu, Tiago H. Falk, Wai-Yip Chan
Speech Commun.3
2010 Scaled factorial hidden Markov models: A new technique for compensating gain differences in model-based single channel speech separation
abstract
In model-based single channel speech separation, factorial hidden Markov models (FHMM) have been successfully applied to model the mixture signal Y(t) = X(t) + V(t) in terms of trained patterns of the speech signals X(t) and V(t). Nonetheless, when the test signals are scaled versions of the trained patterns (i.e. gxX(t) and gvV(t)), the performance of FHMM degrades significantly. In this paper, we introduce a modification to FHMM, called scaled FHMM, which compensates gain difference. In this technique, first, the scale factors are expressed in terms of the target-to-interference ratio (TIR). Then, an iteration quadratic optimization approach is coupled with FHMM to estimate TIR which with the decoded HMM sequences maximize the likelihood of the mixture signal. Experimental results, conducted on 180 mixtures with TIRs from 0 to 15 dB, show that the proposed technique significantly outperforms unscaled FHMM, and scaled/unscaled vector quantization speech separation techniques.
Mohammad H. Radfar, Willy Wong, Richard M. Dansereau, Wai-Yip Chan
ICASSP4
2010 Asynchronous and Reliable Multimedia Multicast with Heterogeneous QoS Constraints
abstract
We present an asynchronous and reliable multicast framework for a scalable multimedia system. An unequal error protection (UEP) transmission scheme employing a layered packetization structure and rateless codes is proposed. With this scheme, asynchronous data reception and guaranteed quality-of-service (QoS) can be achieved at the same time. Furthermore, a novel allocation algorithm is developed within the layered multicast framework, which can minimize the system cost in terms of the number of transmitted packets. Simulation results show that the proposed rateless-codes-based transmission scheme works well for asynchronous multicasting to heterogeneous receivers, and can be applied to a wide range of multimedia applications.
Wei Sheng, Wai-Yip Chan, Steven D. Blostein
ICC2
2010 Unequal error protection rateless coding design for multimedia multicasting
abstract
We study the design and optimization of unequal error protection (UEP) rateless codes for scalable multimedia multicasting. We formulate two general problems of optimizing UEP rateless code for multimedia multicasting to heterogenous users: one focusing on providing guaranteed quality of service (QoS) and the other focusing on providing best-effort QoS. A random interleaved rateless encoder design is proposed. Unlike previous designs, existing standardized raptor codes can be directly applied to this design without degrading performance. For each problem, optimal layer selection parameters are obtained either analytically or numerically. Numerical results demonstrate that the proposed optimized random interleaved UEP rateless code outperforms non-optimized rateless codes and recently proposed UEP rateless codes.
Steven D. Blostein, Wai-Yip Chan
ISIT3
2010 Modulation Spectral Features for Robust Far-Field Speaker Identification
abstract
In this paper, auditory inspired modulation spectral features are used to improve automatic speaker identification (ASI) performance in the presence of room reverberation. The modulation spectral signal representation is obtained by first filtering the speech signal with a 23-channel gammatone filterbank. An eight-channel modulation filterbank is then applied to the temporal envelope of each gammatone filter output. Features are extracted from modulation frequency bands ranging from 3-15 H z and are shown to be robust to mismatch between training and testing conditions and to increasing reverberation levels. To demonstrate the gains obtained with the proposed features, experiments are performed with clean speech, artificially generated reverberant speech, and reverberant speech recorded in a meeting room. Simulation results show that a Gaussian mixture model based ASI system, trained on the proposed features, consistently outperforms a baseline system trained on mel-frequency cepstral coefficients. For multimicrophone ASI applications, three multichannel score combination and adaptive channel selection techniques are investigated and shown to further improve ASI performance.
Tiago H. Falk, Wai-Yip Chan
IEEE Trans. Speech Audio Process.2
2010 A Non-Intrusive Quality and Intelligibility Measure of Reverberant and Dereverberated Speech
abstract
A modulation spectral representation is investigated for non-intrusive quality and intelligibility measurement of reverberant and dereverberated speech. The representation is obtained by means of an auditory-inspired filterbank analysis of critical-band temporal envelopes of the speech signal. Modulation spectral insights are used to develop an adaptive measure termed speech to reverberation modulation energy ratio. Experimental results show the proposed measure outperforming three standard algorithms for tasks involving estimation of multiple dimensions of perceived coloration, as well as quality measurement and intelligibility estimation of reverberant and dereverberated speech.
Tiago H. Falk, Chenxi Zheng, Wai-Yip Chan
IEEE Trans. Speech Audio Process.3
2010 Multiple description quantizer design for multiple-antenna systems with MAP detection
abstract
We study the design of index assignment based multiple description quantizers for multiple-antenna systems with slow Rayleigh fading. A time-interleaver is employed at the transmitter to provide independent channel instances for the multiple descriptions, and a maximum a posteriori (MAP) detector is employed at the receiver to jointly decode the multiple descriptions. The concatenation of index assignment, binary phase-shift keying modulators, space-time orthogonal block coder, multiple-antenna channel, and MAP detector is modeled as a discrete memoryless channel. We optimize multiple description vector quantizers using an upper bound of the channel transition probability achieved by the MAP detector. The proposed scheme of combining multiple description coding and space-time orthogonal block coding is shown to be superior to each individual coding scheme.
Yugang Zhou, Wai-Yip Chan
IEEE Trans. Commun.2
2009 Performance comparison of HMM and VQ based single channel speech separation
Mohammad H. Radfar, Wai-Yip Chan, Richard M. Dansereau, Willy Wong
INTERSPEECH2
2008 Spectro-temporal features for robust far-field speaker identification
abstract
Features derived from an auditory spectro-temporal represen-tation of speech are proposed for robust far-field speaker iden-tification. The auditory representation is obtained by first filtering the speech signal with a gammatone filterbank. A modulation filterbank is then applied to the temporal enve-lope of each gammatone filter output. Compared to com-monly used mel-frequency cepstral coefficients (MFCC), the proposed features are shown to be more robust to mismatched conditions between enrollment and test data and are less sen-sitive to increasing reverberation time (RT). Experiments with simulated and recorded far-field speech show that a Gaus-sian mixture model based identification system, trained on the proposed features, attains an average improvement in identifi-cation accuracy of 15 % relative to a system trained on MFCC. Improvements of up to 85 % are attained for larger RT.
Tiago H. Falk, Wai-Yip Chan
INTERSPEECH2
2008 Long-term spectro-temporal information for improved automatic speech emotion classification
abstract
This paper investigates the contribution of features which con-vey long-term spectro-temporal (ST) information for the pur-pose of automatic emotional speech classification. The ST rep-resentation is obtained by means of a modulation filterbank de-composition of long-term temporal envelopes of the outputs of a gammatone filterbank. The two-dimensional discrete cosine transform is used to reduce the dimensionality of the represen-tation; candidate features are then derived from statistics com-puted from the DCT coefficients. Sequential forward feature selection is used to select the most salient features. Two types of experiments are described which use the Berlin emotional speech database to test the performance of the ST features alone and in combination with prosodic features. In a multi-class experiment, simulation results with a support vector classifier show that a 44 % reduction in classification error is attained once prosodic features are combined with the proposed ST fea-tures. Additionally, in a one-against-all experiment, an average increase in F-score of 33 % is attained when the proposed ST features are included. Index Terms: speech emotion recognition, spectro-temporal features, modulation spectrum, affective computing.
Siqing Wu, Tiago H. Falk, Wai-Yip Chan
INTERSPEECH3
2008 Hybrid Signal-and-Link-Parametric Speech Quality Measurement for VoIP Communications
abstract
A hybrid signal-and-link-parametric approach to speech quality measurement for voice-over-Internet protocol (VoIP) communications is described. Connection parameters are used to determine a base quality representative of the transmission link. Degradation factors, computed from perceptual features extracted from the decoded speech signal, are used to quantify distortions not captured by the connection parameters. The algorithm is tested on speech degraded by acoustic noise, temporal clippings, and noise suppression artifacts, thus simulating degradations present in wireless-VoIP tandem connections. Hybrid measurement is shown to overcome the limitations of pure link parametric and pure signal-based measurement methods, resulting in better measurement accuracy for modern VoIP communications. In addition, the proposed algorithm incurs modest computational overhead relative to pure link parametric measurement and attains up to 88% reduction in processing time relative to the ITU-T standard P.563 signal-based algorithm.
Tiago H. Falk, Wai-Yip Chan
IEEE Trans. Speech Audio Process.2
2007 A Hybrid Signal-and-Link-Parametric Approach to Single-Ended Quality Measurement of Packetized Speech
abstract
A hybrid signal-and-link-parametric approach to single-ended quality measurement of packetized speech is proposed. Transmission link parameters are used to determine a base quality for the test signal. The base quality is adjusted by degradation factors calculated from perceptual features extracted from the test signal. The degradation factors are based on Kullback-Leibler distances between a parametric model trained online for the extracted features and reference models of normative speech behavior. The proposed method overcomes the limitations of pure link parametric and pure signal-based methods.
Tiago H. Falk, Wai-Yip Chan
ICASSP (4)3
2007 Noise suppression based on extending a speech-dominated modulation band
abstract
Previous work on bandpass modulation filtering for noise suppression has resulted in unwanted perceptual artifacts and decreased speech clarity. Artifacts are introduced mainly due to half-wave rectification, which is employed to correct for negative power spectral values resultant from the filtering process. In this paper, modulation frequency estimation (i.e., bandwidth extension) is used to improve perceptual quality. Experiments demonstrate that speech-component lowpass modulation content can be reliably estimated from bandpass modulation content of speech-plus-noise components. Subjective listening tests corroborate that improved quality is attained when the removed speech lowpass modulation content is compensated for by the estimate.
Tiago H. Falk, Svante Stadler, W. Bastiaan Kleijn, Wai-Yip Chan
INTERSPEECH4
2007 Spectro-temporal processing for blind estimation of reverberation time and single-ended quality measurement of reverberant speech
abstract
Auditory spectro-temporal representations of reverberant speech are investigated for blind estimation of reverber-ation time (RT) and for single-ended measurement of speech quality. The auditory representations are obtained from an eight-filter filterbank which is used to extract the modulation spectra from temporal envelopes of the speech signal. Gaussian mixture models (GMM), one for each modulation channel and trained on clean speech signals, serve as reference models of normative speech behavior. Consistency measures, computed between re-verberant test signals and each GMM, are mapped to an estimated RT and to an estimated quality score. Experi-ments show that the proposed measures achieve superior performance relative to current “state-of-art ” algorithms. Index Terms: Reverberation time, quality measurement, GMM, modulation spectrum, consistency.
Tiago H. Falk, Wai-Yip Chan
INTERSPEECH3
2007 Degradation-classification assisted single-ended quality measurement of speech
abstract
We propose an algorithm to classify speech degradations at network endpoints and to estimate the speech quality based on the degradation classification decision. Percep-tual features from degraded speech signals are used to form statistical reference models of different degradation classes. Consistency measures, calculated between de-graded speech signals and the reference models, are used to train a degradation classifier and mean opinion score (MOS) mappings. The quality of a received speech signal is estimated based on its degradation class and the MOS mapping associated with the class. Experimental results show that the proposed algorithm achieves high classifi-cation accuracy, and degradation classification improves the accuracy of the quality estimate. Index Terms: speech communication network, speech degradations, speech transmission impairments, degrada-tion classification, speech quality measurement. 1.
Tiago H. Falk, Wai-Yip Chan
INTERSPEECH3
2007 Architecture for Multiple Reference Frame Variable Block Size Motion Estimation
abstract
This paper proposes a high throughput variable block size motion estimation (VBSME) architecture supporting multiple reference frames (MRF). To enable best rate-distortion performance for different video contents, the architecture allows selection between high spatial resolution motion search over a single reference frame, or MRF search at a lower spatial resolution. Through synthesis of an ASIC implementation, the architecture is shown to be suitable for high definition video resolutions and frame rates. The architecture also provides a higher overall macroblock throughput than other VBSME architectures in the literature.
Stephen Warrington, Subramania Sudharsanan, Wai-Yip Chan
ISCAS3
2007 Multiple Description Quantizer Design for Space-time Orthogonal Block Coded Channels
abstract
We study the design of multiple description quantizers for space-time orthogonal block coded slow Rayleigh fading channels. A time-interleaver is employed at the transmitter to provide independent channel instances for the multiple descriptions, and a maximum a posteriori (MAP) decoder is employed at the receiver to jointly decode the multiple descriptions. We propose a scheme to optimize multiple description vector quantizers using an upper bound of the channel transition probability achieved by the MAP decoder. The scheme furnishes substantial performance gain.
Yugang Zhou, Wai-Yip Chan
ISIT2
2006 Enhanced Non-Intrusive Speech Quality Measurement Using Degradation Models
abstract
The speech quality estimation scheme in [1] is improved with the addition of a reference model of the behavior of speech degraded by different transmission and/or coding schemes. Moreover, via maximization of a mutual information measure, we validate the use of segmental SNR as a measure of the amount of multiplicative noise present in the test signal. These two additions result in an algorithm that is more accurate and more robust to certain distortion conditions. When tested on unseen data, the proposed algorithm outperforms the current "state-of- art" P.563 algorithm while requiring considerably lower computational complexity.
Tiago H. Falk, Wai-Yip Chan
ICASSP (1)2
2006 Low-Complexity Multiple Description Vector Quantization With Constrained Central Codebook
abstract
Conventional multiple description vector quantizers (MDVQ) have high complexity, which limits their practical application. Two central-codebook-constrained MDVQ (CMDVQ) schemes are proposed to reduce the storage and search complexity. Simulation results show that for low channel loss rates, a tradeoff exists between choosing CMDVQ for its low complexity and the conventional MDVQ for its higher signal-to-noise ratio (SNR) performance. For medium to high channel loss rates, CMDVQ is preferred for its low complexity and comparable SNR performance to the conventional MDVQ
Yugang Zhou, Wai-Yip Chan
ICASSP (4)2
2006 Scalable high-throughput architecture for H.264/AVC variable block size motion estimation
abstract
Variable block size motion estimation (VBSME) is a key part of the new H.264/AVC video coding standard. This has increased the demand for high performance VBSME architectures. This paper proposes a VLSI architecture for high throughput VBSME. The VBS calculation is done by combining the results of sub-block calculations to form the results for larger blocks. High motion vector throughput is achieved in two proposed implementations: one performing operations on a 1/spl times/4 set of pixels per cycle, and the second performing operations on a 1/spl times/16 set of pixels per cycle. Using these approaches, the architecture is able to produce motion vector results at a higher throughput than current VBSME designs, while providing a high level of scalability through adjusting the length of the processing element array.
Stephen Warrington, Wai-Yip Chan, Subramania Sudharsanan
ISCAS2
2006 Performance improvement of the H.264/AVC deblocking filter using SIMD instructions
abstract
The H.264/AVC standard defines an in-loop deblocking filter which is used in both the encoder and decoder. This work examines several methods for improving the performance of the H.264/AVC reference software implementation of the deblocking filter. Methods examined include general software optimization, parallelization through standard multimedia SIMD instructions, and augmenting standard SIMD instruction sets with new instructions. Using the above methods, we are able to achieve a large speedup of the deblocking filter computation.
Stephen Warrington, Hassan Shojania, Subramania Sudharsanan, Wai-Yip Chan
ISCAS4
2006 Nonintrusive speech quality estimation using Gaussian mixture models
abstract
An algorithm for nonintrusive speech quality estimation based on Gaussian mixture models (GMMs) is presented. GMMs are used to form an artificial reference model of the behavior of features of undegraded speech. Consistency measures between the degraded speech signal and the reference model serve as indicators of speech quality. Consistency values are mapped to an objective speech quality score using a multivariate adaptive regression splines function. When tested on unseen data, the proposed algorithm generally outperforms ITU-T standard P.563, which is the current "state-of-the-art" algorithm. The algorithm computes objective quality scores roughly twice as fast as P.563.
Tiago H. Falk, Wai-Yip Chan
IEEE Signal Process. Lett.2
2006 Single-Ended Speech Quality Measurement Using Machine Learning Methods
abstract
We describe a novel single-ended algorithm constructed from models of speech signals, including clean and degraded speech, and speech corrupted by multiplicative noise and temporal discontinuities. Machine learning methods are used to design the models, including Gaussian mixture models, support vector machines, and random forest classifiers. Estimates of the subjective mean opinion score (MOS) generated by the models are combined using hard or soft decisions generated by a classifier which has learned to match the input signal with the models. Test results show the algorithm outperforming ITU-T P.563, the current "state-of-art" standard single-ended algorithm. Employed in a distributed double-ended measurement configuration, the proposed algorithm is found to be more effective than P.563 in assessing the quality of noise reduction systems and can provide a functionality not available with P.862 PESQ, the current double-ended standard algorithm
Tiago H. Falk, Wai-Yip Chan
IEEE Trans. Speech Audio Process.2
2005 Performance comparison of layered coding and multiple description coding in packet networks
abstract
We examine the performance of multiple descriptions coding (MDC) with and without the use of automatic retransmission request (ARQ) protocols for packet network communication. The rate-distortion lower bound of MDC and layered coding (LC) are incorporated into a performance measure that accounts for the additional costs of excess rates and delay incurred from using ARQ. Results show that unaided MDC is the best for large packet loss rates and large delay, and LC is the best for small loss and moderate delay. In between these two extremes, MDC aided by ARQ provides the best performance
Yugang Zhou, Wai-Yip Chan
GLOBECOM2
2005 Rate-distortion allocation for time-frequency dependent audio coding
abstract
A stream coding framework is presented for solving the distortion-constrained time-frequency dependent quantization problem that naturally arises when overlapped time-frequency decompositions are used. The main contributions of this paper are: (1) an efficient rate-distortion allocation algorithm for dependent quantization when the neighborhood of dependency is large; and (2) demonstration that a perceptual excitation distortion measure produces better coded audio quality than the conventional noise-to-mask ratio measure.
Ricky Der, Peter Kabal, Wai-Yip Chan
ICASSP (3)3
2005 Non-Intrusive GMM-Based Speech Quality Measurement
abstract
We propose a non-intrusive speech quality measurement algorithm based on using Gaussian-mixture probability models of features of undegraded speech signals as an artificial reference model of "clean" speech behaviour. The consistency between the features of the test speech signal and the reference model serves as an indicator of speech quality. Consistency measures are calculated and mapped to an objective speech quality score using a multivariate adaptive regression splines function. Simulation results show that the proposed method offers accurate and yet low-complexity measurement of speech quality.
Tiago H. Falk, Qingfeng Xu, Wai-Yip Chan
ICASSP (1)3
2005 An improved GMM-based voice quality predictor
abstract
A voice quality prediction method based on Gaussian mixture models (GMMs) is improved by constructing a feature selection algorithm to provide the best GMM-based prediction quality. The proposed sequential se-lection algorithm performs N-survivor search, allowing for trading between design complexity and performance. Simulation shows that predictors designed using the pro-posed algorithm outperform two benchmark selection al-gorithms. Performance improvements over the ITU-T P.862 PESQ standard are also attained. 1.
Tiago H. Falk, Wai-Yip Chan, Peter Kabal
INTERSPEECH2
2004 Bit allocation algorithms for frequency and time spread perceptual coding
abstract
We examine the problem of bit allocation when time spread and frequency spread perceptual distortion criteria are used. For such measures, standard incremental techniques can fail. Two algorithms are introduced for bit allocation; the first a multi-band version of the greedy algorithm, and the second an inverse greedy algorithm initialized by the bit allocation of a forward algorithm driven by a non-spread metric. Experimental results show the second algorithm outperforms the first.
Ricky Der, Peter Kabal, Wai-Yip Chan
ICASSP (4)3
2004 Speech quality estimation using Gaussian mixture models
abstract
Abstract—An algorithm for nonintrusive speech quality esti-mation based on Gaussian mixture models (GMMs) is presented. GMMs are used to form an artificial reference model of the behavior of features of undegraded speech. Consistency measures between the degraded speech signal and the reference model serve as indicators of speech quality. Consistency values are mapped to an objective speech quality score using a multivariate adaptive regression splines function. When tested on unseen data, the proposed algorithm generally outperforms ITU-T standard P.563, which is the current “state-of-the-art ” algorithm. The algorithm computes objective quality scores roughly twice as fast as P.563. Index Terms—Gaussian mixtures, quality assurance, quality measurement, quality of service, speech coding, speech quality, speech transmission, telephony. I.
Tiago H. Falk, Wai-Yip Chan, Peter Kabal
INTERSPEECH2
2003 Towards a new perceptual coding paradigm for audio signals
abstract
A new frequency domain approach to coding audio signals is introduced. The bit assignment strategy is aimed at reducing the perceived loudness difference between the original signal and the coded signal. As such it uses perceptual effects (spread excitation patterns), but does not directly invoke masking results. At low bit rates, examples coded with the new approach sound better than a more traditional bit allocation based on noise-to-mask ratio.
Ricky Der, Peter Kabal, Wai-Yip Chan
ICASSP (5)3
2002 Motion filter vector quantization
abstract
Motion-compensated prediction of video is formulated as a novel vector quantization scheme called motion filter vector quantization (MFVQ). In MFVQ, the motion vector and the pixel-intensity interpolation filter are combined into a motion filter and the entire filter is vector quantized. A codebook design algorithm is proposed for designing unit gain and entropy constrained MFVQ codebooks. The algorithm is tested under two application configurations, MFVQ with static codebook and MFVQ with forward-adaptive codebook, and is shown to furnish up to a dB of PSNR gain.
Dariusz Blasiak, Wai-Yip Chan
ICIP (1)2
2001 A novel code excited pel-recursive motion compensation algorithm
abstract
The conventional scalar-quantizer-based pel-recursive motion compensated coding scheme is enhanced by incorporating vector quantization. The enhanced scheme uses code-excited analysis-by-synthesis to quantize the motion compensated prediction residual. Simulation results show that the enhanced scheme furnishes large rate-distortion performance gains over the conventional scheme, particularly at bit rates below about 0.5 bits per pixel.
Jiandong Shen, Wai-Yip Chan
IEEE Signal Process. Lett.2
2000 Code Excited Pel-Recursive Motion Compensated Video Coding
abstract
We present a novel pel-recursive motion compensation scheme. The new scheme incorporates vector quantization into pel-recursive motion compensation by using a code excited analysis-by-synthesis coding structure. Our experiments demonstrate that the proposed scheme achieves 4-7 dB PSNR gains or equivalently 40-50% bit savings over the conventional scalar quantizer based pel-recursive motion compensation scheme.
Jiandong Shen, Wai-Yip Chan
ICIP2
2000 An experimental assessment of personal speech coding
Wenhui Jia, Wai-Yip Chan
Speech Commun.2
1999 Vector Quantization of Affine Motion Models
abstract
This paper presents an efficient motion compensation algorithm for low bit rate video coding. The motion field is modelled as a block-segmented affine motion field. Affine motion parameter estimation and quantization are performed in one step by searching a static VQ codebook. The codebook is designed using a GLA-like algorithm to minimize the motion compensation error. Furthermore, variable block size is used to allow more flexible bit allocation between motion coding and prediction error coding. Simulations on various test sequences show 0.8-2.0 dB PSNR gain over H.263 with advanced prediction mode enabled.
Jiandong Shen, Wai-Yip Chan
ICIP (2)2
1999 Generalized scalar quantizer design using dynamic programming
abstract
We cast the design of generalized scalar quantizers as a dynamic programming problem. The algorithm enables the design of discrete nonparametric estimators directly from training data and has the advantage of admitting a variety of constraints on the estimator mapping. The utility of the algorithm is illustrated with the design of rate-distortion performance predictors for a video coder.
Dariusz Blasiak, Jiandong Shen, Wai-Yip Chan
IEEE Signal Process. Lett.3
1998 Personal speech coding
abstract
In existing speech coding systems, all quantizer codebooks are designed to suit the statistical and perceptual characteristics of speech signals of a population of speakers. However, an individual's speech signal does not exhibit, even over a long time, the entire range of characteristics of the population. With the advent of the personal communication systems, personal information might become available and be used to improve the rate-distortion performance of speech coders. We assess the potential gain of personal speech coding by designing codebooks for individual speakers. Spectral quantisation, excitation and pitch lag codebooks of existing CELP coders are redesigned. The gains appear to be modest, suggesting that we need to use a different coding framework, which can model personal characteristics explicitly. Amongst the components, the spectral quantizer seems to be most amenable to personalization.
Wenhui Jin, Wai-Yip Chan
ICASSP2
1998 Efficient Wavelet Coding of Motion Compensated Prediction Residuals
abstract
In the coding of motion compensated prediction residuals, the competitiveness of wavelet coding relative to DCT coding is still an open issue. This paper describes an efficient and effective entropy coding strategy for the compression of wavelet transformed, motion compensated prediction residuals. Without the use of arithmetic coding, the proposed coder achieves competitive or better rate-distortion performance in comparison with H.263+, which uses DCT/runlength coding of the prediction residual. Moreover, the proposed coder furnishes better perceptual quality. The use of reasonably sized, static variable-length-codes in the proposed coder also helps to maintain low computational complexity and good error resiliency.
Dariusz Blasiak, Wai-Yip Chan
ICIP (2)2
1998 A Non-Parametric Method for Fast Joint Rate-Distortion Optimization of Motion Estimation and DFD Coding
Jiandong Shen, Wai-Yip Chan
ICIP (3)2
1998 Rate-distortion optimization of hierarchical displacement fields
abstract
The displacement vector field for motion-compensated coding of image sequences is represented and coded using a hierarchy of displacement vector fields over successively finer sampling grids. An efficient method based on the generalized Breiman-Friedman-Olshen-Stone (BFOS) tree-pruning algorithm is used to find a variable-depth hierarchy that satisfies a given bit rate or distortion constraint. The variable-depth hierarchy corresponds to using a nonuniformly sampled displacement field for motion-compensation. The proposed scheme is compared with conventional fixed-block-size block-matching motion-compensation. The motion information bit rate or the motion-compensation residual energy of the conventional scheme serves as the constraint. Simulation results show that the proposed scheme reduces the residual energy by between 0.6-1.5 dB, or the bit rate by 19-54%, while using only 40-100% of the computational complexity of the conventional scheme. The prediction image sequence produced by the hierarchical scheme offers much better viewing quality.
Hao Bi, Wai-Yip Chan
IEEE Trans. Circuits Syst. Video Technol.2
1996 Rate-constrained hierarchical motion estimation using BFOS tree pruning
abstract
The rate-distortion efficiency of motion compensation is improved by employing a nonuniformly sampled displacement field. The field is constructed based on a hierarchy of component displacement fields. Using block matching, the component displacement vectors are determined in a successive refinement manner, from the coarsest level to the finest level of the hierarchy. Based on the hierarchical block-matching motion compensation, a tree is constructed whose nodes are labelled with the distortions and bit rates resulting from the compensation. The BFOS algorithm is used to efficiently find pruned subtrees of the tree. Each pruned subtree corresponds to a nonuniformly sampled displacement field that furnishes the best rate-distortion performance. Compared to the conventional fixed-size block matching algorithm, and under the some bit-rate constraint, our hierarchical algorithm improves the average PSNR by up to 2.2 dB and offers superior subjective video quality.
Hao Bi, Wai-Yip Chan
ICASSP2
1996 Classified nonlinear predictive vector quantization of speech spectral parameters
abstract
Nonlinear predictive split vector quantization (NPSVQ) and classified NPSVQ (CNPSVQ) are introduced to exploit the correlation among the speech spectral parameters from two adjacent analysis frames. By interleaving intraframe SVQ with forward predictive SVQ, the error propagation is limited to at most one adjacent frame. At an overall bit rate of about 21 bits/frame, NPSVQ can provide similar coding quality as intraframe SVQ at 24 bits/frame. Voicing classification as used in CNPSVQ to obtain an additional average gain of 1 bit/frame for unvoiced frames. Therefore, an overall bit rate of 20 bits/frame is obtained for unvoiced frames. The particular form of nonlinear prediction we use incurs virtually no additional encoding computational complexity. We have verified our comparative performance results using subjective listening tests.
James H. Y. Loo, Wai-Yip Chan, Peter Kabal
ICASSP2
1996 Motion compensated transform coding of video using hierarchical displacement field and global rate-distortion optimization
abstract
Seeking a more efficient video coding scheme for low-bit-rate video communication, we enhance the conventional motion-compensated discrete-cosine-transform coding scheme by using (i) a nonuniformly sampled displacement field, constructed from a multiresolution hierarchy of component displacement fields, and (ii) "global" optimization of bit allocation, jointly between the displacement field hierarchy and the motion compensated prediction residual. Our simulation results demonstrate that, for bit rates in the 10-64 kb/s range and for various common test sequences, the enhanced scheme outperforms the conventional scheme by between 1-3 dB of PSNR. The enhanced scheme performs well because it is allowed to judiciously exploit temporal redundancy to a much greater extent than the conventional scheme.
Hao Bi, Wai-Yip Chan
ICIP (3)2
1994 Low-complexity encoding of speech LSF parameters using constrained-storage TSVQ
abstract
Tree structured vector quantization (TSVQ) is employed as a low-complexity approach to performing vector quantization of speech linear prediction coefficients, expressed for the purpose of quantization as line spectral frequency (LSF) parameters. Good tradeoffs between search complexity and distortion-rate performance are obtained using multiple-survivor encoding. The exponential storage-complexity of conventional TSVQ is circumvented by using multiple stages, where one or more tree codebooks may be used in each stage. Experimental results show that for rates between 23-25 bits/frame,the encoding complexity required to achieve "transparent coding" quality ranges from below two hundred to several hundred weighted-squared-error distortion computations per frame.>
Wai-Yip Chan, David Chemla
ICASSP (1)1
1994 Approaches to Layered Coding for Dual-Rate Wireless Video Transmission
abstract
Visual communications over wireless networks require the efficient and robust coding of video signals for transmission over wireless links having time-varying channel capacity. The authors compare several schemes for encoding video data into two priority streams, thereby enabling the transmission of video data over wireless links to be switched between two bit rates. An H.261 (p/spl times/64) algorithm is modified to implement each candidate scheme. The algorithms are evaluated for a microcellular wireless environment and a clear-channel bit rate of 65 kb/s. The results show that by combining layering with automatic-repeat request (wireless-)link control, almost-wireline visual quality can be achieved. >
Masoud R. K. Khansari, Awais Zakauddin, Wai-Yip Chan, Eric Dubois 0002, Paul Mermelstein
ICIP (1)3
1993 Joint Codebook Design for Summation Product-Code Vector Quantizers
abstract
With respect to the generalized product code (GPC) model for structured vector quantization, multistage VQ (MSVQ) and tree-structured VQ are members of a family of summation product codes (SPCs), defined by the prototypical synthesis function x=f/sub 1/+...+f/sub s/, where f/sub i/, i=1, . . ., s are the residual vector features. The authors describe an algorithm paradigm for the joint design of the feature codebooks constituting a GPC. They specialize the paradigm to a joint design algorithm for the SPCs and exhibit experimental results for the MSVQ of simulated sources. The performance improvements over conventional 'greedy' design are essentially 'free' as the only cost is a moderate increase in design complexity.>
Wai-Yip Chan, Allen Gersho, S. Soong
Data Compression Conference1
1992 The design of generalized product-code vector quantizers
abstract
Structured vector quantization (VQ) can achieve superior performance-complexity tradeoffs in comparison with unstructured VQ. Many VQ schemes fall into a class of structured VQ called product codes. A generalization of product codes wherein a feature may have multiple codebooks is considered. For the design of generalized product codes (GPCs) methodologies are devised for achieving independent tradeoffs of codebook storage complexity and encoding complexity versus distortion performance and for the joint optimization of feature codebooks. This GPC design framework makes it possible to attain many intermediate levels of rate-distortion performance between unstructured VQ and conventional product codes. Thus, the performance of unstructured VQ may be approached with a lower-complexity GPC. The framework is illustrated using numerical results from the quantization of speech line-spectral frequency parameters.>
Wai-Yip Chan
ICASSP1
1992 Vector quantization of speech LSF parameters with generalized product codes
Erdal Paksoy, Wai-Yip Chan, Allen Gersho
ICSLP2
1992 Enhanced multistage vector quantization by joint codebook design
abstract
Multistage vector quantization (MSVQ) can achieve very low encoding and storage complexity in comparison to unstructured vector quantization. However, the conventional stage-by-stage design of the codebooks in MSVQ is suboptimal with respect to the overall performance measure. The authors introduce an algorithm for the joint design of the stage codebooks to optimize the overall performance. The performance improvement, although modest, is achieved with no effect on encoding or storage complexity and only a slight increase in design effort.>
Wai-Yip Chan, Smita Gupta, Allen Gersho
IEEE Trans. Commun.1
1991 Constrained-storage vector quantization in high fidelity audio transform coding
abstract
The concept and design methods for efficient use of vector quantization (VQ) in high-fidelity audio coding are presented. It is demonstrated that with constrained-storage VQ (CSVQ), tree-structured codebooks can be constructed for very high rates without incurring an exponential growth in storage complexity and without impairing the rate-distortion performance. Nonlinear interpolative VQ allows efficient coding of the power envelope needed for transform-coefficient normalization and adaptive distortion assignment. These techniques lead to a substantial reduction in the overall bit rate and codebook storage for the audio coder.>
Wai-Yip Chan, Allen Gersho
ICASSP1
1991 Constrained-storage quantization of multiple vector sources by codebook sharing
abstract
A codebook sharing technique, called constrained storage vector quantization (CSVQ), is introduced. This technique offers a convenient and optimal way of trading off performance against storage. The technique can be used in conjunction with tree-structured vector quantization (VQ) and other structured VQ techniques that alleviate the search complexity obstacle. The effectiveness of CSVQ is illustrated for coding transform coefficients of audio signals with multistage VQ.>
Wai-Yip Chan, Allen Gersho
IEEE Trans. Commun.1
1990 High fidelity audio transform coding with vector quantization
abstract
Multi-stage tree-structured vector quantization (MSTVQ) is examined as an alternative to entropy constrained scalar quantization (ECSQ) in transform coding of high fidelity audio signals with a simultaneous masking model for distortion control. Discrete-cosine-transform coefficients are normalized by an interpolated spectral power envelope and groups of adjacent coefficients are vector coded with variable rate to achieve distortion-masking. With the current coder configuration, high fidelity quality for a sampling rate of 32 kHz is achievable with data rates below 64 kbps for some transform and masking model, preliminary results show that MSTVO and ECSQ have a similar rate-distortion performance.>
Wai-Yip Chan, Allen Gersho
ICASSP1
1983 An Integrated Voice/Data System for VHF/UHF Mobile Radio
abstract
This paper presents the basic architecture and performance of a mobile radio multiaccess voice/data system. Natural pauses in conversational speech allow bandwidth saving through interleaving of data packets and talkspurts from different voice sources. A speech detector designed specifically for the mobile environment is presented. Blocking and delay performance of the multiaccess uplink is analyzed for voice traffic, assuming no traffic effects from the low priority data packets. Performance results from simulation are then presented for two downlink strategies in a two-hop virtual circuit in which a base station acts as a relay. The results verify also that the uplink analysis is valid for low voice traffic. For the data traffic, simulation results are presented in terms of data packet transmission delay and probability of collision with talkspurts. The results indicate that data flow may be limited by the collision factor. This work concludes that relative to conventional radio telephoning in which two channels are dedicated to each transmitter/receiver pair, a bandwidth reduction of 30-35 percent can be achieved.
Samy A. Mahmoud, Wai-Yip Chan, J. Spruce Riordon, Salah E. Aidarous
IEEE J. Sel. Areas Commun.2