EDBT 2026 Demo / reviewers in the wild / expert
Kuldip K. Paliwal
dblp:25/2700
· DBLP profile ↗
188ranked-venue papers
50as first author
11since 2021 · last 2025
0000-0002-3553-3662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 44 first-author · 1 since 2021Artificial intelligence and machine learning · 73 · 15 first-authorApplied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UBSTrack: Unified Band Selection and Multimodel Ensemble for Hyperspectral Object TrackingabstractHyperspectral object tracking is notably challenging due to the high-dimensional nature of the data and the necessity of seamlessly integrating spectral, spatial and temporal information. Traditional methods often emphasize detection-based or tracking-based networks, each leveraging their inherent strengths but overlooking the potential advantages of a combined approach, leading to suboptimal performance in complex, real-world scenarios. Furthermore, this challenge is amplified by the variability of spectral bands across datasets, making the maintenance of consistent tracking performance complicated. To address these issues, we propose a novel, unified approach that merges adaptive band selection with a multi-model ensemble strategy. We introduce a local and global attention-based unified band selection (UBS) technique that identifies the most informative three bands from any dataset, significantly reducing data complexity while preserving critical spectral and spatial information. This UBS method employs spectral independence, allowing it to process hyperspectral video frames with any number of bands as input, ultimately generating a three-band pseudocolor image. This is coupled with a multi-model ensemble framework, utilizing a local and global attention-based appearance module. The module selects the optimal candidate by computing the similarity between the proposals generated by the base models and historical frames. Experimental results show that our approach, UBSTrack, achieves state-of-the-art performance, delivering robust and accurate tracking under different real-world challenging scenarios. The code of UBSTrack is available at the following link: source code. Jun Zhou 0001, Wangzhi Xing, Yongsheng Gao 0001, Kuldip K. Paliwal |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Hy-Tracker: A Novel Framework for Enhancing Efficiency and Accuracy of Object Tracking in Hyperspectral VideosabstractHyperspectral images, with their many spectral bands, provide a rich source of material information about an object that can be effectively used for object tracking. However, many trackers in this domain rely on detection-based techniques, which often perform suboptimally in challenging scenarios such as managing occlusions and distinguishing objects in cluttered backgrounds. This underperformance is primarily due to the presence of multiple spectral bands and the inability to leverage this abundance of data for effective tracking. Additionally, the scarcity of annotated hyperspectral videos and the absence of comprehensive temporal information exacerbate these difficulties, further limiting the effectiveness of current tracking methods. To address these challenges, this article introduces the novel Hy-Tracker framework, designed to bridge the gap between hyperspectral data and state-of-the-art object detection methods. Our approach leverages the strengths of YOLOv7 for object tracking in hyperspectral videos, enhancing both accuracy and robustness in complex scenarios. The Hy-Tracker framework comprises two key components. We introduce a hierarchical attention for band selection (HAS-BS) that selectively processes and groups the most informative spectral bands, thereby significantly improving detection accuracy. Additionally, we have developed a refined tracker that refines the initial detections by incorporating a classifier and a temporal network using gated recurrent units (GRUs). The classifier distinguishes similar objects, while the temporal network models temporal dependencies across frames for robust performance despite occlusions and scale variations (SVs). Experimental results on hyperspectral benchmark datasets demonstrate the effectiveness of Hy-Tracker in accurately tracking objects across frames and overcoming the challenges inherent in detection-based hyperspectral object tracking (HOT). Wangzhi Xing, Jun Zhou 0001, Yongsheng Gao 0001, Kuldip K. Paliwal |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Probing RNA structures and functions by solvent accessibility: an overview from experimental and computational perspectivesabstractCharacterizing RNA structures and functions have mostly been focused on 2D, secondary and 3D, tertiary structures. Recent advances in experimental and computational techniques for probing or predicting RNA solvent accessibility make this 1D representation of tertiary structures an increasingly attractive feature to explore. Here, we provide a survey of these recent developments, which indicate the emergence of solvent accessibility as a simple 1D property, adding to secondary and tertiary structures for investigating complex structure-function relations of RNAs. Md. Solayman, Thomas Litfin, Kuldip K. Paliwal, Yaoqi Zhou, Jian Zhan |
Briefings Bioinform. | 4 |
| 2022 | SPOT-Contact-LM: improving single-sequence-based prediction of protein contact map using a transformer language modelabstractMOTIVATION: Accurate prediction of protein contact-map is essential for accurate protein structure and function prediction. As a result, many methods have been developed for protein contact map prediction. However, most methods rely on protein-sequence-evolutionary information, which may not exist for many proteins due to lack of naturally occurring homologous sequences. Moreover, generating evolutionary profiles is computationally intensive. Here, we developed a contact-map predictor utilizing the output of a pre-trained language model ESM-1b as an input along with a large training set and an ensemble of residual neural networks. RESULTS: We showed that the proposed method makes a significant improvement over a single-sequence-based predictor SSCpred with 15% improvement in the F1-score for the independent CASP14-FM test set. It also outperforms evolutionary-profile-based methods trRosetta and SPOT-Contact with 48.7% and 48.5% respective improvement in the F1-score on the proteins without homologs (Neff = 1) in the independent SPOT-2018 set. The new method provides a much faster and reasonably accurate alternative to evolution-based methods, useful for large-scale prediction. AVAILABILITY AND IMPLEMENTATION: Stand-alone-version of SPOT-Contact-LM is available at https://github.com/jas-preet/SPOT-Contact-Single. Direct prediction can also be made at https://sparks-lab.org/server/spot-contact-single. The datasets used in this research can also be downloaded from the GitHub. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Litfin, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 4 |
| 2022 | Predicting RNA distance-based contact maps by integrated deep learning on physics-inferred secondary structure and evolutionary-derived mutational couplingabstractMOTIVATION: Recently, AlphaFold2 achieved high experimental accuracy for the majority of proteins in Critical Assessment of Structure Prediction (CASP 14). This raises the hope that one day, we may achieve the same feat for RNA structure prediction for those structured RNAs, which is as fundamentally and practically important similar to protein structure prediction. One major factor in the recent advancement of protein structure prediction is the highly accurate prediction of distance-based contact maps of proteins. RESULTS: Here, we showed that by integrated deep learning with physics-inferred secondary structures, co-evolutionary information and multiple sequence-alignment sampling, we can achieve RNA contact-map prediction at a level of accuracy similar to that in protein contact-map prediction. More importantly, highly accurate prediction for top L long-range contacts can be assured for those RNAs with a high effective number of homologous sequences (Neff > 50). The initial use of the predicted contact map as distance-based restraints confirmed its usefulness in 3D structure prediction. AVAILABILITY AND IMPLEMENTATION: SPOT-RNA-2D is available as a web server at https://sparks-lab.org/server/spot-rna-2d/ and as a standalone program at https://github.com/jaswindersingh2/SPOT-RNA-2D. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kuldip K. Paliwal, Thomas Litfin, Yaoqi Zhou |
Bioinform. | 2 |
| 2022 | On supervised LPC estimation training targets for augmented Kalman filter-based speech enhancement
Sujan Kumar Roy, Aaron Nicolson, Kuldip K. Paliwal |
Speech Commun. | 3 |
| 2021 | SPOT-1D2: Improving Protein Secondary Structure Prediction using High Sequence Identity Training Set and an Ensemble of Recurrent and Residual-convolutional Neural NetworksabstractProtein secondary structure prediction has been a long-standing problem in computational biology. Recent advances in deep contextual learning have enabled its performance in three-state prediction closer to the theoretical limit at 88–90%. Here, we showed that a large training set with 95% sequence identity cutoff can improve prediction of secondary structures even for those unrelated test sequences (<25% sequence identity cutoff) compared to the use of a non-redundant training dataset with 25% sequence identity cutoff. The three-state prediction edges closer to an accuracy of 87% and eight-state at 76%.The resulting model called SPOT-1D2 is freely available to academic users at https://github.com/jas-preet/SPOT-1D2. Kuldip K. Paliwal, Andrew Busch, Yaoqi Zhou |
CIBCB | 3 |
| 2021 | Single-sequence and profile-based prediction of RNA solvent accessibility using dilated convolutional neural networkabstractMOTIVATION: RNA solvent accessibility, similar to protein solvent accessibility, reflects the structural regions that are accessible to solvents or other functional biomolecules, and plays an important role for structural and functional characterization. Unlike protein solvent accessibility, only a few tools are available for predicting RNA solvent accessibility despite the fact that millions of RNA transcripts have unknown structures and functions. Also, these tools have limited accuracy. Here, we have developed RNAsnap2 that uses a dilated convolutional neural network with a new feature, based on predicted base-pairing probabilities from LinearPartition. RESULTS: Using the same training set from the recent predictor RNAsol, RNAsnap2 provides an 11% improvement in median Pearson Correlation Coefficient (PCC) and 9% improvement in mean absolute errors for the same test set of 45 RNA chains. A larger improvement (22% in median PCC) is observed for 31 newly deposited RNA chains that are non-redundant and independent from the training and the test sets. A single-sequence version of RNAsnap2 (i.e. without using sequence profiles generated from homology search by Infernal) has achieved comparable performance to the profile-based RNAsol. In addition, RNAsnap2 has achieved comparable performance for protein-bound and protein-free RNAs. Both RNAsnap2 and RNAsnap2 (SingleSeq) are expected to be useful for searching structural signatures and locating functional regions of non-coding RNAs. AVAILABILITY AND IMPLEMENTATION: Standalone-versions of RNAsnap2 and RNAsnap2 (SingleSeq) are available at https://github.com/jaswindersingh2/RNAsnap2. Direct prediction can also be made at https://sparks-lab.org/server/rnasnap2. The datasets used in this research can also be downloaded from the GITHUB and the webserver mentioned above. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anil Kumar Hanumanthappa, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 3 |
| 2021 | SPOT-1D-Single: improving the single-sequence-based prediction of protein secondary structure, backbone angles, solvent accessibility and half-sphere exposures using a large training set and ensembled deep learningabstractMOTIVATION: Knowing protein secondary and other one-dimensional structural properties are essential for accurate protein structure and function prediction. As a result, many methods have been developed for predicting these one-dimensional structural properties. However, most methods relied on evolutionary information that may not exist for many proteins due to a lack of sequence homologs. Moreover, it is computationally intensive for obtaining evolutionary information as the library of protein sequences continues to expand exponentially. Here, we developed a new single-sequence method called SPOT-1D-Single based on a large training dataset of 39 120 proteins deposited prior to 2016 and an ensemble of hybrid long-short-term-memory bidirectional neural network and convolutional neural network. RESULTS: We showed that SPOT-1D-Single consistently improves over SPIDER3-Single and ProteinUnet for secondary structure, solvent accessibility, contact number and backbone angles prediction for all seven independent test sets (TEST2018, SPOT-2016, SPOT-2016-HQ, SPOT-2018, SPOT-2018-HQ, CASP12 and CASP13 free-modeling targets). For example, the predicted three-state secondary structure's accuracy ranges from 72.12% to 74.28% by SPOT-1D-Single, compared to 69.1-72.6% by SPIDER3-Single and 70.6-73% by ProteinUnet. SPOT-1D-Single also predicts SS3 and SS8 with 6.24% and 6.98% better accuracy than SPOT-1D on SPOT-2018 proteins with no homologs (Neff = 1), respectively. The new method's improvement over existing techniques is due to a larger training set combined with ensembled learning. AVAILABILITY AND IMPLEMENTATION: Standalone-version of SPOT-1D-Single is available at https://github.com/jas-preet/SPOT-1D-Single. Direct prediction can also be made at https://sparks-lab.org/server/spot-1d-single. The datasets used in this research can also be downloaded from GitHub. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Litfin, Kuldip K. Paliwal, Anil Kumar Hanumanthappa, Yaoqi Zhou |
Bioinform. | 3 |
| 2021 | Improved RNA secondary structure and tertiary base-pairing prediction using evolutionary profile, mutational coupling and two-dimensional transfer learningabstractMOTIVATION: The recent discovery of numerous non-coding RNAs (long non-coding RNAs, in particular) has transformed our perception about the roles of RNAs in living organisms. Our ability to understand them, however, is hampered by our inability to solve their secondary and tertiary structures in high resolution efficiently by existing experimental techniques. Computational prediction of RNA secondary structure, on the other hand, has received much-needed improvement, recently, through deep learning of a large approximate data, followed by transfer learning with gold-standard base-pairing structures from high-resolution 3-D structures. Here, we expand this single-sequence-based learning to the use of evolutionary profiles and mutational coupling. RESULTS: The new method allows large improvement not only in canonical base-pairs (RNA secondary structures) but more so in base-pairing associated with tertiary interactions such as pseudoknots, non-canonical and lone base-pairs. In particular, it is highly accurate for those RNAs of more than 1000 homologous sequences by achieving >0.8 F1-score (harmonic mean of sensitivity and precision) for 14/16 RNAs tested. The method can also significantly improve base-pairing prediction by incorporating artificial but functional homologous sequences generated from deep mutational scanning without any modification. The fully automatic method (publicly available as server and standalone software) should provide the scientific community a new powerful tool to capture not only the secondary structure but also tertiary base-pairing information for building three-dimensional models. It also highlights the future of accurately solving the base-pairing structure by using a large number of natural and/or artificial homologous sequences. AVAILABILITY AND IMPLEMENTATION: Standalone-version of SPOT-RNA2 is available at https://github.com/jaswindersingh2/SPOT-RNA2. Direct prediction can also be made at https://sparks-lab.org/server/spot-rna2/. The datasets used in this research can also be downloaded from the GITHUB and the webserver mentioned above. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kuldip K. Paliwal, Tongchuan Zhang, Thomas Litfin, Yaoqi Zhou |
Bioinform. | 2 |
| 2021 | RNAcmap: a fully automatic pipeline for predicting contact maps of RNAs by evolutionary coupling analysisabstractMOTIVATION: The accuracy of RNA secondary and tertiary structure prediction can be significantly improved by using structural restraints derived from evolutionary coupling or direct coupling analysis. Currently, these coupling analyses relied on manually curated multiple sequence alignments collected in the Rfam database, which contains 3016 families. By comparison, millions of non-coding RNA sequences are known. Here, we established RNAcmap, a fully automatic pipeline that enables evolutionary coupling analysis for any RNA sequences. The homology search was based on the covariance model built by INFERNAL according to two secondary structure predictors: a folding-based algorithm RNAfold and the latest deep-learning method SPOT-RNA. RESULTS: We showed that the performance of RNAcmap is less dependent on the specific evolutionary coupling tool but is more dependent on the accuracy of secondary structure predictor with the best performance given by RNAcmap (SPOT-RNA). The performance of RNAcmap (SPOT-RNA) is comparable to that based on Rfam-supplied alignment and consistent for those sequences that are not in Rfam collections. Further improvement can be made with a simple meta predictor RNAcmap (SPOT-RNA/RNAfold) depending on which secondary structure predictor can find more homologous sequences. Reliable base-pairing information generated from RNAcmap, for RNAs with high effective homologous sequences, in particular, will be useful for aiding RNA structure prediction. AVAILABILITY AND IMPLEMENTATION: RNAcmap is available as a web server at https://sparks-lab.org/server/rnacmap/ and as a standalone application along with the datasets at https://github.com/sparks-lab-org/RNAcmap_standalone. A platform independent and fully configured docker image of RNAcmap is also provided at https://hub.docker.com/r/jaswindersingh2/rnacmap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tongchuan Zhang, Thomas Litfin, Jian Zhan, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 5 |
| 2020 | Deep Residual-Dense Lattice Network for Speech EnhancementabstractConvolutional neural networks (CNNs) with residual links (ResNets) and causal dilated convolutional units have been the network of choice for deep learning approaches to speech enhancement. While residual links improve gradient flow during training, feature diminution of shallow layer outputs can occur due to repetitive summations with deeper layer outputs. One strategy to improve feature re-usage is to fuse both ResNets and densely connected CNNs (DenseNets). DenseNets, however, over-allocate parameters for feature re-usage. Motivated by this, we propose the residual-dense lattice network (RDL-Net), which is a new CNN for speech enhancement that employs both residual and dense aggregations without over-allocating parameters for feature re-usage. This is managed through the topology of the RDL blocks, which limit the number of outputs used for dense aggregations. Our extensive experimental investigation shows that RDL-Nets are able to achieve a higher speech enhancement performance than CNNs that employ residual and/or dense aggregations. RDL-Nets also use substantially fewer parameters and have a lower computational requirement. Furthermore, we demonstrate that RDL-Nets outperform many state-of-the-art deep learning approaches to speech enhancement. Availability: https://github.com/nick-nikzad/RDL-SE. Mohammad Nikzad, Aaron Nicolson, Yongsheng Gao 0001, Jun Zhou 0001, Kuldip K. Paliwal, Fanhua Shang |
AAAI | 5 |
| 2020 | Sum-Product Networks for Robust Automatic Speaker IdentificationabstractWe introduce sum-product networks (SPNs) for robust speech processing through a simple robust automatic speaker identification (ASI) task. SPNs are deep probabilistic graphical models capable of answering multiple probabilistic queries. We show that SPNs are able to remain robust by using the marginal probability density function (PDF) of the spectral features that reliably represent speech. Though current SPN toolkits and learning algorithms are in their infancy, we aim to show that SPNs have the potential to become a useful tool for robust speech processing in the future. SPN speaker models are evaluated here on real-world non-stationary and coloured noise sources at multiple signal-to-noise ratio (SNR) levels. In terms of ASI accuracy, we find that SPN speaker models are more robust than two recent convolutional neural network (CNN)-based ASI systems. Additionally, SPN speaker models consist of significantly fewer parameters than their CNN-based counterparts. The results indicate that SPN speaker models could be a robust, parameter-efficient alternative for ASI. Additionally, this work demonstrates that SPNs have potential in related tasks, such as robust automatic speech recognition (ASR) and automatic speaker verification (ASV). Availability: The SPN ASI system is available at https://github.com/anicolson/SPN-ASI. Aaron Nicolson, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2020 | A Deep Learning-Based Kalman Filter for Speech EnhancementabstractThe existing Kalman filter (KF) suffers from poor estimates of the noise variance and the linear prediction coefficients (LPCs) in real-world noise conditions. This results in a degraded speech enhancement performance. In this paper, a deep learning approach is used to more accurately estimate the noise variance and LPCs, enabling the KF to enhance speech in various noise conditions. Specifically, a deep learning approach to MMSE-based noise power spectral density (PSD) estimation, called DeepMMSE, is used. The estimated noise PSD is used to compute the noise variance. We also construct a whitening filter with its coefficients computed from the estimated noise PSD. It is then applied to the noisy speech, yielding pre-whitened speech for computing the LPCs. The improved noise variance and LPC estimates enable the KF to minimise the residual noise and distortion in the enhanced speech. Experimental results show that the proposed method exhibits higher quality and intelligibility in the enhanced speech than the benchmark methods in various noise conditions for a wide-range of SNR levels. Sujan Kumar Roy, Aaron Nicolson, Kuldip K. Paliwal |
INTERSPEECH | 3 |
| 2020 | Deep Learning with Augmented Kalman Filter for Single-Channel Speech EnhancementabstractThe existing augmented Kalman filter (AKF) suffers from poor LPC estimates in real-world noise conditions, which degrades the speech enhancement performance. In this paper, a deep learning technique exploits the LPC estimates for the AKF to enhance speech in various noise conditions. Specifically, a deep residual network is used to estimate the noise PSD for computing noise LPCs. A whitening filter is also implemented with the noise LPCs to pre-whiten the noisy speech signal prior to estimating the speech LPCs. It is shown that the improved speech and noise LPCs enable the AKF to minimize the residual noise as well as distortion in the enhanced speech. Experimental results show that the enhanced speech produced by the proposed method exhibits higher quality and intelligibility than the benchmark methods in various noise conditions for a wide-range of SNR levels. Sujan Kumar Roy, Aaron Nicolson, Kuldip K. Paliwal |
ISCAS | 3 |
| 2020 | Identifying molecular recognition features in intrinsically disordered regions of proteins by transfer learningabstractMOTIVATION: Protein intrinsic disorder describes the tendency of sequence residues to not fold into a rigid three-dimensional shape by themselves. However, some of these disordered regions can transition from disorder to order when interacting with another molecule in segments known as molecular recognition features (MoRFs). Previous analysis has shown that these MoRF regions are indirectly encoded within the prediction of residue disorder as low-confidence predictions [i.e. in a semi-disordered state P(D)≈0.5]. Thus, what has been learned for disorder prediction may be transferable to MoRF prediction. Transferring the internal characterization of protein disorder for the prediction of MoRF residues would allow us to take advantage of the large training set available for disorder prediction, enabling the training of larger analytical models than is currently feasible on the small number of currently available annotated MoRF proteins. In this paper, we propose a new method for MoRF prediction by transfer learning from the SPOT-Disorder2 ensemble models built for disorder prediction. RESULTS: We confirm that directly training on the MoRF set with a randomly initialized model yields substantially poorer performance on independent test sets than by using the transfer-learning-based method SPOT-MoRF, for both deep and simple networks. Its comparison to current state-of-the-art techniques reveals its superior performance in identifying MoRF binding regions in proteins across two independent testing sets, including our new dataset of >800 protein chains. These test chains share <30% sequence similarity to all training and validation proteins used in SPOT-Disorder2 and SPOT-MoRF, and provide a much-needed large-scale update on the performance of current MoRF predictors. The method is expected to be useful in locating functional disordered regions in proteins. AVAILABILITY AND IMPLEMENTATION: SPOT-MoRF and its data are available as a web server and as a standalone program at: http://sparks-lab.org/jack/server/SPOT-MoRF/index.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jack Hanson, Thomas Litfin, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 3 |
| 2020 | Masked multi-head self-attention for causal speech enhancement
Aaron Nicolson, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2020 | DeepMMSE: A Deep Learning Approach to MMSE-Based Noise Power Spectral Density EstimationabstractAn accurate noise power spectral density (PSD) tracker is an indispensable component of a single-channel speech enhancement system. Bayesian-motivated minimum mean-square error (MMSE)-based noise PSD estimators have been the most prominent in recent time. However, they lack the ability to track highly non-stationary noise sources due to current methods of a priori signal-to-noise (SNR) estimation. This is caused by the underlying assumption that the noise signal changes at a slower rate than the speech signal. As a result, MMSE-based noise PSD trackers exhibit a large tracking delay and produce noise PSD estimates that require bias compensation. Motivated by this, we propose an MMSE-based noise PSD tracker that employs a temporal convolutional network (TCN) a priori SNR estimator. The proposed noise PSD tracker, called DeepMMSE makes no assumptions about the characteristics of the noise or the speech, exhibits no tracking delay, and produces an accurate estimate that requires no bias correction. Our extensive experimental investigation shows that the proposed DeepMMSE method outperforms state-of-the-art noise PSD trackers and demonstrates the ability to track abrupt changes in the noise level. Furthermore, when employed in a speech enhancement framework, the proposed DeepMMSE method is able to outperform state-of-the-art noise PSD trackers, as well as multiple deep learning approaches to speech enhancement. Availability: DeepMMSE is available at: https://github.com/anicolson/DeepXi. Qiquan Zhang, Aaron Nicolson, Mingjiang Wang, Kuldip K. Paliwal, Chenxu Wang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Improving prediction of protein secondary structure, backbone angles, solvent accessibility and contact numbers by using predicted contact maps and an ensemble of recurrent and residual convolutional neural networksabstractMOTIVATION: Sequence-based prediction of one dimensional structural properties of proteins has been a long-standing subproblem of protein structure prediction. Recently, prediction accuracy has been significantly improved due to the rapid expansion of protein sequence and structure libraries and advances in deep learning techniques, such as residual convolutional networks (ResNets) and Long-Short-Term Memory Cells in Bidirectional Recurrent Neural Networks (LSTM-BRNNs). Here we leverage an ensemble of LSTM-BRNN and ResNet models, together with predicted residue-residue contact maps, to continue the push towards the attainable limit of prediction for 3- and 8-state secondary structure, backbone angles (θ, τ, ϕ and ψ), half-sphere exposure, contact numbers and solvent accessible surface area (ASA). RESULTS: The new method, named SPOT-1D, achieves similar, high performance on a large validation set and test set (≈1000 proteins in each set), suggesting robust performance for unseen data. For the large test set, it achieves 87% and 77% in 3- and 8-state secondary structure prediction and 0.82 and 0.86 in correlation coefficients between predicted and measured ASA and contact numbers, respectively. Comparison to current state-of-the-art techniques reveals substantial improvement in secondary structure and backbone angle prediction. In particular, 44% of 40-residue fragment structures constructed from predicted backbone Cα-based θ and τ angles are less than 6 Å root-mean-squared-distance from their native conformations, nearly 20% better than the next best. The method is expected to be useful for advancing protein structure and function prediction. AVAILABILITY AND IMPLEMENTATION: SPOT-1D and its data is available at: http://sparks-lab.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jack Hanson, Kuldip K. Paliwal, Thomas Litfin, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 2 |
| 2019 | Deep learning for minimum mean-square error approaches to speech enhancementabstractRecently, the focus of speech enhancement research has shifted from minimum mean-square error (MMSE) approaches, like the MMSE short-time spectral amplitude (MMSE-STSA) estimator, to state-of-the-art masking- and mapping-based deep learning approaches. We aim to bridge the gap between these two differing speech enhancement approaches. Deep learning methods for MMSE approaches are investigated in this work, with the objective of producing intelligible enhanced speech at a high quality. Since the speech enhancement performance of an MMSE approach improves with the accuracy of the used a priori signal-to-noise ratio (SNR) estimator, a residual long short-term memory (ResLSTM) network is utilised here to accurately estimate the a priori SNR. MMSE approaches utilising the ResLSTM a priori SNR estimator are evaluated using subjective and objective measures of speech quality and intelligibility. The tested conditions include real-world non-stationary and coloured noise sources at multiple SNR levels. MMSE approaches utilising the proposed a priori SNR estimator are able to achieve higher enhanced speech quality and intelligibility scores than recent masking- and mapping-based deep learning approaches. The results presented in this work show that the performance of an MMSE approach to speech enhancement significantly increases when utilising deep learning. Availability: The proposed a priori SNR estimator is available at: https://github.com/anicolson/DeepXi. Aaron Nicolson, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2018 | Bidirectional Long-Short Term Memory Network-based Estimation of Reliable Spectral Component LocationsabstractAn accurate Ideal Binary Mask (IBM) estimate is essential for Missing Feature Theory (MFT)-based speaker identifica-tion, as incorrectly labelled spectral components (where a com-ponent is either reliable or unreliable) will degrade the perfor-mance of an Automatic Speaker Identification (ASI) system ad-versely in the presence of noise. In this work a Bidirectional Re-current Neural Network (BRNN) with Long-Short Term Mem-ory (LSTM) cells is proposed for improved IBM estimation. The proposed system had an average IBM estimate accuracy improvement of 4.5% and an average MFT-based speaker iden-tification accuracy improvement of 3.1% over all tested SNRdB levels, when compared to the previously proposed Multilayer Perceptron (MLP)-IBM estimator. When used for speech en-hancement the proposed system had an average MOS-LQO (ob-jective quality measure) improvement of 0.32 and an average QSTI (objective intelligibility measure) improvement of 0.01 over all tested SNRdB levels, when compared to the MLP-IBM estimator. The results presented in this work highlight the effec-tiveness of the proposed BRNN-IBM estimator for MFT-based speaker identification and IBM-based speech enhancement. Aaron Nicolson, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2018 | Sixty-five years of the long march in protein secondary structure prediction: the final stretch?abstractProtein secondary structure prediction began in 1951 when Pauling and Corey predicted helical and sheet conformations for protein polypeptide backbone even before the first protein structure was determined. Sixty-five years later, powerful new methods breathe new life into this field. The highest three-state accuracy without relying on structure templates is now at 82-84%, a number unthinkable just a few years ago. These improvements came from increasingly larger databases of protein sequences and structures for training, the use of template secondary structure information and more powerful deep learning techniques. As we are approaching to the theoretical limit of three-state prediction (88-90%), alternative to secondary structure prediction (prediction of backbone torsion angles and Cα-atom-based angles and torsion angles) not only has more room for further improvement but also allows direct prediction of three-dimensional fragment structures with constantly improved accuracy. About 20% of all 40-residue fragments in a database of 1199 non-redundant proteins have <6 Å root-mean-squared distance from the native conformations by SPIDER2. More powerful deep learning methods with improved capability of capturing long-range interactions begin to emerge as the next generation of techniques for secondary structure prediction. The time has come to finish off the final stretch of the long march towards protein secondary structure prediction. Yuedong Yang, Jianzhao Gao, Jihua Wang, Rhys Heffernan, Jack Hanson, Kuldip K. Paliwal, Yaoqi Zhou |
Briefings Bioinform. | 6 |
| 2018 | Accurate prediction of protein contact maps by coupling residual two-dimensional bidirectional long short-term memory with convolutional neural networksabstractMotivation: Accurate prediction of a protein contact map depends greatly on capturing as much contextual information as possible from surrounding residues for a target residue pair. Recently, ultra-deep residual convolutional networks were found to be state-of-the-art in the latest Critical Assessment of Structure Prediction techniques (CASP12) for protein contact map prediction by attempting to provide a protein-wide context at each residue pair. Recurrent neural networks have seen great success in recent protein residue classification problems due to their ability to propagate information through long protein sequences, especially Long Short-Term Memory (LSTM) cells. Here, we propose a novel protein contact map prediction method by stacking residual convolutional networks with two-dimensional residual bidirectional recurrent LSTM networks, and using both one-dimensional sequence-based and two-dimensional evolutionary coupling-based information. Results: We show that the proposed method achieves a robust performance over validation and independent test sets with the Area Under the receiver operating characteristic Curve (AUC) > 0.95 in all tests. When compared to several state-of-the-art methods for independent testing of 228 proteins, the method yields an AUC value of 0.958, whereas the next-best method obtains an AUC of 0.909. More importantly, the improvement is over contacts at all sequence-position separations. Specifically, a 8.95%, 5.65% and 2.84% increase in precision were observed for the top L∕10 predictions over the next best for short, medium and long-range contacts, respectively. This confirms the usefulness of ResNets to congregate the short-range relations and 2D-BRLSTM to propagate the long-range dependencies throughout the entire protein contact map 'image'. Availability and implementation: SPOT-Contact server url: http://sparks-lab.org/jack/server/SPOT-Contact/. Supplementary information: Supplementary data are available at Bioinformatics online. Jack Hanson, Kuldip K. Paliwal, Thomas Litfin, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 2 |
| 2018 | Robustness metric-based tuning of the augmented Kalman filter for the enhancement of speech corrupted with coloured noise
Aidan E. W. George, Stephen So, Ratna Ghosh, Kuldip K. Paliwal |
Speech Commun. | 4 |
| 2017 | Improving protein disorder prediction by deep bidirectional long short-term memory recurrent neural networksabstractMotivation: Capturing long-range interactions between structural but not sequence neighbors of proteins is a long-standing challenging problem in bioinformatics. Recently, long short-term memory (LSTM) networks have significantly improved the accuracy of speech and image classification problems by remembering useful past information in long sequential events. Here, we have implemented deep bidirectional LSTM recurrent neural networks in the problem of protein intrinsic disorder prediction. Results: The new method, named SPOT-Disorder, has steadily improved over a similar method using a traditional, window-based neural network (SPINE-D) in all datasets tested without separate training on short and long disordered regions. Independent tests on four other datasets including the datasets from critical assessment of structure prediction (CASP) techniques and >10 000 annotated proteins from MobiDB, confirmed SPOT-Disorder as one of the best methods in disorder prediction. Moreover, initial studies indicate that the method is more accurate in predicting functional sites in disordered regions. These results highlight the usefulness combining LSTM with deep bidirectional recurrent neural networks in capturing non-local, long-range interactions for bioinformatics applications. Availability and Implementation: SPOT-disorder is available as a web server and as a standalone program at: http://sparks-lab.org/server/SPOT-disorder/index.php . Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data is available at Bioinformatics online. Jack Hanson, Yuedong Yang, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 3 |
| 2017 | Capturing non-local interactions by long short-term memory bidirectional recurrent neural networks for improving prediction of protein secondary structure, backbone angles, contact numbers and solvent accessibilityabstractMOTIVATION: The accuracy of predicting protein local and global structural properties such as secondary structure and solvent accessible surface area has been stagnant for many years because of the challenge of accounting for non-local interactions between amino acid residues that are close in three-dimensional structural space but far from each other in their sequence positions. All existing machine-learning techniques relied on a sliding window of 10-20 amino acid residues to capture some 'short to intermediate' non-local interactions. Here, we employed Long Short-Term Memory (LSTM) Bidirectional Recurrent Neural Networks (BRNNs) which are capable of capturing long range interactions without using a window. RESULTS: We showed that the application of LSTM-BRNN to the prediction of protein structural properties makes the most significant improvement for residues with the most long-range contacts (|i-j| >19) over a previous window-based, deep-learning method SPIDER2. Capturing long-range interactions allows the accuracy of three-state secondary structure prediction to reach 84% and the correlation coefficient between predicted and actual solvent accessible surface areas to reach 0.80, plus a reduction of 5%, 10%, 5% and 10% in the mean absolute error for backbone ϕ , ψ , θ and τ angles, respectively, from SPIDER2. More significantly, 27% of 182724 40-residue models directly constructed from predicted C α atom-based θ and τ have similar structures to their corresponding native structures (6Å RMSD or less), which is 3% better than models built by ϕ and ψ angles. We expect the method to be useful for assisting protein structure and function prediction. AVAILABILITY AND IMPLEMENTATION: The method is available as a SPIDER3 server and standalone package at http://sparks-lab.org . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rhys Heffernan, Yuedong Yang, Kuldip K. Paliwal, Yaoqi Zhou |
Bioinform. | 3 |
| 2016 | Highly accurate sequence-based prediction of half-sphere exposures of amino acid residues in proteinsabstractMOTIVATION: Solvent exposure of amino acid residues of proteins plays an important role in understanding and predicting protein structure, function and interactions. Solvent exposure can be characterized by several measures including solvent accessible surface area (ASA), residue depth (RD) and contact numbers (CN). More recently, an orientation-dependent contact number called half-sphere exposure (HSE) was introduced by separating the contacts within upper and down half spheres defined according to the Cα-Cβ (HSEβ) vector or neighboring Cα-Cα vectors (HSEα). HSEα calculated from protein structures was found to better describe the solvent exposure over ASA, CN and RD in many applications. Thus, a sequence-based prediction is desirable, as most proteins do not have experimentally determined structures. To our best knowledge, there is no method to predict HSEα and only one method to predict HSEβ. RESULTS: This study developed a novel method for predicting both HSEα and HSEβ (SPIDER-HSE) that achieved a consistent performance for 10-fold cross validation and two independent tests. The correlation coefficients between predicted and measured HSEβ (0.73 for upper sphere, 0.69 for down sphere and 0.76 for contact numbers) for the independent test set of 1199 proteins are significantly higher than existing methods. Moreover, predicted HSEα has a higher correlation coefficient (0.46) to the stability change by residue mutants than predicted HSEβ (0.37) and ASA (0.43). The results, together with its easy Cα-atom-based calculation, highlight the potential usefulness of predicted HSEα for protein structure prediction and refinement as well as function prediction. AVAILABILITY AND IMPLEMENTATION: The method is available at http://sparks-lab.org CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rhys Heffernan, Abdollah Dehzangi, James G. Lyons, Kuldip K. Paliwal, Alok Sharma, Jihua Wang, Abdul Sattar 0001, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 4 |
| 2016 | Phase distortion resulting in a just noticeable difference in the perceived quality of speech
Roger Chappel, Belinda Schwerin, Kuldip K. Paliwal |
Speech Commun. | 3 |
| 2015 | Gram-positive and gram-negative subcellular localization using rotation forest and physicochemical-based featuresabstractBACKGROUND: The functioning of a protein relies on its location in the cell. Therefore, predicting protein subcellular localization is an important step towards protein function prediction. Recent studies have shown that relying on Gene Ontology (GO) for feature extraction can improve the prediction performance. However, for newly sequenced proteins, the GO is not available. Therefore, for these cases, the prediction performance of GO based methods degrade significantly. RESULTS: In this study, we develop a method to effectively employ physicochemical and evolutionary-based information in the protein sequence. To do this, we propose segmentation based feature extraction method to explore potential discriminatory information based on physicochemical properties of the amino acids to tackle Gram-positive and Gram-negative subcellular localization. We explore our proposed feature extraction techniques using 10 attributes that have been experimentally selected among a wide range of physicochemical attributes. Finally by applying the Rotation Forest classification technique to our extracted features, we enhance Gram-positive and Gram-negative subcellular localization accuracies up to 3.4% better than previous studies which used GO for feature extraction. CONCLUSION: By proposing segmentation based feature extraction method to explore potential discriminatory information based on physicochemical properties of the amino acids as well as using Rotation Forest classification technique, we are able to enhance the Gram-positive and Gram-negative subcellular localization prediction accuracies, significantly. Abdollah Dehzangi, Sohrab Sohrabi, Rhys Heffernan, Alok Sharma, James G. Lyons, Kuldip K. Paliwal, Abdul Sattar 0001 |
BMC Bioinform. | 6 |
| 2015 | A deterministic approach to regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal |
Neurocomputing | 2 |
| 2014 | Improving protein fold recognition using the amalgamation of evolutionary-based and structural based informationabstractDeciphering three dimensional structure of a protein sequence is a challenging task in biological science. Protein fold recognition and protein secondary structure prediction are transitional steps in identifying the three dimensional structure of a protein. For protein fold recognition, evolutionary-based information of amino acid sequences from the position specific scoring matrix (PSSM) has been recently applied with improved results. On the other hand, the SPINE-X predictor has been developed and applied for protein secondary structure prediction. Several reported methods for protein fold recognition have only limited accuracy. In this paper, we have developed a strategy of combining evolutionary-based information (from PSSM) and predicted secondary structure using SPINE-X to improve protein fold recognition. The strategy is based on finding the probabilities of amino acid pairs (AAP). The proposed method has been tested on several protein benchmark datasets and an improvement of 8.9% recognition accuracy has been achieved. We have achieved, for the first time over 90% and 75% prediction accuracies for sequence similarity values below 40% and 25%, respectively. We also obtain 90.6% and 77.0% prediction accuracies, respectively, for the Extended Ding and Dubchak and Taguchi and Gromiha benchmark protein fold recognition datasets widely used for in the literature. Kuldip K. Paliwal, Alok Sharma, James G. Lyons, Abdollah Dehzangi |
BMC Bioinform. | 1 |
| 2014 | A feature selection method using improved regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal, Seiya Imoto, Satoru Miyano |
Mach. Vis. Appl. | 2 |
| 2014 | An educational platform to demonstrate speech processing techniques on Android based smart phones and tablets
Roger Chappel, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2014 | An improved speech transmission index for intelligibility prediction
Belinda Schwerin, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2014 | Using STFT real and imaginary parts of modulation signals for MMSE-based speech enhancement
Belinda Schwerin, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2014 | A Segmentation-Based Method to Extract Structural and Evolutionary Features for Protein Fold RecognitionabstractProtein fold recognition (PFR) is considered as an important step towards the protein structure prediction problem. Despite all the efforts that have been made so far, finding an accurate and fast computational approach to solve the PFR still remains a challenging problem for bioinformatics and computational biology. In this study, we propose the concept of segmented-based feature extraction technique to provide local evolutionary information embedded in position specific scoring matrix (PSSM) and structural information embedded in the predicted secondary structure of proteins using SPINE-X. We also employ the concept of occurrence feature to extract global discriminatory information from PSSM and SPINE-X. By applying a support vector machine (SVM) to our extracted features, we enhance the protein fold prediction accuracy for 7.4 percent over the best results reported in the literature. We also report 73.8 percent prediction accuracy for a data set consisting of proteins with less than 25 percent sequence similarity rates and 80.7 percent prediction accuracy for a data set with proteins belonging to 110 folds with less than 40 percent sequence similarity rates. We also investigate the relation between the number of folds and the number of features being used and show that the number of features should be increased to get better protein fold prediction results when the number of folds is relatively large. Abdollah Dehzangi, Kuldip K. Paliwal, James G. Lyons, Alok Sharma, Abdul Sattar 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | A strategy to select suitable physicochemical attributes of amino acids for protein fold recognitionabstractBACKGROUND: Assigning a protein into one of its folds is a transitional step for discovering three dimensional protein structure, which is a challenging task in bimolecular (biological) science. The present research focuses on: 1) the development of classifiers, and 2) the development of feature extraction techniques based on syntactic and/or physicochemical properties. RESULTS: Apart from the above two main categories of research, we have shown that the selection of physicochemical attributes of the amino acids is an important step in protein fold recognition and has not been explored adequately. We have presented a multi-dimensional successive feature selection (MD-SFS) approach to systematically select attributes. The proposed method is applied on protein sequence data and an improvement of around 24% in fold recognition has been noted when selecting attributes appropriately. CONCLUSION: The MD-SFS has been applied successfully in selecting physicochemical attributes of the amino acids. The selected attributes show improved protein fold recognition performance. Alok Sharma, Kuldip K. Paliwal, Abdollah Dehzangi, James G. Lyons, Seiya Imoto, Satoru Miyano |
BMC Bioinform. | 2 |
| 2012 | Speech Enhancement for Android (SEA): A Speech Processing Demonstration Tool for Android Based Smart Phones and TabletsabstractThis paper presents a speech processing platform which can be used to demonstrate and investigate speech enhancement methods. This platform is called Speech Enhancement for Android (SEA), and has been developed on the Android operating system, available to students, teaching staff and researchers through the Android market. SEA can be used as an additional teaching tool in undergraduate courses as well as a useful tool to aid researchers in gaining an intuitive understanding of speech enhancement methods. The focus of this platform is to present advanced speech processing concepts in a quick and interactive way on a personal phone or tablet. This paper outlines the operation of SEA along with how it can be used to engage students and speech processing professionals. Roger Chappel, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2012 | Improved Pseudoinverse Linear Discriminant Analysis Method for Dimensionality ReductionabstractPseudoinverse linear discriminant analysis (PLDA) is a classical method for solving small sample size problem. However, its performance is limited. In this paper, we propose an improved PLDA method which is faster and produces better classification accuracy when experimented on several datasets. Kuldip K. Paliwal, Alok Sharma |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | A new perspective to null linear discriminant analysis method and its fast implementation using random matrix multiplication with scatter matrices
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. | 2 |
| 2012 | A two-stage linear discriminant analysis for face-recognition
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 2 |
| 2012 | Improving objective intelligibility prediction by combining correlation and coherence based methods with a measure based on the negative distortion ratio
Ángel M. Gómez, Belinda Schwerin, Kuldip K. Paliwal |
Speech Commun. | 3 |
| 2012 | Speech enhancement using a minimum mean-square error short-time spectral modulation magnitude estimator
Kuldip K. Paliwal, Belinda Schwerin, Kamil K. Wójcicki |
Speech Commun. | 1 |
| 2011 | Objective Intelligibility Prediction of Speech by Combining Correlation and Distortion Based TechniquesabstractA number of techniques based on correlation measurements have recently been proposed to provide an objective measure of intelligibility. These techniques are able to detect nonlinear distortions and provide intelligibility scores highly correlated with those given by human listeners. However, the performance of these techniques has not been found satisfactory for measuring the speech intelligibility of speech enhancement algorithms. In this paper we first investigate the different correlation-based methods, in the context of speech enhancement. We then propose to combine these correlation-based techniques with spectral distance based ones. Results presented show that objective intelligibility prediction is significantly improved by this combination. Ángel M. Gómez, Belinda Schwerin, Kuldip K. Paliwal |
INTERSPEECH | 3 |
| 2011 | Single Channel Speech Enhancement Using MMSE Estimation of Short-Time Modulation Magnitude SpectrumabstractIn this paper we investigate the enhancement of speech by applying MMSE short-time spectral magnitude estimation in the modulation domain. For this purpose, the traditional analysismodification- synthesis framework is extended to include modulation domain processing. We compensate the noisy modulation spectrum for additive noise distortion by applying the MMSE short-time spectral magnitude estimation algorithm in the modulation domain. Subjective experiments were conducted to compare the quality of stimuli processed by the MMSE modulation magnitude estimator to those processed using the MMSE acoustic magnitude estimator and the modulation spectral subtraction method. The proposed method is shown to have better noise suppression than MMSE acoustic magnitude estimation, and improved speech quality compared to modulation domain spectral subtraction. Kuldip K. Paliwal, Belinda Schwerin, Kamil K. Wójcicki |
INTERSPEECH | 1 |
| 2011 | Role of modulation magnitude and phase spectrum towards speech intelligibility
Kuldip K. Paliwal, Belinda Schwerin, Kamil K. Wójcicki |
Speech Commun. | 1 |
| 2011 | The importance of phase in speech enhancement
Kuldip K. Paliwal, Kamil K. Wójcicki, Benjamin J. Shannon |
Speech Commun. | 1 |
| 2011 | Suppressing the influence of additive noise on the Kalman gain for low residual noise speech enhancement
Stephen So, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2011 | Modulation-domain Kalman filtering for single-channel speech enhancement
Stephen So, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2011 | Use of speech presence uncertainty with MMSE spectral energy estimation for robust automatic speech recognition
Anthony P. Stark, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2011 | MMSE estimation of log-filterbank energies for robust speech recognition
Anthony P. Stark, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2010 | Fast converging iterative kalman filtering for speech enhancement using long and overlapped tapered windows with large side lobe attenuationabstractIn this paper, we propose an iterative Kalman filtering scheme that has faster convergence and introduces less residual noise, when compared with the iterative scheme of Gibson, et al. This is achieved via the use of long and overlapped frames as well as using a tapered window with a large side lobe attenuation for linear prediction analysis. We show that the Dolph-Chebychev window with a −200 dB side lobe attenuation tends to enhance the dynamic range of the formant structure of speech corrupted with white noise, reduce prediction error variance bias, as well as provide for some spectral smoothing, while the long overlapped frames provide for reliable autocorrelation estimates and temporal smoothing. Speech enhancement experiments on the NOIZEUS corpus show that the proposed method outperformed conventional iterative and non-iterative Kalman filters as well as other enhancement methods such as MMSE-STSA and PSC. Index Terms: speech enhancement, Kalman filtering 1. Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2010 | Single-channel speech enhancement using kalman filtering in the modulation domainabstractIn this paper, we propose the modulation-domain Kalman filter (MDKF) for speech enhancement. In contrast to previous modulation domain enhancement methods based on bandpass filtering, the MDKF is an adaptive and linear MMSE estimator that uses models of the temporal changes of the magnitude spectrum for both speech and noise. Also, because the Kalman filter is a joint magnitude and phase spectrum estimator, under non-stationarity assumptions, it is highly suited for modulation-domain processing, as modulation phase tends to contain more speech information than acoustic phase. Experimental results from the NOIZEUS corpus show the ideal MDKF (with clean speech parameters) to outperform all the acoustic and time-domain enhancement methods that were evaluated, including the conventional time-domain Kalman filter with clean speech parameters. A practical MDKF that uses the MMSE-STSA method to enhance noisy speech in the acoustic domain prior to LPC analysis was also evaluated and showed promising results. Stephen So, Kamil K. Wójcicki, Kuldip K. Paliwal |
INTERSPEECH | 3 |
| 2010 | Improved direct LDA and its application to DNA microarray gene expression data
Kuldip K. Paliwal, Alok Sharma |
Pattern Recognit. Lett. | 1 |
| 2010 | Single-channel speech enhancement using spectral subtraction in the short-time modulation domain
Kuldip K. Paliwal, Kamil K. Wójcicki, Belinda Schwerin |
Speech Commun. | 1 |
| 2009 | Kalman fitler with phase spectrum compensation algorithm for speech enhancementabstractIn this paper, we propose to combine the Kalman filter with a recent speech enhancement technique, called the phase spectrum compensation procedure, or PSC. More specifically, we apply the PSC technique to initialise the Kalman filter, whereby PSC is used to clean the noisy speech prior to LPC estimation for the Kalman recursion. We refer to the combined technique as the Kalman-PSC filter. Using an objective speech quality measure, formal subjective listening tests and spectrogram analysis, we show that the proposed method results in improved speech quality. Stephen So, Kamil K. Wójcicki, James G. Lyons, Anthony P. Stark, Kuldip K. Paliwal |
ICASSP | 5 |
| 2009 | Modulation domain spectral subtraction for speech enhancementabstractIn this paper we investigate the modulation domain as an alternative to the acoustic domain for speech enhancement. More specifically, we wish to determine how competitive the modulation domain is for spectral subtraction as compared to the acoustic domain. For this purpose, we extend the traditional analysis-modification-synthesis framework to include modulation domain processing. We then compensate the noisy modulation spectrum for additive noise distortion by applying the spectral subtraction algorithm in the modulation domain. Using subjective listening tests and objective speech quality evaluation we show that the proposed method results in improved speech quality. Furthermore, applying spectral subtraction in the modulation domain does not introduce the musical noise artifacts that are typically present after acoustic domain spectral subtraction. The proposed method also achieves better background noise reduction than the MMSE method. Index Terms: speech enhancement, spectral subtraction, modulation domain, analysis-modification-synthesis (AMS) Kuldip K. Paliwal, Belinda Schwerin, Kamil K. Wójcicki |
INTERSPEECH | 1 |
| 2009 | Group-delay-deviation based spectral analysis of speechabstractIn this paper, we investigate a new method for extracting useful information from the group delay spectrum of speech. The group delay spectrum is often poorly behaved and noisy. In the literature, various methods have been proposed to address this problem. However, to make the group delay a more tractable function, these methods have typically relied upon some modification of the underlying speech signal. The method proposed in this paper does not require such modifications. To accomplish this, we investigate a new function derived from the group delay spectrum, namely the group delay deviation.We use it for both narrowband analysis and wideband analysis of speech and show that this function exhibits meaningful formant and pitch information. Index Terms: group delay deviation, spectral analysis, speech processing. Anthony P. Stark, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2009 | Special issue on non-linear and non-conventional speech processing
Mohamed Chetouani, Marcos Faúndez-Zanuy, Amir Hussain 0001, Bruno Gas, Jean-Luc Zarader, Kuldip K. Paliwal |
Speech Commun. | 6 |
| 2009 | Speech-Signal-Based Frequency WarpingabstractThe speech signal is used for transmission of linguistic information. High energy portions of the speech spectrum have higher signal-to-noise ratios than the low energy portions. As a result, these regions are more robust to noise. Since the speech signal is known to be very robust to noise, it is expected that the high energy regions of the speech spectrum carry the majority of the linguistic information. This letter tries to derive a frequency warping function directly from the speech signal by sampling the frequency axis nonuniformly with the high energy regions sampled more densely than the low energy regions. To achieve this, an ensemble average short-time power spectrum is computed from a large speech corpus. The speech-signal-based frequency warping is obtained by considering equal area portions of the log spectrum. The proposed frequency warping is shown to be similar to the frequency scales obtained through psycho-acoustic experiments, namely the mel and bark scales. The warping is then used in filterbank design for automatic speech recognition experiments. The results of these experiments show that cepstral features based on the proposed warping achieve performance under clean conditions comparable to that of mel-frequency cepstral coefficients, while outperforming them under noisy conditions. Kuldip K. Paliwal, Benjamin J. Shannon, James G. Lyons, Kamil K. Wójcicki |
IEEE Signal Process. Lett. | 1 |
| 2008 | Effect of compressing the dynamic range of the power spectrum in modulation filtering based speech enhancementabstractGriffith Sciences, Griffith School of Engineering James G. Lyons, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2008 | A long state vector kalman filter for speech enhancementabstractIn this paper, we investigate a long state vector Kalman filter for the enhancement of speech that has been corrupted by white and coloured noise. It has been reported in previous studies that a vector Kalman filter achieves better enhancement than the scalar Kalman filter and it is expected that by increasing the state vector length, one may improve the enhancement performance even further. However, any enhancement improvement that may result from an increase in state vector length is constrained by the typical use of short, non-overlapped speech frames, as the autocorrelation coefficient estimates tend to become less reliable at higher lags. We propose to overcome this problem by incorporating an analysis-modificationsynthesis framework, where long, overlapped frames are used instead. Our enhancement experiments based on the NOIZEUS corpus show that the proposed long state vector Kalman filter achieves higher mean SNR and PESQ scores than the scalar and short state vector Kalman filter, therefore fulfilling the notion that a longer state vector can lead to better enhancement. Index Terms: Kalman filtering, speech enhancement, analysismodification-synthesis 1. Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2008 | Speech analysis using instantaneous frequency deviationabstractGriffith Sciences, Griffith School of Engineering Anthony P. Stark, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2008 | Noise driven short-time phase spectrum compensation procedure for speech enhancementabstractTypical speech enhancement algorithms operate on the shorttime magnitude spectrum, while keeping the short-time phase spectrum unchanged for synthesis. Recently, a novel approach to speech enhancement has been proposed where the noisy magnitude spectrum is recombined with a changed phase spectrum to produce a modified complex spectrum. During synthesis the low energy components of the modified complex spectrum cancel out more than the high energy components, thus reducing background noise. In the present work, a procedure that employs noise estimates to compensate the phase spectrum for additive noise distortion is formulated. The proposed approach is objectively evaluated against several popular speech enhancement methods under various noise conditions and is shown to compare favourably. Index Terms: speech enhancement, magnitude spectrum, phase spectrum, phase spectrum compensation Anthony P. Stark, Kamil K. Wójcicki, James G. Lyons, Kuldip K. Paliwal |
INTERSPEECH | 4 |
| 2008 | A response generation in the Mongolian spoken language system for accessing to multimedia knowledge baseabstractBy using automatic speech recognition (ASR) and text to speech (TTS) systems, which have been available in Mongolian for last few years, this research set out to implement a new version of the Mongolian Virtual Education Environment (VEE) that has not included a speech interface. The spoken language system aims to provide a natural interface between trainees and the environment by using simple and natural dialogues to enable the user to access the multimedia knowledge base of the VEE. We have worked on the response generation part of the system. This paper describes a TTS system for the VEE for university courses held in Mongolian. A concatenative speech synthesizer for Mongolian is applied for the TTS in response generation. A Festvox framework for unit selection speech synthesis was used to build the Mongolian voice. We discuss aspects of the voice development process and the results of a perceptual test of the synthesized voice. Munkhtuya Davaatsagaan, Kuldip K. Paliwal |
SLT | 2 |
| 2008 | Cancer classification by gradient LDA technique using microarray gene expression data
Alok Sharma, Kuldip K. Paliwal |
Data Knowl. Eng. | 2 |
| 2008 | A Gradient Linear Discriminant Analysis for Small Sample Sized Problem
Alok Sharma, Kuldip K. Paliwal |
Neural Process. Lett. | 2 |
| 2008 | Special Issue on Iberian Languages
Isabel Trancoso, Néstor Becerra Yoma, Plínio Almeida Barbosa, Rubén San-Segundo-Hernández, Kuldip K. Paliwal |
Speech Commun. | 5 |
| 2008 | Effect of Analysis Window Duration on Speech IntelligibilityabstractIn this letter, we investigate the effect of the analysis window duration on speech intelligibility in a systematic way. In speech processing, the short-time magnitude spectrum is believed to contain the majority of the intelligible information. Consequently, in our experiments, we construct speech stimuli based purely on the short-time magnitude spectrum. We conduct subjective listening tests in the form of a consonant recognition task to assess intelligibility as a function of analysis window duration. In our investigations, we also employ three objective speech intelligibility measures based on the speech transmission index (STI). The experimental results show that the analysis window duration of 15-35 ms is the optimum choice when speech is reconstructed from the short-time magnitude spectrum. Kuldip K. Paliwal, Kamil K. Wójcicki |
IEEE Signal Process. Lett. | 1 |
| 2008 | Exploiting Conjugate Symmetry of the Short-Time Fourier Spectrum for Speech EnhancementabstractTypical speech enhancement algorithms operate on the short-time magnitude spectrum, while keeping the short-time phase spectrum unchanged for synthesis. We propose a novel approach where the noisy magnitude spectrum is recombined with a changed phase spectrum to produce a modified complex spectrum. During synthesis, the low energy components of the modified complex spectrum cancel out more than the high energy components, thus reducing background noise. Using objective speech quality measures, informal subjective listening tests and spectrogram analysis, we show that the proposed method results in improved speech quality. Kamil K. Wójcicki, Mitar Milacic, Anthony P. Stark, James G. Lyons, Kuldip K. Paliwal |
IEEE Signal Process. Lett. | 5 |
| 2008 | Rotational Linear Discriminant Analysis Technique for Dimensionality ReductionabstractThe linear discriminant analysis (LDA) technique is very popular in pattern recognition for dimensionality reduction. It is a supervised learning technique that finds a linear transformation such that the overlap between the classes is minimum for the projected feature vectors in the reduced feature space. This overlap, if present, adversely affects the classification performance. In this paper, we introduce prior to dimensionality-reduction transformation an additional rotational transform that rotates the feature vectors in the original feature space around their respective class centroids in such a way that the overlap between the classes in the reduced feature space is further minimized. As a result, the classification performance significantly improves, which is demonstrated using several data corpuses. Alok Sharma, Kuldip K. Paliwal |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | Effect of Speech and Noise Cross Correlation on AMFCC Speech Recognition FeaturesabstractWhen designing noise robust speech recognition feature extraction algorithms, it is common to assume that the noise and speech signal are uncorrelated. This assumption allows the cross correlation terms to be ignored in the equations that describe the operation of these algorithms, thus making the mathematics more tractable. In this paper, we investigate the validity of this assumption in the context of the autocorrelation mel frequency cepstral coefficient (AMFCC) feature extraction algorithm. To carry out the investigation, we designed a modified AMFCC algorithm that forces the cross terms in the noisy signal autocorrelation equation to be zero. We then compared the performance of the modified algorithm to the un-modified algorithm in recognition experiments performed using the AURORA II database. From these evaluations, we show that the assumption is fair in 5 out of six tested noise cases. The difference in recognition accuracy between the AMFCC and modified AMFCC for these five noises was less than 5%. Benjamin J. Shannon, Kuldip K. Paliwal |
ICASSP (4) | 2 |
| 2007 | Importance of the Dynamic Range of an Analysis Windowfunction for Phase-Only and Magnitude-Only Reconstruction of SpeechabstractThe short-time Fourier transform (STFT) of a speech signal has two components: the short-time magnitude spectrum and the short-time phase spectrum. It is traditionally believed that the short-time magnitude spectrum plays the dominant role for speech perception at small window durations (20-40 ms). However, recent perceptual studies have shown that the short-time phase spectrum can contribute as much to speech intelligibility as the short-time magnitude spectrum. It was observed that the use of the rectangular (non-tapered) analysis window for the computation of the short-time phase spectrum is more advantageous than the use of the Hamming (tapered) analysis window. This paper investigates the effect that the dynamic range of an analysis window has on the intelligibility of speech for phase-only and magnitude-only stimuli. For this purpose, the Chebyshev analysis window with adjustable equi-ripple side-lobes is employed. Two types of magnitude-only stimuli are investigated: random phase and zero phase. It is shown that the intelligibility of the magnitude-only stimuli constructed with zero phase is independent of the dynamic range of the analysis window, while the random phase stimuli are intelligible only for analysis windows with high dynamic range. This study also shows that for low dynamic range analysis windows, the short-time phase spectrum at small window durations (20-40 ms) contributes as much as to speech intelligibility as the short-time magnitude spectrum. Kamil K. Wójcicki, Kuldip K. Paliwal |
ICASSP (4) | 2 |
| 2007 | The effect of the additivity assumption on time and frequency domain wiener filtering for speech enhancementabstractIn this paper, we investigate the validity of the common assumption made in Wiener filtering that the clean speech and noise signals are uncorrelated under short-time analysis typically used for speech enhancement. In order to achieve this we have performed speech enhancement experiments, where speech corrupted by additive white Gaussian noise is enhanced by a Wiener filter designed in the time as well as the frequency domains. Results of oracle-style experiments confirm that the inclusion of the additivity assumption in Wiener filtering results in negligible degradation of enhanced speech quality. Informal listening tests show that the background noise resulting from time domain enhancement to be more tolerable than the background noise resulting from frequency domain framework. Index Terms: Wiener filtering, speech enhancement 1. Kamil K. Wójcicki, Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 3 |
| 2007 | Intrusion detection using text processing techniques with a kernel based similarity measure
Alok Sharma, Arun K. Pujari, Kuldip K. Paliwal |
Comput. Secur. | 3 |
| 2007 | Iterative reconstruction of speech from short-time Fourier transform phase and magnitude spectra
Leigh D. Alsteris, Kuldip K. Paliwal |
Comput. Speech Lang. | 2 |
| 2007 | Fast principal component analysis using fixed-point algorithm
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 2 |
| 2007 | Special issue on Speech Enhancement
Philipos C. Loizou, Israel Cohen, Sharon Gannot, Kuldip K. Paliwal |
Speech Commun. | 4 |
| 2006 | Multi-Frame GMM-Based Block Quantisation for Distributed Speech Recognition Under Noisy ConditionsabstractIn this paper, we report on the recognition accuracy of the multi-frame GMM-based block quantiser for the coding of MFCC features in a distributed speech recognition framework under varying noise conditions. All experiments were performed using the ETSI Aurora-2 connected-digits recognition task. For comparison, we have also investigated other quantisation schemes such as the memoryless GMM-based block quantiser, the unconstrained vector quantiser, and non-uniform scalar quantisers. The results show that the rate-distortion efficiency of the quantiser is a factor in determining the level of recognition accuracy at low to medium levels of additive noise. For high levels of additive noise, the influence of rate-distortion efficiency diminishes and the recognition accuracy becomes dependent on the recognition features Stephen So, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2006 | Role of phase estimation in speech enhancementabstractTypical speech enhancement algorithms that operate in the Fourier domain only modify the magnitude component. It is commonly understood that the phase component is perceptually unimportant, and thus, it is passed directly to the output.In recent intelligibility experiments, it has been reported that the Short-Time Fourier Transform (STFT) phase spectrum can provide significant intelligibility when estimated using a window function lower in dynamic range than the typical Hamming window. Motivated by this, we investigate the role of the window function for STFT phase estimation in relation to speech enhancement.Using a modified STFT Analysis-Modification-Synthesis (AMS) framework, we show that noise reduction can be achieved by modifying the window function used to estimate the STFT phase spectra. We demonstrate this through spectrogram plots and results from two objective speech quality measures. Benjamin J. Shannon, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2006 | Speech enhancement based on spectral estimation from higher-lag autocorrelationabstractIn this paper, we propose a unique approach to enhance speech signals that have been corrupted by non-stationary noises. This approach is not based on a spectral subtraction algorithm, but on an algorithm that separates the speech signal and noise signal contributions in the autocorrelation domain. We call this technique the AR-HASE speech enhancement algorithm. In this initial study, we evaluate the performance of the new algorithm using the average PESQ score computed from 10 male utterances and 10 female utterances taken from the TIMIT database as a measure of speech quality. We test the algorithm using one broadband stationary noise and two non-stationary noises. We will show that the AR-HASE enhancement algorithm produces near transparent quality for clean speech, gives poor enhancement performance for broadband stationary noises, and gives significantly enhanced quality for the two non-stationary noises. Index Terms: speech enhancement, autocorrelation, impulsive noise. Benjamin J. Shannon, Kuldip K. Paliwal, Climent Nadeu |
INTERSPEECH | 2 |
| 2006 | Subspace independent component analysis using vector kurtosis
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. | 2 |
| 2006 | Class-dependent PCA, MDC and LDA: A combined classifier for pattern classification
Alok Sharma, Kuldip K. Paliwal, Godfrey C. Onwubolu |
Pattern Recognit. | 2 |
| 2006 | Further intelligibility results from human listening tests using the short-time phase spectrum
Leigh D. Alsteris, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2006 | Feature extraction from higher-lag autocorrelation coefficients for robust speech recognition
Benjamin J. Shannon, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2006 | Scalable distributed speech recognition using Gaussian mixture model-based block quantisation
Stephen So, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2006 | Empirical Lower Bound on the Bitrate for the Transparent Memoryless Coding of Wideband LPC ParametersabstractIn this letter, we determine empirical lower bounds on the bitrate required to transparently code linear predictive coding (LPC) parameters derived from wideband speech. This is achieved via extrapolation of the operating distortion-rate curve of an unconstrained vector quantizer that is trained using artificial vectors generated by a Gaussian mixture model. Memoryless coding is considered and two competing LPC parameter representations are investigated. Our results show a lower bound of 31 bits/frame when assuming high-rate linearity in the operating distortion-rate curve and 35 bits/frame for an exponential curve. We also evaluate a recent quantization scheme and compare its performance against this lower bound Stephen So, Kuldip K. Paliwal |
IEEE Signal Process. Lett. | 2 |
| 2006 | Robust speech recognition in noisy environments based on subband spectral centroid histogramsabstractWe investigate how dominant-frequency information can be used in speech feature extraction to increase the robustness of automatic speech recognition against additive background noise. First, we review several earlier proposed auditory-based feature extraction methods and argue that the use of dominant-frequency information might be one of the major reasons for their improved noise robustness. Furthermore, we propose a new feature extraction method, which combines subband power information with dominant subband frequency information in a simple and computationally efficient way. The proposed features are shown to be considerably more robust against additive background noise than standard mel-frequency cepstrum coefficients on two different recognition tasks. The performance improvement increased as we moved from a small-vocabulary isolated-word task to a medium-vocabulary continuous-speech task, where the proposed features also outperformed a computationally expensive auditory-based method. The greatest improvement was obtained for noise types characterized by a relatively flat spectral density. Bojana Gajic, Kuldip K. Paliwal |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Influence of Autocorrelation Lag Ranges on Robust Speech RecognitionabstractIt is generally believed that the lower-lag autocorrelation coefficients carry information about the spectral envelope and the higher-lag autocorrelation coefficients are more related to pitch information. In this paper, we use lower-lag and higher-lag ranges of the autocorrelation function separately for deriving speech recognition features, and investigate their role in terms of speech recognition performance. The state-of-the-art MFCC (mel frequency cepstral coefficient) features use the whole autocorrelation function in their computation and are used here as a benchmark in our experiments. Our recognition results from the Aurora II corpus show that the higher-lag autocorrelation coefficients perform as well as the whole autocorrelation function for clean speech, and provide better performance for noisy speech, while lower-lag autocorrelation coefficients are not as effective in this aspect. Benjamin J. Shannon, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2005 | Multi-Frame GMM-Based Block Quantisation of Line Spectral Frequencies for Wideband Speech CodingabstractWe explore the use of the multi-frame GMM-based block quantiser for quantising line spectral frequencies for wideband speech coding. Its main advantages over vector quantisers are bitrate scalability and bitrate independent complexity. By concatenating multiple frames together, interframe correlation can be exploited by the KLT (Karhunen-Loeve transform), leading to better quantisation. A saving of up to 3 bits/frame can be achieved by switching the quantiser from memoryless mode to jointly quantising two frames, with only a moderate increase in complexity. This quantisation scheme achieves lower spectral distortion than the split-multistage vector quantiser in the AMR-WB speech codec, with transparent coding at 37 bits/frame. Stephen So, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2005 | Some experiments on iterative reconstruction of speech from STFT phase and magnitude spectraabstractGriffith Sciences, Griffith School of Engineering Leigh D. Alsteris, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2005 | Improved noise-robustness in distributed speech recognition via perceptually-weighted vector quantisation of filterbank energiesabstractIn this paper, we examine a coding scheme for quantising feature vectors in a distributed speech recognition environment that is more robust to noise. It consists of a vector quantiser that operates on the logarithmic filterbank energies (LFBEs). Through the use of a perceptually-weighted Euclidean distance measure, which emphasises the LFBEs that represent the spectral peaks, the vector quantiser codebook provides /emph{a priori} knowledge of the spectral characteristics of clean speech and is used to quantise features from noise-corrupted speech. Our comparative results from the ETSI Aurora-2 recognition task show that the perceptually-weighted vector quantisation of LFBEs achieves higher recognition accuracies for noisy speech than the unweighted vector quantisation, memoryless and multi-frame GMM-based block quantisation and scalar quantisation of Mel frequency-warped cepstral coefficients. Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2005 | Switched split vector quantisation of line spectral frequencies for wideband speech codingabstractIn this paper, we investigate the use of the switched split vector quantiser (SSVQ) for coding short-term spectral envelope information for wideband speech coding. The SSVQ is the hybrid of a switch vector quantiser and split vector quantiser, which has been shown in previous studies to be more efficient, in terms of rate-distortion, as well as possessing low computational complexity, than the split vector quantiser. In our experiments, the SSVQ is used to quantise line spectral frequencies from the TIMIT database and its spectral distortion performance is compared with the split vector quantiser, the split-multistage vector quantiser (S-MSVQ) with MA predictor from the AMR-WB speech coder (ITU-T G.722.2), and PDF-optimised scalar quantisers. We show the SSVQ, which is a memoryless scheme, to achieve comparable spectral distortion to the S-MSVQ with MA predictor at 46 bits/frame. The five-part SSVQ requires 42 bits/frame and 17.7 kflops/frame for transparent coding, compared with 46 bits/frame and 40.96 kflops/frame for the five-part split vector quantiser. 1. Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2005 | On the usefulness of STFT phase spectrum in human listening tests
Kuldip K. Paliwal, Leigh D. Alsteris |
Speech Commun. | 1 |
| 2005 | Multi-frame GMM-based block quantisation of line spectral frequencies
Stephen So, Kuldip K. Paliwal |
Speech Commun. | 2 |
| 2005 | Generative factor analyzed HMM for automatic speech recognition
Kaisheng Yao, Kuldip K. Paliwal, Te-Won Lee |
Speech Commun. | 2 |
| 2005 | Maximum likelihood sub-band adaptation for robust speech recognition
Donglai Zhu, Satoshi Nakamura 0001, Kuldip K. Paliwal, Renhua Wang |
Speech Commun. | 3 |
| 2005 | Low-complexity GMM-based block quantisation of images using the discrete cosine transform
Kuldip K. Paliwal, Stephen So |
Signal Process. Image Commun. | 1 |
| 2004 | Importance of window shape for phase-only reconstruction of speechabstractThe authors recently conducted a human perception experiment to measure the intelligibility of speech stimuli synthesised either from short-time magnitude spectra or short-time phase spectra. The results of the experiment indicate that even for small window durations (of relevance for automatic speech recognition applications), the phase spectrum can contribute to speech intelligibility as much as the magnitude spectrum if the analysis-modification-synthesis parameters are properly selected. This intelligibility is significantly more than that reported by Liu et al. (1997), who carried out a similar experiment with the same analysis-modification-synthesis framework. The significant improvement in intelligibility over Liu's results may be attributed to the differences in the parameter settings adopted. In this paper, we review our previous experiment and conduct an additional experiment to determine the contribution that each parameter setting provides towards the intelligibility of stimuli reconstructed from short-time phase spectra. The parameter selection that contributes most to the intelligibility of the phase-only stimuli is that of a rectangular analysis window, as opposed to a Hamming window (which is generally used in speech analysis). Leigh D. Alsteris, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2004 | Multiple frame block quantisation of line spectral frequencies using Gaussian mixture modelsabstractIn this paper, we present a Gaussian mixture model-based block quantiser for coding line spectral frequencies that uses multiple frames and mean squared error as the quantiser selection criterion. The efficiency gained from jointly coding multiple frames permits the use of the mean squared error distortion (MSE) criterion rather than the computationally expensive spectral distortion. The proposed coder encompasses improvements in both distortion performance and complexity with transparency achieved at 23 bits per frame when coding two frames jointly or 21 bits per frame when coding 3 frames. Kuldip K. Paliwal, Stephen So |
ICASSP (1) | 1 |
| 2004 | Product of power spectrum and group delay function for speech recognitionabstractMel-frequency cepstral coefficients (MFCCs) are the most widely used features for speech recognition. These are derived from the power spectrum of the speech signal. Recently, the cepstral features derived from the modified group delay function (MGDF) have been studied by Murthy and Gadde (Proc. ICASSP, vol.1, p.68-71, 2003) for speech recognition. In this paper, we propose to use the product of the power spectrum and the group delay function (GDF), and derive the MFCCs from the product spectrum. This spectrum combines the information from the magnitude spectrum as well as the phase spectrum. The MFCCs of the MGDF are also investigated in this paper. Results show that the cepstral features derived from the power spectrum perform better than that from the MGDF, and the product spectrum based features provide the best performance. Donglai Zhu, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2004 | ASR on speech reconstructed from short-time fourier phase spectraabstractIn our earlier papers, we have measured human intelligibility of speech stimuli reconstructed either from the short-time magnitude spectra (magnitude-only stimuli) or the short-time phase spectra (phase-only stimuli) of a speech stimulus. We demonstrated that, even for small analysis window durations of 20-40 ms (of relevance to automatic speech recognition), the short-time phase spectrum can contribute to speech intelligibility as much as the short-time magnitude spectrum. In this paper, we perform automatic speech recognition on magnitude-only and phase-only stimuli. When employing an MFCC-based front-end, the recognition achieved for these phase-only stimuli is much worse than magnitude-only stimuli at small analysis window durations, which is not consistent with their corresponding human intelligibility results. This implies that the MFCC feature set is not capturing all of the discriminating information present in the speech signal. Leigh D. Alsteris, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2004 | Scalable distributed speech recognition using multi-frame GMM-based block quantizationabstractIn this paper, we propose the use of the multi-frame Gaussian mixture model-based block quantizer for the coding of Mel frequency-warped cepstral coefficient (MFCC) features in distributed speech recognition (DSR) applications. This coding scheme exploits intraframe correlation via the Karhunen-Loeve transform (KLT) and interframe correlation via the joint processing of adjacent frames together with the computational simplicity of scalar quantization. The proposed coder is bit-rate scalable, which means that the bitrate can be adjusted without the need for re-training of the quantizers. Static parameters such as the probability density function (PDF) model and KLT orthogonal matrices are stored at the encoder and decoder and bit allocations are calculated 'on-the-fly' without intensive processing. This coding scheme is evaluated in this paper on the Aurora-2 database in a DSR framework. It is shown that this coding scheme achieves high recognition performance at lower bitrates, with a word error rate (WER) of 2.5% at 800 bps, which is less than 1% degradation from the baseline word recognition accuracy, and graceful degradation down to a WER of 7% at 300 bps. Kuldip K. Paliwal, Stephen So |
INTERSPEECH | 1 |
| 2004 | MFCC computation from magnitude spectrum of higher lag autocorrelation coefficients for robust speech recognitionabstractProcessing of the speech signal in the autocorrelation domain in the context of robust feature extraction is based on the following two properties: 1) pole preserving property (the poles of a given (original) signal are preserved in its autocorrelation function), and 2) noise separation property (the autocorrelation function of a noise signal is confined to lower lags, while the speech signal contribution is spread over all the lags in the autocorrelation function, thus providing a way to eliminate noise by discarding lower-lag autocorrelation coefficients). In this paper, we use these properties to derive robust features for automatic speech recognition. We compute the magnitude spectrum of the one-sided higher-lag autocorrelation sequence, process it through a Mel filter bank and parameterise it in terms of Mel Frequency Cepstral Coefficients (MFCCs). Since the proposed method combines autocorrelation domain processing with Mel filter bank analysis, we call the resulting MFCCs, Autocorrelation Mel Frequency Cepstral Coefficients (AMFCCs). Recognition experiments are conducted on the Aurora II database and it is found that the AMFCC representation performs as well as the MFCC representation in clean conditions and provides more robust performance in the presence of background noise. Benjamin J. Shannon, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2004 | Efficient vector quantisation of line spectral frequencies using the switched split vector quantiserabstractIn this paper, we investigate the use of a switched split vector quantiser (SSVQ) for coding linear predictive coding (LPC) parameters. The SSVQ is applied to quantise the LPC parameters in terms of line spectral frequencies from the TIMIT database and its performance is compared with the split vector quantiser. Experimental results show that the SSVQ provides a better trade-off between bit-rate and distortion performance than the split VQ. In addition, the SSVQ has a lower computational (search) complexity than the split VQ, though this is attained at the expense of an increase in memory requirements. In order to achieve a spectral distortion of 1 dB, the three-part SSVQ with an 8 directional switch requires 23 bits/frame, 4.41 kflops/frame of computations and 8272 floats of memory, while the corresponding values for a traditional three-part split VQ are 25 bits/frame, 13.3 kflops/frame and 3328 floats, respectively. Stephen So, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2004 | Noise adaptive speech recognition based on sequential noise parameter estimation
Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
Speech Commun. | 2 |
| 2004 | Recognition of noisy speech using dynamic spectral subband centroidsabstractDespite their widespread popularity as front-end parameters for speech recognition, the cepstral coefficients derived from either linear prediction analysis or a filter-bank are found to be sensitive to additive noise. In this letter, we discuss the use of spectral subband centroids for robust speech recognition. We show that centroids, if properly selected, can achieve recognition performance comparable to that of the mel-frequency cepstral coefficients (MFCCs) in clean speech, while delivering better performance than MFCC in noisy environments. A procedure is proposed to construct the dynamic centroid feature vector that essentially embodies the transitional spectral information. We discuss some properties of the proposed dynamic features. Jingdong Chen, Yiteng Huang, Kuldip K. Paliwal |
IEEE Signal Process. Lett. | 4 |
| 2003 | Robust speech recognition using features based on zero crossings with peak amplitudesabstractThe paper presents an extensive study of zero crossings with peak amplitudes (ZCPA) features, that have earlier been shown to outperform both conventional and auditory-based features in the presence of additive noise. The study starts by optimizing different parameters involved in ZCPA feature computation, followed by a comparison of ZCPA and MFCC features on two recognition tasks in different background conditions. The main differences between the two feature types are identified, and their individual effects on ASR performance are evaluated. The importance of a proper choice of analysis frame lengths and filter bandwidths in ZCPA feature extraction is demonstrated. Furthermore, the use of dominant frequency information in ZCPA features is found to be a major reason for increased robustness of ZCPA features compared to MFCC features. Bojana Gajic, Kuldip K. Paliwal |
ICASSP (1) | 2 |
| 2003 | Noise resistant audio-visual verification via structural constraintsabstractWe propose a piecewise linear classifier for use as the decision stage in a two-modal verification system, comprised of a face expert and a speech expert. The classifier utilizes a fixed decision boundary that has been specifically designed to account for the effects of noisy audio conditions. Experimental results show that, in clean conditions, the proposed classifier is outperformed by a traditional weighted summation decision stage (using both fixed and adaptive weights); however, in high noise conditions the classifier obtains better performance than the fixed approach and has similar performance as the adaptive approach, with the advantage of having a fixed (non-adaptive) structure. Conrad Sanderson, Kuldip K. Paliwal |
ICASSP (5) | 2 |
| 2003 | Frequency-related representation of speechabstractCepstral features derived from power spectrum are widely used for automatic speech recognition. Very little work, if any, has been done in speech research to explore phase-based representations. In this paper, an attempt is made to investigate the use of phase function in the analytic signal of critical-band filtered speech for deriving a representation of frequencies present in the speech signal. Results are presented which show the validity of this approach. Kuldip K. Paliwal, Bishnu S. Atal |
INTERSPEECH | 1 |
| 2003 | Usefulness of phase spectrum in human speech perceptionabstractShort-time Fourier transform of speech signal has two components: magnitude spectrum and phase spectrum. In this paper, relative importance of short-time magnitudeand phase spectra on speech perception is investigated. Human perception experiments are conducted to measure intelligibility of speech tokens synthesized either from magnitude spectrum or phase spectrum. It is traditionally believed that magnitude spectrum plays a dominant role for shorter windows (20-30 ms); while phase spectrum is more important for longer windows (128-3500 ms). It is shown in this paper that even for shorter windows, phase spectrum can contribute to speech intelligibility as much as the magnitude spectrum if the shape of the window function is properly selected. Kuldip K. Paliwal, Leigh D. Alsteris |
INTERSPEECH | 1 |
| 2003 | Speech recognition with a generative factor analyzed hidden Markov modelabstractWe present a generative factor analyzed hidden Markov model (GFA-HMM) for automatic speech recognition. In a traditional HMM, the observation vectors are represented by mixture of Gaussians (MoG) that are dependent on discrete-valued hidden state sequence. The GFA-HMM introduces a hierarchy of continuous-valued latent representation of observation vectors, where latent vectors in one level are acoustic-unit dependent and the latent vectors in a higher level are acoustic-unit independent. An expectation maximization (EM) algorithm is derived for maximum likelihood parameter estimation of the model. The GFA-HMM can achieve a much more compact representation of the intra-frame statistics of observation vectors than traditional HMM. We conducted an experiment to show that the GFA-HMM can achieve better performances over traditional HMM with the same amount of training data but much smaller number of model parameters. Kaisheng Yao, Kuldip K. Paliwal, Te-Won Lee |
INTERSPEECH | 2 |
| 2003 | Model based noisy speech recognition with environment parameters estimated by noise adaptive speech recognition with priorabstractWe have proposed earlier a noise adaptive speech recognition approach for recognizing speech corrupted by nonstationary noise and channel distortion. In this paper, we extend this approach. Instead of maximum likelihood estimation of environment parameters (as done in our previous work), the present method estimates environment parameters within the Bayesian framework that is capable of incorporating prior knowledge of the environment. Experiments are conducted on a database that contains digit utterances contaminated by channel distortion and nonstationary noise. Results show that this method performs better than the previous methods. Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2003 | Maximum likelihood sub-band weighting for robust speech recognitionabstractSub-band speech recognition approaches have been proposed for robust speech recognition, where full-band power spectra are divided into several sub-bands and then likelihoods or cepstral vectors of the sub-bands are merged depending on their reliability. In conventional sub-band approaches, correlations across the sub-bands are not modeled and the merging weights can only be set experientially or estimated during training procedures, which may not match observed data. The methods further degrade performance for clean speech. We proposed a novel sub-band approach, where frequency sub-bands are multiplied with weighting factors and merged, which considers sub-band dependence and proves to be more robust than both full-band and conventional sub-band approaches. And further the weighting factors can be obtained by using the maximum-likelihood estimation approaches in order to minimize the mismatch between the trained models and the observed features. Finally we evaluated our methods on both the Aurora2 task and the Resource Management task and showed improvement of performance on the two tasks consistently. 1. Donglai Zhu, Satoshi Nakamura 0001, Kuldip K. Paliwal, Renhua Wang |
INTERSPEECH | 3 |
| 2003 | Noise compensation in a person verification system using face and multiple speech feature
Conrad Sanderson, Kuldip K. Paliwal |
Pattern Recognit. | 2 |
| 2003 | Feature extraction and dimensionality reduction algorithms and their applications in vowel recognition
Xuechuan Wang, Kuldip K. Paliwal |
Pattern Recognit. | 2 |
| 2003 | Structurally noise resistant classifier for multi-modal person verification
Conrad Sanderson, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 2 |
| 2003 | Fast features for face authentication under illumination direction changes
Conrad Sanderson, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 2 |
| 2003 | Features for robust face-based identity verification
Conrad Sanderson, Kuldip K. Paliwal |
Signal Process. | 2 |
| 2003 | Cepstrum derived from differentiated power spectrum for robust speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
Speech Commun. | 2 |
| 2002 | Noise adaptive speech recognition in time-varying noise based on sequential kullback proximal algorithmabstractWe present a noise adaptive speech recognition approach, where time-varying noise parameter estimation and Viterbi process are combined together. The Viterbi process provides approximated joint likelihood of active partial paths and observation sequence given the noise parameter sequence estimated till previous frame. The joint likelihood after normalization provides approximation to the posterior probabilities of state sequences for an EM-type recursive process based on sequential Kullback proximal algorithm to estimate the current noise parameter. The combined process can easily be applied to perform continuous speech recognition in presence of non-stationary noise. Experiments were conducted in simulated and real non-stationary noises. Results showed that the noise adaptive system provides significant improvements in word accuracy as compared to the baseline system (without noise compensation) and the normal noise compensation system (which assumes the noise to be stationary). Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
ICASSP | 2 |
| 2002 | Likelihood normalization for face authentication in variable recording conditionsabstractIn this paper we evaluate the effectiveness of two likelihood normalization techniques, the background model set (BMS) and the universal background model (UBM), for improving performance and robustness of four face authentication systems utilizing a Gaussian mixture model (GMM) classifier. The systems differ in the feature extraction method used: eigenfaces (PCA), 2-D DCT, 2-D Gabor wavelets and DCT-mod2. Experiments on the VidTIMIT database, using test images corrupted either by an illumination change or compression artefacts, suggest that likelihood normalization has little effect when using PCA derived features, while providing significant performance improvements when using the remaining features. Kuldip K. Paliwal, Conrad Sanderson |
ICIP (1) | 1 |
| 2002 | Polynomial features for robust face authenticationabstractWe introduce the DCT-mod2 facial feature extraction technique which utilizes polynomial coefficients derived from 2D DCT coefficients of spatially neighbouring blocks. We evaluate its robustness and performance against three popular feature sets for use in an identity verification system subject to illumination changes. Results on the multi-session VidTIMIT database suggest that the proposed feature set is the most robust, followed by (in order of robustness and performance): 2D Gabor wavelets; 2D DCT coefficients; PCA (eigenface) derived features. Moreover, compared to Gabor wavelets, the DCT-mod2 feature set is over 80 times quicker to compute. Conrad Sanderson, Kuldip K. Paliwal |
ICIP (3) | 2 |
| 2002 | Noise adaptive speech recognition with acoustic models trained from noisy speech evaluated on Aurora-2 databaseabstractIn this paper, we apply the noise adaptive speech recognition for noisy speech recognition in non-stationary noise to the situation that acoustic models are trained from noisy speech. We justify it by that the noise adaptive speech recognition includes iterative processes between a noise parameter estimation step and a model adaptation step, which can possibly do non-linear mapping between the original training space and that for recognition. Experiments were performed onAurora-2 task with multi-conditional training set which includes noisy utterances. Through experiments, we observed that the noise adaptive speech recognition can have better performance than the baseline system trained from multiconditional training set without noise adaptive speech recognition. 1. Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2001 | Joint Cohort Normalization in a Multi-Feature Speaker Verification SystemabstractIn this paper we propose a new fusion technique, termed joint cohort normalization fusion, where the information fusion is done prior to the likelihood ratio test in a speaker verification system. The performance of the technique is compared against two popular types of fusion: feature vector concatenation and expert opinion fusion, for fusion of mel frequency cepstral coefficients (MFCC), MFCC with cepstral mean subtraction and maximum auto-correlation values features. In experiments on the NTIMIT database, the proposed technique is shown, in most cases, to outperform the popular methods. Conrad Sanderson, Kuldip K. Paliwal |
FUZZ-IEEE | 2 |
| 2001 | Robust feature extraction using subband spectral centroid histogramsabstractIn this paper we propose a new framework for utilizing frequency information from the short-term power spectrum of speech. Feature extraction is based on the cepstral coefficients derived from the histograms of subband spectral centroids (SSC). Two new feature extraction algorithms are proposed, one based on frequency information alone, and the other which efficiently combines the frequency and amplitude information from the speech power spectrum. Experimental study on an automatic speech recognition task shows that the proposed methods outperform the conventional speech front-ends in the presence of additive white noise, while they perform comparably in the noise-free conditions. Bojana Gajic, Kuldip K. Paliwal |
ICASSP | 2 |
| 2001 | Noise compensation in a multi-modal verification systemabstractIn this paper we propose an adaptive multi-modal verification system comprised of a modified minimum cost Bayesian classifier (MCBC) and a method to find the reliability of the speech expert for various noisy conditions. The modified MCBC takes into account the reliability of each modality expert, allowing the de-emphasis of the contribution of opinions from the expert affected by noise. Reliability of the speech expert is found without directly modeling the noisy speech or finding the reliability a priori for various conditions of the speech signal. Experiments on the digit database show the total error to be reduced by 78% when compared to a non-adaptive system. Conrad Sanderson, Kuldip K. Paliwal |
ICASSP | 2 |
| 2001 | Sub-band based additive noise removal for robust speech recognitionabstractGriffith Sciences, Griffith School of Engineering Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2001 | Robust parameters for speech recognition based on subband spectral centroid histogramsabstractIn this paper we propose a new speech parameterization framework that efficiently combines frequency and magnitude information from the short-term power spectrum of speech. This is achieved through computation of subband spectral centroid histograms (SSCH). Relationship between the proposed method and auditory based speech parameterization methods is discussed. An experimental study on an automatic speech recognition task has shown that the proposed method outperforms the conventional speech front-ends in presence of different types of additive noise, while it performs comparably in the noise-free conditions. In the case of car noise, our method also outperforms the computationally expensive auditory based methods, while having simplicity and low computational cost similar to the conventional front-ends. Bojana Gajic, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2001 | Fast adaptation using constrained affine transformations with hierarchical priorsabstractIn this paper we present an approach to transformation based model adaptation that combines a fast, closed form solution to the MAP estimation of our transforms with robust priors. The robust priors are found using the technique of hierarchical priors, and a closed form solution is achieved by choosing diagonally constrained affine transformations and a suitable family of prior distributions for these transformations. We show that the method gives results comparable to other algorithms, but with significantly reduced computational complexity and memory demands. Experiments are conducted on the SI Recognition Outlier task from the Wall Street Journal corpus, where speaker independent models have to be adapted to handle speech from non-native speakers. Tor André Myrvoll, Kuldip K. Paliwal, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2001 | Information fusion for robust speaker verificationabstractGriffith Sciences, Griffith School of Engineering Conrad Sanderson, Kuldip K. Paliwal |
INTERSPEECH | 2 |
| 2001 | Feature extraction and model-based noise compensation for noisy speech recognition evaluated on AURORA 2 taskabstractNo Full Text Kaisheng Yao, Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2001 | Sequential noise compensation by a sequential kullback proximal algorithmabstractNo Full Text Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2000 | Lossless coding of MPEG-1 Layer III encoded audio streamsabstractIn this paper we describe a lossless coding scheme for the encoding of MPEG-1 Layer III encoded audio bitstreams. Commonly known as MP3, the MPEG-1 Layer III standard has proved widely popular for the transmission of encoded audio files (MP3's) over the Internet. However, the MPEG-1 Layer III standard has been designed with a wide range of applications in mind. As such, the frame sizes are kept small and redundancies between samples in neighboring frames are not exploited. We propose a design which uses a combination of linear predictive coding and arithmetic coding to exploit such redundancies. The proposed coder was tested on a number of Layer III encoded audio (MP3) files and shown to produce an average coding gain of 12.2% over the original Layer III encoded files. Farshid Golchin, Kuldip K. Paliwal |
ICASSP | 2 |
| 2000 | A block cosine transform and its application in speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2000 | Fast nearest-neighbor search algorithms based on approximation-elimination search
V. Ramasubramanian 0001, Kuldip K. Paliwal |
Pattern Recognit. | 2 |
| 2000 | Comments on "modified K-means algorithm for vector quantizer design"abstractPreviously a modified K-means algorithm for vector quantization design has been proposed where the codevector updating step is as follows: new codevector=current codevector+scale factor (new centroid-current codevector). This algorithm uses a fixed value for the scale factor. In this paper, we propose the use of a variable scale factor which is a function of the iteration number. For the vector quantization of image data, we show that it offers faster convergence than the modified K-means algorithm with a fixed scale factor, without affecting the optimality of the codebook. Kuldip K. Paliwal, V. Ramasubramanian 0001 |
IEEE Trans. Image Process. | 1 |
| 1999 | Decorrelated and liftered filter-bank energies for robust speech recognition
Kuldip K. Paliwal |
EUROSPEECH | 1 |
| 1999 | Fast nearest-neighbor search based on Voronoi projections and its application to vector quantization encodingabstractIn this work, we consider two fast nearest-neighbor search methods based on the projections of Voronoi regions, namely, the box-search method and the cell-partition search method. We provide their comprehensive study in the contest of vector quantization encoding. We show that the use of principal component transformation reduces the complexity of Voronoi-projection based search significantly for data with high degree of correlation across their components. V. Ramasubramanian 0001, Kuldip K. Paliwal |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | A lossless image coder with context classification, adaptive prediction and adaptive entropy codingabstractWe combine a context classification scheme with adaptive prediction and entropy coding to produce an adaptive lossless image coder. In this coder, we maximize the benefits of adaptivity using both adaptive prediction and entropy coding. The adaptive prediction is closely tied with the classification of contexts within the image. These contexts are defined with respect to the local edge, texture or gradient characteristics as well as local activity within small blocks of the image. For each context an optimal predictor is found which is used for the prediction of all pixels belonging to that particular context. Once the predicted values have been removed from the original image, a clustering algorithm is used to design a separate, optimal entropy coding scheme for encoding the prediction residual. Blocks of residual pixels are classified into a finite number of classes and members of each class are encoded using the entropy coder designed for that particular class. The combination of these two powerful techniques produces some of the best lossless coding results reported so far. Farshid Golchin, Kuldip K. Paliwal |
ICASSP | 2 |
| 1998 | Spectral subband centroid features for speech recognitionabstractCepstral coefficients derived either through linear prediction (LP) analysis or from filter banks are perhaps the most commonly used features in currently available speech recognition systems. In this paper, we propose spectral subband centroids as new features and use them as a supplement to cepstral features for speech recognition. We show that these features have properties similar to formant frequencies and they are quite robust to noise. Recognition results are reported, justifying the usefulness of these features as supplementary features. Kuldip K. Paliwal |
ICASSP | 1 |
| 1998 | Model parameter estimation for mixture density polynomial segment models
Toshiaki Fukada, Kuldip K. Paliwal, Yoshinori Sagisaka |
Comput. Speech Lang. | 2 |
| 1997 | Model parameter estimation for mixture density polynomial segment modelsabstractIn this paper, we propose parameter estimation techniques for mixture density polynomial segment models (MDPSM) where their trajectories are specified with an arbitrary regression order. MDPSM parameters can be trained in one of three different ways: (1) segment clustering, (2) expectation maximization (EM) training of mean trajectories, or (3) EM training of mean and variance trajectories. These parameter estimation methods were evaluated in TIMIT vowel classification experiments. The experimental results showed that modeling both the mean and variance trajectories are consistently superior to modeling only the mean trajectory. We also found that modeling both trajectories results in significant improvements over the conventional HMM. Toshiaki Fukada, Yoshinori Sagisaka, Kuldip K. Paliwal |
ICASSP | 3 |
| 1997 | Quadtree based classification with arithmetic and trellis coded quantization for subband image codingabstractIn this a paper a quadtree based method is proposed for classifying blocks of samples in image subbands. Classification of blocks of subband samples according to their energy and variable bit allocation within the subsequent classes has demonstrated considerable gains in coding efficiency. The gains due to classification increase as smaller blocks are used; however, so do the overheads for transmitting the classification information. The quadtree based method proposed in this paper allows for more efficient classification by using variable-sized blocks in order to maximize the classification gain, while maintaining a limit on the classification overheads. Using an efficient quantization scheme such as ACTCQ (arithmetic and trellis coded quantization), we have been able to demonstrate competitive coding results at low bit-rates. Farshid Golchin, Kuldip K. Paliwal |
ICASSP | 2 |
| 1997 | Minimum-Entropy Clustering and its Application to Lossless Image CodingabstractThe minimum-entropy clustering (MEC) algorithm proposed in this paper provides an optimal method for addressing the non-stationarity of a source with respect to entropy coding. This algorithm clusters a set of vectors (where each vector consists of a fixed number of contiguous samples from a discrete source) using a minimum entropy criterion. In a manner similar to classified vector quantization (CVQ), a given vector is first classified into the class which leads to the lowest entropy and then its samples are coded by the entropy coder designed for that particular class. The MEC algorithm is used in the design of a lossless, predictive image coder. The MEC-based coder is found to significantly outperform the single entropy coder as well as the other popular lossless coders reported in the literature. Farshid Golchin, Kuldip K. Paliwal |
ICIP (2) | 2 |
| 1997 | Classified adaptive prediction and entropy coding for lossless coding of imagesabstractNatural images often consist of many distinct regions with individual characteristics. Adaptive image coders exploit this feature of natural images to obtain better compression results. In this paper, we propose a classification-based scheme for both adaptive prediction and entropy coding in a lossless image coder. In the proposed coder, blocks of image samples (in the PCM domain) are classified to select an appropriate linear predictor from finite set of predictors. Once the predictors have been determined, the image is DPCM coded. A second classification is then performed to select a suitable entropy coder for each block of DPCM samples. These classification schemes are designed using two separate clustering procedures which attempt to minimize the bit-rate of the encoded image. The coder was tested on a set of monochrome images and was found to produce very promising results. Farshid Golchin, Kuldip K. Paliwal |
ICIP (3) | 2 |
| 1997 | Cyclic autocorrelation-based linear prediction analysis of speech
Kuldip K. Paliwal, Yoshinori Sagisaka |
EUROSPEECH | 1 |
| 1997 | Reducing the complexity of the LPC vector quantizer using the k-d tree search algorithm
V. Ramasubramanian 0001, Kuldip K. Paliwal |
EUROSPEECH | 2 |
| 1996 | Design of a speech recognition system based on acoustically derived segmental unitsabstractThe design of a speech recognition system based on acoustically-derived, segmental units can be divided in three steps: unit design, lexicon building and pronunciation modeling. We formulate an iterative unit design procedure which consistently uses a maximum likelihood (ML) objective in successive application of resegmentation and model re-estimation. The lexicon building allows multi-word entries in the lexicon but restricts the number of these entries in order to avoid a too costly search. Selected multi-word lexical entries are those with high frequency (such as function words) and those which consistently exhibit cross-word phone assimilation. The stochastic pronunciation model represents the likelihood of a particular acoustic segment sequence given the phonetic baseform of a lexical item, where the sequence of baseform phones are treated as a Markov state sequence and each state can emit multiple segments. Michiel Bacchiani, Mari Ostendorf, Yoshinori Sagisaka, Kuldip K. Paliwal |
ICASSP | 4 |
| 1996 | Speech recognition based on acoustically derived segment units
Toshiaki Fukada, Michiel Bacchiani, Kuldip K. Paliwal, Yoshinori Sagisaka |
ICSLP | 3 |
| 1996 | Effect of speech coders on speech recognition performance
B. T. Lilly, Kuldip K. Paliwal |
ICSLP | 2 |
| 1995 | Interpolation properties of linear prediction parametric representations
Kuldip K. Paliwal |
EUROSPEECH | 1 |
| 1995 | A maximum likelihood equalization technique for robust speech recognition in adverse environments
Kuldip K. Paliwal |
EUROSPEECH | 1 |
| 1995 | Minimum classification error training algorithm for feature extractor and pattern classifier in speech recognition
Kuldip K. Paliwal, Michiel Bacchiani, Yoshinori Sagisaka |
EUROSPEECH | 1 |
| 1995 | Effect of rasta-type processing for speech recognition with speaking-rate mismatches
Harald Singer, Kuldip K. Paliwal, Tomohiko Beppu, Yoshinori Sagisaka |
EUROSPEECH | 2 |
| 1994 | A comparative study of feature representations for robust speech recognition in adverse environments
Kuldip K. Paliwal, Bishnu S. Atal |
ICSLP | 1 |
| 1993 | Use of temporal correlation between successive frames in a hidden Markov model based speech recognizer
Kuldip K. Paliwal |
ICASSP (2) | 1 |
| 1993 | Efficient vector quantization of LPC parameters at 24 bits/frameabstractFor low bit rate speech coding applications, it is important to quantize the LPC parameters accurately using as few bits as possible. Though vector quantizers are more efficient than scalar quantizers, their use for accurate quantization of linear predictive coding (LPC) information (using 24-26 bits/frame) is impeded by their prohibitively high complexity. A split vector quantization approach is used here to overcome the complexity problem. An LPC vector consisting of 10 line spectral frequencies (LSFs) is divided into two parts, and each part is quantized separately using vector quantization. Using the localized spectral sensitivity property of the LSF parameters, a weighted LSF distance measure is proposed. With this distance measure, it is shown that the split vector quantizer can quantize LPC information in 24 bits/frame with an average spectral distortion of 1 dB and less than 2% of the frames having spectral distortion greater than 2 dB. The effect of channel errors on the performance of this quantizer is also investigated and results are reported.> Kuldip K. Paliwal, Bishnu S. Atal |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | Vector equalization in hidden Markov models for noisy speech recognitionabstractSpeech recognizers often experience serious performance degradation when deployed in an unknown acoustic (particularly, noise contaminated) environment. To combat this problem, the authors proposed in a previous study a distortion measure that takes into account the norm shrinkage bias in the noisy cepstrum. The authors incorporate a first-order equalization mechanism, specifically aimed at avoiding the norm shrinkage problem, in a hidden Markov model (HMM) framework to model the speech cepstral sequence. Such a modeling technique requires special care as the formulation inevitably involves parameter estimation from a set of data with singular dispersion. The authors provide solutions to this HMM stochastic modeling problem and give algorithms for estimating the necessary model parameters. They experimentally show that incorporation of the first-order normal equalization model makes the HMM-based speech recognizer robust to noise. With respect to a conventional HMM recognizer, this leads to an improvement in recognition performance which is equivalent to about 15-20-dB gain in signal-to-noise ratio.> Biing-Hwang Juang, Kuldip K. Paliwal |
ICASSP | 2 |
| 1992 | An efficient approximation-elimination algorithm for fast-nearest-neighbour search (speech coding)abstractThe authors present an efficient algorithm for fast nearest-neighbour search in multidimensional space under a so called approximation-elimination framework. The algorithm is based on an approximation procedure which selects codevectors for distance computation in the close proximity of the test vector and eliminates codevectors using the triangle inequality based elimination. The algorithm is studied in the context of vector quantization of speech and compared with related algorithms proposed earlier. It is shown to be more efficient in terms of reducing the main search complexity, overhead costs and storage.> V. Ramasubramanian 0001, Kuldip K. Paliwal |
ICASSP | 2 |
| 1992 | An efficient approximation-elimination algorithm for fast nearest-neighbour search based on a spherical distance coordinate formulation
V. Ramasubramanian 0001, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 2 |
| 1991 | Speech coding at 4 kb/s and lower using single-pulse and stochastic models of LPC excitationabstractThe authors present an LPC (linear predictive coding) speech coder that classifies speech into periodic and nonperiodic intervals. The coder uses a stochastic codebook, as in CELP (code excited linear prediction), to synthesize nonperiodic speech and single-pulse excitation (SPE) to synthesize periodic speech. This coder is called SPE-CELP. The optimization of the single-pulse excitation is based on a new algorithm that determines the time instants of pitch periods within a short interval of periodic speech of approximately 32 ms. It thus causes a coding delay that is acceptable for many applications. Results using a fixed delta-impulse shape and time-varying pulse shapes obtained from codebooks are reported.> W. Granzow, Bishnu S. Atal, Kuldip K. Paliwal, Juergen Schroeter |
ICASSP | 3 |
| 1991 | Cell-conditioned multistage vector quantizationabstractTwo-stage vector quantization (2VQ) reduces the complexity of single-stage VQ at the cost of reduced performance. Cell-conditioned (CC) 2VQ is proposed to improve the performance of 2VQ, while retaining the advantage of reduced complexity. In CC2VQ, the error vector from the first stage is transformed prior to its quantization by second stage. The effect of transformation is to cause the size and orientation of different first-stage cells to be as similar as possible. It is shown theoretically that CC2VQ performs as well as single-stage VQ under asymptotic conditions. Some experimental results for the vector quantization of speech waveforms are presented to support these theoretical results.> David L. Neuhoff, Kuldip K. Paliwal |
ICASSP | 3 |
| 1991 | Efficient vector quantization of LPC parameters at 24 bits/frameabstractThough vector quantizers are more efficient than scalar quantizers, their use for fine quantization of linear predictive coding (LPC) information (using 24-26 b/frame) is impeded due to their prohibitively high complexity. In the present work, a split vector quantization approach is used to overcome the complexity problem. The LPC vector, consisting of ten line spectral frequencies (LSFs), is divided into two parts and each part is quantized separately using vector quantization. Using the localized spectral sensitivity property of the LSF parameters, a weighted LSF distance measure is proposed. Using this distance measure, it is shown that the split vector quantizer can quantize LPC information in 24 b/frame with 1-dB average spectral distortion and> Kuldip K. Paliwal, Bishnu S. Atal |
ICASSP | 1 |
| 1991 | Recognition of noisy speech using cumulant-based linear prediction analysisabstractThe use of cumulant-based LP (linear prediction) analysis for speech recognition in the presence of noise is proposed. This method assumes the speech signal to be non-Gaussian. It is shown that cepstral coefficients derived by this method are quite insensitive to additive Gaussian noise which can be white or colored. The performance of a recognizer based on these estimates is compared to the performance of one that uses LP estimates derived from the autocorrelation function. It is found that at low SNR (below about 20 dB) the cumulant-based estimates outperform the autocorrelation-based estimates. At higher SNRs the reverse is true. The reasons for this behavior are not yet understood. However, it is shown that, by combining the two estimates, one can achieve recognition accuracy that is better than that of the conventional recognizer at all SNRs.> Kuldip K. Paliwal, Man Mohan Sondhi |
ICASSP | 1 |
| 1991 | Performance study of stochastic speech coders
Bernt Ribbum, Andrew Perkis, Kuldip K. Paliwal, Tor A. Ramstad |
Speech Commun. | 3 |
| 1990 | Neural net classifiers for robust speech recognition under noisy environmentsabstractThe multilayer perceptron (MLP) classifier is studied for the recognition of noisy speech, and its performance is compared with that of conventional pattern classifiers such as the maximum-likelihood (ML) classifier and the k-nearest-neighbor (kNN) classifier. The linear prediction (LP) parameters derived through tenth-order LP analysis are used as the recognition parameters. Different LP parametric representations are compared as to their recognition performance with the MLP classifier, and the cepstral coefficient representation is found to be the best parametric representation. When ten cepstral coefficients are used as recognition parameters, the performance of the MLP classifier is found to be significantly better than that of the ML and the kNN classifiers for noisy speech. Use of 15 cepstral coefficients (obtained by extrapolating the ten cepstral coefficients) improves the recognition performance of the MLP classifier for noisy speech further.> Kuldip K. Paliwal |
ICASSP | 1 |
| 1990 | Lexicon-building methods for an acoustic sub-word based speech recognizerabstractThe use of an acoustic subword unit (ASWU)-based speech recognition system for the recognition of isolated words is discussed. Some methods are proposed for generating the deterministic and the statistical types of word lexicon. It is shown that the use of a modified k-means algorithm on the likelihoods derived through the Viterbi algorithm provides the best deterministic-type of word lexicon. However, the ASWU-based speech recognizer leads to better performance with the statistical type of word lexicon than with the deterministic type. Improving the design of the word lexicon makes it possible to narrow the gap in the recognition performances of the whole word unit (WWU)-based and the ASWU-based speech recognizers considerably. Further improvements are expected by designing the word lexicon better.> Kuldip K. Paliwal |
ICASSP | 1 |
| 1990 | A study of LSF representation for speaker-dependent and speaker-independent HMM-based speech recognition systemsabstractThe line spectral-pair frequency (LSF) representation is used as the parametric representation for speech recognition. Its performance is compared with that of the cepstral coefficient (CC) representation for the speaker-dependent and speaker-independent hidden Markov model (HMM)-based isolated work recognition systems. It is shown that the CC and the LSF representations result in comparable recognition performances for the full covariance matrix case. For the diagonal covariance matrix case, the LSF representation provides significantly better recognition performance than the CC representation.> Kuldip K. Paliwal |
ICASSP | 1 |
| 1989 | An improved sub-word based speech recognizerabstractThe authors describe a system for speaker-dependent speech recognition based on acoustic subword units. Several strategies for automatic generation of an acoustic lexicon are outlined. Preliminary tests have been performed on a small vocabulary. In these tests, the proposed system showed results comparable to those of whole-word-based systems.> Torbjørn Svendsen, Kuldip K. Paliwal, Erik Harborg, P. O. Husoy |
ICASSP | 2 |
| 1989 | A study of line spectrum pair frequencies for vowel recognition
Kuldip K. Paliwal |
Speech Commun. | 1 |
| 1989 | Effect of ordering the codebook on the efficiency of the partial distance search algorithm for vector quantizationabstractRecently, C.D. Bei and R.M. Gray (1985) used a partial distance search algorithm that reduces the computational complexity of the minimum distortion encoding for vector quantization. The effect of ordering the codevectors on the computational complexity of the algorithm is studied. It is shown that the computational complexity of this algorithm can be reduced further by ordering the codevectors according to the sizes of their corresponding clusters.> Kuldip K. Paliwal, V. Ramasubramanian 0001 |
IEEE Trans. Commun. | 1 |
| 1988 | A comparative performance evaluation of adaptive ARMA spectral estimation methods for noisy speechabstractThe problem of adaptive estimation of linear prediction (LP) coefficients from noisy speech is considered. Performance of three adaptive ARMA spectral estimation algorithms are studied for this purpose: the recursive extended least squares (RELS) algorithm, the recursive maximum likelihood (RML) algorithm, and the overdetermined recursive instrumental variable (ORIV) algorithm. To put them in proper perspective, the normalized LMS (NLMS) has also been considered. The ORIV algorithm is found to be the best in terms of Itakura distance from the ideal LP coefficients and the power spectral density estimation. The RML algorithm is found to be robust in highly noisy cases.> Anjan Basu, Kuldip K. Paliwal |
ICASSP | 2 |
| 1988 | A study of line spectrum pair frequencies for speech recognitionabstractThe line spectrum pair (LSP) frequency representation has recently been proposed as an alternative linear prediction (LP) parametric representation. In the context of speech coding, this representation shows better quantization properties than other LP parametric representations. The LSP representation is studied for speech recognition. Several distance measures based on this representation are investigated The weighted LSP distance measure is found to result in the best performance. The performance of the weighted LSP distance measure is compared with that of the other popular LP distance measures (such as the Itakura, cepstral, weighted cepstral, root-power-sum, log area ratio and reflection coefficient distance measures). The weighted LSP distance measure is found to perform significantly better than these popular LP distance measures.> Kuldip K. Paliwal |
ICASSP | 1 |
| 1988 | Low bit-rate image coding using 2-D linear prediction and 2-D stochastic excitationabstractThe use of stochastic excited linear predictive coding method is investigated for image coding and its parameters are studied. It is shown that it is not necessary to transmit local bias values of the image frames. It is also shown that the stochastic excitation is not adequate to represent the prediction residual signal. In order to get good performance from this coder, it is necessary to generate the codebook from the actual prediction residual signal.> Kuldip K. Paliwal |
ICASSP | 1 |
| 1987 | Estimation of noise variance from the noisy AR signal and its application in speech enhancementabstractIn a number of applications involving the processing of noisy signals, it is desirable to know a priori the noise variance. We propose here a method of estimating the noise variance from the autoregressive (AR) signal corrupted by the additive white noise. This method first estimates the AR parameters from the high-order Yule-Walker equations and then uses these AR parameters to estimate the noise variance from the low-order Yule-Walker equations. The method is studied for a number of examples of noisy AR signals and its performance is found to be close to the Cramer-Rao lower bound for high signal-to-noise ratios. It is also used in a speech enhancement application where its performance is studied for stationary as well as nonstationary noise conditions. The results are found to be encouraging. Kuldip K. Paliwal |
ICASSP | 1 |
| 1987 | A speech enhancement method based on Kalman filteringabstractIn this paper, the problem of speech enhancement when only corrupted speech signal is available for processing is considered. For this, the Kalman filtering method is studied and compared with the Wiener filtering method. Its performance is found to be significantly better than the Wiener filtering method. A delayed-Kalman filtering method is also proposed which improves the speech enhancement performance of Kalman filter further. Kuldip K. Paliwal, Anjan Basu |
ICASSP | 1 |
| 1986 | Speech enhancement using multi-pulse excited linear prediction systemabstractThe multi-pulse excited linear prediction (MPELP) system is proposed for speech enhancement. It is shown that for successful enhancement of speech the error-weighting filter should not be used in the MPELP system. A new method (the constrained forward-backward correlation prediction method) is proposed for accurate estimation of LP coefficients from noisy speech. This method guarantees the stability of the estimated all-pole filter which is an important prerequirement for the MPELP system. It is shown that the MPELP system can improve the SNR of 0 dB speech by as much as 5.4 dB. Kuldip K. Paliwal |
ICASSP | 1 |
| 1986 | A noise-compensated long correlation matching method for AR spectral estimation of noisy signalsabstractA noise-compensated long correlation matching (NCLCM) method is proposed for autoregressive (AR) spectral estimation of the noisy AR signals. This method first computes the AR parameters from the high-order Yule-Walker equations. Next, it employs these AR parameters and uses the low-order Yule-Walker equations to compensate the zeroth autocorrelation coefficient for the additive white noise. Finally, it solves the low- as well as high-order Yule-Walker equations in a least-squares sense to determine the AR parameters. It is shown that for the noisy AR signals the NCLCM method performs better than the conventional Burg method and the high-order Yule-Walker method. Kuldip K. Paliwal |
ICASSP | 1 |
| 1985 | A study of three coders (sub-band, RELP and MPE) for speech with additive white noiseabstractThe following three speech coders are implemented for a bitrate of 9.6 kbits/s 1) Sub-band coder, 2) Residual Excited Linear Predictive (RELP) coder, and 3) Multi-Pulse Excited linear predictive (MPE) coder. Performance of these coders is evaluated for speech corrupted by additive white noise. Evaluation of speech coders is done both subjectively and objectively. The MPE coder is found to give the best performance among the three coders. It is also shown that the MPE coder can be used for noisy speech with signal-to-noise ratio as low as -10 dB giving reasonably good quality speech provided 1) one does not use the error weighting filter and 2) one can use a better LP analysis algorithm which can estimate LP coefficients correctly from noisy speech. Kuldip K. Paliwal, Torbjørn Svendsen |
ICASSP | 1 |
| 1984 | Effect of preemphasis on vowel recognition performance
Kuldip K. Paliwal |
Speech Commun. | 1 |
| 1984 | Performance of the weighted burg methods of ar spectral estimation for pitch-synchronous analysis of voiced speech
Kuldip K. Paliwal |
Speech Commun. | 1 |
| 1984 | A comparative performance evaluation of pitch estimation methods for TDHS/sub-band coding of speech
Kuldip K. Paliwal, A. I. Aarskog |
Speech Commun. | 1 |
| 1983 | Application of k-Nearest-Neighbor Decision Rule in Vowel RecognitionabstractThe k-nearest-neighbor decision rule is known to provide a useful nonparametric procedure for pattern classification. This rule is applied here to a vowel recognition problem and the effect of the number (k) of nearest neighbors, the size of the trained set and the type of the distance measure on vowel recognition performance is studied. It is shown that the vowel recognition performance remains approximately constant for all the values of k. The recognition performance initially improves with the size of the training set and then converges to an asymptotic value. Selection of a better distance measure leads to a significant improvement in vowel recognition performance. Kuldip K. Paliwal, P. V. S. Rao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1983 | A study of two-formant models for vowel identification
Kuldip K. Paliwal, William A. Ainsworth, D. Lindsay |
Speech Commun. | 1 |
| 1983 | A synthesis-based method for pitch extraction
Kuldip K. Paliwal, P. V. S. Rao |
Speech Commun. | 1 |
| 1982 | A modification over Sakoe and Chiba's dynamic time warping algorithm for isolated word recognitionabstractA modification over Sakoe and Chiba's dynamic time warping algorithm for isolated word recognition is proposed. It is shown that this modified algorithm works better without any slope constraint. Also, this algorithm not only consumes less computation time but also improves the word recognition accuracy. Kuldip K. Paliwal, Anant Agarwal, Sarvajit S. Sinha |
ICASSP | 1 |
| 1982 | On the performance of the quefrency-weighted cepstral coefficients in vowel recognition
Kuldip K. Paliwal |
Speech Commun. | 1 |