Tieran Zheng

dblp:18/8675 · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
13since 2021 · last 2025
0009-0006-5179-0278ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dual Orthogonality Sub-center Loss for Enhanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Tieran Zheng, Guibin Zheng, Yongjun He 0002
INTERSPEECH3
2025 Adaptive Across-Subcenter Representation Learning for Imbalanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Guibin Zheng, Tieran Zheng, Yongjun He 0002
INTERSPEECH4
2025 Knowledge Distillation Method for Pruned RNN-T Models via Pruning Bounds Sharing and Losses Confusion
Xiaocan Zhang, Guibin Zheng, Chenhao Jing, Jiqing Han 0001, Tieran Zheng
INTERSPEECH6
2025 Joint Energy-Based Model for Semi-Supervised Respiratory Sound Classification: A Method of Insensitive to Distribution Mismatch
abstract
Semi-supervised learning effectively mitigates the lack of labeled data by introducing extensive unlabeled data. Despite achieving success in respiratory sound classification, in practice, it usually takes years to acquire a sufficiently sizeable unlabeled set, which consequently results in an extension of the research timeline. Considering that there are also respiratory sounds available in other related tasks, like breath phase detection and COVID-19 detection, it might be an alternative manner to treat these external samples as unlabeled data for respiratory sound classification. However, since these external samples are collected in different scenarios via different devices, there inevitably exists a distribution mismatch between the labeled and external unlabeled data. For existing methods, they usually assume that the labeled and unlabeled data follow the same data distribution. Therefore, they cannot benefit from external samples. To utilize external unlabeled data, we propose a semi-supervised method based on Joint Energy-based Model (JEM) in this paper. During training, the method attempts to use only the essential semantic components within the samples to model the data distribution. When non-semantic components like recording environments and devices vary, as these non-semantic components have a small impact on the model training, a relatively accurate distribution estimation is obtained. Therefore, the method exhibits insensitivity to the distribution mismatch, enabling the model to leverage external unlabeled data to mitigate the lack of labeled data. Taking ICBHI 2017 as the labeled set, HF_Lung_V1 and COVID-19 Sounds as the external unlabeled sets, the proposed method exceeds the baseline by 12.86.
Wenjie Song 0003, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng, Yongjun He 0002
IEEE J. Biomed. Health Informatics4
2024 Modeling Quasi-Periodic Dependency via Self-Supervised Pre-Training for Respiratory Sound Classification
abstract
Despite the success of self-supervised respiratory sound classification methods, they do not consider that respiratory sounds are quasi-periodic signals with repetitive patterns in successive breaths, which is vital for distinguishing respiratory sounds from non-quasi-periodic sounds like noises. Therefore, the existing methods may achieve limited improvement due to ignoring the quasi-periodic dependency. To this end, considering that the segments containing the same respiratory sound pattern should be similar in a sample, we extract the segment-wise representations and evaluate the similarity between the periodic-dependent representations via a sparse self-relation matrix. By defining a periodic consistency loss, we push the sparse self-relation matrixes of two clips of the same sample closer, encouraging a larger similarity between the representations. In this manner, the method can focus more on the respiratory sound-related quasi-periodic patterns that repeatedly recur in the periodic-dependent segments. Taking HF_Lung_V1 and COVID-19 Sounds as pre-training sets, the method exceeds the baseline by 7.67% on the ICBHI 2017 classification task.
Wenjie Song 0003, Jiqing Han 0001, Jianchen Li, Guibin Zheng, Tieran Zheng, Yongjun He 0002
ICASSP5
2024 Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event Detection
abstract
Overlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events using shared and entangled frame-wise features, which degrades the feature discrimination. To solve the problem, we propose a disentangled feature learning framework to learn a category-specific representation. Specifically, we employ different projectors to learn the frame-wise features for each category. To ensure that these feature does not contain information of other categories, we maximize the common information between frame-wise features within the same category and propose a frame-wise contrastive loss. In addition, considering that the labeled data used by the proposed method is limited, we propose a semi-supervised frame-wise contrastive loss that can leverage large amounts of unlabeled data to achieve feature disentanglement. The experimental results demonstrate the effectiveness of our method.
Yadong Guan, Jiqing Han 0001, Wenjie Song 0003, Guibin Zheng, Tieran Zheng, Yongjun He 0002
ICASSP6
2024 Sound Activity-Aware Based Cross-Task Collaborative Training for Semi-Supervised Sound Event Detection
abstract
The training of sound event detection (SED) models remains a challenge of insufficient supervision due to limited frame-wise labeled data. Mainstream research on this problem has adopted semi-supervised training strategies that generate pseudo-labels for unlabeled data and use these data for the training of a model. Recent works further introduce multi-task training strategies to impose additional supervision. However, the auxiliary tasks employed in these methods either lack frame-wise guidance or exhibit unsuitable task designs. Furthermore, they fail to exploit inter-task relationships effectively, which can serve as valuable supervision. In this paper, we introduce a novel task, sound occurrence and overlap detection (SOD), which detects predefined sound activity patterns, including non-overlapping and overlapping cases. On the basis of SOD, we propose a cross-task collaborative training framework that leverages the relationship between SED and SOD to improve the SED model. Firstly, by jointly optimizing the two tasks in a multi-task manner, the SED model is encouraged to learn features sensitive to sound activity. Subsequently, the cross-task consistency regularization is proposed to promote consistent predictions between SED and SOD. Finally, we propose a pseudo-label selection method that uses inconsistent predictions between the two tasks to identify potential wrong pseudo-labels and mitigate their confirmation bias. In the inference phase, only the trained SED model is used, thus no additional computation and storage costs are incurred. Extensive experiments on the DESED dataset demonstrate the effectiveness of our method.
Yadong Guan, Jiqing Han 0001, Shiwen Deng, Guibin Zheng, Tieran Zheng, Yongjun He 0002
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Distance Metric-Based Open-Set Domain Adaptation for Speaker Verification
abstract
Domain shift poses a significant challenge in speaker verification, especially in open-set scenarios where the speaker categories are disjoint between the source and target domains. To alleviate the domain shift, traditional domain adaptation methods typically align the source and target distributions in the speaker embedding space, but this may cause the overlap of embeddings from different speakers. To address this problem, this paper proposes to perform the domain alignment in a novel distance metric space, where the source and target domains exhibit the shared within-speaker and between-speaker categories. Thus, the discrepancy between the source and target domains arises only from the domain shift. We refer to the proposed method as Cross-Domain Distance Metric Adaptation (CDMA), in which the within- and between-speaker distance distributions in the target domain are aligned with the source distance distributions and further separated to minimize their overlap. This alignment and separation require estimating the within- and between-speaker distance distributions based on speaker labels, which are unavailable in the unlabeled target domain. Thus, we further propose a learnable speaker clustering method called Graph Convolutional Network with Graph Pruning (GCN-GP). This method generates high-quality pseudo-labels to estimate the two distance distributions in the target domain. Experimental results demonstrate that our method achieves state-of-the-art performance on the FFSVC2022 and VOiCES datasets.
Jianchen Li, Jiqing Han 0001, Fan Qian, Tieran Zheng, Yongjun He 0002, Guibin Zheng
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Graph-Based Spectro-Temporal Dependency Modeling for Anti-Spoofing
abstract
A great deal of recent research reveals that artifacts introduced by spoofing algorithms reside in specific frequency subbands or temporal segments. Therefore, the performance of spoofing detection can be improved by focusing on these regions. However, it is difficult for the detection system to choose an appropriate region when it encounters an unknown spoofing algorithm, resulting in poor generalization. Actually, there is a noticeable difference in the inter-region relationship between the bonafide and spoofed speeches. We name the inter-region relationship spectro-temporal dependency and design a method to model it for anti-spoofing. By focusing on the general dependency difference rather than specific regions, the generalization ability of the detection system can be improved. We employ a graph neural network to model the dependency and incorporate prior knowledge into the graph by designing the graph structure and edge weight, which forces the network to pay more attention to potential relationships. In addition, an attention mechanism is introduced in the graph pooling to focus on more critical nodes. The proposed method achieves an equal error rate of 0.58% on the ASVspoof 2019 LA dataset and outperforms all competing systems.
Shiwen Deng, Tieran Zheng, Yongjun He 0002, Jiqing Han 0001
ICASSP3
2023 Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
abstract
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g. the log-Mel spectrogram’s average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
Jian Guan 0001, Youde Liu, Qiaoxi Zhu, Tieran Zheng, Jiqing Han 0001, Wenwu Wang 0001
ICASSP4
2023 Mutual Information-based Embedding Decoupling for Generalizable Speaker Verification
Jianchen Li, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Yongjun He 0002, Guibin Zheng
INTERSPEECH4
2021 Model-Agnostic Fast Adaptive Multi-Objective Balancing Algorithm for Multilingual Automatic Speech Recognition Model Training
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
Interspeech2
2021 Exploring attention mechanisms based on summary information for end-to-end automatic speech recognition
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
Neurocomputing2
2020 Structured Sparse Attention for end-to-end Automatic Speech Recognition
abstract
The Softmax normalization function-based attention mechanism is often employed by End-to-End Automatic Speech Recognition (E2E ASR) models to tell the network where to focus within the input. However, this mechanism leads to the attention distribution becoming increasingly flatter as the input sequence length increases, since the output probability of this function is dense and nonnegative, which makes it unable to highlight the important information in speech. In this paper, we present two sparse attention mechanisms for ASR tasks with long utterances, which try to improve the attention mechanism by introducing the sparse transformation. First, we propose to replace the Softmax with the Sparsemax that normalizes the attention weight by finding the closest point in the probability simplex. Then, considering the structured characteristics, the pronunciation has a relatively stable duration. Therefore, we further present a structured sparse transformation that forces the networks to pay attention to a continuous segment of speech by applying the l2penalty. A noniterative solution algorithm that can be used in the backpropagation is designed here. The experiments show that our methods achieve better ASR results compared to a well-tuned attention-based baseline system on a character ASR task.
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
ICASSP2
2020 Error Heuristic Based Text-Only Error Correction Method for Automatic Speech Recognition
Linhan Zhang, Tieran Zheng, Jiabin Xue
ICONIP (1)2
2019 Convolutional Grid Long Short-Term Memory Recurrent Neural Network for Automatic Speech Recognition
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
ICONIP (5)2
2018 Deep Neural Network Based Discriminative Training for I-Vector/PLDA Speaker Verification
abstract
In the studies of i-vector based speaker verification, the discriminative training of probabilistic linear discriminative analysis (PLDA) model has been proven to be an effective way to improve performance. This paper focuses on using a deep neural network (DNN) to strengthen the original discriminatively trained classifiers by its strong capability of nonlinear modeling representation. We first propose a deep neural network based dimensionality reduction model to replace the linear discriminant analysis (LDA) process, and then a discriminative training algorithm is also proposed to jointly optimize the network and PLDA scoring function under single discriminative criterion. Our experiments show that performance improvements are achieved in the male trials of short2-short3 core data set of NIST SRE08.
Tieran Zheng, Jiqing Han 0001, Guibin Zheng
ICASSP1
2017 Learning Deep Neural Network Based Kernel Functions for Small Sample Size Classification
Tieran Zheng, Jiqing Han 0001, Guibin Zheng
ICONIP (1)1
2016 Speaker Verification via Modeling Kurtosis Using Sparse Coding
abstract
This paper proposes a new model for speaker verification by employing kurtosis statistical method based on sparse coding of human auditory system. Since only a small number of neurons in primary auditory cortex are activated in encoding acoustic stimuli and sparse independent events are used to represent the characteristics of the neurons. Each individual dictionary is learned from individual speaker samples where dictionary atoms correspond to the cortex neurons. The neuron responses possess statistical properties of acoustic signals in auditory cortex so that the activation distribution of individual speaker’s neurons is approximated as the characteristics of the speaker. Kurtosis is an efficient approach to measure the sparsity of the neuron from its activation distribution, and the vector composed of the kurtosis of every neuron is obtained as the model to characterize the speaker’s voice. The experimental results demonstrate that the kurtosis model outperforms the baseline systems and an effective identity validation function is achieved desirably.
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
Int. J. Pattern Recognit. Artif. Intell.3
2015 Soft Margin Based Low-Rank Audio Signal Classification
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
Neural Process. Lett.3
2014 Robust minimum statistics project coefficients feature for acoustic environment recognition
abstract
Acoustic environment recognition has been widely used in many applications, and is a considerable difficult problem for the real-life and complex environment. This paper proposes a novel feature, named minimum statistics project coefficients (MSPC), and intents to solve this problem. The MSPC feature is extracted from the background sound which is more robust than the foreground sound for the task of acoustic environment recognition. Experimental results show the outstanding performance of the MSPC feature compared with the conventional acoustic features, especially in very complex acoustic environments.
Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP4
2013 Upper and lower bounds for approximation of the Kullback-Leibler divergence between Hidden Markov models
abstract
The Kullback-Leibler (KL) divergence is often used for a similarity comparison between two Hidden Markov models (HMMs). However, there is no closed form expression for computing the KL divergence between HMMs, and it can only be approximated. In this paper, we propose two novel methods for approximating the KL divergence between the left-to-right transient HMMs. The first method is a product approximation which can be calculated recursively without introducing extra parameters. The second method is based on the upper and lower bounds of KL divergence, and the mean of these bounds provides an available approximation of the divergence. We demonstrate the effectiveness of the proposed methods through experiments including the deviations to the numerical approximation and the task of predicting the confusability of phone pairs. Experimental results show that the proposed product approximation is comparable with the current variational approximation, and the proposed approximation based on bounds performs better than current methods in the experiments.
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP3
2013 Case based reasoning solution to the problem of sustained learning in keyword spotting
abstract
In some practical keyword spotting applications, users or service providers are willing to provide spotting-result feedback to help improve system performance. To do so, they require a keyword spotting technique with a sustained learning ability. This paper presents a new Chinese keyword spotting method based on a case based reasoning framework. Two level keyword case representations are adopted based on a set of symbols that are discriminative both in acoustic feature vector space and in semantic space. Then case bases are indexed with a tree structure and searched for test speech based on an elastic matching strategy. Finally, the feedback is used to adjust the statistics attached to the cases or to append new cases. Two experiments were conducted to compare our approach with a syllable lattice based method and to test the sustained learning ability.
Tieran Zheng, Jiqing Han 0001, Guibin Zheng, Shiwen Deng
ICASSP1
2013 Guarantees of Augmented Trace Norm Models in Tensor Recovery
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
IJCAI3
2013 Audio Segment Classification Using Online Learning Based Tensor Representation Feature Discrimination
abstract
In order to naturally combine audio information from different dimensions and build robust audio processing system, a novel framework based on low-rank tensor representation features for audio segment classification is proposed in this paper. The audio signal is first transformed into tensor format data, and then these tensor data are mapped to a low-rank space which is insensitive under certain noises, especially white Gaussian noise and gross corruptions. For these low-rank tensor based features, tensor classification via a linear classifier based on minimization a smooth loss function regularized by the trace norm proposed recently is used. Most previous methods find the weight tensor and bias in batch-mode learning, which makes them inefficient for large-scale problems. In this paper, we propose to address this problem with an online learning algorithm based on the accelerated proximal gradient (APG) method, which scales up gracefully to large data sets. Experiments on simulation and real audio data demonstrate the efficiency of the methods.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng, Shiwen Deng
IEEE Trans. Speech Audio Process.3
2013 Identification of Objectionable Audio Segments Based on Pseudo and Heterogeneous Mixture Models
abstract
In this paper, we generalize the Gaussian Mixture Model (GMM) in two ways: a) by introducing novel distance measures between two vectors based on nonlinear maps to give more general mixture models; b) by building mixture models based on multiple different kinds of distributions. These two generalizations cope with different problems arisen in feature modeling. Mixture model obtained by first method is called pseudo Gaussian Mixture Model (pseudo GMM). Compared to the traditional GMM, pseudo GMM with nonlinear maps have better performance on nonlinear problems, while the computational complexity is almost the same as the Expectation-Maximization (EM) algorithm for traditional GMM according to the iteration procedures. The second generalization considers that in practice the practical learning problem often involves multiple, heterogeneous data sources, while classical mixture models are based on a single kind of distribution. In this work, we consider heterogeneous mixture models (hetMM) based on multiple different kinds of distributions. Different types of distributions in hetMM may have quite different properties and may capture different features of the data. Component classifiers including pseudo and hetMM based classifiers are employed in our task of erotic audio recognition. Experimental results with classifiers built based on pseudo GMM and hetMM for erotic audio recognition demonstrate the effectiveness of the proposed model. Online and off-line experiments show that the proposed approach is highly effective for erotic audio recognition.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
IEEE Trans. Speech Audio Process.3
2013 Audio classification with low-rank matrix representation features
abstract
In this article, a novel framework based on trace norm minimization for audio classification is proposed. In this framework, both the feature extraction and classification are obtained by solving corresponding convex optimization problem with trace norm regularization. For feature extraction, robust principle component analysis (robust PCA) via minimization a combination of the nuclear norm and the ℓ 1 -norm is used to extract low-rank matrix features which are robust to white noise and gross corruption for audio signal. These low-rank matrix features are fed to a linear classifier where the weight and bias are learned by solving similar trace norm constrained problems. For this linear classifier, most methods find the parameters, that is the weight matrix and bias in batch-mode, which makes it inefficient for large scale problems. In this article, we propose a parallel online framework using accelerated proximal gradient method. This framework has advantages in processing speed and memory cost. In addition, as a result of the regularization formulation of matrix classification, the Lipschitz constant was given explicitly, and hence the step size estimation of the general proximal gradient method was omitted, and this part of computing burden is saved in our approach. Extensive experiments on real data sets for laugh/non-laugh and applause/non-applause classification indicate that this novel framework is effective and noise robust.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
ACM Trans. Intell. Syst. Technol.3
2012 A solution to residual noise in speech denoising with sparse representation
abstract
As a promising technique, sparse representation has been extensively investigated in signal processing community. Recently, sparse representation is widely used for speech processing in noisy environments; however, many problems need to be solved because of the particularity of speech. One assumption for speech denoising with sparse representation is that the representation of speech over the dictionary is sparse, while that of the noise is dense. Unfortunately, this assumption is not sustained in speech denoising scenario. We find that many noises, e.g., the babble and white noises, are also sparse over the dictionary trained with clean speech, resulting in severe residual noise in sparse enhancement. To solve this problem, we propose a novel residual noise reduction (RNR) method which first finds out the atoms which represents the noise sparely, and then ignores them in the reconstruction of speech. Experimental results show that the proposed method can reduce residual noise substantially.
Yongjun He 0002, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng
ICASSP4
2012 Sparse power spectrum based robust voice activity detector
abstract
This paper presents a robust approach to improve the performance of voice activity detector (VAD) in low signal-to-noise ratio (SNR) noisy environments. To this end, we first generate sparse representations by Bregman Iteration based sparse decomposition with a learned over-complete dictionary, and derive a kind of audio feature called sparse power spectrum from the sparse representations. we then propose a method to calculate the short segment average spectrum and long segment average spectrum from sparse power spectrum. Finally, we design a criterion to detect speech region and non-speech region based on the above average spectrum. Experiments show that the proposed approach further improves the performance of VAD in low SNR noisy environments.
Datao You, Jiqing Han 0001, Guibin Zheng, Tieran Zheng
ICASSP4
2012 A Novel Confidence Measure Based on Context Consistency for Spoken Term Detection
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
INTERSPEECH3
2012 Low-rank Audio Signal Classification Under Soft Margin and Trace Norm Constraints
Ziqiang Shi, Tieran Zheng, Jiqing Han 0001, Shiwen Deng
INTERSPEECH2
2012 Sparse-Based auditory Model for robust speaker Recognition
abstract
The mismatch between the training and the testing environments greatly degrades the performance of speaker recognition. Although many robust techniques have been proposed, speaker recognition in mismatch condition is still a challenge. To solve this problem, we propose a sparse-based auditory model as the front-end of speaker recognition by simulating auditory processing of speech signal. To this end, we introduce narrow-band filter-bank instead of the widely used wide-band filter-bank to simulate the basilar membrane filter-bank, use sparse representation as the approximation of basilar membrane coding strategy, and incorporate the frequency selectivity enhance mechanism between tectorial membrane and basilar membrane by practical engineering approximation. Compared with the standard Mel-frequency cepstral coefficient approach, our preliminary experimental results indicate that the sparse-based auditory model consistently improve the robustness of speaker recognition in mismatched condition.
Datao You, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
Int. J. Pattern Recognit. Artif. Intell.3
2011 A modified MAP criterion based on hidden Markov model for voice activity detecion
abstract
The maximum a posteriori (MAP) criterion is broadly used in the statistical model-based voice activity detection (VAD) approaches. In the conventional MAP criterion, however, the inter-frame correlation of the voice activity is not taken into consideration. In this paper, we proposes a novel modified MAP criterion based on a two-state hidden Markov model (HMM) to improve the performance of the VAD, and the the inter-frame correlation of the voice activity is modeled. With the proposed MAP criterion, the decision rule is derived by explicitly incorporating the a priori, a posteriori, and inter-frame correlation information into the likelihood ratio test (LRT). In the LRT, a compensation factor for the hypothesis of speech presence is used to regulate the trade-off between the probability of detection and the false alarm probability. Experimental results show the superiority of the VAD algorithm based on the proposed MAP criterion in comparison with that based on the recent conditional MAP criterion (CMAP) under various noise conditions.
Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP3
2011 Compensation of partly reliable components for band-limited speech recognition with missing data techniques
abstract
Mismatch in speech bandwidth between training and real operation greatly degrades the performance of automatic speech recognition (ASR) systems. Missing feature technique (MFT) is effective in handling bandwidth mismatch. However, current MFT-based methods ignore the mismatch in the filter bank channels which cover the upper and lower limit cutoff frequencies. To solve this problem, we propose to partition the feature into reliable, unreliable and partly reliable parts, and then modify the probability density functions (PDFs) of the partly reliable part to match band-limited features. Experiments showed that such compensation further improved the performances of MFT-based methods under band-limited conditions.
Yongjun He 0002, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP3
2011 A cochlear neuron based robust feature for speaker recognition
abstract
In this paper, a robust feature for text-independent speaker recognition is proposed, which simulate the response mode of cochlear neurons in processing acoustic signal. The feature is derived from sparse coding coefficient which is computed on a learned over-complete dictionary, and the dictionary is considered similar to part of speech sensitive cochlear neurons. Furthermore, the feature is generated without dimension reducing and de-correlation. The robust feature is implemented to address the problem of mismatch situation between training and testing. Experiments show that the proposed feature outperforms the Mel-frequency cepstral coefficients (MFCC) feature, especially under noisy environments, the equal error rate (EER) of the MFCC drops to 21.6% (10 dB) from 10.3% (25 dB), while the EER of the proposed feature is also 6.6% (10 dB) with no degradation.
Datao You, Jiqing Han 0001, Tieran Zheng
ICASSP4
2011 A Novel Framework Based on Trace Norm Minimization for Audio Event Detection
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
ICONIP (2)3
2011 AUC Optimization Based Confidence Measure for Keyword Spotting
Jiqing Han 0001, Tieran Zheng
INTERSPEECH3
2011 Real-World Speech/Non-Speech Audio Classification Based on Sparse Representation Features and GPCs
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
INTERSPEECH3
2010 Study on the Recognition of Objectionable Audio
abstract
In this paper, a novel method from the feature — porno-sounds recognition — point of view is proposed to detect adult video sequences automatically which may serve as a verification step, a supplementary method or an independent detector. To the specificity of erotic sound, its feature analysis is given. Based on the popular features, histograms and contours are introduced as new sets of features. At the same time due to the complexity of outside data, a general framework called in-class clustering is proposed which selects the most representative subclass for training and classification. All these efforts increase the recall rate and decrease the false positive rate. Experiments on real data from the Internet indicate that the proposed method yields superior performance with 89.17% recall rate and 10.78% false positive rate being achieved.
Ziqiang Shi, Boyang Gao, Tieran Zheng, Jiqing Han 0001
Int. J. Pattern Recognit. Artif. Intell.3