Qianhua He

dblp:02/4984 · also Qian-Hua He · DBLP profile ↗
← Back
48ranked-venue papers
2as first author
24since 2021 · last 2025
0000-0002-9079-4566ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 20 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments
abstract
Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features produced during enhancement. In this paper, we design a low-complexity network to address the challenges of KWS in noisy environments. We integrate the tasks of both SE and KWS into a unified network that learns a shared representation from both tasks. The proposed network features two blocks: a Residual Full-band and Sub-band Fusion (RFSF) block, and a Deformable Transition (DT) block. Our dual-task network surpasses existing KWS models in accuracy with low complexity, making it suitable for deployment on edge devices.
Qianhua He, Yanxiong Li, Zunxian Liu, Mingru Yang, Jinxin Huang
ICASSP2
2025 An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty
abstract
Deep learning has achieved success across many domains when sufficient training samples are available. However, the commonly used mini-batch stochastic gradient descent (SGD) training paradigm treats each sample equally, resulting in massive computational waste on samples that are easily identifiable. In contrast, low-quality samples, such as those with erroneous labels, can negatively impact the training process. To address these issues, we propose a training sample utilization method based on sample uncertainty. Once the model has acquired preliminary decision-making abilities, the class uncertainty for each sample can be evaluated within a training epoch. Subsequently, the samples are probabilistically selected based on their uncertainty for the next epoch. Experiments conducted on the GSC v2 and CIFAR-10 datasets demonstrate that the proposed method can reduce training time by over 32% and 58%, respectively, with only a loss of 1% performance. Additionally, the method has the capability to mitigate the adverse effects of samples with erroneous labels.
Jinxin Huang, Qianhua He, Jiezhi Xu, Sam Kwong, Mingru Yang
ICASSP2
2025 Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection
abstract
Personalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and target domains. In this paper, we formulate a KWS task of Cross-Domain Few-Shot Open-Set (CD-FSOS) and propose a dedicated framework Adapt-KWS to bridge the distribution gap between the source domain and target open-set domain with quite limited enrollment data. The proposed Adapt-KWS consists of a set of Custom-Keyword Adapters (CKAs) and a Prototype Reprojection Module (PRM). CKAs enable the efficient adaptation to new target tasks with limited training samples, aiming to improve cross-domain generalization. PRM reprojects the support prototypes into the query embedding space to enhance their alignment, mitigating the potential covariate shift between open-set queries and enrollments. Experimental results demonstrate the effectiveness of our framework and proposed modules on multiple datasets. Code will be available at: https://github.com/Raynaming/CD-FSOS-KWS.
Mingru Yang, Qianhua He, Jinxin Huang, Zunxian Liu, Yanxiong Li
ICASSP2
2025 Fully Few-shot Class-incremental Audio Classification Using Multi-level Embedding Extractor and Ridge Regression Classifier
Yongjie Si, Yanxiong Li, Jiaxin Tan, Qianhua He, Il-Youp Kwak
INTERSPEECH4
2025 Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
Mingru Yang, Yanmei Gu, Qianhua He, Yanxiong Li, Peirong Zhang 0001, Huijia Zhu, Weiqiang Wang 0002
INTERSPEECH3
2025 Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk Minimization
abstract
With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.
Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 0001, Haolin He, Huijia Zhu, Weiqiang Wang 0002
ACM Multimedia3
2025 Noise-robust feature extraction for keyword spotting based on supervised adversarial domain adaptation training strategies
Qianhua He, Zunxian Liu, Mingru Yang, Wenwu Wang 0001
Speech Commun.2
2024 Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
Yanxiong Li, Jiaxin Tan, Jialong Li 0002, Yongjie Si, Qianhua He
INTERSPEECH6
2024 Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
Yongjie Si, Yanxiong Li, Jialong Li 0002, Jiaxin Tan, Qianhua He
INTERSPEECH5
2024 Detecting video anomalies by jointly utilizing appearance and skeleton information
Wenfeng Pang, Qianhua He, Yanxiong Li, Noman Ahmed
Expert Syst. Appl.2
2024 Few-Shot Class-Incremental Audio Classification With Adaptive Mitigation of Forgetting and Overfitting
abstract
Few-shot Class-incremental Audio Classification (FCAC) is a task to continuously identify incremental classes with only few training samples after training the model on base classes with abundant samples. The key to solving the FCAC problem is to ensure that the model has good stability (without forgetting base classes) and strong plasticity (without overfitting incremental classes). In this paper, we propose a FCAC method which is able to adaptively mitigate the model's forgetting of base classes and overfitting of incremental classes. Our model consists of an embedding extractor and an expandable classifier. The former is the backbone of a residual network and is frozen after being trained using sufficient samples of base classes, whereas the latter can be expandable and is updated using few training samples of incremental classes in each incremental session. The expandable classifier consists of two branches and one fusion module. The two branches are designed to mitigate the model's forgetting of base classes and overfitting of incremental classes, respectively. The fusion module is designed to adaptively fuse predictions output by the above two branches. In addition, we define two losses for model training in the base and incremental sessions, respectively. Three experimental datasets (NSynth-100, FSC-89 and LS-100) are created by randomly choosing samples from audio corpora of NSynth, FSD-MIX-CLIP and LibriSpeech, respectively. Experimental results demonstrate that our proposed method outperforms all previous methods in accuracy and has advantage over most previous methods in computational load. The code is available athttps://github.com/Jialongdustin/AMFO.
Yanxiong Li, Jialong Li 0002, Yongjie Si, Jiaxin Tan, Qianhua He
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Clean Sample Guided Self-Knowledge Distillation for Image Classification
abstract
For two-stage knowledge distillation, the combination with Data Augmentation (DA) is straightforward and effective. Yet, for online Self-knowledge Distillation (SD), DA is not always beneficial because of the absence of a trustworthy teacher model. To address this issue, this paper proposes an SD method named Clean sample guided Self-knowledge Distillation (CleanSD), in which the original clean sample is used as a guide when the model is trained with the augmented samples. The implementation of the CleanSD comes with two DA techniques, namely Mixup (for label-mixing) and Cutout (for label-preserving). Results on CIFAR-100 demonstrate that error rates obtained by the proposed CleanSD are reduced by 2.59%, 1.39%, and 0.47-1.20%, compared to that obtained by the baseline, the vanilla DA techniques, and other peer SD methods, respectively. In addition, the effectiveness and robustness of the CleanSD are verified across multiple DA methods and datasets.
Jiyue Wang, Yanxiong Li, Qianhua He, Wei Xie 0013
ICASSP3
2023 Few-shot Class-incremental Audio Classification Using Stochastic Classifier
Yanxiong Li, Wenchang Cao, Jialong Li 0002, Wei Xie 0013, Qianhua He
INTERSPEECH5
2023 Few-shot Class-incremental Audio Classification Using Adaptively-refined Prototypes
abstract
New classes of sounds constantly emerge with a few samples, making it challenging for models to adapt to dynamic acoustic environments. This challenge motivates us to address the new problem of few-shot class-incremental audio classification. This study aims to enable a model to continuously recognize new classes of sounds with a few training samples of new classes while remembering the learned ones. To this end, we propose a method to generate discriminative prototypes and use them to expand the model's classifier for recognizing sounds of new and learned classes. The model is first trained with a random episodic training strategy, and then its backbone is used to generate the prototypes. A dynamic relation projection module refines the prototypes to enhance their discriminability. Results on two datasets (derived from the corpora of Nsynth and FSD-MIX-CLIPS) show that the proposed method exceeds three state-of-the-art methods in average accuracy and performance dropping rate.
Wei Xie 0013, Yanxiong Li, Qianhua He, Wenchang Cao, Tuomas Virtanen
INTERSPEECH3
2023 Few-shot class-incremental audio classification via discriminative prototype learning
Wei Xie 0013, Yanxiong Li, Qianhua He, Wenchang Cao
Expert Syst. Appl.3
2023 Few-Shot Speaker Identification Using Lightweight Prototypical Network With Feature Grouping and Interaction
abstract
Existing methods for few-shot speaker identification (FSSI) obtain high accuracy, but their computational complexities and model sizes need to be reduced for lightweight applications. In this work, we propose a FSSI method using a lightweight prototypical network with the final goal to implement the FSSI on intelligent terminals with limited resources, such as smart watches and smart speakers. In the proposed prototypical network, an embedding module is designed to perform feature grouping for reducing the memory requirement and computational complexity, and feature interaction for enhancing the representational ability of the learned speaker embedding. In the proposed embedding module, audio feature of each speech sample is split into several low-dimensional feature subsets that are transformed by a recurrent convolutional block in parallel. Then, the operations of averaging, addition, concatenation, element-wise summation and statistics pooling are sequentially executed to learn a speaker embedding for each speech sample. The recurrent convolutional block consists of a block of bidirectional long short-term memory, and a block of de-redundancy convolution in which feature grouping and interaction are conducted too. Our method is compared to baseline methods on three datasets that are selected from three public speech corpora (VoxCeleb1, VoxCeleb2, and LibriSpeech). The results show that our method obtains higher accuracy under several conditions, and has advantages over all baseline methods in computational complexity and model size.
Yanxiong Li, Wenchang Cao, Qisheng Huang, Qianhua He
IEEE Trans. Multim.5
2023 Audiovisual Dependency Attention for Violence Detection in Videos
abstract
Violence detection in videos can help maintain public order, detect crimes, or provide timely assistance. In this paper, we aim to leverage multimodal information to determine whether successive frames contain violence. Specifically, we propose an audiovisual dependency attention (AVD-attention) module modified from the co-attention architecture to fuse visual and audio information, unlike commonly used methods such as the feature concatenation, addition, and score fusion. Because the AVD-attention module’s dependency map contains sufficient fusion information, we argue that it should be applied more sufficiently. A combination pooling method is utilized to convert the dependency map to an attention vector, which can be considered a new feature that includes fusion information or a mask of the attention feature map. Since some information in the input feature might be lost after processing by attention modules, we employ a multimodal low-rank bilinear method that considers all pairwise interactions among two features in each time step to complement the original information for output features of the module. AVD-attention outperformed co-attention in experiments on the XD-Violence dataset. Our system outperforms state-of-the-art systems.
Wenfeng Pang, Wei Xie 0013, Qianhua He, Yanxiong Li
IEEE Trans. Multim.3
2022 Predicting skeleton trajectories using a Skeleton-Transformer for video anomaly detection
Wenfeng Pang, Qianhua He, Yanxiong Li
Multim. Syst.2
2022 Fall event detection with global and temporal local information in real-world videos
Wenfeng Pang, Qianhua He, Yuanfeng Chen, Yanxiong Li
Multim. Tools Appl.2
2021 Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder Network
abstract
Recent efforts have been made on domestic activities classification from audio recordings, especially the works submitted to the challenge of DCASE (Detection and Classification of Acoustic Scenes and Events) since 2018. In contrast, few studies were done on domestic activities clustering, which is a newly emerging problem. Domestic activities clustering from audio recordings aims at merging audio clips which belong to the same class of domestic activity into a single cluster. Domestic activities clustering is an effective way for unsupervised estimation of daily activities performed in home environment. In this study, we propose a method for domestic activities clustering using a convolutional capsule autoencoder network (CCAN). In the method, the deep embeddings are learned by the autoencoder in the CCAN, while the deep embeddings which belong to the same class of domestic activities are merged into a single cluster by a clustering layer in the CCAN. Evaluated on a public dataset adopted in DCASE- 2018 Task 5, the results show that the proposed method outperforms state-of-the-art methods in terms of the metrics of clustering accuracy and normalized mutual information.
Ziheng Lin, Yanxiong Li, Zhangjin Huang, Yufeng Tan, Yichun Chen, Qianhua He
ICASSP7
2021 Violence Detection in Videos Based on Fusing Visual and Audio Information
abstract
Determining whether given video frames contain violent content is a basic problem in violence detection. Visual and audio information are useful for detecting violence included in a video, and are usually complementary; however, violence detection studies focusing on fusing visual and audio information are relatively rare. Therefore, we explored methods for fusing visual and audio information. We proposed a neural network containing three modules for fusing multimodal information: 1) attention module for utilizing weighted features to generate effective features based on the mutual guidance between visual and audio information; 2) fusion module for integrating features by fusing visual and audio information based on the bilinear pooling mechanism; and 3) mutual Learning module for enabling the model to learn visual information from another neural network with a different architecture. Experimental results indicated that the proposed neural network outperforms existing state-of-the-art methods on the XD-Violence dataset.
Wen-Feng Pang, Qianhua He, Yongjian Hu, Yanxiong Li
ICASSP2
2021 A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of Query
abstract
The state-of-the-art of query-by-example spoken term detection (QbE-STD) strategies are usually based on segmental dynamic time warping (S-DTW). However, the sliding window in S-DTW may separate signal of a word into different segments and produce many illegal candidates required to be compared with the query, which significantly reduce the accuracy and efficiency of detection. In this paper, we propose a stage match strategy based on the structure information of the query, represented with the unvoiced-voiced attribute of the portions in itself. The strategy first locates potential candidates with similar structure against the query in utterances, and further matches the query with Type-Location DTW (TL-DTW), which is a modified DTW with the constraints of pronunciation types and relative positions of paired frames in the voiced sub-segments. Experiments on AISHELL-1 Corpus showed that the proposed approach achieved a relative improvement of 30.51% in AUC against S-DTW and speeded up the retrieval.
Junyao Zhan, Qianhua He, Jianbin Su, Yanxiong Li
ICASSP2
2021 LIS-Net: An end-to-end light interior search network for speech command recognition
Nguyen Tuan Anh, Yongjian Hu, Qianhua He, Tran Thi Ngoc Linh, Hoang Thi Kim Dung, Chen Guang
Comput. Speech Lang.3
2021 Speaker Clustering by Co-Optimizing Deep Representation Learning and Cluster Estimation
abstract
Speaker clustering is a task to merge speech segments uttered by the same speaker into a single cluster, which is an effective tool for alleviating the management of massive amount of audio documents. In this paper, we present a work for co-optimizing the two main steps of speaker clustering, namely, feature learning and cluster estimation. In our method, the deep representation feature is learned by a deep convolutional autoencoder network (DCAN), while the cluster estimation is realized by a softmax layer that is combined with the DCAN. We devise an integrated loss function to simultaneously minimize the reconstruction loss (for deep representation learning) and the clustering loss (for cluster estimation). Many state-of-the-art audio features and clustering methods are evaluated on experimental datasets selected from two publicly available speech corpora (the AISHELL-2 and the VoxCeleb1). The results show that the proposed method exceeds other speaker clustering methods in regard to the normalized mutual information (NMI) and the clustering accuracy (CA). Additionally, the proposed deep representation feature outperforms other features that were widely used in previous works, in terms of both NMI and CA.
Yanxiong Li, Wucheng Wang, Mingle Liu, Zhongjie Jiang, Qianhua He
IEEE Trans. Multim.5
2020 Crnn-Ctc Based Mandarin Keywords Spotting
abstract
Deep learning based approaches have greatly improved the performance of spoken keyword spotting (KWS). However, KWS of different languages should have their own corresponding modeling units to optimize the performance. In this paper, we propose an end-to-end Mandarin KWS system using Convolutional Recurrent Neural Network with the Connectionist Temporal Classification (CTC) loss function (CRNN-CTC). The tonal syllables are adopted as modeling units. Experiments on AISHELL-2 datasets showed that the proposed approach on the tasks of 13 keywords and 20 keywords can achieve a false rejection rate of 5.35% with 0.26 FA/hour and 6.37% with 0.17 FA/hour, respectively.
Haikang Yan, Qianhua He, Wei Xie 0013
ICASSP2
2020 Acoustic Scene Clustering Using Joint Optimization of Deep Embedding Learning and Clustering Iteration
abstract
Recent efforts have been made on acoustic scene classification in the audio signal processing community. In contrast, few studies have been conducted on acoustic scene clustering, which is a newly emerging problem. Acoustic scene clustering aims at merging the audio recordings of the same class of acoustic scene into a single cluster without using prior information and training classifiers. In this study, we propose a method for acoustic scene clustering that jointly optimizes the procedures of feature learning and clustering iteration. In the proposed method, the learned feature is a deep embedding that is extracted from a deep convolutional neural network (CNN), while the clustering algorithm is the agglomerative hierarchical clustering (AHC). We formulate a unified loss function for integrating and optimizing these two procedures. Various features and methods are compared. The experimental results demonstrate that the proposed method outperforms other unsupervised methods in terms of the normalized mutual information and the clustering accuracy. In addition, the deep embedding outperforms many state-of-the-art features.
Yanxiong Li, Mingle Liu, Wucheng Wang, Qianhua He
IEEE Trans. Multim.5
2018 Frontal Face Generation from Multiple Pose-Variant Faces with CGAN in Real-World Surveillance Scene
abstract
It is well known that frontal face is much easier to be recognized than pose-variant face for both human and machine perception. However, it is not easy to acquire a frontal face in real-world video surveillance. This paper proposes a method to synthetize a frontal face for recognition in video surveillance scene, which is based on Conditional Generative Adversarial Networks (cGAN) with input of multiple pose-variant faces from a video. Experimental results show that the proposed approach can generate suitable frontal faces and improve face recognition by around 20% on a dataset of 43276 face images from 19 persons, collected from the real-world video surveillance scene. The effectiveness of multiple frames against single frame as input is demonstrated. Moreover, we investigate the generator with different depth for synthetizing frontal faces, in which an up-down sampling trick is designed for synthetizing higher quality frontal face images and boosts the performance of the generator.
Zhu-Liang Chen, Qianhua He, Wen-Feng Pang, Yanxiong Li
ICASSP2
2018 Feature with Complementarity of Statistics and Principal Information for Spoofing Detection
Ji-Chen Yang, Chang Huai You, Qianhua He
INTERSPEECH3
2018 Dictionary learning based on M-PCA-N for audio signal sparse representation
abstract
The current popular dictionary learning algorithms for sparse representation of signals are K‐means Singular Value Decomposition (K‐SVD) and K‐SVD‐extended. Only rank‐1 approximation is used to update one atom at a time and it is unable to cope with large dictionary efficiently. In order to tackle these two problems, this study proposes M‐Principal Component Analysis‐N (M‐PCA‐N), which is an algorithm for dictionary learning and sparse representation. First, M‐Principal Component Analysis (M‐PCA) utilised information from the top M ranks of SVD decomposition to update M atoms at a time. Then, in order to further utilise the information from remaining ranks, M‐PCA‐N is proposed on the basis of M‐PCA, by transforming information from the following N non‐principal ranks onto the top M principal ranks. The mathematic formula indicates that M‐PCA may be seen as a generalisation of K‐SVD. Experimental results on the BBC Sound Effects Library show that M‐PCA‐N not only lowers the MSE between original signal and approximation signal in audio signal sparse representation, but also obtains higher audio signal classification precision than K‐SVD.
Ji-Chen Yang, Qianhua He, Yanxiong Li, Lei-an Liu, Xiaohui Feng
IET Signal Process.2
2018 Using multi-stream hierarchical deep neural network to extract deep audio feature for acoustic event detection
Yanxiong Li, Hai Jin 0008, Xianku Li, Qin Wang 0014, Qianhua He
Multim. Tools Appl.6
2018 Mobile Phone Clustering From Speech Recordings Using Deep Representation and Spectral Clustering
abstract
Considerable attention has been paid to acquisition device recognition over the past decade in the forensic community, especially in digital image forensics. In contrast, acquisition device clustering from speech recordings is a new problem that aims to merge the recordings acquired by the same device into a single cluster without having prior information about the recordings and training classifiers in advance. In this paper, we propose a method for mobile phone clustering from speech recordings by using a new feature of deep representation and a spectral clustering algorithm. The new feature is learned by a deep auto-encoder network for representing the intrinsic trace left behind by each phone in the recordings, and spectral clustering is used to merge recordings acquired by the same phone into a single cluster. The impacts of the structures of the deep auto-encoder network on the performance of the new feature are discussed. Different features are compared with one another. The proposed method is compared with others and evaluated under special conditions. The results show that the proposed method is effective under these conditions and the new feature outperforms other features.
Yanxiong Li, Xianku Li, Ji-Chen Yang, Qianhua He
IEEE Trans. Inf. Forensics Secur.6
2017 Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clustering
abstract
Acquisition device clustering from speech recordings is a new and critical problem in the field of speech forensic, which aims at merging speech recordings acquired by the same device into one cluster without both pre-knowing prior information of the processed data and pre-training classifier. We propose a mobile phone clustering method, in which deep Gaussian supervector learned by deep neural network is used to represent the intrinsic trace left behind by mobile phone in speech recordings, and then spectral clustering technique is adopted to merge speech recordings acquired by the same mobile phone into one cluster. The performance of the proposed method is evaluated on a public corpus of speech recordings acquired by mobile phones. The results show that the proposed method is effective for mobile phone clustering from acquired speech recordings.
Yanxiong Li, Xianku Li, Xiaohui Feng, Ji-Chen Yang, Aiwu Chen, Qianhua He
ICASSP7
2017 Unsupervised classification of speaker roles in multi-participant conversational speech
Yanxiong Li, Qin Wang 0014, Xinchao Li, Ji-Chen Yang, Xiaohui Feng, Qianhua He
Comput. Speech Lang.9
2017 Sparse representation-based quasi-clean speech construction for speech quality assessment under complex environments
abstract
A non‐intrusive speech quality assessment method for complex environments was proposed. In the proposed approach, a new sparse representation‐based speech reconstruction algorithm was presented to acquire the quasi‐clean speech from the noisy degraded signal. Firstly, an over‐complete dictionary of the clean speech power spectrum was learned by the K‐singular value decomposition algorithm. Then in the sparse representation stage, the stopping residue error was adaptively achieved according to the estimated cross‐correlation and the noise spectrum which was adjusted by a posteriori SNR‐weighted factor, and the orthogonal matching pursuit approach was applied to reconstruct the clean speech spectrum from the noisy speech. The quasi‐clean speech was considered as the reference to a modified PESQ perceptual model, and the mean opinion score of the noisy degraded speech was achieved via the distortions estimation between the quasi‐clean speech and the degraded speech. Experimental results show that the proposed approach obtains a correlation coefficient of 0.925 on NOIZEUS complex environment database, which is 99% similar to the performance of the intrusive standard ITU‐T PESQ, and 7.1% outperforms non‐intrusive standard ITU‐T P.563.
Weili Zhou, Qianhua He, Yalou Wang, Yanxiong Li
IET Signal Process.2
2016 Source cell phone matching from speech recordings by sparse representation and KISS metric
abstract
Source recording device matching from two speech recordings is a new and important problem of digital media forensics. It aims to answer the question that whether or not two speech recordings are recorded by the same recording device. In this study we propose a source cell phone matching scheme. The Gaussian supervector (GSV) based on Mel-frequency cepstral coefficients (MFCCs) is extracted from the speech recording and is sparse represented with respect to a dictionary learned by K-SVD algorithm. The reduced-dimensional sparse representation coefficient is utilized to characterize the intrinsic fingerprint of the recording device. Then, KISS metric learning based similarity matching is conducted on a pair of fingerprints extracted from the two speech recordings. Evaluation experiments were conducted on a database of speech recordings recorded by 14 cell phones. The experimental results demonstrated the feasibility of the proposed scheme.
Ling Zou 0003, Qianhua He, Ji-Chen Yang, Yanxiong Li
ICASSP2
2015 Acoustic feature extraction by tensor-based sparse representation for sound effects classification
abstract
This paper describes a method to extract time-frequency (TF) audio features by tensor-based sparse approximation for sound effects classification. In the proposed method, the observed data is encoded as a higher-order tensor and discriminative features are extracted in spectrotemporal domain. Firstly, audio signals are represented by a joint time-frequency-duration tensor based on sparse approximation; then tensor factorization is applied to calculate feature vectors. The three arrays of the proposed tensor are used to represent frequency, time and duration of transient TF atoms respectively. Experimental results show that exploiting tensor representation allows to characterize distinctive transient TF atoms, yielding an average accuracy improvement of 9.7% and 12.5% compared with matching pursuit (MP) and MFCC features.
Xueyuan Zhang, Qianhua He, Xiaohui Feng
ICASSP2
2015 Cell phone verification from speech recordings using sparse representation
abstract
Source recording device recognition is an important emerging research field of digital media forensic. Most of the prior literature focus on the recording device identification problem. In this study we propose a source cell phone verification scheme based on sparse representation. We employed Gaussian supervectors (GSVs) based on Mel-frequency cepstral coefficients (MFCCs) extracted from the speech recordings to characterize the intrinsic fingerprint of the cell phone. For the sparse representation, both exemplar based dictionary and dictionary learned by K-SVD algorithm were examined to this problem. Evaluation experiments were conducted on a corpus consists of speech recording recorded by 14 cell phones. The achieved equal error rate (EER) demonstrated the feasibility of the proposed scheme.
Ling Zou 0003, Qianhua He, Xiaohui Feng
ICASSP2
2015 Quasi-clean Speech Construction Based Speech Quality Evaluation under Complex Environments
abstract
Objective evaluation of speech quality for complex environments is a main part of the quality of communications service. Usually, intrusive approaches outperform non-intrusive approaches because reference is used in the intrusive ones. This paper presented a new non-intrusive evaluation method for complex environments in which an improved noise tracking and subtraction algorithm is used to obtain the quasi-clean speech from the noisy speech. The quasi-clean speech is regarded as the reference speech to a perceptual model, which is a modified version of ITU-T PESQ. The perceptual model acquires the Mean Opinion Score (MOS) via the measurement of the distortions between the noisy speech and quasi clean speech. Experimental results demonstrate that the presented method got a correlation coefficient of 0.92 on NOIZEUS noisy dataset, which is 99% similar to that of PESQ, and 6.4% superior to ITUT P.563.
Weili Zhou, Qianhua He
SMC2
2014 Fast speaker clustering using distance of feature matrix mean and adaptive convergence threshold
abstract
The authors propose a method of fast speaker clustering in which a distance (distance of feature matrix mean, DFMM) is first defined for characterising the similarities between any two clusters, and then an adaptive convergence threshold is introduced for terminating the procedure of speaker clustering. If the minimum of the DFMMs between any two clusters is smaller than the threshold, then they are merged. The above mergence of clusters is repeated until the minimum of the DFMMs between any two clusters is larger than the threshold. They conduct experiments on both shorter voice segments (≤ 3 s) and longer voice segments (> 3 s) to compare their method with state‐of‐the‐art methods, agglomerative hierarchical clustering with Bayesian information criterion (AHC + BIC) and vector quantisation with spectral clustering. Experiments show that their method achieves the best results for clustering shorter voice segments, and also obtains satisfactory results for clustering longer voice segments in comparison with other two methods. What is more, their method is faster than other methods in all experimental cases. The initial results show that the hybrid methods by combining their method with the AHC + BIC obtain further improvement in terms of the F score.
Yanxiong Li, Hai Jin 0008, Qianhua He, Xiaohui Feng
IET Signal Process.4
2009 Characteristics-based effective applause detection for meeting speech
Yanxiong Li, Qianhua He, Sam Kwong, Ji-Chen Yang
Signal Process.2
2008 A survey on emotional semantic image retrieval
abstract
Emotional semantic image retrieval is a new and promising research direction in recent years. This paper attempts to introduce this emerging area to researchers, give them a brief overview of current research progress and framework. In this field three key issues, emotional semantic representation, image feature extraction, and emotion recognition, are discussed in detail with approaches and challenges. In addition, future research directions are suggested.
Weining Wang 0003, Qianhua He
ICIP2
2004 Adaptation of hidden Markov models using maximum model distance algorithm
abstract
This paper presents a new approach that uses the maximum model distance (MMD) method for the adaptation of hidden Markov models (HMMs). This method has the same framework as it is used for constructing speech recognizers with abundant data, and work effectively with any amount of adaptation data. All parameters of the HMMs with or without the adaptation data could be adapted. If the adaptation data is sufficient, then the adapted models will gradually become a speaker-dependent one. Both the dialect and the speaker adaptation experiments were conducted to investigate the effectiveness of the proposed algorithm. In the speaker adaptation experiments, up to 65.55% phoneme error reduction was achieved, and the MMD could reduce the phoneme error by 16.91% even only one adaptation utterance is available.
Qianhua He, Sam Kwong, Q. Y. Hong
IEEE Trans. Syst. Man Cybern. Part A1
2002 A genetic classification error method for speech recognition
Sam Kwong, Qianhua He, K. W. Ku, Tak-Ming Chan, Kim-Fung Man, Wallace Kit-Sang Tang
Signal Process.2
2000 Deformation Transformation for Handwritten Chinese Character Shape Correction
Jiancheng Huang, Junxun Yin, Qianhua He
ICMI4
2000 An improved maximum model distance approach for HMM-based speech recognition systems
Qianhua He, Sam Kwong, Kim-Fung Man, Wallace Kit-Sang Tang
Pattern Recognit.1
1998 Parallel Genetic-Based Hybrid Pattern Matching Algorithm for Isolated Word Recognition
abstract
Dynamic Time Warping (DTW) is a common technique widely used for nonlinear time normalization of different utterances in many speech recognition systems. Two major problems are usually encountered when the DTW is applied for recognizing speech utterances: (i) the normalization factors used in a warping path; and (ii) finding the K-best warping paths. Although DTW is modified to compute multiple warping paths by using the Tree-Trellis Search (TTS) algorithm, the use of actual normalization factor still remains a major problem for the DTW. In this paper, a Parallel Genetic Time Warping (PGTW) is proposed to solve the above said problems. A database extracted from the TIMIT speech database of 95 isolated words is set up for evaluating the performance of the PGTW. In the database, each of the first 15 words had 70 different utterances, and the remaining 80 words had only one utterance. For each of the 15 words, one utterance is arbitrarily selected as the test template for recognition. Distance measure for each test template to the utterances of the same word and to those of the 80 words is calculated with three different time warping algorithms: TTS, PGTW and Sequential Genetic Time Warping (SGTW). A Normal Distribution Model based on Rabiner23 is used to evaluate the performance of the three algorithms analytically. The analyzed results showed that the PGTW had performed better than the TTS. It also showed that the PGTW had very similar results as the SGTW, but about 30% CPU time is saved in the single processor system.
Sam Kwong, Qianhua He, Kim-Fung Man, Chak-Wai Chau, Wallace Kit-Sang Tang
Int. J. Pattern Recognit. Artif. Intell.2
1998 A maximum model distance approach for HMM-based speech recognition
Sam Kwong, Qianhua He, Kim-Fung Man, Wallace Kit-Sang Tang
Pattern Recognit.2
1996 Genetic Time Warping for Isolated Word Recognition
abstract
In this paper, a Genetic Time Warping (GTW) algorithm for isolated word recognition was proposed. Relative representation techniques, fitness techniques and reproduction techniques were described and genetic operators were also discussed in detail. Different from the conventional genetic algorithms with fixed genes, every chromosome has its own number of genes. A modified order-based crossover operator was introduced in order to deal with the chromosomes with a different number of genes. Besides the mutation and crossover operators, a new heuristic local optimum operator was also built and it could alter part of a chromosome based on a function of local distance and average distortion of the paths. Finally, experimental investigations were carried out to test the performance of GTW. Based on Rabiner's normal assumptions23 on the distributions of the distances, the overall probability of making a word error could be calculated experimentally. Results demonstrated that GTW performed better or much better than the DTW method for most of the tested words.
Sam Kwong, Qianhua He, Kim-Fung Man
Int. J. Pattern Recognit. Artif. Intell.2