Jiqing Han 0001

dblp:h/JiqingHan · also Ji-Qing Han 0001 · DBLP profile ↗
← Back
105ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0002-4297-4300ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 2 first-author · 31 since 2021Artificial intelligence and machine learning · 58 · 2 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Dual-Path Conditional Chain for CTC-Based Multi-Talker Speech Recognition
Ying Shi 0001, Jiqing Han 0001
IEEE Signal Process. Lett.2
2026 AS-EVNorm: Tail-Aware Extreme Value Normalization for Speaker Verification
abstract
Score normalization is a key back-end technique in speaker verification for improving score comparability across trials. As a convenient and widely used normalization method, Adaptive Symmetric Normalization (AS-Norm) standardizes raw scores using mean-and-standard-deviation normalization parameters estimated from an adaptive cohort comprising the most similar impostor scores. However, since the adaptive cohort retains only the top-ranked impostor scores, these central-moment statistics may not optimally characterize the empirical distribution of these scores, resulting in suboptimal speaker verification performance. In this letter, we propose Adaptive Symmetric Extreme Value Normalization (AS-EVNorm), which treats the adaptive cohort as upper-tail samples from the impostor-score distribution and models them under extreme-value theory for more accurate normalization parameters. Experiments on VoxCeleb and CN-Celeb show that AS-EVNorm consistently reduces both EER and minDCF compared with AS-Norm across a broad range of adaptive cohort configurations, while maintaining competitive normalization time.
Zekai Su, Jiqing Han 0001, Jianchen Li, Yikun Jiang, Zhifeng Jiang 0007, Tianhong Ding, Yongjun He 0002
IEEE Signal Process. Lett.2
2025 Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
abstract
Language Model (LM)-based Text-to-Speech (TTS) systems often generate hallucinated speech that deviates from input text.Existing mitigation strategies either demand excessive training resources or introduce significant inference latency.In this paper, we propose GFlOwNet-guided distribution AlignmenT (GOAT) for LM-based TTS, a posttraining framework that mitigates hallucinations without relying on massive resources or inference cost.Specifically, we first conduct an uncertainty analysis, revealing a strong positive correlation between hallucination and model uncertainty.Based on this, we reformulate TTS generation as a trajectory flow optimization problem and introduce an enhanced Subtrajectory Balance objective together with a sharpened internal reward as target distribution.We further integrate reward temperature decay and learning rate optimization for stability and performance balance.Extensive experiments show that GOAT reduce over 50% character error rates on challenging test cases and lowering uncertainty by up to 58%, demonstrating its strong generalization ability and effectiveness.Code:
Chenlin Liu, Minghui Fang 0002, Patrick Zhang, Jiqing Han 0001
EMNLP6
2025 InfoMin-based Query Embedding Optimization For Query-based Universal Sound Separation
abstract
The query-based universal sound separation (QUSS) has been addressed, aiming to perform the separation of specific sound sources based on a given query. Most of existed methods focus on the improvement of separation models, ignoring the influence of category-conditioned query embedding distribution on separation performance. To address this issue, we propose an optimization method for query embedding that reduces mutual information (MI) between query embeddings while keeping task-related information intact, named the InfoMin principle. In addition, we propose the Frequency-varying Feature-wise Linear Modulation (FFiLM), which leverages frequency band differences in acoustic events to enhance the modulation capability of query embedding and improve the performance of the separation model. Experimental results show that our method achieves considerable improvements over the existing SoTA method.
Jiqing Han 0001, Liwen Zhang 0001, Youcheng Zhang
ICASSP2
2025 Dual Orthogonality Sub-center Loss for Enhanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Tieran Zheng, Guibin Zheng, Yongjun He 0002
INTERSPEECH2
2025 Adaptive Across-Subcenter Representation Learning for Imbalanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Guibin Zheng, Tieran Zheng, Yongjun He 0002
INTERSPEECH2
2025 Knowledge Distillation Method for Pruned RNN-T Models via Pruning Bounds Sharing and Losses Confusion
Xiaocan Zhang, Guibin Zheng, Chenhao Jing, Jiqing Han 0001, Tieran Zheng
INTERSPEECH5
2025 Knowledge-Decoupled Functionally Invariant Path With Synthetic Personal Data for Personalized ASR
abstract
Fine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic personal data without forgetting real knowledge, and in adapting to personal data without forgetting generic knowledge. Considering that the functionally invariant path (FIP) framework enables model adaptation while preserving prior knowledge, in this letter, we introduce FIP into synthetic-data-augmented personalized ASR models. However, the model still struggles to balance the learning of synthetic, personalized, and generic knowledge when applying FIP to train the model on all three types of data simultaneously. To decouple this learning process and further address the above two challenges, we integrate a gated parameter-isolation strategy into FIP and propose a knowledge-decoupled functionally invariant path (KDFIP) framework, which stores generic and personalized knowledge in separate modules and applies FIP to them sequentially. Specifically, KDFIP adapts the personalized module to synthetic and real personal data and the generic module to generic data. Both modules are updated along personalization-invariant paths, and their outputs are dynamically fused through a gating mechanism. With augmented synthetic data, KDFIP achieves a 29.38% relative character error rate reduction on target speakers and maintains comparable generalization performance to the unadapted ASR baseline.
Zhihao Du, Ying Shi 0001, Jiqing Han 0001, Yongjun He 0002
IEEE Signal Process. Lett.4
2025 Joint Energy-Based Model for Semi-Supervised Respiratory Sound Classification: A Method of Insensitive to Distribution Mismatch
abstract
Semi-supervised learning effectively mitigates the lack of labeled data by introducing extensive unlabeled data. Despite achieving success in respiratory sound classification, in practice, it usually takes years to acquire a sufficiently sizeable unlabeled set, which consequently results in an extension of the research timeline. Considering that there are also respiratory sounds available in other related tasks, like breath phase detection and COVID-19 detection, it might be an alternative manner to treat these external samples as unlabeled data for respiratory sound classification. However, since these external samples are collected in different scenarios via different devices, there inevitably exists a distribution mismatch between the labeled and external unlabeled data. For existing methods, they usually assume that the labeled and unlabeled data follow the same data distribution. Therefore, they cannot benefit from external samples. To utilize external unlabeled data, we propose a semi-supervised method based on Joint Energy-based Model (JEM) in this paper. During training, the method attempts to use only the essential semantic components within the samples to model the data distribution. When non-semantic components like recording environments and devices vary, as these non-semantic components have a small impact on the model training, a relatively accurate distribution estimation is obtained. Therefore, the method exhibits insensitivity to the distribution mismatch, enabling the model to leverage external unlabeled data to mitigate the lack of labeled data. Taking ICBHI 2017 as the labeled set, HF_Lung_V1 and COVID-19 Sounds as the external unlabeled sets, the proposed method exceeds the baseline by 12.86.
Wenjie Song 0003, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng, Yongjun He 0002
IEEE J. Biomed. Health Informatics2
2024 Modeling Quasi-Periodic Dependency via Self-Supervised Pre-Training for Respiratory Sound Classification
abstract
Despite the success of self-supervised respiratory sound classification methods, they do not consider that respiratory sounds are quasi-periodic signals with repetitive patterns in successive breaths, which is vital for distinguishing respiratory sounds from non-quasi-periodic sounds like noises. Therefore, the existing methods may achieve limited improvement due to ignoring the quasi-periodic dependency. To this end, considering that the segments containing the same respiratory sound pattern should be similar in a sample, we extract the segment-wise representations and evaluate the similarity between the periodic-dependent representations via a sparse self-relation matrix. By defining a periodic consistency loss, we push the sparse self-relation matrixes of two clips of the same sample closer, encouraging a larger similarity between the representations. In this manner, the method can focus more on the respiratory sound-related quasi-periodic patterns that repeatedly recur in the periodic-dependent segments. Taking HF_Lung_V1 and COVID-19 Sounds as pre-training sets, the method exceeds the baseline by 7.67% on the ICBHI 2017 classification task.
Wenjie Song 0003, Jiqing Han 0001, Jianchen Li, Guibin Zheng, Tieran Zheng, Yongjun He 0002
ICASSP2
2024 Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event Detection
abstract
Overlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events using shared and entangled frame-wise features, which degrades the feature discrimination. To solve the problem, we propose a disentangled feature learning framework to learn a category-specific representation. Specifically, we employ different projectors to learn the frame-wise features for each category. To ensure that these feature does not contain information of other categories, we maximize the common information between frame-wise features within the same category and propose a frame-wise contrastive loss. In addition, considering that the labeled data used by the proposed method is limited, we propose a semi-supervised frame-wise contrastive loss that can leverage large amounts of unlabeled data to achieve feature disentanglement. The experimental results demonstrate the effectiveness of our method.
Yadong Guan, Jiqing Han 0001, Wenjie Song 0003, Guibin Zheng, Tieran Zheng, Yongjun He 0002
ICASSP2
2024 Serialized Output Training by Learned Dominance
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001
INTERSPEECH5
2024 Personality-memory Gated Adaptation: An Efficient Speaker Adaptation for Personalized End-to-end Automatic Speech Recognition
Zhihao Du, Shiliang Zhang, Jiqing Han 0001, Yongjun He 0002
INTERSPEECH4
2024 Capturing High-Level Semantic Correlations via Graph for Multimodal Sentiment Analysis
abstract
Modeling intra-modal and cross-modal interactions poses significant challenges in multimodal sentiment analysis. Currently, graph-based methods like HGraph-CL achieve promising performance, which rely on two different levels of graph contrastive learning within and between modalities to explore sentiment correlations. However, HGraph-CL still faces the following drawbacks in graph construction: 1) nodes of the graph are represented at the frame level, only containing low-level information, neglecting the correlations among high-level semantics; 2) edges of the graph are based on the fixed dependency relations between words in the text sequence and the adjacent relations between frame-level nodes in the non-verbal sequences, failing to effectively capture implicit and long-distance correlations. To this end, this letter introduces capsule networks to construct high-level semantic nodes in a graph, uncovering deep sentimental structures. Furthermore, the learnable adjacency matrices are employed to construct edges of graph, thus adaptively learning the relations between nodes. Experimental results on several benchmark datasets for multimodal sentiment analysis demonstrate the effectiveness of the proposed method.
Fan Qian, Jiqing Han 0001, Yadong Guan, Wenjie Song 0003, Yongjun He 0002
IEEE Signal Process. Lett.2
2024 Keyword Guided Target Speech Recognition
abstract
This letter presents a new target speech recognition problem, where the target speech is defined by a keyword. For instance, when a person speaks “Hey Google” or “Help Me”, we hope the model can recognize the entire contextual speech of that person, even with strong interference speech from other people. The new problem is denoted by target content ASR (TC-ASR). The core challenge of TC-ASR is that the model needs to simultaneously detect the existence of the keyword from heavily mixed speech and recognize the target speech component using the information of the detected keyword segment. Surprisingly, our experiments show that an attention encoder-decoder (AED) model augmented with a keyword encoder can solve this problem pretty well. We also defined a key content spotting (KCS) task and tested the proposed model on it. Our experiments on the LibriMix dataset demonstrated that our approach could address the KCS task with a promising accuracy, outperforming two baseline models by a large margin. Further analysis shows that the proposed model identifies the target speech by a timbre cue, i.e., ensuring that the identified speech is coherent in speaker trait.
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001
IEEE Signal Process. Lett.4
2024 Sound Activity-Aware Based Cross-Task Collaborative Training for Semi-Supervised Sound Event Detection
abstract
The training of sound event detection (SED) models remains a challenge of insufficient supervision due to limited frame-wise labeled data. Mainstream research on this problem has adopted semi-supervised training strategies that generate pseudo-labels for unlabeled data and use these data for the training of a model. Recent works further introduce multi-task training strategies to impose additional supervision. However, the auxiliary tasks employed in these methods either lack frame-wise guidance or exhibit unsuitable task designs. Furthermore, they fail to exploit inter-task relationships effectively, which can serve as valuable supervision. In this paper, we introduce a novel task, sound occurrence and overlap detection (SOD), which detects predefined sound activity patterns, including non-overlapping and overlapping cases. On the basis of SOD, we propose a cross-task collaborative training framework that leverages the relationship between SED and SOD to improve the SED model. Firstly, by jointly optimizing the two tasks in a multi-task manner, the SED model is encouraged to learn features sensitive to sound activity. Subsequently, the cross-task consistency regularization is proposed to promote consistent predictions between SED and SOD. Finally, we propose a pseudo-label selection method that uses inconsistent predictions between the two tasks to identify potential wrong pseudo-labels and mitigate their confirmation bias. In the inference phase, only the trained SED model is used, thus no additional computation and storage costs are incurred. Extensive experiments on the DESED dataset demonstrate the effectiveness of our method.
Yadong Guan, Jiqing Han 0001, Shiwen Deng, Guibin Zheng, Tieran Zheng, Yongjun He 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Distance Metric-Based Open-Set Domain Adaptation for Speaker Verification
abstract
Domain shift poses a significant challenge in speaker verification, especially in open-set scenarios where the speaker categories are disjoint between the source and target domains. To alleviate the domain shift, traditional domain adaptation methods typically align the source and target distributions in the speaker embedding space, but this may cause the overlap of embeddings from different speakers. To address this problem, this paper proposes to perform the domain alignment in a novel distance metric space, where the source and target domains exhibit the shared within-speaker and between-speaker categories. Thus, the discrepancy between the source and target domains arises only from the domain shift. We refer to the proposed method as Cross-Domain Distance Metric Adaptation (CDMA), in which the within- and between-speaker distance distributions in the target domain are aligned with the source distance distributions and further separated to minimize their overlap. This alignment and separation require estimating the within- and between-speaker distance distributions based on speaker labels, which are unavailable in the unlabeled target domain. Thus, we further propose a learnable speaker clustering method called Graph Convolutional Network with Graph Pruning (GCN-GP). This method generates high-quality pseudo-labels to estimate the two distance distributions in the target domain. Experimental results demonstrate that our method achieves state-of-the-art performance on the FFSVC2022 and VOiCES datasets.
Jianchen Li, Jiqing Han 0001, Fan Qian, Tieran Zheng, Yongjun He 0002, Guibin Zheng
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Graph-Based Spectro-Temporal Dependency Modeling for Anti-Spoofing
abstract
A great deal of recent research reveals that artifacts introduced by spoofing algorithms reside in specific frequency subbands or temporal segments. Therefore, the performance of spoofing detection can be improved by focusing on these regions. However, it is difficult for the detection system to choose an appropriate region when it encounters an unknown spoofing algorithm, resulting in poor generalization. Actually, there is a noticeable difference in the inter-region relationship between the bonafide and spoofed speeches. We name the inter-region relationship spectro-temporal dependency and design a method to model it for anti-spoofing. By focusing on the general dependency difference rather than specific regions, the generalization ability of the detection system can be improved. We employ a graph neural network to model the dependency and incorporate prior knowledge into the graph by designing the graph structure and edge weight, which forces the network to pay more attention to potential relationships. In addition, an attention mechanism is introduced in the graph pooling to focus on more critical nodes. The proposed method achieves an equal error rate of 0.58% on the ASVspoof 2019 LA dataset and outperforms all competing systems.
Shiwen Deng, Tieran Zheng, Yongjun He 0002, Jiqing Han 0001
ICASSP5
2023 Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
abstract
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g. the log-Mel spectrogram’s average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
Jian Guan 0001, Youde Liu, Qiaoxi Zhu, Tieran Zheng, Jiqing Han 0001, Wenwu Wang 0001
ICASSP5
2023 Subband Dependency Modeling for Sound Event Detection
abstract
In the domain of sound event detection (SED), Convolutional Recurrent Neural Network (CRNN) has become the most successful architecture, which adopts Recurrent Neural Network (RNN) to model temporal dependencies from the output of Convolutional Neural Network (CNN). However, CRNN does not fully use the subband dependencies that have been proved critical for human perception of sound events. In this paper, we propose a subband dependency model (SDM) to enhance the capability of CRNN in modeling subband dependencies from the input spectrogram. To select prominent subband dependencies, we propose a novel SoftSparsemax transformation. It can select the salient parts by comparing all dependencies and further strengthen them by projecting them onto a probability simplex. Furthermore, since subband dependencies of different sound events may be prominent in different timescales, multi-timescale subband dependency is considered. The experiment results demonstrate the effectiveness of our method.
Yadong Guan, Guibin Zheng, Jiqing Han 0001, Huanliang Wang
ICASSP3
2023 Using Auxiliary Tasks In Multimodal Fusion of Wav2vec 2.0 And Bert for Multimodal Emotion Recognition
abstract
The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pre-trained models as upstream network, wav2vec 2.0 for audio modality and BERT for text modality, and finetune them in downstream task of MER to cope with the lack of data. For the difficulty of multimodal fusion, we use a K-layer multi-head attention mechanism as a downstream fusion module. Starting from the MER task itself, we design two auxiliary tasks to alleviate the insufficient fusion between modalities and guide the network to capture and align emotion-related features. Compared to the previous state-of-the-art models, we achieve a better performance by 78.42% Weighted Accuracy (WA) and 79.71% Unweighted Accuracy (UA) on the IEMOCAP dataset.
Dekai Sun, Yancheng He, Jiqing Han 0001
ICASSP3
2023 Spot Keywords From Very Noisy and Mixed Speech
Ying Shi 0001, Dong Wang 0013, Lantian Li, Jiqing Han 0001
INTERSPEECH4
2023 Personality-aware Training based Speaker Adaptation for End-to-end Speech Recognition
Zhihao Du, Shiliang Zhang, Qian Chen 0003, Jiqing Han 0001
INTERSPEECH5
2023 Mutual Information-based Embedding Decoupling for Generalizable Speaker Verification
Jianchen Li, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Yongjun He 0002, Guibin Zheng
INTERSPEECH2
2023 Task-driven common subspace learning based semantic feature extraction for acoustic event recognition
Qiuying Shi, Shiwen Deng, Jiqing Han 0001
Expert Syst. Appl.3
2022 Sparse Self-Attention for Semi-Supervised Sound Event Detection
abstract
Self-attention mechanism has been widely employed in semi-supervised sound event detection (SS-SED). In self-attention, since dependencies between pairwise features at all moments are captured, the irrelevant features of different classes of sounds and background sounds at other moments are inevitably mixed in the current embedding when self-attention performs weighted summation. These irrelevant features will weaken the ability of the aggregated embedding to describe sound events. In this paper, we propose a sparse self-attention mechanism to alleviate the impact. Specifically, the Sparsemax function is introduced for attention weights normalization, which uses Euclidean projection to project attention weights onto a probability simplex. After the normalization, the attention weights of the irrelevant features are projected onto the boundary of the simplex and then removed. Furthermore, to solve the excessive sparsity problem of the Sparsemax, we further propose the Sparsemax with adjustable sparsity. Experimental results demonstrate the effectiveness of the proposed method.
Yadong Guan, Jiabin Xue, Guibin Zheng, Jiqing Han 0001
ICASSP4
2022 CDMA: Cross-Domain Distance Metric Adaptation for Speaker Verification
abstract
To solve the domain shift problem in speaker verification, one effective domain adaptation approach is to learn domain-invariant embeddings via aligning the source and target distributions in the embedding space. However, this approach could be problematic when the source and target domains are from the disjoint speaker label spaces as the embedding distributions of different speakers cannot be aligned. In this paper, we propose a Cross-domain Distance Metric Adaptation (CDMA) approach to alleviate the domain shift in the distance metric space, where the source and target domains share the same classes, i.e., within- and between-speaker. Specifically, the two target pairwise distance distributions are aligned with the source pairwise distance distributions and further separated to learn a domain-invariant metric, which is more suitable for speaker verification based on metric learning. Experiments indicate that CDMA significantly outperforms the approach proposed in the embedding space.
Jianchen Li, Jiqing Han 0001
ICASSP2
2022 Exploring Transformer's Potential on Automatic Piano Transcription
abstract
Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic notation. Based on a high-resolution piano transcription system, we explore the possibility of incorporating another powerful sequence transformation tool—the Transformer—to deal with the AMT problem. We argue that the properties of the Transformer make it more suitable for certain AMT subtasks. We confirm the Transformer’s superiority on the velocity detection task by experiments on the MAESTRO dataset and a cross-dataset evaluation on the MAPS dataset. We observe a performance improvement on both frame-level and note-level metrics after introducing the Transformer network.
Longshen Ou, Emmanouil Benetos, Jiqing Han 0001, Ye Wang 0007
ICASSP4
2022 Word-wise Sparse Attention for Multimodal Sentiment Analysis
Fan Qian, Jiqing Han 0001
INTERSPEECH3
2022 Exploring Inter-Node Relations in CNNs for Environmental Sound Classification
abstract
For environmental sound classification, CNNs have become the most successful architecture. By regarding the CNN features as a collection of nodes arranged on a 2D time-frequency grid, typical CNN layers process nodes within a limited local region. However, the rich relation information between nodes, especially the non-local relations, is mostly ignored. For environmental sound, these inter-node relations carry rich information about the existence of repetitive sound event patterns and the complex interactions between different sound events in acoustic scenes, which are valuable for categorizing environmental sound. In this letter, we propose a relation module, named the R-Block, to explore the relation information in an explicit and comprehensive way. The R-Block is designed to not only capture and utilize the inter-node relations, but also explore the structure of the learned relations, which leads to a more expressive representation. Experimental results reveal that, by augmenting a powerful ResNeXt backbone with the R-Block, our model is able to achieve competitive performance on ESC-50 and US8K sound event classification dataset and state-of-the-art result on DCASE2018 acoustics scene classification dataset.
Shiwen Deng, Jiqing Han 0001
IEEE Signal Process. Lett.3
2021 Contrastive Embeddind Learning Method for Respiratory Sound Classification
abstract
Respiratory sound classification refers to identifying adventitious sounds from given recordings automatically. Due to the difficulty of collection and the expensive manual annotation, there are only limited samples available, which impacts on learning better models. Meanwhile, a majority of these models do not explicitly encourage intra-class compactness and inter-class separability between the learned embeddings, leading to the difficulty of identifying several samples and a reduced generalization performance. To address the problems, we propose a contrastive embedding learning method, where the input is a contrastive tuple. And the composite input strategy provides more possible network inputs. By the comparison among the samples in the tuple, we can learn the slight differences among the similar samples, and the easily-confused samples are more likely to be identified. In the embedding space, we explicitly promote the intra-class compactness and inter-class separability, thereby the generalization performance is improved. Our method is evaluated on ICBHI 2017, and the classification score is increased from 75.61% of a conventional cross-entropy network to 78.18%, outperforming the state-of-the-art methods.
Wenjie Song 0003, Jiqing Han 0001
ICASSP2
2021 Capturing Temporal Dependencies Through Future Prediction for CNN-Based Audio Classifiers
abstract
This paper focuses on the problem of temporal dependency modeling in the CNN-based models for audio classification tasks. To capture audio temporal dependencies using CNNs, we take a different approach from the purely architecture-induced method and explicitly encode temporal dependencies into the CNN-based audio classifiers. More specifically, in addition to the classification objective, we require the CNN model to solve an auxiliary task of predicting the future features, which is formulated by leveraging the Contrastive Predictive Coding (CPC) loss. Furthermore, a novel hierarchical CPC (HCPC) model is proposed for capturing multi-level temporal dependencies at the same time. The proposed model is evaluated on a wide range of non-speech audio signals, including musical and in-the-wild environmental audio signals. We show that the proposed approach improves the backbone CNNs consistently on all tested benchmark datasets and outperforms a DenseNet model trained from scratch.
Jiqing Han 0001, Shiwen Deng, Zhihao Du
ICASSP2
2021 Gradient Regularization for Noise-Robust Speaker Verification
Jianchen Li, Jiqing Han 0001
Interspeech2
2021 Multimodal Sentiment Analysis with Temporal Modality Attention
Fan Qian, Jiqing Han 0001
Interspeech2
2021 Model-Agnostic Fast Adaptive Multi-Objective Balancing Algorithm for Multilingual Automatic Speech Recognition Model Training
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
Interspeech3
2021 Can We Trust Deep Speech Prior?
abstract
Recently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) architecture. Compared to conventional approaches that represent clean speech by shallow models such as Gaussians with a low-rank covariance, the new approach employs deep generative models to represent the clean speech, which often provides a better prior. Despite the clear advantage in theory, we argue that deep priors must be used with much caution, since the likelihood produced by a deep generative model does not always coincide with the speech quality. We designed a comprehensive study on this issue and demonstrated that based on deep speech priors, a reasonable SE performance can be achieved, but the results might be suboptimal. A careful analysis showed that this problem is deeply rooted in the disharmony between the flexibility of deep generative models and the nature of the maximum-likelihood (ML) training.
Ying Shi 0001, Zhiyuan Tang, Lantian Li, Dong Wang 0013, Jiqing Han 0001
SLT6
2021 Exploring attention mechanisms based on summary information for end-to-end automatic speech recognition
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
Neurocomputing3
2020 TDMF: Task-Driven Multilevel Framework for End-to-End Speaker Verification
abstract
In this paper, a task-driven multilevel framework (TDMF) is proposed for end-to-end speaker verification. The TDMF has four layers, and each layer has different effects on speaker models or representations to implement the functions of universal background model (UBM), Gaussian mixture model (GMM), total variability model (TVM) and probabilistic linear discriminant analysis (PLDA). Unlike the typical i-vector method, the proposed TDMF can supervise the optimal solution of each phase (layer) towards the direction required by the PLDA classifier. Moreover, different from most end-to-end neural network approaches, which extract embeddings first and then additionally calculate the distance between two embeddings as the verification score, the TDMF can directly provide scores via the fourth-layer PLDA. The experimental results show that the TDMF can achieve better performance than that of the typical i-vector framework and VGG-M convolutional neural networks (CNN) framework.
Chen Chen 0086, Jiqing Han 0001
ICASSP2
2020 Pan: Phoneme-Aware Network for Monaural Speech Enhancement
abstract
Current methods for monaural speech enhancement only utilize acoustic information but seldom consider the phonetic information of an utterance. In the voice conversion community, significant progress has been achieved by using the phonetic information via the phonetic posteriorgrams (PPGs). Inspired by the progress, we propose a phoneme-aware network (PAN) to utilize the noisy PPGs for speech enhancement. Since the PPG prediction and speech enhancement benefit from each other, a PPG predictor is involved into the PAN and an iterative training algorithm is proposed for PAN. Experimental results show that the enhancement performance is improved by using the phonetic information in terms of speech intelligibility, perceptual quality and character error rate. To the best of our knowledge, this is the first time to introduce the PPG into speech enhancement.
Zhihao Du, Jiqing Han 0001, Shiliang Zhang
ICASSP3
2020 Structured Sparse Attention for end-to-end Automatic Speech Recognition
abstract
The Softmax normalization function-based attention mechanism is often employed by End-to-End Automatic Speech Recognition (E2E ASR) models to tell the network where to focus within the input. However, this mechanism leads to the attention distribution becoming increasingly flatter as the input sequence length increases, since the output probability of this function is dense and nonnegative, which makes it unable to highlight the important information in speech. In this paper, we present two sparse attention mechanisms for ASR tasks with long utterances, which try to improve the attention mechanism by introducing the sparse transformation. First, we propose to replace the Softmax with the Sparsemax that normalizes the attention weight by finding the closest point in the probability simplex. Then, considering the structured characteristics, the pronunciation has a relatively stable duration. Therefore, we further present a structured sparse transformation that forces the networks to pay attention to a continuous segment of speech by applying the l2penalty. A noniterative solution algorithm that can be used in the backpropagation is designed here. The experiments show that our methods achieve better ASR results compared to a well-tuned attention-based baseline system on a character ASR task.
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
ICASSP3
2020 Double Adversarial Network Based Monaural Speech Enhancement for Robust Speech Recognition
Zhihao Du, Jiqing Han 0001, Xueliang Zhang 0001
INTERSPEECH2
2020 Self-Supervised Adversarial Multi-Task Learning for Vocoder-Based Monaural Speech Enhancement
Zhihao Du, Jiqing Han 0001, Shiliang Zhang
INTERSPEECH3
2020 Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity Loss
abstract
Deep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation. This work investigates how to extend dual-path BiLSTM to result in a new state-of-the-art approach, called TasTas, for multi-talker monaural speech separation (a.k.a cocktail party problem). TasTas introduces two simple but effective improvements, one is an iterative multi-stage refinement scheme, and the other is to correct the speech with imperfect separation through a loss of speaker identity consistency between the separated speech and original speech, to boost the performance of dual-path BiLSTM based networks. TasTas takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. Our experiments on the notable benchmark WSJ0-2mix data corpus result in 20.55dB SDR improvement, 20.35dB SI-SDR improvement, 3.69 of PESQ, and 94.86\% of ESTOI, which shows that our proposed networks can lead to big performance improvement on the speaker separation task. We have open sourced our re-implementation of the DPRNN-TasNet here (this https URL), and our TasTas is realized based on this implementation of DPRNN-TasNet, it is believed that the results in this paper can be reproduced with ease.
Ziqiang Shi, Rujie Liu, Jiqing Han 0001
INTERSPEECH3
2020 ATReSN-Net: Capturing Attentive Temporal Relations in Semantic Neighborhood for Acoustic Scene Classification
Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi
INTERSPEECH2
2020 FurcaNeXt: End-to-End Monaural Speech Separation with Dynamic Gated Dilated Temporal Convolutional Networks
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001, Anyan Shi, Ding Ma 0001
MMM (1)3
2020 Learning Temporal Relations from Semantic Neighbors for Acoustic Scene Classification
abstract
Convolutional networks have achieved the state-of-the-art performance on Acoustic Scene Classification (ASC). Given the Log Mel-Spectrogram of an audio sample, the network can extract useful semantic contents in a certain range receptive field by stacking local convolutional operations. However, the temporal relations between different receptive fields are not captured explicitly. In this letter, we propose an end-to-end 3D Convolutional Neural Network (CNN) for ASC, named SeNoT-Net, which can generate effective audio representations by capturing temporal relations from semantic neighbors of different receptive fields over time. The SeNoT-Net treats the Log-Mel spectrogram as an ordered segment-level sequence. For each segment, the residual block can produce the semantic feature maps, then the semantic neighbors over time (SeNoT) module is applied to capture the relations between each feature point in the feature maps and its top-k semantic neighbors. The proposed SeNoT-Net outperforms most of the state-of-the-art CNN models on both DCASE 2018 and 2019 ASC datasets.
Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi
IEEE Signal Process. Lett.2
2020 A Joint Framework of Denoising Autoencoder and Generative Vocoder for Monaural Speech Enhancement
abstract
Conventional monaural speech enhancement methods usually enhance the magnitude spectrum of noisy speech and leave the phase unchanged. Recent studies suggest that phase is also important for both speech intelligibility and perceptual quality. Although deep learning exhibits great potential on enhancing the magnitude and phase spectra in complex spectrogram domain and waveform domain, complex spectrogram and waveform are always more difficult to predict than the magnitude spectrum due to lack of clear structure in them. In this study, a Mel-domain denoising autoencoder and a deep generative vocoder are stacked to form a joint framework for monaural speech enhancement, in which the clean speech waveform is reconstructed without using the phase. Specifically, a convolutional recurrent network (CRN) is employed as the denoising autoencoder to enhance the Mel power spectrum of noisy speech. Then, the enhanced Mel power spectrum is fed to a deep generative vocoder to synthesize the speech waveform. Furthermore, the denoising autoencoder and generative vocoder are jointly fine-tuned. Experimental results show that the proposed method significantly improves speech intelligibility and perceptual quality. More importantly, our method achieves much better generalization ability for untrained noises than previous methods.
Zhihao Du, Xueliang Zhang 0001, Jiqing Han 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Nonnegative Matrix Factorization Based Transfer Subspace Learning for Cross-Corpus Speech Emotion Recognition
abstract
This article focuses on the cross-corpus speech emotion recognition (SER) task. To overcome the problem that the distribution of training (source) samples is inconsistent with that of testing (target) samples, we propose a non-negative matrix factorization based transfer subspace learning method (NMFTSL). Our method tries to find a shared feature subspace for the source and target corpora, in which the discrepancy between the two distributions is eliminated as much as possible and their individual components are excluded, thus the knowledge of the source corpus can be transferred to the target corpus. Specifically, in this induced subspace, we minimize the distances not only between the marginal distributions but also between the conditional distributions, where both distances are measured by the maximum mean discrepancy criterion. To estimate the conditional distribution of the target corpus, we propose to integrate the prediction of target label and the learning of feature representation into a joint learning model. Meanwhile, we introduce a difference loss to exclude the individual components from the shared subspace, which can further reduce the mutual interference between the source and target individual components. Moreover, we propose a discrimination loss to introduce the labels into the shared subspace, which can improve the discrimination ability of the feature representation. We also provide the solution for the corresponding optimization problem. To evaluate the performance of our method, we construct 30 cross-corpus SER schemes using 6 popular speech emotion corpora. Experimental results show that our approach achieves better overall performance than state-of-the-art methods.
Hui Luo 0008, Jiqing Han 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Pyramidal Temporal Pooling With Discriminative Mapping for Audio Classification
abstract
Audio signals are temporally-structured data, and learning their discriminative representations containing temporal information is crucial for the audio classification. In this article, we propose an audio representation learning method with a hierarchical pyramid structure called pyramidal temporal pooling (PTP) which aims to capture the temporal information of an entire audio sample. By stacking a global temporal pooling layer on multiple local temporal pooling layers, the PTP can capture the high-level temporal dynamics of the input feature sequence in an unsupervised way. Furthermore, in the top global temporal pooling layer, we jointly optimize a learnable discriminative mapping (DM) and a softmax classifier. Such that, a joint learning method for the discriminative audio representations and the classifier called DM-PTP is also presented. By treating the temporal encoding as a low-level constraint of a bi-level optimization problem, the DM-PTP can produce the discriminative representation while maintaining the temporal information of the whole sequence. For an audio sample with an arbitrary time duration, both our PTP and DM-PTP can encode the input feature sequence with arbitrary length into a fixed-length representation. Without using any data augmentation and ensemble learning methods, both PTP and DM-PTP outperform the state-of-the-art CNNs on the audio event recognition (AER) dataset, and can achieve comparable performance on the DCASE 2018 acoustic scene classification (ASC) dataset compared with other best models in the challenge.
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example Training
abstract
Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, such as signal-to-distortion ratio (SDR) upper bound of separated utterances and also fail to exploit an end-to-end framework. In this paper we present an integrated simple and effective end-to-end approach called FurcaX1to monaural speech separation, which consists of deep gated (de)convolutional neural networks (GCNN) that takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. For the objective, we propose to train the network by directly optimizing utterance level SDR in a permutation invariant training (PIT) style. We execute generative adversarial training (GAT) throughout the training, which makes the separated speech indistinguishable from the real one. Our experiments on the the public WSJ0-2mix data corpus demonstrate that this new scheme can produce more discriminative separated utterances and leading to performance improvement on the speaker separation task.
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Jiqing Han 0001
ICASSP6
2019 Convolutional Grid Long Short-Term Memory Recurrent Neural Network for Automatic Speech Recognition
Jiabin Xue, Tieran Zheng, Jiqing Han 0001
ICONIP (5)3
2019 Cross-Corpus Speech Emotion Recognition Using Semi-Supervised Transfer Non-Negative Matrix Factorization with Adaptation Regularization
Jiqing Han 0001
INTERSPEECH2
2019 Subspace Pooling Based Temporal Features Extraction for Audio Event Recognition
Qiuying Shi, Jiqing Han 0001
INTERSPEECH3
2019 End-to-End Monaural Speech Separation with Multi-Scale Dynamic Weighted Gated Dilated Convolutional Pyramid Network
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Shoji Hayakawa, Shouji Harada, Jiqing Han 0001
INTERSPEECH7
2019 Deep Attention Gated Dilated Temporal Convolutional Networks with Intra-Parallel Convolutional Modules for End-to-End Monaural Speech Separation
Ziqiang Shi, Huibin Lin, Liu Liu 0020, Rujie Liu, Jiqing Han 0001, Anyan Shi
INTERSPEECH5
2019 Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events
abstract
In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from existing strategies, which focus on characterizing global acoustical distributions of audio or the temporal evolution of short-term audio features, without analysis down to the level of sound events. To identify distinct sound events for each scene, we formulate ASC in a multi-instance learning (MIL) framework, where each audio recording is mapped into a bag-of-instances representation. Here, instances can be seen as high-level representations for sound events inside a scene. We also propose a MIL neural networks model, which implicitly identifies distinct instances (i.e., sound events). Furthermore, we propose two specially designed modules that model the multi-temporal scale and multi-modal natures of the sound events respectively. The experiments were conducted on the official development set of the DCASE2018 Task1 Subtask B, and our best-performing model improves over the official baseline by 9.4% (68.3% vs 58.9%) in terms of classification accuracy. This study indicates that recognizing acoustic scenes by identifying distinct sound events is effective and paves the way for future studies that combine this strategy with previous ones.
Jiqing Han 0001, Shiwen Deng, Zhihao Du
INTERSPEECH2
2018 Deep Neural Network Based Discriminative Training for I-Vector/PLDA Speaker Verification
abstract
In the studies of i-vector based speaker verification, the discriminative training of probabilistic linear discriminative analysis (PLDA) model has been proven to be an effective way to improve performance. This paper focuses on using a deep neural network (DNN) to strengthen the original discriminatively trained classifiers by its strong capability of nonlinear modeling representation. We first propose a deep neural network based dimensionality reduction model to replace the linear discriminant analysis (LDA) process, and then a discriminative training algorithm is also proposed to jointly optimize the network and PLDA scoring function under single discriminative criterion. Our experiments show that performance improvements are achieved in the male trials of short2-short3 core data set of NIST SRE08.
Tieran Zheng, Jiqing Han 0001, Guibin Zheng
ICASSP2
2018 A Compact and Discriminative Feature Based on Auditory Summary Statistics for Acoustic Scene Classification
abstract
One of the biggest challenges of acoustic scene classification (ASC) is to find proper features to better represent and characterize environmental sounds. Environmental sounds generally involve more sound sources while exhibiting less structure in temporal spectral representations. However, the background of an acoustic scene exhibits temporal homogeneity in acoustic properties, suggesting it could be characterized by distribution statistics rather than temporal details. In this work, we investigated using auditory summary statistics as the feature for ASC tasks. The inspiration comes from a recent neuroscience study, which shows the human auditory system tends to perceive sound textures through time-averaged statistics. Based on these statistics, we further proposed to use linear discriminant analysis to eliminate redundancies among these statistics while keeping the discriminative information, providing an extreme com-pact representation for acoustic scenes. Experimental results show the outstanding performance of the proposed feature over the conventional handcrafted features.
Jiqing Han 0001, Shiwen Deng
INTERSPEECH2
2018 Unsupervised Temporal Feature Learning Based on Sparse Coding Embedded BoAW for Acoustic Event Recognition
Liwen Zhang 0001, Jiqing Han 0001, Shiwen Deng
INTERSPEECH2
2017 Learning Deep Neural Network Based Kernel Functions for Small Sample Size Classification
Tieran Zheng, Jiqing Han 0001, Guibin Zheng
ICONIP (1)2
2017 Speaker Verification via Estimating Total Variability Space Using Probabilistic Partial Least Squares
Chen Chen 0086, Jiqing Han 0001, Yilin Pan
INTERSPEECH2
2017 Heart sound classification based on scaled spectrogram and tensor decomposition
Wenjie Zhang 0008, Jiqing Han 0001, Shiwen Deng
Expert Syst. Appl.2
2016 Realistic human action recognition: When deep learning meets VLAD
abstract
Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action recognition from realistic datsets. We adopt stack convolu-tional independent subspace analysis (ISA) networks to learn 3D cuboid representation directly from spatio-temporal video data; we propose an improved VLAD by incorporating the spatio-temporal geometrical information to encode the deep learned local features. On two challenging realistic datasets: the YouTube action and HMDB51 datasets, the proposed method achieves state-of-the-art performance with an efficient linear SVM classifier, which is competitive with and even better than existing sophisticated algorithms.
Lei Zhang 0093, Yangyang Feng, Jiqing Han 0001, Xiantong Zhen
ICASSP3
2016 Towards optimal vlad for human action recognition from still images
abstract
Human action recognition from still image has recently drawn increasing attention in human behavior analysis vision and also poses great challenges due to the huge inter ambiguity and intra variability. Vector of locally aggregated descriptors (VLAD) has achieved state-of-the-art performance in many image classification tasks based on local features. The great success of VLAD is largely due to its high descriptive ability and computational efficiency. In this paper, towards optimal VLAD representations for human action recognition from still images, we improve VLAD by tackling two important issues in VLAD including empty cavity and assignment ambiguity. The empty cavity issue severely compromises the performance of VLAD and has long been overlooked. We investigate the empty cavity and provide an effective solution to deal with it, which largely improves the performance of VLAD; we propose middle level assignments to conquer the assignment ambiguity, which are more reliable and can provide more useful information for realistic activity. We have conducted extensive experiments on two widely-used benchmarks to validate the proposed method for human action recognition from still images. Our method produces competitive performance with state-of-the-art algorithms.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001
ICASSP3
2016 Towards heart sound classification without segmentation via autocorrelation feature and diffusion maps
Shiwen Deng, Jiqing Han 0001
Future Gener. Comput. Syst.2
2016 Optimization of learned dictionary for sparse coding in speech processing
Yongjun He 0002, Guanglu Sun, Jiqing Han 0001
Neurocomputing3
2016 Speaker Verification via Modeling Kurtosis Using Sparse Coding
abstract
This paper proposes a new model for speaker verification by employing kurtosis statistical method based on sparse coding of human auditory system. Since only a small number of neurons in primary auditory cortex are activated in encoding acoustic stimuli and sparse independent events are used to represent the characteristics of the neurons. Each individual dictionary is learned from individual speaker samples where dictionary atoms correspond to the cortex neurons. The neuron responses possess statistical properties of acoustic signals in auditory cortex so that the activation distribution of individual speaker’s neurons is approximated as the characteristics of the speaker. Kurtosis is an efficient approach to measure the sparsity of the neuron from its activation distribution, and the vector composed of the kurtosis of every neuron is obtained as the model to characterize the speaker’s voice. The experimental results demonstrate that the kurtosis model outperforms the baseline systems and an effective identity validation function is achieved desirably.
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
Int. J. Pattern Recognit. Artif. Intell.2
2016 Sparse Decomposition for Signal Periodic Model Over Complex Exponential Dictionary
abstract
In this letter, we propose an algorithm of the sparse decomposition based on the signal periodic model (SDSPM) to decompose a signal into a series of periodic components and residual or noise. Instead of directly using the complex dictionary, the SDSPM is alternatively performed over two real-valued dictionaries with weighted ℓ1regularization to yield a complex sparse representation of a signal. The sparse representation is the complex projection coefficient vector of a signal into the complex conjugate subspaces and reveals the periodic structure of the signal. Being different from previous methods with the assumption that the signals are noise free, the proposed method can also perform the decomposition of the noisy signals. Moreover, the block version of the SDSPM is also used to decompose a signal with local or time-varying periodicities. In addition, several examples are presented to illustrate the effectiveness of the proposed method.
Shiwen Deng, Jiqing Han 0001
IEEE Signal Process. Lett.2
2015 Noise-robust speaker recognition based on morphological component analysis
Yongjun He 0002, Chen Chen 0086, Jiqing Han 0001
INTERSPEECH3
2015 Dictionary evaluation and optimization for sparse coding based speech processing
Yongjun He 0002, Guanglu Sun, Jiqing Han 0001
Inf. Sci.4
2015 Soft Margin Based Low-Rank Audio Signal Classification
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
Neural Process. Lett.2
2014 Robust minimum statistics project coefficients feature for acoustic environment recognition
abstract
Acoustic environment recognition has been widely used in many applications, and is a considerable difficult problem for the real-life and complex environment. This paper proposes a novel feature, named minimum statistics project coefficients (MSPC), and intents to solve this problem. The MSPC feature is extracted from the background sound which is more robust than the foreground sound for the task of acoustic environment recognition. Experimental results show the outstanding performance of the MSPC feature compared with the conventional acoustic features, especially in very complex acoustic environments.
Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP2
2014 Learning semantic kernels for scene classification
abstract
In this paper we propose to learn semantic kernels for scene classification. We first decompose the Object Bank representation into subspaces associated with each object, Anchor Objects are then created by clustering for each scene class separately. The Anchor Distances are computed to measure the distance between objects to scene classes. In order to take the advantage of the discriminative information from different scene classes, we propose semantic kernels based on the anchor distances to different classes for scene classification. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed Semantic Kernels can significantly improve the original Object Bank and achieve state-of-the-art performance.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001, Xuezhi Xiang
ICASSP3
2014 Evaluation of dictionary for sparse coding in speech processing
Yongjun He 0002, Guanglu Sun, Guibin Zheng, Jiqing Han 0001
INTERSPEECH4
2013 Upper and lower bounds for approximation of the Kullback-Leibler divergence between Hidden Markov models
abstract
The Kullback-Leibler (KL) divergence is often used for a similarity comparison between two Hidden Markov models (HMMs). However, there is no closed form expression for computing the KL divergence between HMMs, and it can only be approximated. In this paper, we propose two novel methods for approximating the KL divergence between the left-to-right transient HMMs. The first method is a product approximation which can be calculated recursively without introducing extra parameters. The second method is based on the upper and lower bounds of KL divergence, and the mean of these bounds provides an available approximation of the divergence. We demonstrate the effectiveness of the proposed methods through experiments including the deviations to the numerical approximation and the task of predicting the confusability of phone pairs. Experimental results show that the proposed product approximation is comparable with the current variational approximation, and the proposed approximation based on bounds performs better than current methods in the experiments.
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP2
2013 Case based reasoning solution to the problem of sustained learning in keyword spotting
abstract
In some practical keyword spotting applications, users or service providers are willing to provide spotting-result feedback to help improve system performance. To do so, they require a keyword spotting technique with a sustained learning ability. This paper presents a new Chinese keyword spotting method based on a case based reasoning framework. Two level keyword case representations are adopted based on a set of symbols that are discriminative both in acoustic feature vector space and in semantic space. Then case bases are indexed with a tree structure and searched for test speech based on an elastic matching strategy. Finally, the feedback is used to adjust the statistics attached to the cases or to append new cases. Two experiments were conducted to compare our approach with a syllable lattice based method and to test the sustained learning ability.
Tieran Zheng, Jiqing Han 0001, Guibin Zheng, Shiwen Deng
ICASSP2
2013 Guarantees of Augmented Trace Norm Models in Tensor Recovery
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
IJCAI2
2013 Audio Segment Classification Using Online Learning Based Tensor Representation Feature Discrimination
abstract
In order to naturally combine audio information from different dimensions and build robust audio processing system, a novel framework based on low-rank tensor representation features for audio segment classification is proposed in this paper. The audio signal is first transformed into tensor format data, and then these tensor data are mapped to a low-rank space which is insensitive under certain noises, especially white Gaussian noise and gross corruptions. For these low-rank tensor based features, tensor classification via a linear classifier based on minimization a smooth loss function regularized by the trace norm proposed recently is used. Most previous methods find the weight tensor and bias in batch-mode learning, which makes them inefficient for large-scale problems. In this paper, we propose to address this problem with an online learning algorithm based on the accelerated proximal gradient (APG) method, which scales up gracefully to large data sets. Experiments on simulation and real audio data demonstrate the efficiency of the methods.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng, Shiwen Deng
IEEE Trans. Speech Audio Process.2
2013 Identification of Objectionable Audio Segments Based on Pseudo and Heterogeneous Mixture Models
abstract
In this paper, we generalize the Gaussian Mixture Model (GMM) in two ways: a) by introducing novel distance measures between two vectors based on nonlinear maps to give more general mixture models; b) by building mixture models based on multiple different kinds of distributions. These two generalizations cope with different problems arisen in feature modeling. Mixture model obtained by first method is called pseudo Gaussian Mixture Model (pseudo GMM). Compared to the traditional GMM, pseudo GMM with nonlinear maps have better performance on nonlinear problems, while the computational complexity is almost the same as the Expectation-Maximization (EM) algorithm for traditional GMM according to the iteration procedures. The second generalization considers that in practice the practical learning problem often involves multiple, heterogeneous data sources, while classical mixture models are based on a single kind of distribution. In this work, we consider heterogeneous mixture models (hetMM) based on multiple different kinds of distributions. Different types of distributions in hetMM may have quite different properties and may capture different features of the data. Component classifiers including pseudo and hetMM based classifiers are employed in our task of erotic audio recognition. Experimental results with classifiers built based on pseudo GMM and hetMM for erotic audio recognition demonstrate the effectiveness of the proposed model. Online and off-line experiments show that the proposed approach is highly effective for erotic audio recognition.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
IEEE Trans. Speech Audio Process.2
2013 Audio classification with low-rank matrix representation features
abstract
In this article, a novel framework based on trace norm minimization for audio classification is proposed. In this framework, both the feature extraction and classification are obtained by solving corresponding convex optimization problem with trace norm regularization. For feature extraction, robust principle component analysis (robust PCA) via minimization a combination of the nuclear norm and the ℓ 1 -norm is used to extract low-rank matrix features which are robust to white noise and gross corruption for audio signal. These low-rank matrix features are fed to a linear classifier where the weight and bias are learned by solving similar trace norm constrained problems. For this linear classifier, most methods find the parameters, that is the weight matrix and bias in batch-mode, which makes it inefficient for large scale problems. In this article, we propose a parallel online framework using accelerated proximal gradient method. This framework has advantages in processing speed and memory cost. In addition, as a result of the regularization formulation of matrix classification, the Lipschitz constant was given explicitly, and hence the step size estimation of the general proximal gradient method was omitted, and this part of computing burden is saved in our approach. Extensive experiments on real data sets for laugh/non-laugh and applause/non-applause classification indicate that this novel framework is effective and noise robust.
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
ACM Trans. Intell. Syst. Technol.2
2012 A solution to residual noise in speech denoising with sparse representation
abstract
As a promising technique, sparse representation has been extensively investigated in signal processing community. Recently, sparse representation is widely used for speech processing in noisy environments; however, many problems need to be solved because of the particularity of speech. One assumption for speech denoising with sparse representation is that the representation of speech over the dictionary is sparse, while that of the noise is dense. Unfortunately, this assumption is not sustained in speech denoising scenario. We find that many noises, e.g., the babble and white noises, are also sparse over the dictionary trained with clean speech, resulting in severe residual noise in sparse enhancement. To solve this problem, we propose a novel residual noise reduction (RNR) method which first finds out the atoms which represents the noise sparely, and then ignores them in the reconstruction of speech. Experimental results show that the proposed method can reduce residual noise substantially.
Yongjun He 0002, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng
ICASSP2
2012 Sparse power spectrum based robust voice activity detector
abstract
This paper presents a robust approach to improve the performance of voice activity detector (VAD) in low signal-to-noise ratio (SNR) noisy environments. To this end, we first generate sparse representations by Bregman Iteration based sparse decomposition with a learned over-complete dictionary, and derive a kind of audio feature called sparse power spectrum from the sparse representations. we then propose a method to calculate the short segment average spectrum and long segment average spectrum from sparse power spectrum. Finally, we design a criterion to detect speech region and non-speech region based on the above average spectrum. Experiments show that the proposed approach further improves the performance of VAD in low SNR noisy environments.
Datao You, Jiqing Han 0001, Guibin Zheng, Tieran Zheng
ICASSP2
2012 A Novel Confidence Measure Based on Context Consistency for Spoken Term Detection
Jiqing Han 0001, Tieran Zheng, Guibin Zheng
INTERSPEECH2
2012 Low-rank Audio Signal Classification Under Soft Margin and Trace Norm Constraints
Ziqiang Shi, Tieran Zheng, Jiqing Han 0001, Shiwen Deng
INTERSPEECH3
2012 Likelihood ratio sign test for voice activity detection
abstract
Voice activity detection (VAD) plays an important role on the performance of speech processing systems in adverse environments. Recently, statistical model-based VADs have demonstrated impressive performance. The study presents a novel decision test (named likelihood ratio sign test, LRST) for VAD by using sign test and Neyman–Pearson criterion to improve the performance of statistical model-based VAD. The proposed LRST is derived based on the likelihood ratios (LRs) calculated from multiple independent observations by incorporating the long-term speech information into the decision rule. An implementation of the LRST VAD is introduced by defining the LRST over a sliding window and calculating the LRs based on complex Gaussian distribution for an input signal. For experiments, the multiple-observation LRT (MO-LRT) VAD based on multiple observations is used as a reference owing to its outstanding performance compared with conventional VADs. The experimental results show that the proposed approach outperforms the MO-LRT VAD in various noise environments.
Shiwen Deng, Jiqing Han 0001
IET Signal Process.2
2012 Sparse-Based auditory Model for robust speaker Recognition
abstract
The mismatch between the training and the testing environments greatly degrades the performance of speaker recognition. Although many robust techniques have been proposed, speaker recognition in mismatch condition is still a challenge. To solve this problem, we propose a sparse-based auditory model as the front-end of speaker recognition by simulating auditory processing of speech signal. To this end, we introduce narrow-band filter-bank instead of the widely used wide-band filter-bank to simulate the basilar membrane filter-bank, use sparse representation as the approximation of basilar membrane coding strategy, and incorporate the frequency selectivity enhance mechanism between tectorial membrane and basilar membrane by practical engineering approximation. Compared with the standard Mel-frequency cepstral coefficient approach, our preliminary experimental results indicate that the sparse-based auditory model consistently improve the robustness of speaker recognition in mismatched condition.
Datao You, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
Int. J. Pattern Recognit. Artif. Intell.2
2011 A modified MAP criterion based on hidden Markov model for voice activity detecion
abstract
The maximum a posteriori (MAP) criterion is broadly used in the statistical model-based voice activity detection (VAD) approaches. In the conventional MAP criterion, however, the inter-frame correlation of the voice activity is not taken into consideration. In this paper, we proposes a novel modified MAP criterion based on a two-state hidden Markov model (HMM) to improve the performance of the VAD, and the the inter-frame correlation of the voice activity is modeled. With the proposed MAP criterion, the decision rule is derived by explicitly incorporating the a priori, a posteriori, and inter-frame correlation information into the likelihood ratio test (LRT). In the LRT, a compensation factor for the hypothesis of speech presence is used to regulate the trade-off between the probability of detection and the false alarm probability. Experimental results show the superiority of the VAD algorithm based on the proposed MAP criterion in comparison with that based on the recent conditional MAP criterion (CMAP) under various noise conditions.
Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP2
2011 Compensation of partly reliable components for band-limited speech recognition with missing data techniques
abstract
Mismatch in speech bandwidth between training and real operation greatly degrades the performance of automatic speech recognition (ASR) systems. Missing feature technique (MFT) is effective in handling bandwidth mismatch. However, current MFT-based methods ignore the mismatch in the filter bank channels which cover the upper and lower limit cutoff frequencies. To solve this problem, we propose to partition the feature into reliable, unreliable and partly reliable parts, and then modify the probability density functions (PDFs) of the partly reliable part to match band-limited features. Experiments showed that such compensation further improved the performances of MFT-based methods under band-limited conditions.
Yongjun He 0002, Jiqing Han 0001, Tieran Zheng, Guibin Zheng
ICASSP2
2011 A cochlear neuron based robust feature for speaker recognition
abstract
In this paper, a robust feature for text-independent speaker recognition is proposed, which simulate the response mode of cochlear neurons in processing acoustic signal. The feature is derived from sparse coding coefficient which is computed on a learned over-complete dictionary, and the dictionary is considered similar to part of speech sensitive cochlear neurons. Furthermore, the feature is generated without dimension reducing and de-correlation. The robust feature is implemented to address the problem of mismatch situation between training and testing. Experiments show that the proposed feature outperforms the Mel-frequency cepstral coefficients (MFCC) feature, especially under noisy environments, the equal error rate (EER) of the MFCC drops to 21.6% (10 dB) from 10.3% (25 dB), while the EER of the proposed feature is also 6.6% (10 dB) with no degradation.
Datao You, Jiqing Han 0001, Tieran Zheng
ICASSP3
2011 A Novel Framework Based on Trace Norm Minimization for Audio Event Detection
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
ICONIP (2)2
2011 AUC Optimization Based Confidence Measure for Keyword Spotting
Jiqing Han 0001, Tieran Zheng
INTERSPEECH2
2011 Real-World Speech/Non-Speech Audio Classification Based on Sparse Representation Features and GPCs
Ziqiang Shi, Jiqing Han 0001, Tieran Zheng
INTERSPEECH2
2011 Gaussian Specific Compensation for Channel Distortion in Speech Recognition
abstract
Channel distortion is one of the major factors degrading the performance of automatic speech recognition (ASR) systems. Most of the current compensation methods rely on the assumption that the channel distortion remains unchanged within an utterance or globally. However, we show in this letter that the distortion varies over speech frames even if the channel response is unchanged. To address this problem, we relax the above-mentioned assumption and propose a new method to compensate the channel distortion for each Gaussian of the acoustic models. Firstly, we derive the relationship between the clean and distorted models, and then estimate the channel magnitude response with the expectation-maximization (EM) algorithm. Finally, we obtain the matched models with the estimated magnitude response and the clean models. Experiments were conducted on the TIMIT/NTIMIT databases and the results confirmed the effectiveness of the proposed method.
Yongjun He 0002, Jiqing Han 0001
IEEE Signal Process. Lett.2
2010 Voice Activity Detection Based on Complex Exponential Atomic Decomposition and Likelihood Ratio Test
abstract
The voice activity detection (VAD) algorithms by using Discrete Fourier Transform (DFT) coefficients are widely found in literature. However, some shortcomings for modeling a signal in the DFT can easily degrade the performance of a VAD in noise environment. To overcome the problem, this paper presents a novel approach by using the complex coefficients derived from complex exponential atomic decomposition of a signal. Those coefficients are modeled by a complex Gaussian probability distribution and a statistical model is employed to derive the decision rule from the likelihood ratio test. According to the experimental results, the proposed VAD method shows better performance than the VAD based on DFT coefficients in various noise environments.
Shiwen Deng, Jiqing Han 0001
ICPR2
2010 Robust statistical voice activity detection using a likelihood ratio sign test
Shiwen Deng, Jiqing Han 0001
INTERSPEECH2
2010 Model synthesis for band-limited speech recognition
Yongjun He 0002, Jiqing Han 0001
INTERSPEECH2
2010 Study on the Recognition of Objectionable Audio
abstract
In this paper, a novel method from the feature — porno-sounds recognition — point of view is proposed to detect adult video sequences automatically which may serve as a verification step, a supplementary method or an independent detector. To the specificity of erotic sound, its feature analysis is given. Based on the popular features, histograms and contours are introduced as new sets of features. At the same time due to the complexity of outside data, a general framework called in-class clustering is proposed which selects the most representative subclass for training and classification. All these efforts increase the recall rate and decrease the false positive rate. Experiments on real data from the Internet indicate that the proposed method yields superior performance with 89.17% recall rate and 10.78% false positive rate being achieved.
Ziqiang Shi, Boyang Gao, Tieran Zheng, Jiqing Han 0001
Int. J. Pattern Recognit. Artif. Intell.4
2010 Particle-based realistic simulation of fluid-solid interaction
abstract
Abstract In this paper a novel method for simulating incompressible viscous fluid and solid coupling is presented. In the coupling model, a rigid object is treated as a special fluid constrained to rigid body motion. To animate the coupling model, the Smoothed Particle Hydrodynamics method is used for solving the fluid motion equations. For keeping the rigidity of rigid objects, the total force and total torque exerted on solids is first worked out according to the impulse–momentum theorem, and then the movement of these rigid bodies is restricted to translations and rotations. Moreover, in order to prevent the fluids particles leaking into solids, a detection and correction procedure is presented, and the velocities of fluid particles will be tuned if the penetration is detected in this procedure. The proposed method can be implemented easily by extending the existing fluid solvers, the experimental results show that this method is capable of animating the realistic solid and fluid coupling. Copyright © 2010 John Wiley & Sons, Ltd.
Hongquan Sun, Jiqing Han 0001
Comput. Animat. Virtual Worlds2
2006 A multi-space distribution (MSD) approach to speech recognition of tonal languages
Huanliang Wang, Yao Qian, Frank K. Soong, Jian-Lai Zhou, Jiqing Han 0001
INTERSPEECH5
2005 Modifying Spectral Envelope to Synthetically Adjust Voice Quality and Articulation Parameters for Emotional Speech Synthesis
Yanqiu Shao, Jiqing Han 0001, Ting Liu 0001
ACII3
2001 Robust Speech Recognition Method Based on Discriminative Environment Feature Extraction
Jiqing Han 0001, Wen Gao 0001
J. Comput. Sci. Technol.1
2000 An environment model-based robust speech recognition
Lei Zhang 0093, Jiqing Han 0001, Chengguo Lv, Chengfa Wang
INTERSPEECH2
1999 Robust telephone speech recognition based on channel compensation
Jiqing Han 0001, Wen Gao 0001
Pattern Recognit.1
1998 Discriminative learning of additive noise and channel distortions for robust speech recognition
abstract
Learning the influence of additive noise and channel distortions from training data is an effective approach for robust speech recognition. Most of the previous methods are based on maximum likelihood estimation criterion. We propose a new method of discriminative learning environmental parameters, which is based on the minimum classification error (MCE) criterion. By using a simple classifier defined by ourselves and the generalized probabilistic descent (GPD) algorithm, we iteratively learn environmental parameters. After getting the parameters, we estimate the clean speech features from the observed speech features and then use the estimation of the clean speech features to train or test the back-end HMM classifier. The best error rate reduction of 32.1% is obtained, tested on a Korean 18 isolated confusion words task, relative to the conventional HMM system.
Jiqing Han 0001, Munsung Han, Gyu-Bong Park, Jeongue Park, Wen Gao 0001, Doosung Hwang
ICASSP1
1997 Relative mel-frequency cepstral coefficients compensation for robust telephone speech recognition
abstract
It is a crucial factor to find the robust and simple computation methods for the actual application of telephone speech recognition. In this paper, we propose a new channel compensation method, which uses a RASTA-like band-pass filter on the mel-frequency cepstral coefficients for robust telephone speech recognition. It is shown from the experiments that the proposed method, comparing with the RASTA processing, reduces the computational complexity without losing performance, and it is also better than CMS and two level CMS on the performance. We also verify that it is an effective approach to suppress very low modulation frequencies for robust telephone speech recognition.
Jiqing Han 0001, Munsung Han, Gyu-Bong Park, Jeongue Park, Wen Gao 0001
EUROSPEECH1