EDBT 2026 Demo / reviewers in the wild / expert
Guibin Zheng
dblp:z/GuibinZheng
· DBLP profile ↗
24ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0002-4280-5224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual Orthogonality Sub-center Loss for Enhanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Tieran Zheng, Guibin Zheng, Yongjun He 0002 |
INTERSPEECH | 4 |
| 2025 | Adaptive Across-Subcenter Representation Learning for Imbalanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
INTERSPEECH | 3 |
| 2025 | Knowledge Distillation Method for Pruned RNN-T Models via Pruning Bounds Sharing and Losses Confusion
Xiaocan Zhang, Guibin Zheng, Chenhao Jing, Jiqing Han 0001, Tieran Zheng |
INTERSPEECH | 3 |
| 2025 | Joint Energy-Based Model for Semi-Supervised Respiratory Sound Classification: A Method of Insensitive to Distribution MismatchabstractSemi-supervised learning effectively mitigates the lack of labeled data by introducing extensive unlabeled data. Despite achieving success in respiratory sound classification, in practice, it usually takes years to acquire a sufficiently sizeable unlabeled set, which consequently results in an extension of the research timeline. Considering that there are also respiratory sounds available in other related tasks, like breath phase detection and COVID-19 detection, it might be an alternative manner to treat these external samples as unlabeled data for respiratory sound classification. However, since these external samples are collected in different scenarios via different devices, there inevitably exists a distribution mismatch between the labeled and external unlabeled data. For existing methods, they usually assume that the labeled and unlabeled data follow the same data distribution. Therefore, they cannot benefit from external samples. To utilize external unlabeled data, we propose a semi-supervised method based on Joint Energy-based Model (JEM) in this paper. During training, the method attempts to use only the essential semantic components within the samples to model the data distribution. When non-semantic components like recording environments and devices vary, as these non-semantic components have a small impact on the model training, a relatively accurate distribution estimation is obtained. Therefore, the method exhibits insensitivity to the distribution mismatch, enabling the model to leverage external unlabeled data to mitigate the lack of labeled data. Taking ICBHI 2017 as the labeled set, HF_Lung_V1 and COVID-19 Sounds as the external unlabeled sets, the proposed method exceeds the baseline by 12.86. Wenjie Song 0003, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng, Yongjun He 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Modeling Quasi-Periodic Dependency via Self-Supervised Pre-Training for Respiratory Sound ClassificationabstractDespite the success of self-supervised respiratory sound classification methods, they do not consider that respiratory sounds are quasi-periodic signals with repetitive patterns in successive breaths, which is vital for distinguishing respiratory sounds from non-quasi-periodic sounds like noises. Therefore, the existing methods may achieve limited improvement due to ignoring the quasi-periodic dependency. To this end, considering that the segments containing the same respiratory sound pattern should be similar in a sample, we extract the segment-wise representations and evaluate the similarity between the periodic-dependent representations via a sparse self-relation matrix. By defining a periodic consistency loss, we push the sparse self-relation matrixes of two clips of the same sample closer, encouraging a larger similarity between the representations. In this manner, the method can focus more on the respiratory sound-related quasi-periodic patterns that repeatedly recur in the periodic-dependent segments. Taking HF_Lung_V1 and COVID-19 Sounds as pre-training sets, the method exceeds the baseline by 7.67% on the ICBHI 2017 classification task. Wenjie Song 0003, Jiqing Han 0001, Jianchen Li, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
ICASSP | 4 |
| 2024 | Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event DetectionabstractOverlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events using shared and entangled frame-wise features, which degrades the feature discrimination. To solve the problem, we propose a disentangled feature learning framework to learn a category-specific representation. Specifically, we employ different projectors to learn the frame-wise features for each category. To ensure that these feature does not contain information of other categories, we maximize the common information between frame-wise features within the same category and propose a frame-wise contrastive loss. In addition, considering that the labeled data used by the proposed method is limited, we propose a semi-supervised frame-wise contrastive loss that can leverage large amounts of unlabeled data to achieve feature disentanglement. The experimental results demonstrate the effectiveness of our method. Yadong Guan, Jiqing Han 0001, Wenjie Song 0003, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
ICASSP | 5 |
| 2024 | Sound Activity-Aware Based Cross-Task Collaborative Training for Semi-Supervised Sound Event DetectionabstractThe training of sound event detection (SED) models remains a challenge of insufficient supervision due to limited frame-wise labeled data. Mainstream research on this problem has adopted semi-supervised training strategies that generate pseudo-labels for unlabeled data and use these data for the training of a model. Recent works further introduce multi-task training strategies to impose additional supervision. However, the auxiliary tasks employed in these methods either lack frame-wise guidance or exhibit unsuitable task designs. Furthermore, they fail to exploit inter-task relationships effectively, which can serve as valuable supervision. In this paper, we introduce a novel task, sound occurrence and overlap detection (SOD), which detects predefined sound activity patterns, including non-overlapping and overlapping cases. On the basis of SOD, we propose a cross-task collaborative training framework that leverages the relationship between SED and SOD to improve the SED model. Firstly, by jointly optimizing the two tasks in a multi-task manner, the SED model is encouraged to learn features sensitive to sound activity. Subsequently, the cross-task consistency regularization is proposed to promote consistent predictions between SED and SOD. Finally, we propose a pseudo-label selection method that uses inconsistent predictions between the two tasks to identify potential wrong pseudo-labels and mitigate their confirmation bias. In the inference phase, only the trained SED model is used, thus no additional computation and storage costs are incurred. Extensive experiments on the DESED dataset demonstrate the effectiveness of our method. Yadong Guan, Jiqing Han 0001, Shiwen Deng, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Distance Metric-Based Open-Set Domain Adaptation for Speaker VerificationabstractDomain shift poses a significant challenge in speaker verification, especially in open-set scenarios where the speaker categories are disjoint between the source and target domains. To alleviate the domain shift, traditional domain adaptation methods typically align the source and target distributions in the speaker embedding space, but this may cause the overlap of embeddings from different speakers. To address this problem, this paper proposes to perform the domain alignment in a novel distance metric space, where the source and target domains exhibit the shared within-speaker and between-speaker categories. Thus, the discrepancy between the source and target domains arises only from the domain shift. We refer to the proposed method as Cross-Domain Distance Metric Adaptation (CDMA), in which the within- and between-speaker distance distributions in the target domain are aligned with the source distance distributions and further separated to minimize their overlap. This alignment and separation require estimating the within- and between-speaker distance distributions based on speaker labels, which are unavailable in the unlabeled target domain. Thus, we further propose a learnable speaker clustering method called Graph Convolutional Network with Graph Pruning (GCN-GP). This method generates high-quality pseudo-labels to estimate the two distance distributions in the target domain. Experimental results demonstrate that our method achieves state-of-the-art performance on the FFSVC2022 and VOiCES datasets. Jianchen Li, Jiqing Han 0001, Fan Qian, Tieran Zheng, Yongjun He 0002, Guibin Zheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Subband Dependency Modeling for Sound Event DetectionabstractIn the domain of sound event detection (SED), Convolutional Recurrent Neural Network (CRNN) has become the most successful architecture, which adopts Recurrent Neural Network (RNN) to model temporal dependencies from the output of Convolutional Neural Network (CNN). However, CRNN does not fully use the subband dependencies that have been proved critical for human perception of sound events. In this paper, we propose a subband dependency model (SDM) to enhance the capability of CRNN in modeling subband dependencies from the input spectrogram. To select prominent subband dependencies, we propose a novel SoftSparsemax transformation. It can select the salient parts by comparing all dependencies and further strengthen them by projecting them onto a probability simplex. Furthermore, since subband dependencies of different sound events may be prominent in different timescales, multi-timescale subband dependency is considered. The experiment results demonstrate the effectiveness of our method. Yadong Guan, Guibin Zheng, Jiqing Han 0001, Huanliang Wang |
ICASSP | 2 |
| 2023 | Mutual Information-based Embedding Decoupling for Generalizable Speaker Verification
Jianchen Li, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Yongjun He 0002, Guibin Zheng |
INTERSPEECH | 6 |
| 2022 | Sparse Self-Attention for Semi-Supervised Sound Event DetectionabstractSelf-attention mechanism has been widely employed in semi-supervised sound event detection (SS-SED). In self-attention, since dependencies between pairwise features at all moments are captured, the irrelevant features of different classes of sounds and background sounds at other moments are inevitably mixed in the current embedding when self-attention performs weighted summation. These irrelevant features will weaken the ability of the aggregated embedding to describe sound events. In this paper, we propose a sparse self-attention mechanism to alleviate the impact. Specifically, the Sparsemax function is introduced for attention weights normalization, which uses Euclidean projection to project attention weights onto a probability simplex. After the normalization, the attention weights of the irrelevant features are projected onto the boundary of the simplex and then removed. Furthermore, to solve the excessive sparsity problem of the Sparsemax, we further propose the Sparsemax with adjustable sparsity. Experimental results demonstrate the effectiveness of the proposed method. Yadong Guan, Jiabin Xue, Guibin Zheng, Jiqing Han 0001 |
ICASSP | 3 |
| 2018 | Deep Neural Network Based Discriminative Training for I-Vector/PLDA Speaker VerificationabstractIn the studies of i-vector based speaker verification, the discriminative training of probabilistic linear discriminative analysis (PLDA) model has been proven to be an effective way to improve performance. This paper focuses on using a deep neural network (DNN) to strengthen the original discriminatively trained classifiers by its strong capability of nonlinear modeling representation. We first propose a deep neural network based dimensionality reduction model to replace the linear discriminant analysis (LDA) process, and then a discriminative training algorithm is also proposed to jointly optimize the network and PLDA scoring function under single discriminative criterion. Our experiments show that performance improvements are achieved in the male trials of short2-short3 core data set of NIST SRE08. Tieran Zheng, Jiqing Han 0001, Guibin Zheng |
ICASSP | 3 |
| 2017 | Learning Deep Neural Network Based Kernel Functions for Small Sample Size Classification
Tieran Zheng, Jiqing Han 0001, Guibin Zheng |
ICONIP (1) | 3 |
| 2016 | Speaker Verification via Modeling Kurtosis Using Sparse CodingabstractThis paper proposes a new model for speaker verification by employing kurtosis statistical method based on sparse coding of human auditory system. Since only a small number of neurons in primary auditory cortex are activated in encoding acoustic stimuli and sparse independent events are used to represent the characteristics of the neurons. Each individual dictionary is learned from individual speaker samples where dictionary atoms correspond to the cortex neurons. The neuron responses possess statistical properties of acoustic signals in auditory cortex so that the activation distribution of individual speaker’s neurons is approximated as the characteristics of the speaker. Kurtosis is an efficient approach to measure the sparsity of the neuron from its activation distribution, and the vector composed of the kurtosis of every neuron is obtained as the model to characterize the speaker’s voice. The experimental results demonstrate that the kurtosis model outperforms the baseline systems and an effective identity validation function is achieved desirably. Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2014 | Robust minimum statistics project coefficients feature for acoustic environment recognitionabstractAcoustic environment recognition has been widely used in many applications, and is a considerable difficult problem for the real-life and complex environment. This paper proposes a novel feature, named minimum statistics project coefficients (MSPC), and intents to solve this problem. The MSPC feature is extracted from the background sound which is more robust than the foreground sound for the task of acoustic environment recognition. Experimental results show the outstanding performance of the MSPC feature compared with the conventional acoustic features, especially in very complex acoustic environments. Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
ICASSP | 5 |
| 2014 | Evaluation of dictionary for sparse coding in speech processing
Yongjun He 0002, Guanglu Sun, Guibin Zheng, Jiqing Han 0001 |
INTERSPEECH | 3 |
| 2013 | Upper and lower bounds for approximation of the Kullback-Leibler divergence between Hidden Markov modelsabstractThe Kullback-Leibler (KL) divergence is often used for a similarity comparison between two Hidden Markov models (HMMs). However, there is no closed form expression for computing the KL divergence between HMMs, and it can only be approximated. In this paper, we propose two novel methods for approximating the KL divergence between the left-to-right transient HMMs. The first method is a product approximation which can be calculated recursively without introducing extra parameters. The second method is based on the upper and lower bounds of KL divergence, and the mean of these bounds provides an available approximation of the divergence. We demonstrate the effectiveness of the proposed methods through experiments including the deviations to the numerical approximation and the task of predicting the confusability of phone pairs. Experimental results show that the proposed product approximation is comparable with the current variational approximation, and the proposed approximation based on bounds performs better than current methods in the experiments. Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
ICASSP | 4 |
| 2013 | Case based reasoning solution to the problem of sustained learning in keyword spottingabstractIn some practical keyword spotting applications, users or service providers are willing to provide spotting-result feedback to help improve system performance. To do so, they require a keyword spotting technique with a sustained learning ability. This paper presents a new Chinese keyword spotting method based on a case based reasoning framework. Two level keyword case representations are adopted based on a set of symbols that are discriminative both in acoustic feature vector space and in semantic space. Then case bases are indexed with a tree structure and searched for test speech based on an elastic matching strategy. Finally, the feedback is used to adjust the statistics attached to the cases or to append new cases. Two experiments were conducted to compare our approach with a syllable lattice based method and to test the sustained learning ability. Tieran Zheng, Jiqing Han 0001, Guibin Zheng, Shiwen Deng |
ICASSP | 3 |
| 2012 | A solution to residual noise in speech denoising with sparse representationabstractAs a promising technique, sparse representation has been extensively investigated in signal processing community. Recently, sparse representation is widely used for speech processing in noisy environments; however, many problems need to be solved because of the particularity of speech. One assumption for speech denoising with sparse representation is that the representation of speech over the dictionary is sparse, while that of the noise is dense. Unfortunately, this assumption is not sustained in speech denoising scenario. We find that many noises, e.g., the babble and white noises, are also sparse over the dictionary trained with clean speech, resulting in severe residual noise in sparse enhancement. To solve this problem, we propose a novel residual noise reduction (RNR) method which first finds out the atoms which represents the noise sparely, and then ignores them in the reconstruction of speech. Experimental results show that the proposed method can reduce residual noise substantially. Yongjun He 0002, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng |
ICASSP | 5 |
| 2012 | Sparse power spectrum based robust voice activity detectorabstractThis paper presents a robust approach to improve the performance of voice activity detector (VAD) in low signal-to-noise ratio (SNR) noisy environments. To this end, we first generate sparse representations by Bregman Iteration based sparse decomposition with a learned over-complete dictionary, and derive a kind of audio feature called sparse power spectrum from the sparse representations. we then propose a method to calculate the short segment average spectrum and long segment average spectrum from sparse power spectrum. Finally, we design a criterion to detect speech region and non-speech region based on the above average spectrum. Experiments show that the proposed approach further improves the performance of VAD in low SNR noisy environments. Datao You, Jiqing Han 0001, Guibin Zheng, Tieran Zheng |
ICASSP | 3 |
| 2012 | A Novel Confidence Measure Based on Context Consistency for Spoken Term Detection
Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
INTERSPEECH | 4 |
| 2012 | Sparse-Based auditory Model for robust speaker RecognitionabstractThe mismatch between the training and the testing environments greatly degrades the performance of speaker recognition. Although many robust techniques have been proposed, speaker recognition in mismatch condition is still a challenge. To solve this problem, we propose a sparse-based auditory model as the front-end of speaker recognition by simulating auditory processing of speech signal. To this end, we introduce narrow-band filter-bank instead of the widely used wide-band filter-bank to simulate the basilar membrane filter-bank, use sparse representation as the approximation of basilar membrane coding strategy, and incorporate the frequency selectivity enhance mechanism between tectorial membrane and basilar membrane by practical engineering approximation. Compared with the standard Mel-frequency cepstral coefficient approach, our preliminary experimental results indicate that the sparse-based auditory model consistently improve the robustness of speaker recognition in mismatched condition. Datao You, Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2011 | A modified MAP criterion based on hidden Markov model for voice activity detecionabstractThe maximum a posteriori (MAP) criterion is broadly used in the statistical model-based voice activity detection (VAD) approaches. In the conventional MAP criterion, however, the inter-frame correlation of the voice activity is not taken into consideration. In this paper, we proposes a novel modified MAP criterion based on a two-state hidden Markov model (HMM) to improve the performance of the VAD, and the the inter-frame correlation of the voice activity is modeled. With the proposed MAP criterion, the decision rule is derived by explicitly incorporating the a priori, a posteriori, and inter-frame correlation information into the likelihood ratio test (LRT). In the LRT, a compensation factor for the hypothesis of speech presence is used to regulate the trade-off between the probability of detection and the false alarm probability. Experimental results show the superiority of the VAD algorithm based on the proposed MAP criterion in comparison with that based on the recent conditional MAP criterion (CMAP) under various noise conditions. Shiwen Deng, Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
ICASSP | 4 |
| 2011 | Compensation of partly reliable components for band-limited speech recognition with missing data techniquesabstractMismatch in speech bandwidth between training and real operation greatly degrades the performance of automatic speech recognition (ASR) systems. Missing feature technique (MFT) is effective in handling bandwidth mismatch. However, current MFT-based methods ignore the mismatch in the filter bank channels which cover the upper and lower limit cutoff frequencies. To solve this problem, we propose to partition the feature into reliable, unreliable and partly reliable parts, and then modify the probability density functions (PDFs) of the partly reliable part to match band-limited features. Experiments showed that such compensation further improved the performances of MFT-based methods under band-limited conditions. Yongjun He 0002, Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
ICASSP | 4 |