EDBT 2026 Demo / reviewers in the wild / expert
Yongjun He 0002
dblp:48/1117-2
· DBLP profile ↗
34ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mamba-UNet for reference-based super-resolution reconstruction
Bo Ding 0003, Yongjun He 0002, Jun Zhou 0001 |
Appl. Intell. | 4 |
| 2026 | Hierarchical progressive fusion: A novel explainability method for point cloud deep neural networks
Bo Ding 0003, Guangzhen Li, Jun Zhou 0001, Yongjun He 0002 |
Pattern Recognit. Lett. | 5 |
| 2026 | AS-EVNorm: Tail-Aware Extreme Value Normalization for Speaker VerificationabstractScore normalization is a key back-end technique in speaker verification for improving score comparability across trials. As a convenient and widely used normalization method, Adaptive Symmetric Normalization (AS-Norm) standardizes raw scores using mean-and-standard-deviation normalization parameters estimated from an adaptive cohort comprising the most similar impostor scores. However, since the adaptive cohort retains only the top-ranked impostor scores, these central-moment statistics may not optimally characterize the empirical distribution of these scores, resulting in suboptimal speaker verification performance. In this letter, we propose Adaptive Symmetric Extreme Value Normalization (AS-EVNorm), which treats the adaptive cohort as upper-tail samples from the impostor-score distribution and models them under extreme-value theory for more accurate normalization parameters. Experiments on VoxCeleb and CN-Celeb show that AS-EVNorm consistently reduces both EER and minDCF compared with AS-Norm across a broad range of adaptive cohort configurations, while maintaining competitive normalization time. Zekai Su, Jiqing Han 0001, Jianchen Li, Yikun Jiang, Zhifeng Jiang 0007, Tianhong Ding, Yongjun He 0002 |
IEEE Signal Process. Lett. | 7 |
| 2026 | Multimodal Local Global Interaction Networks for Automatic Depression Severity EstimationabstractPhysiological studies have shown that differences between depressed and healthy individuals are manifested in the audio and video modalities. Hence, some researchers have combined local and global information from audio or video modality to obtain the unimodal representation. Attention mechanisms or Multi-Layer Perceptrons (MLPs) are then used to complete the fusion of different representations. However, attention mechanisms or MLPs is essentially a linear aggregation manner, and lacks the ability to explore the element-wise interaction between local and global representations within and across modalities, which affects the accuracy of estimating the depression severity. To this end, we propose a Representation Interaction (RI) module, which uses the mutual linear adjustment to achieve element-wise interaction between representations. Thus, the RI module can be seen as an mutual observation of two representations, which helps to achieve complementary advantages and improve the model’s ability to characterize depression cues. Furthermore, since the interaction process generates multiple representations, we propose a Multi-representation Prediction (MP) module. This module implements multi-representation vectorization in a hierarchical manner from summarizing a single representation to aggregating multiple representations, and adopts the attention mechanism to obtain the estimation of an individual depression severity. In this way, we use the RI and MP modules to construct the Multimodal Local Global Interaction (MLGI) network. The experimental performance on AVEC 2013 and AVEC 2014 depression datasets demonstrates the effectiveness of our method. Mingyue Niu, Zhuhong Shao, Yongjun He 0002, Jianhua Tao 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Dual Orthogonality Sub-center Loss for Enhanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Tieran Zheng, Guibin Zheng, Yongjun He 0002 |
INTERSPEECH | 5 |
| 2025 | Adaptive Across-Subcenter Representation Learning for Imbalanced Anomalous Sound Detection
Dong Wang 0013, Jiqing Han 0001, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
INTERSPEECH | 5 |
| 2025 | Knowledge-Decoupled Functionally Invariant Path With Synthetic Personal Data for Personalized ASRabstractFine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic personal data without forgetting real knowledge, and in adapting to personal data without forgetting generic knowledge. Considering that the functionally invariant path (FIP) framework enables model adaptation while preserving prior knowledge, in this letter, we introduce FIP into synthetic-data-augmented personalized ASR models. However, the model still struggles to balance the learning of synthetic, personalized, and generic knowledge when applying FIP to train the model on all three types of data simultaneously. To decouple this learning process and further address the above two challenges, we integrate a gated parameter-isolation strategy into FIP and propose a knowledge-decoupled functionally invariant path (KDFIP) framework, which stores generic and personalized knowledge in separate modules and applies FIP to them sequentially. Specifically, KDFIP adapts the personalized module to synthetic and real personal data and the generic module to generic data. Both modules are updated along personalization-invariant paths, and their outputs are dynamically fused through a gating mechanism. With augmented synthetic data, KDFIP achieves a 29.38% relative character error rate reduction on target speakers and maintains comparable generalization performance to the unadapted ASR baseline. Zhihao Du, Ying Shi 0001, Jiqing Han 0001, Yongjun He 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | A Dual-Path Multiple Instance Learning Network Guided by Image Quality Assessment for Cervical Whole Slide Image ClassificationabstractThe existing cervical whole slide image classification methods ignore the influence of image quality, resulting in low classification accuracy. To address this, we propose a dual-path multiple instance learning classification method guided by image quality assessment. Specifically, a pre-trained quality assessment model assigns quality scores to patches, splitting them into high- and low-quality paths. In the high-quality path, patch features are weighted by their quality scores to emphasize reliable diagnostic regions. In the low-quality path, a key instance is selected using clustering and feature distance matching. Finally, a cross-attention module fuses features across quality levels. Our method achieves 94.64% accuracy and 91.74% AUC on a dataset of 2,434 WSIs collected from five medical centers, outperforming state-of-the-art methods. Lanlan Kang, Jian Wang 0039, Yongjun He 0002, Bo Ding 0003 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Examining the Fourier Spectrum of Speech Signal From a Time-Frequency Perspective for Automatic Depression Level PredictionabstractCurrently, many studies use Fourier amplitude spectra of speech signals to predict depression levels. However, those works often treat Fourier amplitude spectra as images or sequences to capture depression cues using convolutional neural networks or multilayer perceptrons. Therefore, they ignore the complex element composition and time-frequency attributes of Fourier spectra, which is not conducive to capturing the differences among individuals with different depression levels. For this reason, we construct a Time-Frequency Self-Embedding (TFSE) module, which not only stores the correlation relationship among real (imaginary) parts of Fourier spectra of different subjects from the time-frequency perspective, but also maintain the physical properties of data through the weight embedding process. Besides, Global Average Pooling (GAP) or linear layers are difficult to balance both temporal and frequency dimensions in the vectorization process. Therefore, we construct a Time-Frequency Tensor Vectorization (TFTV) module, which summarizes each channel along time and frequency dimensions, and then generates the vectorization result by integrating various channels. In this way, we combine TFSE and TFTV modules to form our SpectrumFormer model for predicting depression levels. Evaluation indicators on AVEC 2013 and AVEC 2014 depression databases imply the progressiveness of our model. Mingyue Niu, Jianhua Tao 0001, Yongjun He 0002, Shiqing Zhang, Ming Li 0065 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Joint Energy-Based Model for Semi-Supervised Respiratory Sound Classification: A Method of Insensitive to Distribution MismatchabstractSemi-supervised learning effectively mitigates the lack of labeled data by introducing extensive unlabeled data. Despite achieving success in respiratory sound classification, in practice, it usually takes years to acquire a sufficiently sizeable unlabeled set, which consequently results in an extension of the research timeline. Considering that there are also respiratory sounds available in other related tasks, like breath phase detection and COVID-19 detection, it might be an alternative manner to treat these external samples as unlabeled data for respiratory sound classification. However, since these external samples are collected in different scenarios via different devices, there inevitably exists a distribution mismatch between the labeled and external unlabeled data. For existing methods, they usually assume that the labeled and unlabeled data follow the same data distribution. Therefore, they cannot benefit from external samples. To utilize external unlabeled data, we propose a semi-supervised method based on Joint Energy-based Model (JEM) in this paper. During training, the method attempts to use only the essential semantic components within the samples to model the data distribution. When non-semantic components like recording environments and devices vary, as these non-semantic components have a small impact on the model training, a relatively accurate distribution estimation is obtained. Therefore, the method exhibits insensitivity to the distribution mismatch, enabling the model to leverage external unlabeled data to mitigate the lack of labeled data. Taking ICBHI 2017 as the labeled set, HF_Lung_V1 and COVID-19 Sounds as the external unlabeled sets, the proposed method exceeds the baseline by 12.86. Wenjie Song 0003, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng, Yongjun He 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Semantic-Enhanced ULIP for Zero-Shot 3D Shape RecognitionabstractIn recent years, the zero-shot image recognition with semantic knowledge has achieved good performance due to vision-language models. However, because of the complexity of 3D shapes, the model cannot fully use the semantic knowledge of 3D shapes, which results in low accuracy of zero-shot 3D shape recognition. To address this problem, we propose a Semantic-enhanced ULIP for Zero-shot 3D Shape Recognition (SE-ULIP). This method utilizes the contrastive learning to fine-tune the text encoder in two stages, including the domain adaptation fine-tuning and the triplets-based text encoder fine-tuning. In the domain adaptation fine-tuning, we fine-tune the image encoder and the text encoder using the views and the Semantic Descriptive Text (SDT) of each view generated by the Visual Question Answering (VQA) model, which aims to align the view features with the semantic knowledge. In the triplets-based text encoder fine-tuning, we propose an Adaptive Conditional Adjustment Context Optimization (ACACoOp) to learn the optimal context vectors. The optimal context vectors are used as the input to fine-tune the text encoder again, which enhance SE-ULIP to understand the semantic knowledge of 3D shapes. Experiments show that our method achieves the state-of-the-art performance through the fine-tuned text encoder on three 3D backbone networks for both zero-shot and standard 3D shape recognition. Bo Ding 0003, Libao Zhang, Yongjun He 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Modeling Quasi-Periodic Dependency via Self-Supervised Pre-Training for Respiratory Sound ClassificationabstractDespite the success of self-supervised respiratory sound classification methods, they do not consider that respiratory sounds are quasi-periodic signals with repetitive patterns in successive breaths, which is vital for distinguishing respiratory sounds from non-quasi-periodic sounds like noises. Therefore, the existing methods may achieve limited improvement due to ignoring the quasi-periodic dependency. To this end, considering that the segments containing the same respiratory sound pattern should be similar in a sample, we extract the segment-wise representations and evaluate the similarity between the periodic-dependent representations via a sparse self-relation matrix. By defining a periodic consistency loss, we push the sparse self-relation matrixes of two clips of the same sample closer, encouraging a larger similarity between the representations. In this manner, the method can focus more on the respiratory sound-related quasi-periodic patterns that repeatedly recur in the periodic-dependent segments. Taking HF_Lung_V1 and COVID-19 Sounds as pre-training sets, the method exceeds the baseline by 7.67% on the ICBHI 2017 classification task. Wenjie Song 0003, Jiqing Han 0001, Jianchen Li, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
ICASSP | 6 |
| 2024 | Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event DetectionabstractOverlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events using shared and entangled frame-wise features, which degrades the feature discrimination. To solve the problem, we propose a disentangled feature learning framework to learn a category-specific representation. Specifically, we employ different projectors to learn the frame-wise features for each category. To ensure that these feature does not contain information of other categories, we maximize the common information between frame-wise features within the same category and propose a frame-wise contrastive loss. In addition, considering that the labeled data used by the proposed method is limited, we propose a semi-supervised frame-wise contrastive loss that can leverage large amounts of unlabeled data to achieve feature disentanglement. The experimental results demonstrate the effectiveness of our method. Yadong Guan, Jiqing Han 0001, Wenjie Song 0003, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
ICASSP | 7 |
| 2024 | Personality-memory Gated Adaptation: An Efficient Speaker Adaptation for Personalized End-to-end Automatic Speech Recognition
Zhihao Du, Shiliang Zhang, Jiqing Han 0001, Yongjun He 0002 |
INTERSPEECH | 5 |
| 2024 | Capturing High-Level Semantic Correlations via Graph for Multimodal Sentiment AnalysisabstractModeling intra-modal and cross-modal interactions poses significant challenges in multimodal sentiment analysis. Currently, graph-based methods like HGraph-CL achieve promising performance, which rely on two different levels of graph contrastive learning within and between modalities to explore sentiment correlations. However, HGraph-CL still faces the following drawbacks in graph construction: 1) nodes of the graph are represented at the frame level, only containing low-level information, neglecting the correlations among high-level semantics; 2) edges of the graph are based on the fixed dependency relations between words in the text sequence and the adjacent relations between frame-level nodes in the non-verbal sequences, failing to effectively capture implicit and long-distance correlations. To this end, this letter introduces capsule networks to construct high-level semantic nodes in a graph, uncovering deep sentimental structures. Furthermore, the learnable adjacency matrices are employed to construct edges of graph, thus adaptively learning the relations between nodes. Experimental results on several benchmark datasets for multimodal sentiment analysis demonstrate the effectiveness of the proposed method. Fan Qian, Jiqing Han 0001, Yadong Guan, Wenjie Song 0003, Yongjun He 0002 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Sound Activity-Aware Based Cross-Task Collaborative Training for Semi-Supervised Sound Event DetectionabstractThe training of sound event detection (SED) models remains a challenge of insufficient supervision due to limited frame-wise labeled data. Mainstream research on this problem has adopted semi-supervised training strategies that generate pseudo-labels for unlabeled data and use these data for the training of a model. Recent works further introduce multi-task training strategies to impose additional supervision. However, the auxiliary tasks employed in these methods either lack frame-wise guidance or exhibit unsuitable task designs. Furthermore, they fail to exploit inter-task relationships effectively, which can serve as valuable supervision. In this paper, we introduce a novel task, sound occurrence and overlap detection (SOD), which detects predefined sound activity patterns, including non-overlapping and overlapping cases. On the basis of SOD, we propose a cross-task collaborative training framework that leverages the relationship between SED and SOD to improve the SED model. Firstly, by jointly optimizing the two tasks in a multi-task manner, the SED model is encouraged to learn features sensitive to sound activity. Subsequently, the cross-task consistency regularization is proposed to promote consistent predictions between SED and SOD. Finally, we propose a pseudo-label selection method that uses inconsistent predictions between the two tasks to identify potential wrong pseudo-labels and mitigate their confirmation bias. In the inference phase, only the trained SED model is used, thus no additional computation and storage costs are incurred. Extensive experiments on the DESED dataset demonstrate the effectiveness of our method. Yadong Guan, Jiqing Han 0001, Shiwen Deng, Guibin Zheng, Tieran Zheng, Yongjun He 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2024 | Distance Metric-Based Open-Set Domain Adaptation for Speaker VerificationabstractDomain shift poses a significant challenge in speaker verification, especially in open-set scenarios where the speaker categories are disjoint between the source and target domains. To alleviate the domain shift, traditional domain adaptation methods typically align the source and target distributions in the speaker embedding space, but this may cause the overlap of embeddings from different speakers. To address this problem, this paper proposes to perform the domain alignment in a novel distance metric space, where the source and target domains exhibit the shared within-speaker and between-speaker categories. Thus, the discrepancy between the source and target domains arises only from the domain shift. We refer to the proposed method as Cross-Domain Distance Metric Adaptation (CDMA), in which the within- and between-speaker distance distributions in the target domain are aligned with the source distance distributions and further separated to minimize their overlap. This alignment and separation require estimating the within- and between-speaker distance distributions based on speaker labels, which are unavailable in the unlabeled target domain. Thus, we further propose a learnable speaker clustering method called Graph Convolutional Network with Graph Pruning (GCN-GP). This method generates high-quality pseudo-labels to estimate the two distance distributions in the target domain. Experimental results demonstrate that our method achieves state-of-the-art performance on the FFSVC2022 and VOiCES datasets. Jianchen Li, Jiqing Han 0001, Fan Qian, Tieran Zheng, Yongjun He 0002, Guibin Zheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | 3D shape classification based on global and local features extraction with collaborative learning
Bo Ding 0003, Libao Zhang, Yongjun He 0002 |
Vis. Comput. | 3 |
| 2023 | Graph-Based Spectro-Temporal Dependency Modeling for Anti-SpoofingabstractA great deal of recent research reveals that artifacts introduced by spoofing algorithms reside in specific frequency subbands or temporal segments. Therefore, the performance of spoofing detection can be improved by focusing on these regions. However, it is difficult for the detection system to choose an appropriate region when it encounters an unknown spoofing algorithm, resulting in poor generalization. Actually, there is a noticeable difference in the inter-region relationship between the bonafide and spoofed speeches. We name the inter-region relationship spectro-temporal dependency and design a method to model it for anti-spoofing. By focusing on the general dependency difference rather than specific regions, the generalization ability of the detection system can be improved. We employ a graph neural network to model the dependency and incorporate prior knowledge into the graph by designing the graph structure and edge weight, which forces the network to pay more attention to potential relationships. In addition, an attention mechanism is introduced in the graph pooling to focus on more critical nodes. The proposed method achieves an equal error rate of 0.58% on the ASVspoof 2019 LA dataset and outperforms all competing systems. Shiwen Deng, Tieran Zheng, Yongjun He 0002, Jiqing Han 0001 |
ICASSP | 4 |
| 2023 | Mutual Information-based Embedding Decoupling for Generalizable Speaker Verification
Jianchen Li, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Yongjun He 0002, Guibin Zheng |
INTERSPEECH | 5 |
| 2023 | Enhanced VAEGAN: a zero-shot image classification method
Bo Ding 0003, Yufei Fan, Yongjun He 0002 |
Appl. Intell. | 3 |
| 2022 | A Multi-Task Feature Fusion Model for Cervical Cell ClassificationabstractCervical cell classification is a crucial technique for automatic screening of cervical cancer. Although deep learning has greatly improved the accuracy of cell classification, the performance still cannot meet the needs of practical applications. To solve this problem, we propose a multi-task feature fusion model that consists of one auxiliary task of manual feature fitting and two main classification tasks. The auxiliary task enhances the main tasks in a manner of low-layer feature fusion. The main tasks, i.e., a 2-class classification task and a 5-class classification task, are learned together to realize their mutual reinforcement and alleviate the influence of unreliable labels. In addition, a label smoothing method based on cell category similarity is designed to bring inter-cell class information into the label. Comparative experimental results with other state-of-the-art models on the HUSTC and SIPaKMeD datasets prove the effectiveness of the proposed method. With a high sensitivity of 99.82% and a specificity of 98.12% for the 2-class classification task on the HUSTC dataset, our method shows potential to reduce cytologist workload. Yongjun He 0002, Jinping Ge, Yiqin Liang |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | AL-Net: Attention Learning Network Based on Multi-Task Learning for Cervical Nucleus SegmentationabstractCervical nucleus segmentation is a crucial and challenging issue in automatic pathological diagnosis due to uneven staining, blurry boundaries, and adherent or overlapping nuclei in nucleus images. To overcome the limitation of current methods, we propose a multi-task network based on U-Net for cervical nucleus segmentation. This network consists of a primary task and an auxiliary task. The primary task is employed to predict nuclei regions. The auxiliary task, which predicts the boundaries of nuclei, is designed to improve the feature extraction of the main task. Furthermore, a context encoding layer is added behind each encoding layer of the U-Net. The output of each context encoding layer is processed by an attention learning module and then fused with the features of the decoding layer. In addition, a codec block is used in the attention learning module to obtain saliency-based attention and focused attention simultaneously. Experiment results show that the proposed network performs better than the state-of-the-art methods on the 2014 ISBI dataset, BNS, MoNuSeg, and our nucluesSeg dataset. Yongjun He 0002, Si-Qi Zhao, Jinjie Huang, Wangmeng Zuo |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Illumination compensation for microscope images based on illumination difference estimation
Huili Shao, Yongjun He 0002, Lejun Zhang |
Vis. Comput. | 5 |
| 2021 | Attention-Guided Digital Adversarial Patches on Visual DetectionabstractDeep learning has been widely used in the field of image classification and image recognition and achieved positive practical results. However, in recent years, a number of studies have found that the accuracy of deep learning model based on classification greatly drops when making only subtle changes to the original examples, thus realizing the attack on the deep learning model. The main methods are as follows: adjust the pixels of attack examples invisible to human eyes and induce deep learning model to make the wrong classification; by adding an adversarial patch on the detection target, guide and deceive the classification model to make it misclassification. Therefore, these methods have strong randomness and are of very limited use in practical application. Different from the previous perturbation to traffic signs, our paper proposes a method that is able to successfully hide and misclassify vehicles in complex contexts. This method takes into account the complex real scenarios and can perturb with the pictures taken by a camera and mobile phone so that the detector based on deep learning model cannot detect the vehicle or misclassification. In order to improve the robustness, the position and size of the adversarial patch are adjusted according to different detection models by introducing the attachment mechanism. Through the test of different detectors, the patch generated in the single target detection algorithm can also attack other detectors and do well in transferability. Based on the experimental part of this paper, the proposed algorithm is able to significantly lower the accuracy of the detector. Affected by the real world, such as distance, light, angles, resolution, etc., the false classification of the target is realized by reducing the confidence level and background of the target, which greatly perturbs the detection results of the target detector. In COCO Dataset 2017, it reveals that the success rate of this algorithm reaches 88.7%. Dapeng Lang, Yongjun He 0002 |
Secur. Commun. Networks | 4 |
| 2021 | Overlapping region reconstruction in nuclei image segmentation
Yining Xie, Yongjun He 0002 |
Vis. Comput. | 4 |
| 2016 | Optimization of learned dictionary for sparse coding in speech processing
Yongjun He 0002, Guanglu Sun, Jiqing Han 0001 |
Neurocomputing | 1 |
| 2015 | Noise-robust speaker recognition based on morphological component analysis
Yongjun He 0002, Chen Chen 0086, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2015 | Dictionary evaluation and optimization for sparse coding based speech processing
Yongjun He 0002, Guanglu Sun, Jiqing Han 0001 |
Inf. Sci. | 1 |
| 2014 | Evaluation of dictionary for sparse coding in speech processing
Yongjun He 0002, Guanglu Sun, Guibin Zheng, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2012 | A solution to residual noise in speech denoising with sparse representationabstractAs a promising technique, sparse representation has been extensively investigated in signal processing community. Recently, sparse representation is widely used for speech processing in noisy environments; however, many problems need to be solved because of the particularity of speech. One assumption for speech denoising with sparse representation is that the representation of speech over the dictionary is sparse, while that of the noise is dense. Unfortunately, this assumption is not sustained in speech denoising scenario. We find that many noises, e.g., the babble and white noises, are also sparse over the dictionary trained with clean speech, resulting in severe residual noise in sparse enhancement. To solve this problem, we propose a novel residual noise reduction (RNR) method which first finds out the atoms which represents the noise sparely, and then ignores them in the reconstruction of speech. Experimental results show that the proposed method can reduce residual noise substantially. Yongjun He 0002, Jiqing Han 0001, Shiwen Deng, Tieran Zheng, Guibin Zheng |
ICASSP | 1 |
| 2011 | Compensation of partly reliable components for band-limited speech recognition with missing data techniquesabstractMismatch in speech bandwidth between training and real operation greatly degrades the performance of automatic speech recognition (ASR) systems. Missing feature technique (MFT) is effective in handling bandwidth mismatch. However, current MFT-based methods ignore the mismatch in the filter bank channels which cover the upper and lower limit cutoff frequencies. To solve this problem, we propose to partition the feature into reliable, unreliable and partly reliable parts, and then modify the probability density functions (PDFs) of the partly reliable part to match band-limited features. Experiments showed that such compensation further improved the performances of MFT-based methods under band-limited conditions. Yongjun He 0002, Jiqing Han 0001, Tieran Zheng, Guibin Zheng |
ICASSP | 1 |
| 2011 | Gaussian Specific Compensation for Channel Distortion in Speech RecognitionabstractChannel distortion is one of the major factors degrading the performance of automatic speech recognition (ASR) systems. Most of the current compensation methods rely on the assumption that the channel distortion remains unchanged within an utterance or globally. However, we show in this letter that the distortion varies over speech frames even if the channel response is unchanged. To address this problem, we relax the above-mentioned assumption and propose a new method to compensate the channel distortion for each Gaussian of the acoustic models. Firstly, we derive the relationship between the clean and distorted models, and then estimate the channel magnitude response with the expectation-maximization (EM) algorithm. Finally, we obtain the matched models with the estimated magnitude response and the clean models. Experiments were conducted on the TIMIT/NTIMIT databases and the results confirmed the effectiveness of the proposed method. Yongjun He 0002, Jiqing Han 0001 |
IEEE Signal Process. Lett. | 1 |
| 2010 | Model synthesis for band-limited speech recognition
Yongjun He 0002, Jiqing Han 0001 |
INTERSPEECH | 1 |