VLDB 2026 Research / reviewers in the wild / expert
Ying Hu 0005
dblp:92/4882-5
· DBLP profile ↗
30ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-7505-1767ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 23 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dual Consistency Training (DCT) strategy for polyphonic sound event detection
Ying Hu 0005, Xinchun Ma, Zhijian Ou |
Neurocomputing | 3 |
| 2025 | A Singing Melody Extraction Network Via Self-Distillation and Multi-Level SupervisionabstractExtracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper layers generally contain more contextual information than shallower layers. To help the shallower layers enhance the ability of task-relevant feature extraction, we propose a self-distillation and multi-level supervision (SD-MS) method, which leverages the feature distillation from the deepest layer to the shallower one and multi-level supervision to guide network training. Visualization analysis shows that by introducing SD-MS, the same-level layer in the network can obtain a clearer representation of fundamental frequency components, while the shallower layers can even learn more task-relevant semantic information. Ablation study results indicate that SD-MS applies to existing melody extraction models and can consistently improve performance. Experimental results show that our proposed method, MF-TFA with SD-MS, outperforms six compared state-of-the-art methods, achieving overall accuracy (OA) scores of 87.1%, 89.9%, and 76.6% on the ADC 2004, MIREX 05, and MEDLEY DB datasets, respectively. The main code will be available at https://github.com/SmoothJing/MF-TFA_SD-MS. Ying Hu 0005, Jiabo Jing, Fan Li 0003, Lijun He 0001, Wenzhong Yang |
ICASSP | 1 |
| 2025 | TDE-VC: Timbre Disentanglement and Extraction Via Consistency for Zero-Shot Voice ConversionabstractVoice conversion (VC) transforms certain characteristics of speech from a source to a target while preserving the original linguistic content. This paper focuses on timbre conversion, a key type of VC. Current VC methods face two challenges: retaining source speaker information in the extracted content and inadequately capturing timbre features, often leading to suboptimal speaker similarity in the converted speech. To address these issues, we propose the TDE-VC model, a zero-shot voice conversion framework that incorporates a phased-trained content extractor, combining the strengths of adversarial speaker classifier and data perturbation to extract cleaner content. Critically, we introduce a timbre disentanglement and extraction strategy, based on a multi-level consistency constraint, which effectively disentangles timbre from content and guides the timbre encoder to focus solely on timbre extraction. Additionally, we present an effective multi-scale timbre encoder. Experimental results demonstrate that TDE-VC significantly improves speaker similarity, especially for unseen target speakers, while maintaining competitive naturalness compared to existing methods. The demo page is publicly available.1. Ying Hu 0005, Shangkun Tu, Fan Li 0003, Lijun He 0001, Hai Yan |
ICME | 1 |
| 2025 | VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
Zijing Zhao 0008, Hao Huang 0009, Ying Hu 0005, Liang He 0003 |
INTERSPEECH | 4 |
| 2025 | A Joint Network for Singing Melody Extraction from Polyphonic Music with Attention Aggregation and Self-Consistency Training
Jiabo Jing, Ying Hu 0005, Hao Huang 0009, Liang He 0003, Zhijian Ou |
INTERSPEECH | 2 |
| 2024 | SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual AttentionabstractWe propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on the WSJ0-2mix-extr dataset. The ablation and comparison studies verified the effectiveness of SM strategy and MA block. The experimental results show that our proposed method outperforms the state-of-the-art methods by a sizable margin of 1.3 dB on the metric of scale-invariant signal-to-distortion ratio improvement. Additionally, SMMA-Net achieved that the performance of model for TSE task exceeds that for speaker separation task under the similar architecture. The main code will be available at https://github.com/Ht-Xu/SMMA-Net. Ying Hu 0005, Zhongcun Guo, Hao Huang 0009, Liang He 0003 |
ICASSP | 1 |
| 2024 | Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker VerificationabstractIncorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic speech recognition (ASR). Considering that speaker verification datasets typically consist of multiple languages, there are instances where speakers are proficient in multiple languages, resulting in discrepancies between the languages used in the enrolled and test utterances. To address these challenges, we employ a pre-trained multilingual ASR Conformer encoder to initialize the MFA-Conformer network for speaker verification. Experimental results on the VoxCeleb dataset demonstrate a significant improvement in the performance of the system that incorporates multilingual phonetic information across different evaluation sets, including VoxCeleb1-O, E, and H, as well as the VoxSRC21 validation set, which focuses on multilingual verification. The source code is released at https://github.com/zds-potato/multilingual-phonetic-sv. Zhida Song, Liang He 0003, Ying Hu 0005, Hao Huang 0009 |
ICASSP | 4 |
| 2024 | Speaker Recognition Based on Pre-Trained Model and Deep ClusteringabstractIn this paper, we propose a novel loss by integrating a deep clustering (DC) loss at the frame-level and a speaker recognition loss at the segment-level into a single network without additional data requirements and exhaustive computation. The DC loss implicitly generates soft pseudo-phoneme labels for each frame-level feature, which facilitates extracting more discriminant speaker representation by suppressing phonetic content information. We study the DC loss not only on the acoustic feature, but also on the features extracted by the pre-trained models, such as wav2vec 2.0, HuBERT and WavLM. Experimental results on the VoxCeleb dataset shows that the overall system performance based on the pre-trained model features are better than the one on the acoustic feature. The proposed loss is significantly effective for systems on the acoustic feature and has a marginal improvement for systems on the pre-trained model feature. Liang He 0003, Zhida Song, Shuanghong Liu, Mengqi Niu, Ying Hu 0005, Hao Huang 0009 |
ICME | 5 |
| 2024 | Cross-modal Features Interaction-and-Aggregation Network with Self-consistency Training for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 1 |
| 2024 | YOLOPitch: A Time-Frequency Dual-Branch YOLO Model for Pitch Estimation
Hao Huang 0009, Ying Hu 0005, Liang He 0003, Yuyi Wang 0007 |
INTERSPEECH | 3 |
| 2024 | IIFC-Net: A Monaural Speech Enhancement Network With High-Order Information Interaction and Feature CalibrationabstractRecently, many Transformer-style dual-path models have achieved impressive performance for speech enhancement. However, their high parameters and computational complexity hinder their practical application. In this letter, we propose a monaural speech enhancement network with lower parameter count and complexity based on high-order information interaction and feature calibration (IIFC-Net). The network includes high-order information interaction Transformer (HOIIFormer) with high-order information interaction (HOII) block instead of a multi-head self-attention (MHSA) in Transformer. IIFC-Net leverages dual-path HOIIFormer (DPH) to model the distant dependency relation along time and frequency dimensions, respectively, and effectively captures deep-level information through the HOII block. We also design a feature calibration (FC) block to enhance the frequency components of target speech, which can be verified by a visualization analysis. The outcomes of experiments conducted on the VoiceBank+DEMAND and WHAMR! Datasets demonstrate that IIFC-Net achieves comparable performance in terms of denoising, dereverberation, and simultaneous denoising & dereverberation. Wenbing Wei, Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Improving Speaker Verification With Noise-Aware Label Ensembling and Sample Selection: Learning and Correcting Noisy Speaker LabelsabstractSupervised deep learning has achieved tremendous success in speaker verification. However, deep speaker models tend to overfit noisy labels when they are present in the speaker datasets. To mitigate the detrimental effects of noisy labels, in this paper, we propose a novelLabel Ensembling and Sample Selectionframework. Firstly, we select labels with high confidence rankings as clean samples. Additionally, we use predictions from different epochs during training to smoothly correct the noisy labels. Our method does not require staged training and achieves integration of learning from noisy labels, selecting clean labels, and correcting noisy labels. A significant number of experimental results demonstrate the robustness of our method under noisy labels. Even when the training data contains 50% noisy labels, our method can mitigate an average of 86.54% of the performance degradation compared to the standard training method. Furthermore, further ablation experiments and analysis validate the effectiveness of High Confidence Ranking for sample selection and the correctness of Label Ensembling for noisy label correction. Zhihua Fang, Liang He 0003, Lin Li 0032, Ying Hu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter ManipulationabstractExisting speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to mitigate speaker mismatch of domain mismatch. The SA consists of two sub-policies: (1) SA-Vocoder, which uses a vocoder to manipulate pitch and formants parameters of speakers. (2) SA-Spectrum, which directly performs pitch-shift and time-stretch on the spectrum of each speech signal. The SA is simple and effective. Experimental results show that using SA can significantly improve the generalization ability of models, especially for: 1) The training set with fewer speakers, e.g., WSJ0-2mix, or 2) The target test set with complex linguistic conditions, e.g., the TIMIT based test set. Moreover, as a data augmentation method, SA has good potential to be applicable to other speech related tasks. We validate this by applying SA in speech recognition, and experimental results show that the generalization ability is also improved. Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 4 |
| 2023 | A Joint Network Based on Interactive Attention for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) has played a vital role in human-machine interaction. In this paper, we propose a separate spectrum-based SER model and a joint network combining pre-trained and spectrum-based models. In the joint network, we design an interactive attention module to effectively fuse the intermediate features from two models. Our proposed separate spectrum-based model is superior to four compared spectrum-based methods under the speaker-dependent setting. For the application in real scenarios, we compared our proposed joint network with six methods utilizing the pre-trained model under the speaker-independent setting. Experimental results show that our proposed joint network achieves the best performance among four unimodal models on the unweighted accuracy (UA) of 73.32 % and weighted accuracy (WA) of 72.48 %, respectively. Ying Hu 0005, Shijing Hou, Hao Huang 0009, Liang He 0003 |
ICME | 1 |
| 2023 | Speech Topic Classification Based on Pre-trained and Graph NetworksabstractSpeech Topic Classification (STC) automatically classifies audio clips into predefined categories, which is widely used in short video, personalized recommendation and other fields. At present, the common system is composed of two parts: first, the speech is converted into text by automatic speech recognition (ASR), and then the text topic is classified by natural language processing (NLP). Most of them have problems such as error propagation and lack of global structure. So in this paper, we propose a new end-to-end framework based on a pre-trained model and graph network. The pre-trained model is used to extract the semantic features with sequential structure instead of acoustic features, and the combination with the global features of conversational context constructed by graph network has achieved good results on the Fisher dataset. Fangjing Niu, Ying Hu 0005, Hao Huang 0009, Liang He 0003 |
ICME | 3 |
| 2023 | CRA-DIFFUSE: Improved Cross-Domain Speech Enhancement Based on Diffusion Model with T-F Domain Pre-DenoisingabstractSpeech enhancement (SE) methods in both the Time-Frequency (T-F) domain and time-domain domains have their own advantages. Leveraging both T-F domain and time-domain (cross domain) inputs has shown to be successful in the speech enhancement task. Recent SE methods based on diffusion models have shown promising results. However, little research effort has been made in the cross-domain speech enhancement using a diffusion model. We propose CRA-DiffuSE, a cross-domain SE model that uses a diffusion-based enhancement model as a refinement module after initial enhancement to achieve better results. For pre-enhance stage, we design CRANet, a T-F domain enhancement model combining channel attention and spatial attention. For the post-enhance stage, we design DiffuNet, a conditional generation model based on Denoising Diffusion Implicit Model (DDIM) for speech enhancement. Experiments demonstrate that the proposed CRA-DiffuSE is significantly superior to the baselines. Zhibin Qiu, Yachao Guo, Mengfan Fu, Hao Huang 0009, Ying Hu 0005, Liang He 0003, Fuchun Sun 0001 |
ICME | 5 |
| 2023 | MTANet: Multi-band Time-frequency Attention Network for Singing Melody Extraction from Polyphonic Music
Ying Hu 0005, Liusong Wang, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 2 |
| 2022 | Mining Hard Samples Locally And Globally For Improved Speech SeparationabstractSpeech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently good, improving hard samples may contribute more to back-end tasks. In this paper, we propose two methods to alleviate data imbalance in speech separation task, based on local and global hard sample mining. For the local, we propose weighted loss to compensate for hard samples by increasing their weights in each batch. For the global, we perform global hard sample mining and re-sample to increase the proportion of hard samples in the training set. Because hard sample mining using objective loss in dynamic mixing leads to local results, we propose an indirect method using speaker-specific parameters, based on the fact that pitch median difference and x-vector cosine distance of two speakers in a mixture are closely correlated with separation SI-SNRi. Experimental results show that both methods decrease the percentage of hard samples in the test set than using dynamic mixing only while keeping the average SI-SNRi comparable, and the global method shows more promising results than the local one. Yizhou Peng, Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 4 |
| 2022 | A Graph Isomorphism Network with Weighted Multiple Aggregators for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) is an essential part of human-computer interaction. In this paper, we propose an SER network based on a Graph Isomorphism Network with Weighted Multiple Aggregators (WMA-GIN), which can effectively handle the problem of information confusion when neighbour nodes' features are aggregated together in GIN structure. Moreover, a Full-Adjacent (FA) layer is adopted for alleviating the over-squashing problem, which is existed in all Graph Neural Network (GNN) structures, including GIN. Furthermore, a multi-phase attention mechanism and multi-loss training strategy are employed to avoid missing the useful emotional information in the stacked WMA-GIN layers. We evaluated the performance of our proposed WMA-GIN on the popular IEMOCAP dataset. The experimental results show that WMA-GIN outperforms other GNN-based methods and is comparable to some advanced non-graph-based methods by achieving 72.48% of weighted accuracy (WA) and 67.72% of unweighted accuracy (UA). Ying Hu 0005, Yuwu Tang, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 1 |
| 2022 | A Multi-grained based Attention Network for Semi-supervised Sound Event DetectionabstractSound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. This paper presents a multi-grained based attention network (MGA-Net) for semi-supervised sound event detection. To obtain the feature representations related to sound events, a residual hybrid convolution (RH-Conv) block is designed to boost the vanilla convolution's ability to extract the time-frequency features. Moreover, a multi-grained attention (MGA) module is designed to learn temporal resolution features from coarse-level to fine-level. With the MGA module,the network could capture the characteristics of target events with short- or long-duration, resulting in more accurately determining the onset and offset of sound events. Furthermore, to effectively boost the performance of the Mean Teacher (MT) method, a spatial shift (SS) module as a data perturbation mechanism is introduced to increase the diversity of data. Experimental results show that the MGA-Net outperforms the published state-of-the-art competitors, achieving 53.27% and 56.96% event-based macro F1 (EB-F1) score, 0.709 and 0.739 polyphonic sound detection score (PSDS) on the validation and public set respectively. Ying Hu 0005, Xiujuan Zhu, Hao Huang 0009, Liang He 0003 |
INTERSPEECH | 1 |
| 2022 | How to Boost Anti-Spoofing with X-VectorsabstractWith the development of speech synthesis or voice conversion, speech spoofing countermeasures are increasingly required for protecting automatic speaker verification system. In our daily life, if we are familiar with the speaker, we tend to seek her/his traits in our memory to distinguish between bona fide and spoofed speech of her/him. Speaker label can not be directly used to guide the training of anti-spoofing network because it is difficult to obtain in real scenes. Motivated by this, we use x-vectors to represent speaker information and propose two novel methods by introducing x-vectors on the acoustic feature and embedding level into the two mainstream anti-spoofing methods (LightCNN and SeNet). An attention module is also added on the embedding level for further improvement. Experimental results on the ASVspoof 2019 logical access (LA) database show that the best EER and mintDCF in our methods are 0.98% and 0.0294, outperforming state-of-the-art single systems as far as we know. Shen Huang, Ji Gao, Ying Hu 0005, Liang He 0003 |
SLT | 5 |
| 2022 | Multi-stage music separation network with dual-branch attention and hybrid convolution
Yadong Chen 0003, Ying Hu 0005, Liang He 0003, Hao Huang 0009 |
J. Intell. Inf. Syst. | 2 |
| 2022 | A bimodal network based on Audio-Text-Interactional-Attention with ArcFace loss for speech emotion recognition
Yuwu Tang, Ying Hu 0005, Liang He 0003, Hao Huang 0009 |
Speech Commun. | 2 |
| 2022 | Hierarchic Temporal Convolutional Network With Cross-Domain Encoder for Music Source SeparationabstractRecently, the time-domain-based methods (i.e., the method of modeling the raw waveform directly) for audio source separation have shown tremendous potential. In this paper, we propose a model which combines the complexed spectrogram domain feature and time-domain feature by a cross-domain encoder (CDE) and adopts the hierarchic temporal convolutional network (HTCN) for multiple music sources separation. The CDE is designed to enable the network to code the interactive information of the time-domain and complexed spectrogram domain features. HTCN enables it to learn the long-time series dependence effectively. We also designed a feature calibration unit (FCU) to be applied in the HTCN and adopted the multi-stage training strategy during the training stage. The ablation study demonstrates the effectiveness of each designed component in the model. We conducted the experiments on the MUSDB18 dataset. The experimental results indicate that our proposed CDE-HTCN model outperforms the top-of-the-line methods and, compared with the state-of-the-art method, DEMUCS, achieves the improvement of the average SDR score of 0.61 dB. Significantly, the improvement of the SDR score for the$\ bass$source has a sizable margin of 0.91 dB. Ying Hu 0005, Yadong Chen 0003, Wenzhong Yang, Liang He 0003, Hao Huang 0009 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone ClassificationabstractWe pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms. We propose RNN based Encoder-Decoder framework with gating mechanism which underlying models both the state cost estimation and Viterbi back-tracing pass implemented in the RAPT algorithm. Then we apply the pitch extractor to a down-stream Mandarin tone classification task. The basic motivation is to combine together the two conventional components in tone classification (i.e., the pitch extractor and tone classifier) and then the whole network are trained simultaneously in an end-to-end fashion. Various cascade methods are evaluated. We carry out pitch extraction and tone classification experiments on Mandarin continuous speech database to show the superiority of the proposed models. Experimental results on pitch extraction show proposed pitch tracking model outperforms the DNN-RNN and bi-directional variants. Tone classification experimental results show the composite model outperforms the traditional cascade tone classification framework which makes use of pitch related feature and a back-end classifier. Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
ICASSP | 3 |
| 2021 | End-to-End Speech Separation Using Orthogonal Representation in Complex and Real Time-Frequency Domain
Hao Huang 0009, Ying Hu 0005, Sheng Li 0010 |
Interspeech | 3 |
| 2020 | A Lightweight Model Based on Separable Convolution for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Wushour Slamu |
INTERSPEECH | 2 |
| 2018 | Defect characterization of amorphous silicon thin film solar cell based on low frequency noise
Linna Hu, Liang He 0003, Xiaofei Jia, Ying Hu 0005, Hongmei Ma, Dandan Guo |
Sci. China Inf. Sci. | 5 |
| 2016 | Monaural Singing Voice Separation by Non-negative Matrix Partial Co-Factorization with Temporal Continuity and Sparsity Criteria
Ying Hu 0005, Hao Huang 0009 |
ICIC (3) | 1 |
| 2016 | Scene Text Detection Based on Text Probability and Pruning Algorithm
Ying Hu 0005 |
ICIC (3) | 4 |