VLDB 2026 Research / reviewers in the wild / expert
Dingding Yao
dblp:195/9235
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-9610-8782ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Teacher-Student Interactive Cycle: Joint Optimization With Inner-Loop Self-Distillation in Prompted Foundation Models for Efficient Semantic SegmentationabstractIn the field of semantic segmentation, the high computational cost of deep models poses a major barrier to deployment on edge devices. Among various efficiency-oriented methods, knowledge distillation has emerged as a promising technique for transferring knowledge from large models to lightweight networks. However, current knowledge distillation methods for efficient semantic segmentation still face two key challenges: (1) they often rely on large offline pre-trained teacher networks that remain fixed during training, and (2) they lack joint optimization mechanisms that enable effective teacher-student interaction in pixel-wise dense prediction. As a result, mutual learning strategies originally designed for image-level classification often fail to capture the fine-grained consistency required for semantic segmentation. To address these two challenges, we propose a novel training framework termed Teacher-Student Interactive Cycle (TSIC), which performs efficient semantic segmentation. Specifically, TSIC integrates a lightweight student network into a prompt-based foundation model as a prompted segmentor to assist an online-trained teacher. The student provides coarse mask prompts to guide the teacher, while the teacher offers fine-grained supervision through posterior probabilities and intermediate feature maps. This loop enables joint online optimization without relying on offline pre-trained teachers and fosters effective bidirectional communication. Extensive experiments conducted on several benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, demonstrate the effectiveness of TSIC. Compared to previous methods, TSIC achieves superior segmentation mIoU in most scenarios. Our code will be made publicly available at https://github.com/CV-ShuchangLyu/TSIC. Qi Zhao 0037, Shuchang Lyu, Longhao Zou, Dingding Yao, Chenguang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Hybrid dual-path network: Singing voice separation in the waveform domain by combining Conformer and Transformer architectures
Chunxi Wang, Mao-shen Jia, Meiran Li, Dingding Yao |
Speech Commun. | 5 |
| 2025 | GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech EnhancementabstractSpeech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion models have gained lots of attention in speech enhancement area, achieving competitive results. Current diffusion-based methods blur the distribution of the signal with isotropic Gaussian noise and recover clean speech distribution from the prior. However, these methods often suffer from a substantial computational burden. We argue that the computational inefficiency partially stems from the oversight that speech enhancement is not purely a generative task; it primarily involves noise reduction and completion of missing information, while the clean clues in the original mixture do not need to be regenerated. In this paper, we propose a method that introduces noise with anisotropic guidance during the diffusion process, allowing the neural network to preserve clean clues within noisy recordings. This approach substantially reduces computational complexity while exhibiting robustness against various forms of noise and speech distortion. Experiments demonstrate that the proposed method achieves state-of-the-art results with only approximately 4.5 million parameters, a number significantly lower than that required by other diffusion methods. This effectively narrows the model size disparity between diffusion-based and predictive speech enhancement approaches. Additionally, the proposed method performs well in very noisy scenarios, demonstrating its potential for applications in highly challenging environments. Chengzhong Wang, Jianjun Gu 0005, Dingding Yao, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | A novel semi-blind source separation framework towards maximum signal-to-interference ratio
Jianjun Gu 0005, Dingding Yao, Yonghong Yan 0002 |
Signal Process. | 2 |
| 2023 | Exploring Auditory Attention Decoding using Speaker Features
Zelin Qiu, Jianjun Gu 0005, Dingding Yao |
INTERSPEECH | 3 |
| 2023 | The effect of source sparsity on independent vector analysis for blind source separation
Jianjun Gu 0005, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002 |
Signal Process. | 3 |
| 2023 | Multi-Source Localization Using Optimized Time-Frequency Representation and Sparsity Component AnalysisabstractThis paper aims to address the multi-source localization problem by exploiting the sparsity of the speech signal in the time-frequency domain, where the challenge mainly lies in extracting the sparse component. An optimized time-frequency representation and sparsity component analysis-based multi-source localization method is proposed to overcome this challenge. Firstly, extracting the sparse components relies on the accurate representation in the time-frequency domain. However, the energy leakage problem caused by linear time-frequency transformation limits the accuracy of sparse component extraction. To tackle this problem, inspired by empirical mode decomposition, the proposed method classifies all the points in the time-frequency domain into four categories based on their phase feature and mode characteristics. Each type of the point is modeled separately, and a point-by-point analysis is conducted to remove all the points affected by energy leakage. Then, based on the optimized time-frequency representation, the phase coherence criterion is used to detect the sparse component in the point level. Following that, guided by the mode consistency characteristic of sparse components, an extension scheme is proposed to recover the falsely removed sparse components. Finally, the detected sparse components are applied for the multiple source localization. The objective evaluation is performed in both simulation and actual recording environments, and the proposed method can achieve better localization accuracy compared to several existing methods. Mao-shen Jia, Dingding Yao, Jing Wang 0037 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | A Secondary Path-Decoupled Active Noise Control Algorithm Based on Deep LearningabstractActive noise control (ANC) systems are widely used to cancel unwanted noise. However, for high-level noise, the residual error signal cannot be fully eliminated because of the nonlinearity of the secondary path, resulting in the diverging of the adaptive filter. In this letter, we propose a secondary path-decoupled ANC (SPD-ANC) algorithm based on deep learning. Specifically, the secondary path decoupled module consisting of two time-domain convolutional recurrent networks, one for modeling the nonlinear secondary path and the other for modeling the reverse process, is employed to calculate the secondary path-decoupled (SPD) error signal. The control signal is then generated by an adaptive filter that is optimized towards minimizing the SPD error signal. Simulation results indicate that the proposed method outperforms the conventional ANC methods under different conditions. Daocheng Chen, Longbiao Cheng, Dingding Yao, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Estimation Reliability Function Assisted Sound Source Localization With Enhanced Steering Vector Phase DifferenceabstractThe performance of the traditional direction-of-arrival (DOA) estimation algorithms greatly degrades in noisy and reverberant environments. Recently, deep learning has been applied to sound source localization and provided the substantial improvement in robustness for DOA estimation. In this paper, we propose a sound source localization approach using the deep learning-based steering vector phase difference enhancement. The steering vectors and their estimation reliability functions (ERFs) are first estimated under the guidance of the time-frequency masks that are predicted using deep neural network (DNN). The phase difference of the steering vectors is further enhanced with a second DNN model, which is trained with the ERF-weighted mean square error (MSE) loss. The DOA of the sound source is finally determined by the ERF-weighted histogram analysis. Experimental results with various types and levels of noise and various reverberant conditions show that the proposed approach outperforms the state-of-the-art sound source localization algorithms in utterance and frame-level DOA estimation. Longbiao Cheng, Xingwei Sun, Dingding Yao, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | A Subband Energy Modification Method for Elevation Control in Median PlaneabstractElevation perception is crucial for binaural reproduction. A recent study proposed an elevation control method by modifying the energy of HRTFs in each auditory scale subband, such as the ERB and Mel subband. However, this subband division is designed based on auditory excitation patterns and may not be consistent with the elevation localization cues. To this end, this study proposes a novel subband division strategy which emphasizes the physiological information involved in elevation localization based on a statistical analysis of the HRTF. Then, the elevation controlled HRTFs are constructed by modifying the energy of the HRTF magnitudes in each subband. Results of the listening test demonstrate that our method with the proposed subband division strategy outperforms the method with ERB scale subdivision in terms of the accuracy for controlling the perceived elevation of sound image. Dingding Yao, Huaxing Xu, Risheng Xia, Yonghong Yan 0002 |
ICASSP | 1 |