EDBT 2026 Demo / reviewers in the wild / expert
Yulin Wu 0003
dblp:49/9858-3
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-1940-2141ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WavGateMamba: A Frequency-Enhanced and Gated Mamba Model for Multimodal Depression Detection
Haiyang Ye, Dengshi Li, Yulin Wu 0003 |
MMM (1) | 4 |
| 2026 | CMDiff: Clip-guided multi-dimension mamba diffusion model for low light image enhancement
Chengbo Yu, Dengshi Li, Yulin Wu 0003, Aolei Chen |
Image Vis. Comput. | 4 |
| 2025 | Multimodal and multichannel speech separation using location-guided speech feature mapping network
Yulin Wu 0003, Xiaochen Wang 0001, Dengshi Li, Ruimin Hu |
Neurocomputing | 1 |
| 2024 | Acoustic scene classification: A comprehensive survey
Biyun Ding, Tao Zhang 0025, Chao Wang 0135, Ganjun Liu, Jinhua Liang, Ruimin Hu, Yulin Wu 0003, Difei Guo |
Expert Syst. Appl. | 7 |
| 2024 | Adaptive subband partition encoding scheme for multiple audio objects using CNN and residual dense blocks mixture network
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001 |
Expert Syst. Appl. | 1 |
| 2023 | Multi-speaker Direction of Arrival Estimation Using Audio and Visual Modalities with Convolutional Neural NetworkabstractIn reality, audible and visible sound sources are closely aligned, and they can help humans locate sources exactly. To exploit the complementarity between audio and visual data in multi-speaker direction of arrival (DoA) estimation, we propose a novel network consisting of 3D convolution neural networks (3D-CNNs) and 2D-CNNs mixture networks with residual dense blocks. It has two main advantages: 1) both input audio and visual features are low-level signal representation: the real and imaginary parts of STFT coefficients for the audio feature and pixel coordinates for the visual feature, which can allow the network to learn to extract the most informative high-level features. 2) 3D-CNNs with the residual dense block are used for audio and visual feature mapping along the time and frequency axis. The following 2D-CNNs are to ensemble the high-level features along the DoA axis. Experimental results demonstrate promising SSL performance. Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001 |
ICME | 1 |
| 2023 | Perceptual Audio Object Coding Using Adaptive Subband Grouping with CNN and Residual BlockabstractSpatial audio content is becoming increasingly popular and is regarded as a set of object signals with associated metadata. The object-based content representation is independent of loudspeaker layouts and provides high spatial resolution when reproduced on more loudspeakers. The audio quality of the traditional spatial audio object coding (SAOC) method has severe aliasing distortion, which impairs the immersive listening experience. In this study, we reduce aliasing distortion by perceptual adaptive subband grouping strategy and use the convolutional neural network (CNN) and residual block to build the side information compressing model. Both objective and subjective experiments on benchmark datasets with different bitrates show that the proposed method achieves favorable performance against state-of-the-art methods. Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001 |
ICME | 1 |
| 2023 | Multi-speaker DoA Estimation Using Audio and Visual Modality
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke |
Neural Process. Lett. | 1 |
| 2022 | High Parameter Frequency Resolution Encoding Scheme for Spatial Audio Objects Using Stacked Sparse Autoencoder
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu, Shanfa Ke |
Neural Process. Lett. | 1 |
| 2021 | Spatial Audio Object Coding Based on Time-Frequency Shifting and SchedulingabstractSpatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. It divides the full frequency band into 28 subbands and extracts spatial parameters for de-coding. In this way, objects can be encoded into a downmix signal with a few parameters. However, using the same parameters in one subband will cause frequency aliasing distortion, which severely impacts the listening experience. Existing studies to enhance SAOC cannot eliminate the aliasing distortion of all objects effectively. This paper describes a new structure to balance the bit-rate and decode quality based on time-frequency (TF) shifting and scheduling. In this structure, a TF shifting strategy (contains global shifting and local shifting) is proposed to reduce frequency aliasing distortion. Furthermore, a scheduling strategy is used to decide which part should be shifted according to the degree of aliasing. From the experiment results, the performance of the proposed method is better than SAOC and other enhanced methods. Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003 |
ICME | 4 |
| 2021 | Efficient Multi-Step Audio Object Coding with Limited Residual InformationabstractSpatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. Audio systems can provide personalized services under this framework. However, this method causes frequency aliasing distortion, which severely impacts the listening experience. The multi-step SAOC (MS-SAOC) scheme was proposed to enhance the sound quality of each audio object by using residual information. Compared with SAOC, the bit-rate increases three times due to the residual data of multiple objects. In this paper, an efficient multi-step residual coding method is proposed to reduce the residual bit-rate of MS-SAOC. A two-level filter is designed to remove redundant residual information, and the limited residual information can efficiently compensate for frequency aliasing distortion. From experiment results, the residual bit-rate is half of MS-SAOC, and the sound quality is maintained at the Good-Excellent level. Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003, Wenke Liu |
ICME | 4 |
| 2021 | Low Bitrates Audio Object Coding Using Convolutional Auto-Encoder and Densenet Mixture ModelabstractThe efficient transmission of the audio objects can be achieved by spatial audio object coding (SAOC) method that conveys a mono downmix signal together with side information parameters that enable object reconstruction in the decoder. To allow the transmission of audio objects at low bitrates, we present a new audio coding method with convolutional auto-encoder (CAE) and dense convolutional network (DenseNet) mixture model, optimizing the compression of side information parameters of audio objects. It has two main advantages: 1) Different from the linear transform methods, CAE can dig the nonlinear relationship of side information parameters and can effectively reduce the dimension of side information parameters; 2) DenseNet is adding in the decoder to make full use of low dimensional features of side information parameters, which improves the audio quality at low bitrate. Experiments show that our method outperforms base-line methods permitting bitrates as low as 1 kbps per object. Yulin Wu 0003, Ruimin Hu, Chenhao Hu, Shanfa Ke, Xiaochen Wang 0001 |
ICME | 1 |
| 2021 | Stacked Sparse Autoencoder for Audio Object Coding
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu |
MMM (1) | 1 |
| 2021 | Audio object coding based on N-step residual compensating
Chenhao Hu, Xiaochen Wang 0001, Ruimin Hu, Yulin Wu 0003 |
Multim. Tools Appl. | 4 |