EDBT 2026 Demo / reviewers in the wild / expert
Rilin Chen
dblp:285/0216
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0004-6873-1683ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveabstractExisting neural vocoders have demonstrated promising performance by leveraging Mel-spectrum as an acoustic feature for conditional audio generation. Nonetheless, they remain constrained by an inherent ``performance-cost'' dilemma that significantly hinders the development of this field. This paper revisits this foundational task from a novel degradation perspective, where Mel-spectrum is regarded as a special signal degradation process from the target spectrum. Drawing inspiration from traditional sparse signal recovery problems, we propose DegVoC, a GAN-based neural vocoder with a two-step solution procedure. First, by exploiting degradation priors, we attempt to retrieve the initial spectral structure from Mel-domain representations as an initial solution via a simple linear transformation. Based on that, we introduce a deep prior solver that accounts for the heterogeneous distribution of sub-bands in the time-frequency domain. A convolution-style attention module with a large kernel size is specially devised for efficient inter-frame and inter-band contextual modeling. With 3.89 M parameters and substantially reduced inference complexity, DegVoC achieves state-of-the-art performance across objective and subjective evaluations, outperforming existing GAN-, DDPM- and flow-matching-based baselines. Andong Li, Lingling Dai, Rilin Chen, Meng Yu 0003, Xiaodong Li 0002, Dong Yu 0001, Chengshi Zheng |
AAAI | 5 |
| 2025 | STA-V2A: Video-to-Audio Generation with Semantic and Temporal AlignmentabstractVisual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A. Yong Ren 0006, Chenxing Li, Manjie Xu, Rilin Chen, Dong Yu 0001 |
ICASSP | 6 |
| 2025 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisabstractIn this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released. Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001 |
ICASSP | 6 |
| 2025 | BridgeVoC: Neural Vocoder with Schrödinger BridgeabstractWhile previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restoration task, this paper proposes a time-frequency (T-F) domain-based neural vocoder with the Schrödinger Bridge, called BridgeVoC, which is the first to follow the data-to-data generation paradigm. Specifically, the mel-spectrogram can be projected into the target linear-scale domain and regarded as a degraded spectral representation with a deficient rank distribution. Based on this, the Schrödinger Bridge is leveraged to establish a connection between the degraded and target data distributions. During the inference stage, starting from the degraded representation, the target spectrum can be gradually restored rather than generated from a Gaussian noise process. Quantitative experiments on LJSpeech and LibriTTS show that BridgeVoC achieves faster inference and surpasses existing diffusion-based vocoder baselines, while also matching or exceeding non-diffusion state-of-the-art methods across evaluation metrics. Rilin Chen, Meng Yu 0003, Chengshi Zheng, Dong Yu 0001, Andong Li |
IJCAI | 3 |
| 2025 | Learning Neural Vocoder from Range-Null Space DecompositionabstractDespite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned challenges. To be specific, we bridge the connection between the classical signal range-null decomposition (RND) theory and vocoder task, and the reconstruction of target spectrogram can be decomposed into the superimposition between the range-space and null-space, where the former is enabled by a linear domain shift from the original mel-scale domain to the target linear-scale domain, and the latter is instantiated via a learnable network for further spectral detail generation. Accordingly, we propose a novel dual-path framework, where the spectrum is hierarchically encoded/decoded, and the cross- and narrow-band modules are elaborately devised for efficient sub-band and sequential modeling. Comprehensive experiments are conducted on the LJSpeech and LibriTTS benchmarks. Quantitative and qualitative results show that while enjoying lightweight network parameters, the proposed approach yields state-of-the-art performance among existing advanced methods. Our code and the pretrained model weights are available at https://github.com/Andong-Li-speech/RNDVoC. Andong Li, Zhihang Sun, Rilin Chen, Erwei Yin, Xiaodong Li 0002, Chengshi Zheng |
IJCAI | 4 |
| 2025 | SMRU-Lite: Efficient Low-Complexity Speech Enhancement Model with Uncertainty EstimationabstractAlthough neural network-based speech enhancement models perform much better than their traditional counterparts, their substantial computational demands make it challenging for real-time applications on edge devices. Moreover, compact models often exhibit weak generalization in complex and out-of-domain scenarios. In this paper, we propose an efficient model based on our previous work Split-and-Merge Recurrent-based UNet (SMRU). The proposed model achieves a significant reduction in computational load through the incorporation of Skip-RNN layers and an attention-based sub-band compression module. Moreover, the employment of a two-stage uncertainty-driven loss function for aleatoric uncertainty capture leads to enhanced generalization and denoising performance without increasing the computational complexity during inference. Experimental results demonstrate that our model not only surpasses the original SMRU but also outperforms recently proposed lightweight models with similar computational cost (approximately 200M MACs). Furthermore, our model exhibits strong generalization in cross-corpus test sets, making it a promising solution for real-time speech enhancement applications. Zhihang Sun, Feiran Yang 0001, Rilin Chen, Chunguo Li |
IJCNN | 5 |
| 2025 | Video-to-Audio Generation with Fine-grained Temporal Semantics
Chenxing Li, Rilin Chen, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2025 | Scaling beyond Denoising: Submitted System and Findings in URGENT Challenge 2025
Zhihang Sun, Andong Li, Rilin Chen, Meng Yu 0003, Chengshi Zheng, Yi Zhou 0014, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2025 | From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language ModelsabstractThis paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications. Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu 0003, Xinyu Yang 0001, Dong Yu 0001 |
ACM Multimedia | 2 |
| 2024 | Opine: Leveraging a Optimization-Inspired Deep Unfolding Method for Multi-Channel Speech EnhancementabstractProximal gradient theory has demonstrated its superiority in the compressive sensing field for complex signal recovery. As an early trial in the speech front-end field, we propose OPINE, an optimization-inspired deep unfolding framework to simulate traditional iterative optimization process for multi-channel speech enhancement. Specifically, we formulate the joint optimization of beamforming weights and target speech using the Bayesian maximum a posteriori (MAP) criterion. By splitting and introducing the proximal gradient descent method, the original problem can be formulated into the alternating target solving of two sub-problems. Furthermore, we propose to formulate the proximal function into a more generalized NN-based modules, enabling the end-to-end learning from massive training data. The experiments are conducted on the spatialized LibriSpeech dataset, and quantitative results show that the proposed method can achieve comparable performance over existing advanced baselines. Andong Li, Rilin Chen, Chao Weng |
ICASSP | 2 |
| 2024 | LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
Shihao Chen, Jie Zhang 0042, Rilin Chen, Li-Rong Dai 0001 |
INTERSPEECH | 5 |
| 2024 | SMRU: Split-And-Merge Recurrent-Based UNet For Acoustic Echo Cancellation And Noise SuppressionabstractThe proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down-/upsampling layer with varying compression ratios and the dualpath structure for inter- and intra-band modeling. The model is configured from $50 \mathrm{M} / \mathrm{s}$ to $6.8 \mathrm{G} / \mathrm{s}$ in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements. Zhihang Sun, Andong Li, Rilin Chen, Hao Zhang 0112, Meng Yu 0003, Yi Zhou 0014, Dong Yu 0001 |
SLT | 3 |