VLDB 2026 Research / reviewers in the wild / expert
Lingling Dai
dblp:384/4993
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksabstractIn the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our well-chosen combination of different loss functions further optimizes the overall model capability. Lingling Dai, Andong Li, Yifan Liang, Xiaodong Li 0002, Chengshi Zheng |
AAAI | 1 |
| 2026 | DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveabstractExisting neural vocoders have demonstrated promising performance by leveraging Mel-spectrum as an acoustic feature for conditional audio generation. Nonetheless, they remain constrained by an inherent ``performance-cost'' dilemma that significantly hinders the development of this field. This paper revisits this foundational task from a novel degradation perspective, where Mel-spectrum is regarded as a special signal degradation process from the target spectrum. Drawing inspiration from traditional sparse signal recovery problems, we propose DegVoC, a GAN-based neural vocoder with a two-step solution procedure. First, by exploiting degradation priors, we attempt to retrieve the initial spectral structure from Mel-domain representations as an initial solution via a simple linear transformation. Based on that, we introduce a deep prior solver that accounts for the heterogeneous distribution of sub-bands in the time-frequency domain. A convolution-style attention module with a large kernel size is specially devised for efficient inter-frame and inter-band contextual modeling. With 3.89 M parameters and substantially reduced inference complexity, DegVoC achieves state-of-the-art performance across objective and subjective evaluations, outperforming existing GAN-, DDPM- and flow-matching-based baselines. Andong Li, Lingling Dai, Rilin Chen, Meng Yu 0003, Xiaodong Li 0002, Dong Yu 0001, Chengshi Zheng |
AAAI | 3 |
| 2026 | SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisabstractAlthough lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations. Yifan Liang, Andong Li, Guochen Yu, Fangkun Liu, Lingling Dai, Xiaodong Li 0002, Chengshi Zheng |
AAAI | 6 |
| 2026 | Rethinking the Joint Estimation of Magnitude and Phase for Time-Frequency Domain Neural VocodersabstractTime-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders. Lingling Dai, Andong Li, Xiaodong Li 0002, Chengshi Zheng |
IEEE Signal Process. Lett. | 1 |
| 2025 | Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing AidsabstractWide dynamic range compression (WDRC) provides level-dependent amplification, intended to make the output of a hearing aid fall between the hearing threshold and the highest comfortable level of the listener. Hearing aids often combine noise reduction with WDRC, applied sequentially. Unfortunately, this can result in across-source modulation, especially when fast-acting compression is used. However, fast-acting compression is theoretically preferable to compensate for the loss of compression in the cochlea. To apply fast-acting compression to speech while avoiding across-source modulation, we propose a deep-learning-based method that integrates noise reduction and source-independent WDRC in a two-stage low-complexity framework. The method applies fast-acting compression to speech and slow-acting compression to noise by an amount depending on the audiogram, with a controllable residual noise level. Objective measurements using simulated hearing-impaired listeners showed that the proposed method reduced negative interactions of speech and noise in highly non-stationary noise scenarios and generalized well across various hearing losses. Huiyong Zhang, Brian C. J. Moore, Lingling Dai, Fengyuan Hao, Xiaodong Li 0002, Chengshi Zheng |
ICASSP | 3 |
| 2025 | BAPEN: Towards Versatile Audio Phase Retrieval
Lingling Dai, Andong Li, Chengshi Zheng, Xiaodong Li 0002 |
ACM Multimedia | 1 |