EDBT 2026 Demo / reviewers in the wild / expert
Li Xiao 0007
dblp:14/5505-7
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-0287-4663ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FreeCodec: A Disentangled Neural Speech Codec with Fewer Tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang, Li Xiao 0007, Yuhong Yang 0001 |
INTERSPEECH | 6 |
| 2024 | SnoreOxiNet: Non-contact Diagnosis of Nocturnal Hypoxemia Using Cross-Domain Acoustic Features
Weiyan Yi, Xiuping Yang, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001 |
ICANN (8) | 3 |
| 2024 | Srcodec: Split-Residual Vector Quantization for Neural Speech CodecabstractEnd-to-end neural speech coding achieves state-of-the-art performance by using residual vector quantization. However, it is a challenge to quantize the latent variables with as few bits as possible. In this paper, we propose SRCodec, a neural speech codec that relies on a fully convolutional encoder/decoder network with specifically proposed split-residual vector quantization. In particular, it divides the latent representation into two parts with the same dimensions. We utilize two different quantizers to quantize the low-dimensional features and the residual between the low- and high-dimensional features. Meanwhile, we propose a dual attention module in split-residual vector quantization to improve information sharing along both dimensions. Both subjective and objective evaluations demonstrate that the effectiveness of our proposed method can achieve a higher quality of reconstructed speech at 0.95 kbps than Lyra-v1 at 3 kbps and Encodec at 3 kbps. Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu |
ICASSP | 3 |
| 2024 | SuperCodec: A Neural Speech Codec with Selective Back-Projection NetworkabstractNeural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction, especially at low bitrates. In this study, we introduce SuperCodec, a neural speech codec that achieves state-of-the-art performance at low bitrates. It employs a novel back projection method with selective feature fusion for augmented representation. Specifically, we propose to use Selective Up-sampling Back Projection (SUBP) and Selective Down-sampling Back Projection (SDBP) modules to replace the standard up- and down-sampling layers at the encoder and decoder, respectively. Experimental results show that our method outperforms the existing neural speech codecs operating at various bitrates. Specifically, our proposed method can achieve higher quality reconstructed speech at 1 kbps than Lyra V2 at 3.2 kbps and Encodec at 6 kbps. Youqiang Zheng, Weiping Tu, Li Xiao 0007, Xinmeng Xu |
ICASSP | 3 |
| 2024 | LungAdapter: Efficient Adapting Audio Spectrogram Transformer for Lung Sound Classification
Li Xiao 0007, Lucheng Fang, Yuhong Yang 0001, Weiping Tu |
INTERSPEECH | 1 |
| 2024 | SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
Xiuping Yang, Li Xiao 0007, Weiyan Yi, Yuhong Yang 0001, Weiping Tu |
INTERSPEECH | 3 |
| 2023 | MBMS-GAN: Multi-Band Multi-Scale Adversarial Learning for Enhancement of Coded Speech at Very Low Rate
Weiping Tu, Yong Luo 0002, Xin Zhou 0003, Li Xiao 0007, Youqiang Zheng |
ICANN (7) | 5 |
| 2023 | Freevc: Towards High-Quality Text-Free One-Shot Voice ConversionabstractVoice conversion (VC) can be achieved by first extracting source content information and target speaker information, and then reconstructing waveform with these information. However, current approaches normally either extract dirty content information with speaker information leaked in, or demand a large amount of annotated data for training. Besides, the quality of reconstructed waveform can be degraded by the mismatch between conversion model and vocoder. In this paper, we adopt the end-to-end framework of VITS for high-quality waveform reconstruction, and propose strategies for clean content information extraction without text annotation. We disentangle content information by imposing an information bottleneck to WavLM features, and propose the spectrogram-resize based data augmentation to improve the purity of extracted content information. Experimental results show that the proposed method outperforms the latest VC models trained with annotated data and has greater robustness. Weiping Tu, Li Xiao 0007 |
ICASSP | 3 |
| 2023 | Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with TransformerabstractWe propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single microphone recording and sent to the far end, and consists of two key components. First, we apply a pyramid encoder, which adopts the coarse-to-fine structure, to extract the latent correlations between double-end signals and to fuse them in a multiscale manner. Second, we propose a mix-transformer, a combination of local and global attention in a parallel way, to leverage local and global speech information for separation. Experimental results show that the proposed method outperforms recent AEC methods in terms of objective evaluation metrics. In addition, exploring the correlation between speech local and global features by using the mix-transformer significantly improves the system performance and shows more robustness than the conventional transformer. Xinmeng Xu, Weiping Tu, Yuhong Yang 0001, Li Xiao 0007 |
ICASSP | 5 |
| 2023 | ONEI: Unveiling Route and Phase of Breathing from Snoring Sounds
Baoai Han, Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren |
ICONIP (9) | 3 |
| 2023 | A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren |
INTERSPEECH | 1 |
| 2023 | CQNV: A Combination of Coarsely Quantized Bitstream and Neural Vocoder for Low Rate Speech Coding
Youqiang Zheng, Li Xiao 0007, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu |
INTERSPEECH | 2 |