VLDB 2026 Research / reviewers in the wild / expert
Tingwei Guo
dblp:246/4989
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language ModelsabstractEnd-to-end Large Speech Language Models (LSLMs) have demonstrated impressive conversational generation abilities, yet consistently fall short of traditional pipeline systems on semantic understanding benchmarks.In this work, we reveal through systematic experimentation that although LSLMs lose some text input performance after speech-text alignment training, the performance gap between speech and text inputs is more pronounced, which we refer to as the modality gap.To understand this gap, we analyze both coarse-and fine-grained text and speech representations.At the coarse-grained level, representations of speech and text in deeper layers are found to be increasingly aligned in direction (cosine similarity), while concurrently diverging in magnitude (Euclidean distance).We further find that representation similarity is strongly correlated with the modality gap.At the fine-grained level, a spontaneous token-level alignment pattern between text and speech representations is observed.Based on this, we introduce the Alignment Path Score to quantify token-level alignment quality, which exhibits stronger correlation with the modality gap.Building on these insights, we design targeted interventions on critical tokens through angle projection and length normalization.These strategies demonstrate the potential to improve correctness for speech inputs.Our study provides the first systematic empirical analysis of the modality gap and alignment mechanisms in LSLMs, offering both theoretical and methodological guidance for future optimization. Bajian Xiang, Shuaijiang Zhao, Tingwei Guo |
EMNLP | 3 |
| 2023 | VISinger2: High-Fidelity End-to-End Singing Voice Synthesis Enhanced by Digital Signal Processing Synthesizer
Yongmao Zhang, Heyang Xue, Hanzhao Li, Lei Xie 0001, Tingwei Guo, Ruixiong Zhang, Caixia Gong |
INTERSPEECH | 5 |
| 2022 | Time Domain Adversarial Voice Conversion for ADD 2022abstractIn this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language content into target speaker’s fake speech. Then the converted speech generated from VC is post-processed in time-domain to improve the deception ability. The experimental results show that our system has adversarial ability against anti-spoofing detectors with a little compromise in audio quality and speaker similarity. This system ranks top in Track 3.1 in the ADD 2022, showing that our method could also gain good generalization ability against different detectors. Cheng Wen 0004, Tingwei Guo, Xingjun Tan, Shuran Zhou, Chuandong Xie, Xiangang Li |
ICASSP | 2 |
| 2022 | Audio-Visual Wake Word Spotting System for MISP Challenge 2021abstractThis paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the environmental robustness of far-field wake word spotting. In the proposed system, firstly, we take advantage of speech enhancement algorithms such as beamforming and weighted prediction error (WPE) to address the multi-microphone conversational audio. Secondly, several data augmentation techniques are applied to simulate a more realistic far-field scenario. For the video information, the provided region of interest (ROI) is used to obtain visual representation. Then the multi-layer CNN is proposed to learn audio and visual representations, and these representations are fed into our two-branch attention-based net-work which can be employed for fusion, such as transformer and conformer. The focal loss is used to fine-tune the model and improve the performance significantly. Finally, multiple trained models are integrated by casting vote to achieve our final 0.091 score. Yanguang Xu, Shuaijiang Zhao, Chaoyang Mei, Tingwei Guo, Shuran Zhou, Chuandong Xie, Xiangang Li |
ICASSP | 6 |
| 2022 | Audio Deepfake Detection System with Neural Stitching for ADD 2022abstractThis paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge[1]. The very same system was used for both two rounds of evaluation in Track 3.2 with similar training methodology. The first round of Track 3.2 data is generated from Text-to-Speech(TTS) or voice conversion (VC) algorithms, while the second round of data consists of generated fake audio from other participants in Track 3.1, aming to spoof our systems. Our systems uses a standard 34-layer ResNet [2], with multi-head attention pooling [3] to learn the discriminative embedding for fake audio and spoof detection. We further utilize neural stitching to boost the model’s generalization capability in order to perform equally well in different tasks, and more details will be explained in the following sessions. The experiments show that our proposed method outperforms all other systems with 10.1% equal error rate(EER) in Track 3.2. Cheng Wen 0004, Shuran Zhou, Tingwei Guo, Xiangang Li |
ICASSP | 4 |
| 2022 | Improving GAN-based vocoder for fast and high-quality speech synthesis
Mengnan He, Tingwei Guo, Zhenxing Lu, Ruixiong Zhang, Caixia Gong |
INTERSPEECH | 2 |
| 2021 | Didispeech: A Large Scale Mandarin Speech CorpusabstractThis paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech processing tasks, such as voice conversion, multi-speaker text-to-speech and automatic speech recognition. We conduct experiments with multiple speech tasks and evaluate the performance, showing that it is promising to use the corpus for both academic research and practical application. The corpus is available at https://outreach.didichuxing.com/research/opendata/. Tingwei Guo, Cheng Wen 0004, Dongwei Jiang, Ne Luo, Ruixiong Zhang, Shuaijiang Zhao, Wubo Li, Xiangang Li |
ICASSP | 1 |