Tingle Li

dblp:248/9136 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-1654-8030ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SPACE: Speaker Adaptation for Acoustic Eavesdropping Using mmWave Radio Signals
abstract
The prevalence of voice-related interaction and communication has raised concerns about privacy leakage and security. For example, millimeter-wave (mmWave) radio signals have been exploited as a potential attacker for acoustic eavesdropping. However, speaker variability and low-quality input pose significant challenges for the practical deployment of mmWave-based eavesdropping. In this paper, we proposeSPACE, an acoustic eavesdropping system to recover intelligible speech from low-quality mmWave signals, which can adapt to numerous different speakers and unseen ones.SPACEis a two-stage system that first reconstructs the spectrogram using a novelRadio TransUNetand then synthesizes the waveform through a neural vocoder. Specifically, to alleviate the negative effect of speaker variability, we introduce a speaker encoder to capture speaker features and a fusion network to condition the spectrogram reconstruction based on the extracted speaker characteristics. Further, to facilitate intelligible speech recovery from low-quality input, we design a Frequency Transformation Layer to exploit the correlation among all frequency harmonics and incorporate the neural vocoder to synthesize the speech waveform from the reconstructed spectrogram without using the contaminated phase. The experimental results show thatSPACEoutperforms existing mmWave-based approaches in scenarios with numerous different speakers and unseen speakers.
Running Zhao, Jiang-Tao Yu 0001, Tingle Li, Zhihan Jiang 0001, Chenshu Wu, Hang Zhao 0021, Edith C. H. Ngai
IEEE Trans. Mob. Comput.3
2025 Full-Duplex-Bench: A Benchmark to Evaluate Full-Duplex Spoken Dialogue Models on Turn-taking Capabilities
abstract
Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode—processing one turn at a time—emerging full-duplex SDMs can listen and speak simultaneously, enabling more natural conversations. However, current evaluations remain limited, focusing mainly on turn-based metrics or coarse corpus-level analyses. To address this, we introduce Full-Duplex-Bench, a benchmark that systematically evaluates key interactive behaviors: pause handling, backchanneling, turn-taking, and interruption management. Our framework uses automatic metrics for consistent, reproducible assessment and provides a fair, fast evaluation setup. By releasing our benchmark and code, we aim to advance spoken dialogue modeling and foster the development of more natural and engaging SDMs.
Guan-Ting Lin, Jiachen Lian, Tingle Li, Gopala Krishna Anumanchipalli, Alexander H. Liu, Hung-yi Lee
ASRU3
2025 EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
abstract
Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still lacking. To address this, we introduce EMO-Reasoning, a benchmark for assessing emotional coherence in dialogue systems. It leverages a curated dataset generated via text-to-speech to simulate diverse emotional states, overcoming the scarcity of emotional speech data. We further propose the Cross-turn Emotion Reasoning Score to assess the emotion transitions in multi-turn dialogues. Evaluating seven dialogue systems through continuous, categorical, and perceptual metrics, we show that our framework effectively detects emotional inconsistencies, providing insights for improving current dialogue systems. By releasing a systematic evaluation benchmark, we aim to advance emotion-aware spoken dialogue modeling toward more natural and adaptive interactions.
Kan Jen Cheng, Jiachen Lian, Akshay Anand, Faith Qiao, Robert Netzorg, Huang-Cheng Chou, Tingle Li, Guan-Ting Lin, Gopala Krishna Anumanchipalli
ASRU9
2025 Audio Texture Manipulation by Exemplar-Based Analogy
abstract
Audio texture manipulation involves modifying the perceptual characteristics of a sound to achieve specific transformations, such as adding, removing, or replacing auditory elements. In this paper, we propose an exemplar-based analogy model for audio texture manipulation. Instead of conditioning on text-based instructions, our method uses paired speech examples, where one clip represents the original sound and another illustrates the desired transformation. The model learns to apply the same transformation to new input, allowing for the manipulation of sound textures. We construct a quadruplet dataset representing various editing tasks, and train a latent diffusion model in a self-supervised manner. We show through quantitative evaluations and perceptual studies that our model outperforms text-conditioned baselines and generalizes to real-world, out-of-distribution, and non-speech scenarios.
Kan Jen Cheng, Tingle Li, Gopala Krishna Anumanchipalli
ICASSP2
2025 Sounding that Object: Interactive Object-Aware Image to Audio Generation
abstract
Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an interactive object-aware audio generation model that grounds sound generation in user-selected visual objects within images. Our method integrates object-centric learning into a conditional latent diffusion model, which learns to associate image regions with their corresponding sounds through multi-modal attention. At test time, our model employs image segmentation to allow users to interactively generate sounds at the object level. We theoretically validate that our attention mechanism functionally approximates test-time segmentation masks, ensuring the generated audio aligns with selected objects. Quantitative and qualitative evaluations show that our model outperforms baselines, achieving better alignment between objects and their associated sounds.
Tingle Li, Baihe Huang, Xiaobin Zhuang, Dongya Jia, Yuping Wang 0005, Zhuo Chen 0006, Gopala Krishna Anumanchipalli, Yuxuan Wang 0002
ICML1
2024 Self-Supervised Audio-Visual Soundscape Stylization
Tingle Li, Renhao Wang, Po-Yao Huang 0001, Andrew Owens, Gopala Krishna Anumanchipalli
ECCV (80)1
2023 Unconstrained Dysfluency Modeling for Dysfluent Speech Transcription and Detection
abstract
Dysfluent speech modeling requires time-accurate and silence-aware transcription at both the word-level and phonetic-level. However, current research in dysfluency modeling primarily focuses on either transcription or detection, and the performance of each aspect remains limited. In this work, we present an unconstrained dysfluency modeling (UDM) approach that addresses both transcription and detection in an automatic and hierarchical manner. UDM eliminates the need for extensive manual annotation by providing a comprehensive solution. Furthermore, we introduce a simulated dysfluent dataset called VCTK++to enhance the capabilities of UDM in phonetic transcription. Our experimental results demonstrate the effectiveness and robustness of our proposed methods in both transcription and detection tasks.
Jiachen Lian, Carly Feng, Naasir Farooqi, Steve Li, Anshul Kashyap, Cheol Jun Cho, Peter Wu, Robert Netzorg, Tingle Li, Gopala Krishna Anumanchipalli
ASRU9
2023 On Uni-Modal Feature Learning in Supervised Multi-Modal Learning
abstract
We abstract the features (i.e. learned representations) of multi-modal data into 1) uni-modal features, which can be learned from uni-modal training, and 2) paired features, which can only be learned from cross-modal interactions. Multi-modal models are expected to benefit from cross-modal interactions on the basis of ensuring uni-modal feature learning. However, recent supervised multi-modal late-fusion training approaches still suffer from insufficient learning of uni-modal features on each modality. We prove that this phenomenon does hurt the model’s generalization ability. To this end, we propose to choose a targeted late-fusion learning method for the given supervised multi-modal task from Uni-Modal Ensemble (UME) and the proposed Uni-Modal Teacher (UMT), according to the distribution of uni-modal and paired features. We demonstrate that, under a simple guiding strategy, we can achieve comparable results to other complex late-fusion or intermediate-fusion methods on various multi-modal datasets, including VGG-Sound, Kinetics-400, UCF101, and ModelNet40.
Chenzhuang Du, Jiaye Teng, Tingle Li, Tianyuan Yuan, Yang Yuan 0010, Hang Zhao 0021
ICML3
2023 Deep Speech Synthesis from MRI-Based Articulatory Representations
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W. Black, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli
INTERSPEECH2
2022 Learning Visual Styles from Audio-Visual Associations
Tingle Li, Andrew Owens, Hang Zhao 0021
ECCV (37)1
2022 Radio2Speech: High Quality Speech Recovery from Radio Frequency Signals
abstract
Considering the microphone is easily affected by noise and soundproof materials, the radio frequency (RF) signal is a promising candidate to recover audio as it is immune to noise and can traverse many soundproof objects.In this paper, we introduce Radio2Speech, a system that uses RF signals to recover high quality speech from the loudspeaker.Radio2Speech can recover speech comparable to the quality of the microphone, advancing from recovering only single tone music or incomprehensible speech in existing approaches.We use Radio UNet to accurately recover speech in time-frequency domain from RF signals with limited frequency band.Also, we incorporate the neural vocoder to synthesize the speech waveform from the estimated time-frequency representation without using the contaminated phase.Quantitative and qualitative evaluations show that in quiet, noisy and soundproof scenarios, Radio2Speech achieves state-of-the-art performance and is on par with the microphone that works in quiet scenarios.
Running Zhao, Jiang-Tao Yu 0001, Tingle Li, Hang Zhao 0021, Edith C. H. Ngai
INTERSPEECH3
2021 CVC: Contrastive Learning for Non-Parallel Voice Conversion
abstract
Cycle consistent generative adversarial network (CycleGAN) and variational autoencoder (VAE) based models have gained popularity in non-parallel voice conversion recently.However, they often suffer from difficult training process and unsatisfactory results.In this paper, we propose CVC, a contrastive learning-based adversarial approach for voice conversion.Compared to previous CycleGAN-based methods, CVC only requires an efficient one-way GAN training by taking the advantage of contrastive learning.When it comes to nonparallel one-to-one voice conversion, CVC is on par or better than CycleGAN and VAE while effectively reducing training time.CVC further demonstrates superior performance in manyto-one voice conversion, enabling the conversion from unseen speakers.
Tingle Li, Chenxu Hu, Hang Zhao 0021
Interspeech1
2021 Neural Dubber: Dubbing for Videos According to Scripts
abstract
Dubbing is a post-production process of re-recording actors’ dialogues, which is extensively used in filmmaking and video production. It is usually performed manually by professional voice actors who read lines with proper prosody, and in synchronization with the pre-recorded videos. In this work, we propose Neural Dubber, the first neural network model to solve a novel automatic video dubbing (AVD) task: synthesizing human speech synchronized with the given video from the text. Neural Dubber is a multi-modal text-to-speech (TTS) model that utilizes the lip movement in the video to control the prosody of the generated speech. Furthermore, an image-based speaker embedding (ISE) module is developed for the multi-speaker setting, which enables Neural Dubber to generate speech with a reasonable timbre according to the speaker’s face. Experiments on the chemistry lecture single-speaker dataset and LRS2 multi-speaker dataset show that Neural Dubber can generate speech audios on par with state-of-the-art TTS models in terms of speech quality. Most importantly, both qualitative and quantitative evaluations show that Neural Dubber can control the prosody of synthesized speech by the video, and generate high-fidelity speech temporally synchronized with the video.
Chenxu Hu, Qiao Tian 0001, Tingle Li, Yuping Wang 0005, Yuxuan Wang 0002, Hang Zhao 0021
NeurIPS3
2020 Atss-Net: Target Speaker Separation via Attention-Based Neural Network
abstract
Recently, Convolutional Neural Network (CNN) and Long short-term memory (LSTM) based models have been introduced to deep learning-based target speaker separation. In this paper, we propose an Attention-based neural network (Atss-Net) in the spectrogram domain for the task. It allows the network to compute the correlation between each feature parallelly, and using shallower layers to extract more features, compared with the CNN-LSTM architecture. Experimental results show that our Atss-Net yields better performance than the VoiceFilter, although it only contains half of the parameters. Furthermore, our proposed model also demonstrates promising performance in speech enhancement.
Tingle Li, Qingjian Lin, Yuanyuan Bao, Ming Li 0026
INTERSPEECH1
2020 The DKU Speech Activity Detection and Speaker Identification Systems for Fearless Steps Challenge Phase-02
Qingjian Lin, Tingle Li, Ming Li 0026
INTERSPEECH2