VLDB 2026 Research / reviewers in the wild / expert
Tyler Vuong
dblp:227/2337
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human InteractionabstractAdvait Gosai, Tyler Vuong, Utkarsh Tyagi, Steven Li, Wenjia You, Miheer Bavare, Arda Uçar, Zhongwang Fang, Brian Jang, Bing Liu, Yunzhong He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Advait Gosai, Tyler Vuong, Utkarsh Tyagi, Steven Li, Wenjia You, Miheer Bavare, Arda Uçar, Zhongwang Fang, Brian Jang, Yunzhong He |
ACL (1) | 2 |
| 2025 | Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze |
INTERSPEECH | 4 |
| 2023 | Unsupervised Voice Type Discrimination Score Adaptation Using X-Vector ClustersabstractVoice type discrimination (VTD) is the task of automatically detecting speech produced in the same room as a recording device ("live speech") among other speech and non-speech noises, such as traffic noises or radio broadcasts ("distractor audio"). Existing work has described methods for performing the VTD task. This paper presents a method for adapting the output of these existing methods in an unsupervised manner via x-vector clustering and correlation. This adaptation method can be applied to the output of any VTD algorithm, requires no additional training data, and has been shown to yield a relative decrease in decision cost function (DCF) score of up to 47% on a standardized database collected for the task. Mark Lindsey, Tyler Vuong, Richard M. Stern |
ICASSP | 2 |
| 2022 | Improved Modulation-Domain Loss for Neural-Network-based Speech Enhancement
Tyler Vuong, Richard M. Stern |
INTERSPEECH | 1 |
| 2022 | Investigating the Important Temporal Modulations for Deep-Learning-Based Speech Activity DetectionabstractWe describe a learnable modulation spectrogram feature for speech activity detection (SAD). Modulation features capture the temporal dynamics of each frequency subband. We compute learnable modulation spectrogram features by first calculating the log-mel spectrogram. Next, we filter each frequency subband with a bandpass filter that contains a learnable center frequency. The resulting SAD system was evaluated on the Fearless Steps Phase-04 SAD challenge. Experimental results showed that temporal modulations around the 4–6 Hz range are crucial for deep-learning-based SAD. These experimental results align with previous studies that found slow temporal modulation to be most important for speech-processing tasks and speech intelligibility. Additionally, we found that the learnable modulation spectrogram feature outperforms both the standard log-mel and fixed modulation spectrogram features on the Fearless Steps Phase-04 SAD test set. Tyler Vuong, Nikhil Madaan, Rohan Panda, Richard M. Stern |
SLT | 1 |
| 2021 | A Modulation-Domain Loss for Neural-Network-Based Real-Time Speech EnhancementabstractWe describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then used to calculate a weighted mean-squared error (MSE) in the modulation domain for training a speech enhancement system. Experiments showed that adding the modulation-domain MSE to the MSE in the spectro-temporal domain substantially improved the objective prediction of speech quality and intelligibility for real-time speech enhancement systems without incurring additional computation during inference. Tyler Vuong, Yangyang Xia, Richard M. Stern |
ICASSP | 1 |
| 2021 | Generalized Spoofing Detection Inspired from Audio Generation ArtifactsabstractState-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains.Such artifacts could be well captured by the frequency domain analysis over the spectrogram.Thus, we propose a novel use of long-range spectro-temporal modulation feature -2D DCT over log-Mel spectrogram for the audio deepfake detection.We show that this feature works better than log-Mel spectrogram, CQCC, MFCC, as a suitable candidate to capture such artifacts.We employ spectrum augmentation and feature normalization to decrease overfitting and bridge the gap between training and test dataset along with this novel feature introduction.We developed a CNN-based baseline that achieved a 0.0849 t-DCF and outperformed the previously top single systems reported in the ASVspoof 2019 challenge.Finally, by combining our baseline with our proposed 2D DCT spectro-temporal feature, we decrease the t-DCF score down by 14% to 0.0737, making it a state-of-the-art system for spoofing detection.Furthermore, we evaluate our model using two external datasets, showing the proposed feature's generalization ability.We also provide analysis and ablation studies for our proposed feature and results. Tyler Vuong, Mahsa Elyasi, Gaurav Bharaj, Rita Singh |
Interspeech | 2 |
| 2021 | The Application of Learnable STRF Kernels to the 2021 Fearless Steps Phase-03 SAD Challenge
Tyler Vuong, Yangyang Xia, Richard M. Stern |
Interspeech | 1 |
| 2020 | Learnable Spectro-Temporal Receptive Fields for Robust Voice Type DiscriminationabstractVoice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio that were played back such as traffic noise and television broadcasts ("Distractor Audio"). In this work, we propose a deep-learning-based VTD system that features an initial layer of learnable spectro-temporal receptive fields (STRFs). Our approach is also shown to provide very strong performance on a similar spoofing detection task in the ASVspoof 2019 challenge. We evaluate our approach on a new standardized VTD database that was collected to support research in this area. In particular, we study the effect of using learnable STRFs compared to static STRFs or unconstrained kernels. We also show that our system consistently improves a competitive baseline system across a wide range of signal-to-noise ratios on spoofing detection in the presence of VTD distractor noise. Tyler Vuong, Yangyang Xia, Richard M. Stern |
INTERSPEECH | 1 |