VLDB 2026 Research / reviewers in the wild / expert
Duc-Tuan Truong
dblp:317/1170
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0002-1767-7598ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Audio Deepfake Detection using Ensemble Confidence CalibrationabstractModel ensembles using linear interpolation are commonly employed to improve classification performance, with higher weights assigned to better-performing models in the ensemble. However, prior methods use fixed weights across all test samples, which is suboptimal as different models may perform better in different subsets of the samples, especially in out-of-domain (OOD) scenarios. This is a key challenge in Audio Deepfake Detection (ADD) due to variations between training and testing domains. To address this, we propose using EOW-Softmax, a method for modeling open-world uncertainties, to calibrate the magnitudes of OOD classification scores at the sample level. This dynamic adjustment improves ensemble predictions on OOD samples. When tested on the ASVspoof 2021 dataset, our calibrated ensemble reduced the equal error rate (EER) from 2.66% to 2.03%. Kwok Chin Yuen, Duc-Tuan Truong, Jia Qi Yip |
ICASSP | 2 |
| 2025 | Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-SpoofingabstractSpeech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net’s superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net. Tianchi Liu 0004, Duc-Tuan Truong, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker VerificationabstractKnowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from non-target speakers, particularly their classification probabilities, which can be crucial for automatic speaker verification. In this paper, we first demonstrate that leveraging a larger number of training non-target speakers improves the performance of automatic speaker verification models. Inspired by this finding about the importance of non-target speakers’ knowledge, we modified the conventional label-level KD by disentangling and emphasizing the classification probabilities of non-target speakers during knowledge distillation. The proposed method is applied to three different student model architectures and achieves an average of 13.67% improvement in EER on the VoxCeleb dataset compared to embedding-level and conventional label-level KD methods.1 Duc-Tuan Truong, Ruijie Tao, Jia Qi Yip, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 1 |
| 2024 | Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech DetectionabstractRecent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts.This improvement could be due to the powerful modeling ability of the multi-head selfattention (MHSA) in the Transformer model, which learns the temporal relationship of each input token.However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence.In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporalchannel dependencies.Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER.Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech 1 . Duc-Tuan Truong, Ruijie Tao, Hieu-Thi Luong, Kong-Aik Lee, Chng Eng Siong |
INTERSPEECH | 1 |
| 2024 | Multi-Stage Face-Voice Association Learning with Keynote Speaker DiarizationabstractThe human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as "cross-modal speaker verification''. This task poses significant challenges due to the complex relationship between the modalities. In this paper, we propose a "Multi-stage Face-voice Association Learning with Keynote Speaker Diarization''(MFV-KSD) framework. MFV-KSD contains a keynote speaker diarization front-end to effectively address the noisy speech inputs issue. To balance and enhance the intra-modal feature learning and inter-modal correlation understanding, MFV-KSD utilizes a novel three-stage training strategy. Our experimental results demonstrated robust performance, achieving the first rank in the 2024 Face-voice Association in Multilingual Environments (FAME) challenge with an overall Equal Error Rate (EER) of 19.9%. Details can be found in https://github.com/TaoRuijie/MFV-KSD. Ruijie Tao, Yidi Jiang, Duc-Tuan Truong, Chng Eng Siong, Massimo Alioto, Haizhou Li 0001 |
ACM Multimedia | 4 |
| 2024 | Room Impulse Responses Help Attackers to Evade Deep Fake DetectionabstractThe ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system’s EER. To counter this type of attack, We augmented training data with a large-scale synthetic/simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%. Hieu-Thi Luong, Duc-Tuan Truong, Kong-Aik Lee, Chng Eng Siong |
SLT | 2 |
| 2023 | ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2022 | Estimation of speaker age and height from speech signal using bi-encoder transformer mixture model
Duc-Tuan Truong, Chng Eng Siong |
INTERSPEECH | 2 |