VLDB 2026 Research / reviewers in the wild / expert
Van Hai Do
dblp:125/2874
· DBLP profile ↗
18ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-9554-5171ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Emotion-Complexity Guided Expert Activation for Speech Emotion Recognition
Ta Bao Thang, Huynh Thi Thanh Binh, Van Hai Do |
IEEE Signal Process. Lett. | 3 |
| 2025 | Exploring Non-Matching Multiple References for Speech Quality AssessmentabstractNon-Matching Reference-based Speech Quality Assessment models typically require numerous references during inference to ensure stable and accurate predictions. However, this dependency introduces significant computational overhead, limiting their suitability for real-time applications. In this paper, we propose a novel training paradigm that directly addresses prediction instability at its source by integrating multiple references during training rather than during inference, as in existing approaches. This method allows the model to capture the inherent variability of reference signals, thereby enhancing prediction reliability. Additionally, we introduce an auxiliary variance loss function to minimize inconsistencies across predictions, ensuring stable assessments regardless of the number of references used. Experiments on the NISQA datasets demonstrate that, with the same training time, our method achieves consistent predictions with a single reference during inference, resulting in a 100-fold reduction in computational time while maintaining high accuracy. Ta Bao Thang, Nhat Minh Le, Huynh Thi Thanh Binh, Van Hai Do |
IEEE Signal Process. Lett. | 4 |
| 2024 | Human Behavior Modeling in Speech Transcribing Process via Pretrained Speech Recognition ModelsabstractThe process of transcribing speech plays a critical role in developing automatic speech recognition (ASR) systems. During this process, human transcribers listen to spoken utterances and convert them into accurately written texts. However, our initial investigations using telephone conversation datasets have shown that approximately 40% of the recorded utterances are challenging to transcribe and often disregarded by transcribers due to various factors, such as unintelligibility. As a result, transcribers waste significant time and effort listening to such utterances, significantly slowing down the overall transcribing speed. In this paper, we explore the potential of pretrained speech recognition models to model transcribers’ behavior in the speech transcribing process. By leveraging the knowledge and insights encoded in these models, we aim to filter out problematic utterances that transcribers typically disregard before they reach human transcribers, thus optimizing their time and effort. Through experiments on practical telephone datasets, our approach utilizing pretrained speech recognition models successfully reduces unwanted utterances by 50% while preserving the high quality of the collected dataset. Ta Bao Thang, Minh Khang Pham, Nhat Minh Le, Van Hai Do |
IJCNN | 4 |
| 2024 | Enhancing Non-Matching Reference Speech Quality Assessment through Dynamic Weight Adaptation
Ta Bao Thang, Van Hai Do, Huynh Thi Thanh Binh |
INTERSPEECH | 2 |
| 2024 | Enhancing No-Reference Speech Quality Assessment with Pairwise, Triplet Ranking Losses, and ASR Pretraining
Ta Bao Thang, Minh Tu Le, Van Hai Do, Huynh Thi Thanh Binh |
INTERSPEECH | 3 |
| 2024 | Transfer learning methods for low-resource speech accent recognition: A case study on Vietnamese language
Ta Bao Thang, Nhat Minh Le, Van Hai Do |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | LightVoc: An Upsampling-Free GAN Vocoder Based On Conformer And Inverse Short-time Fourier Transform
Dinh Son Dang, Tung Lam Nguyen 0002, Ta Bao Thang, Tien Thanh Nguyen, Thi Ngoc Anh Nguyen, Dang Linh Le, Nhat Minh Le, Van Hai Do |
INTERSPEECH | 8 |
| 2023 | Probing Speech Quality Information in ASR Systems
Ta Bao Thang, Minh Tu Le, Nhat Minh Le, Van Hai Do |
INTERSPEECH | 4 |
| 2018 | Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched TranscriptionabstractIt is challenging to obtain large amounts of native (matched) labels for speech audio in underresourced languages. This challenge is often due to a lack of literate speakers of the language, or in extreme cases, a lack of universally acknowledged orthography as well. One solution is to increase the amount of labeled data by using mismatched transcription, which employs transcribers who do not speak the underresourced language of interest called the target language (in place of native speakers), to transcribe what they hear as nonsense speech in their own annotation language (≠ target language). Previous uses of mismatched transcription converted it to a probabilistic transcription (PT), but PT is limited by the errors of nonnative perception. This paper proposes, instead, a multitask learning framework in which one deep neural network (DNN) is trained to optimize two separate tasks: acoustic modeling of a small number of matched transcription with matched target-language graphemes; and acoustic modeling of a large number of mismatched transcription with mismatched annotation-language graphemes. We find that: first, the multitask learning framework gives significant improvement over monolingual, semisupervised learning, multilingual DNN training, and transfer learning baselines; second, a Gaussian Mixture Model-Hidden-Markov Model (GMM-HMM) model adapted using PT improves alignments, thereby improving training; and third, bottleneck features trained on the mismatched transcriptions lead to even better alignments, resulting in further performance gains of the multitask DNN. Our experiments are conducted on the IARPA Georgian and Vietnamese BABEL corpora as well as on our newly collected speech corpus of Singapore Hokkien, an underresourced language with no standard written form. Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Multi-Task Learning Using Mismatched Transcription for Under-Resourced Speech Recognition
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2016 | Improving Efficiency of Sentence Boundary Detection by Feature Selection
Thi-Nga Ho, Tze Yuang Chong, Van Hai Do, Van Tung Pham, Chng Eng Siong |
ACIIDS (2) | 3 |
| 2016 | Exemplar-inspired strategies for low-resource spoken keyword search in SwahiliabstractWe present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples. Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 5 |
| 2016 | Approximate search of audio queries by using DTW with phone time boundary and data augmentationabstractDynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DTW is sensitive to the mismatch of signal conditions between the query and the speech search data. To allow approximate search, we propose a partial template matching strategy using phone time boundary information generated by a phone recognizer. To have more invariant representation of audio signals, we use bottleneck features (BNF) as the input of DTW. The BNF network is trained from augmented data, which is generated by adding reverberation and additive noises to the clean training data. Experimental results on QUESST 2015 task shows the effectiveness of the proposed methods for QbE-STD when the queries and search data are both distorted by reverberation and noises. Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Cheung-Chi Leung, Lei Wang 0020, Van Hai Do, Hang Lv 0001, Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 7 |
| 2016 | Analysis of Mismatched Transcriptions Generated by Humans and Machines for Under-Resourced Languages
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2015 | A comparative study of BNF and DNN multilingual training on cross-lingual low-resource speech recognition
Haihua Xu 0001, Van Hai Do, Chng Eng Siong |
INTERSPEECH | 2 |
| 2014 | Kernel density-based acoustic model with cross-lingual bottleneck features for resource limited LVCSR
Van Hai Do, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2013 | Context-dependent phone mapping for LVCSR of under-resourced languages
Van Hai Do, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2013 | A study on LVCSR and keyword search for tagalog
Korbinian Riedhammer, Van Hai Do, James Hieronymus |
INTERSPEECH | 2 |