VLDB 2026 Research / reviewers in the wild / expert
Duc-Quang Vu
dblp:260/4485
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-5458-3713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLSS: Mandarin English Code-Switching Speech Recognition via Mutual Learning-Based Semi-Supervised MethodabstractCode-switching is a phenomenon of alternating use of two or more languages within or between utterances in communication that often occurs in multilingual communities. Recently, code-switching natural language processing and automatic speech recognition (ASR) have attracted numerous studies. However, a major obstacle affecting the results of these studies is the lack of transcribed data. In this letter, we propose a novel semi-supervised learning (SSL) approach to deal with this problem, namely Mutual Learning-Based Semi-Supervised Method (MLSS). The MLSS method involves the utilization of two networks for interleaved fine-tuning on a combination of transcribed dataset and pseudo-labeled data generated from another network. This iterative fine-tuning process repeats until all unlabeled data are selected for training or reaches a certain number of iterations. By incorporating mutual learning between the two networks, our approach effectively leverages the knowledge acquired from previous iterations during the training stage and combines the knowledge from both networks during the decoding process, resulting in a more robust and effective approach. To evaluate the effectiveness of our proposed method, we conduct experiments on the SEAME Mandarin-English code-switching corpus. The experimental results clearly illustrate that our approach outperforms other state-of-the-art methods, as evidenced by achieving a Mixed Error Rate (MER) of 15.6% /21.1% on test$_{man}$/test$_{sge}$sets. Cao Hong Nga, Duc-Quang Vu, Phuong Le Thi, Huong Hoang Luong, Jia-Ching Wang |
IEEE Signal Process. Lett. | 2 |
| 2024 | LCSL: Long-Tailed Classification via Self-LabelingabstractDuring the last decades, deep learning (DL) has been proven to be a very powerful and successful technique in many real-world applications, e.g., video surveillance or object detection. However, when class label distributions are highly skewed, DL classifiers tend to be biased towards majority classes during training phases. This leads to poor generalization of minority classes and consequently reduces the overall accuracy. How to effectively deal with this long-tailed class distribution in DL, i.e., deep long-tailed classification (DLC), remains a challenging problem despite many research efforts. Among various approaches, data augmentation, which aims at generating more samples for reducing label imbalance, is the most common and practical one. However, simply relying on existing class-agnostic augmentation strategies without properly considering the label differences would worsen the problem since more head-class samples can be inevitably augmented than tail-class ones. Moreover, none of the existing works consider the quality and suitability of augmented samples during the training process. Our proposed approach, called Long-tailed Classification via Self-Labeling (LCSL), is specifically designed to address these limitations. LCSL fundamentally differs from existing works by the way it iteratively exploits the preceding network during the training process to re-label the labeled augmented samples and uses the output confidence to decide whether new samples belong to minority classes before adding them to the data. Not only does this help to reduce imbalance ratios among classes, but this also helps to reduce the uncertainty of class prediction problems for minority classes by selecting more confident samples to the data. This incremental learning and generating scheme thus provide a new robust approach for decreasing model over-fitting, thus enhancing the overall accuracy, especially for minority classes. Extensive experiments have demonstrated that LCSL acquires better performance than state-of-the-art long-tailed learning techniques on various standard benchmark datasets. More specifically, our LCSL obtains 85.8%, 54.4%, and 56.2% in terms of accuracy on CIFAR10-LT, CIFAR100-LT, and ImageNet-LT (with moderate to extreme imbalance ratios), respectively. The source code is available athttps://github.com/vdquang1991/lcsl/. Duc-Quang Vu, Trang T. T. Phung, Jia-Ching Wang, Son T. Mai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Selinet: A Lightweight Model for Single Channel Speech SeparationabstractThe time-domain speech separation methods adopting deep learning have obtained impressive performance. However, the computational complexity, model size, and performance are still the challenges for the implementation on real-time low-resource devices. In this paper, we introduce a lightweight yet effective network for speech separation, namely SeliNet. The SeliNet is the one-dimensional convolutional architecture that employs bottleneck modules, and atrous temporal pyramid pooling. In bottleneck modules, the depth-wise separable convolution significantly decreases the model size and computational cost meanwhile the squeeze excitation uses a context vector to interact with the entire hidden state vector. Specifically, the atrous temporal pyramid pooling recognizes long-time sequences of various lengths and extracts context at different field-of-views. This helps SeliNet to obtain impressive performance while still maintaining the small computational cost and model size. Ha Minh Tan, Duc-Quang Vu, Jia-Ching Wang |
ICASSP | 2 |
| 2023 | Cyclic Transfer Learning for Mandarin-English Code-Switching Speech RecognitionabstractTransfer learning is a common method to improve the performance of the model on a target task via pre-training the model on pretext tasks. Different from the methods using monolingual corpora for pre-training, in this study, we propose a Cyclic Transfer Learning method (CTL) that utilizes both code-switching (CS) and monolingual speech resources as the pretext tasks. Moreover, the model in our approach is always alternately learned among these tasks. This helps our model can improve its performance via maintaining CS features during transferring knowledge. The experiment results on the standard SEAME Mandarin-English CS corpus have shown that our proposed CTL approach achieves the best performance with Mixed Error Rate (MER) of 16.3% on test$_{man}$, 24.1% on test$_{sge}$. In comparison to the baseline model that was pre-trained with monolingual data, our CTL method achieves 11.4% and 8.7% relative MER reduction on the test$_{man}$and test$_{sge}$sets, respectively. Besides, the CTL approach also outperforms compared to other state-of-the-art methods. The source code of the CTL method can be found athttps://github.com/caohongnga/CTL-CSSR. Cao Hong Nga, Duc-Quang Vu, Huong Hoang Luong, Chien-Lin Huang, Jia-Ching Wang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Selective Mutual Learning: An Efficient Approach for Single Channel Speech SeparationabstractMutual learning, the related idea to knowledge distillation, is a group of untrained lightweight networks, which simultaneously learn and share knowledge to perform tasks together during training. In this paper, we propose a novel mutual learning approach, namely selective mutual learning. This is the simple yet effective approach to boost the performance of the networks for speech separation. There are two networks in the selective mutual learning method, they are like a pair of friends learning and sharing knowledge with each other. Especially, the high-confidence predictions are used to guide the remaining network while the low-confidence predictions are ignored. This helps to remove poor predictions of the two networks during sharing knowledge. The experimental results have shown that our proposed selective mutual learning method significantly improves the separation performance compared to existing training strategies including independently training, knowledge distillation, and mutual learning with the same network architecture. Ha Minh Tan, Duc-Quang Vu, Chung-Ting Lee, Yung-Hui Li, Jia-Ching Wang |
ICASSP | 2 |
| 2022 | (2+1)D Distilled ShuffleNet: A Lightweight Unsupervised Distillation Network for Human Action RecognitionabstractWhile most existing deep neural networks (DNN) architectures are proposed for increasing performance, they also raise overall model complexity. However, practical applications require lightweight DNN models, that are able to run real-time in edge computing devices. In this work, we present a simple and elegant unsupervised distillation learning paradigm to train a lightweight network to human action recognition called (2+1)D Distilled ShuffleNet. Leveraging the distilling technique, the proposed method allows us to create a lightweight DNN model that achieves high accuracy and real-time speed. Our lightweight (2+1)D Distilled ShuffleNet is designed as an unsupervised paradigm; it does not require labelled data during distilling knowledge from the teacher to the student. Furthermore, to help the student be more "intelligent", we propose to distill the knowledge from two different teachers, i.e., 2D teacher and 3D teacher. The experimental results have shown that our lightweight (2+1)D Distilled ShuffleNet outperforms other state-of-the-art distillation networks with 86.4% and 59.9% top-1 accuracy on UCF101 and HMDB51 datasets, respectively, whereas the inference running time is at 47.16 FPS on CPU with only 17.1M parameters and 12.07 GFLOPs. Duc-Quang Vu, T. Hoang Ngan Le, Jia-Ching Wang |
ICPR | 1 |
| 2021 | A Novel Self-Knowledge Distillation Approach with Siamese Representation Learning for Action RecognitionabstractKnowledge distillation is an effective transfer of knowledge from a heavy network (teacher) to a small network (student) to boost students' performance. Self-knowledge dis-tillation, the special case of knowledge distillation, has been proposed to remove the large teacher network training process while preserving the student's performance. This paper intro-duces a novel Self-knowledge distillation approach via Siamese representation learning, which minimizes the difference between two representation vectors of the two different views from a given sample. Our proposed method, SKD-SRL, utilizes both soft label distillation and the similarity of representation vectors. Therefore, SKD-SRL can generate more consistent predictions and representations in various views of the same data point. Our benchmark has been evaluated on various standard datasets. The experimental results have shown that SKD-SRL significantly improves the accuracy compared to existing supervised learning and knowledge distillation methods regardless of the networks. Duc-Quang Vu, Thi-Thu-Trang Phung, Jia-Ching Wang |
VCIP | 1 |