Thanh Duc Ngo

dblp:65/3565 · also Thanh-Duc Ngo · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-6882-0070ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 NII-UIT at VBS2026: Towards Effective Visual Question Answering for Interactive and Multimodal Video Retrieval
Bao Tran, Tien Do, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001
MMM (4)3
2026 ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
abstract
Vision Language Models (VLMs) have rapidly advanced and show strong promise for text-based person search (TBPS), a task that requires capturing fine-grained relationships between images and text to distinguish individuals. Previous methods address these challenges through local alignment, yet they are often prone to shortcut learning and spurious correlations, yielding misalignment. Moreover, injecting prior knowledge can distort intra-modality structure. Motivated by our finding that encoder attention surfaces spatially precise evidence from the earliest training epochs, and to alleviate these issues, we introduce ITSELF, an attention-guided framework for implicit local alignment. At its core, Guided Representation with Attentive Bank (GRAB) converts the model’s own attention into an Attentive Bank of high-saliency tokens and applies local objectives on this bank, learning fine-grained correspondences without extra supervision. To make the selection reliable and non-redundant, we introduce Multi-Layer Attention for Robust Selection (MARS), which aggregates attention across layers and performs diversity-aware top-k selection; and Adaptive Token Scheduler (ATS), which schedules the retention budget from coarse to fine over training, preserving context early while progressively focusing on discriminative details. Extensive experiments on three widely used TBPS benchmarks show state-of-the-art performance and strong cross-dataset generalization, confirming the effectiveness and robustness of our approach without additional prior supervision. Our project is publicly available at https://trhuuloc.github.io/itself
Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo
WACV3
2026 Skeleton-guided artistic text recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo
Int. J. Document Anal. Recognit.5
2026 Towards scalable and context-aware multimodal interactive video retrieval
Bao Tran, Khiem Le 0001, Thanh Duc Ngo
Multim. Syst.3
2025 Towards Understanding the Logical Layout of Scene Text in Signboard Images
Giang Tran Thi Cam, Cam-Nguyen Tran-Nhu, Thuyen Tran 0001, Thanh Duc Ngo
ICDAR (3)4
2025 Skeleton-Guided Artistic Text Recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo
ICDAR (5)5
2025 Multi-Perspective Data Augmentation for Few-shot Object Detection
abstract
Recent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of $17.5\%$ in nAP50 over the baseline on PASCAL VOC.
Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002
ICLR4
2025 NII-UIT at VBS2025: Multimodal Video Retrieval with LLM Integration and Dynamic Temporal Search
Bao Tran Gia, Tuong Bui Cong Khanh, Tam Le Thi Thanh, Thuyen Tran 0001, Khiem Le 0001, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001
MMM (5)8
2025 Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
abstract
Unsupervised domain adaptation (UDA) has become increasingly prevalent in scene text recognition (STR), especially where training and testing data reside in different domains. The efficacy of existing UDA approaches tends to degrade when there is a large gap between the source and target domains. To deal with this problem, gradually shifting or progressively learning to shift from domain to domain is the key issue. In this paper, we introduce the Stratified Domain Adaptation (StrDA) approach, which examines the gradual escalation of the domain gap for the learning process. The objective is to partition the target data into subsets so that the progressively self-trained model can adapt to gradual changes. We stratify the target data by evaluating the proximity of each data sample to both the source and target domains. We propose a novel method for employing domain discriminators to estimate the out-of-distribution and domain discriminative levels of data samples. Extensive experiments on benchmark scene-text datasets show that our approach significantly improves the performance of baseline (source-trained) STR models. The source code is available at https://github.com/KhaLee2307/StrDA.
Kha Nhat Le, Hoang-Tuan Nguyen, Hung Tien Tran, Thanh Duc Ngo
WACV4
2023 Unsupervised Domain Adaptation with Imbalanced Character Distribution for Scene Text Recognition
abstract
Recent deep learning based methods have demonstrated promising results in scene text recognition. One of the major difficulty is the lack of manually annotated data. Synthetic data are then used to eliminate the requirement for human annotation. However, the domain gap between synthetic and real-world data remains a challenging issue. To bridge the gap, unsupervised domain adaptation (UDA) was introduced to transfer knowledge from a labeled source domain to a target domain. In this work, we introduce an unsupervised domain adaptation method based on a sequence-to-sequence attention model. We take into account imbalanced distribution of characters to optimize the adaptation process. We propose to use focal loss as the classification loss for the labeled source domain and focal entropy as the entropy loss for the unlabeled target domain. Our proposed method, named ICD-DA, outperforms other UDA methods on official benchmarks.
Hung Tran Tien, Thanh Duc Ngo
ICIP2
2023 Abstraction-perception preserving cartoon face synthesis
Sy-Tuyen Ho, Manh-Khanh Ngo Huu, Thanh-Danh Nguyen, Nguyen Phan, Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Tam V. Nguyen 0002
Multim. Tools Appl.6
2023 Instance-Level Few-Shot Learning With Class Hierarchy Mining
abstract
Few-shot learning is proposed to tackle the problem of scarce training data in novel classes. However, prior works in instance-level few-shot learning have paid less attention to effectively utilizing the relationship between categories. In this paper, we exploit the hierarchical information to leverage discriminative and relevant features of base classes to effectively classify novel objects. These features are extracted from abundant data of base classes, which could be utilized to reasonably describe classes with scarce data. Specifically, we propose a novel superclass approach that automatically creates a hierarchy considering base and novel classes as fine-grained classes for few-shot instance segmentation (FSIS). Based on the hierarchical information, we design a novel framework called Soft Multiple Superclass (SMS) to extract relevant features or characteristics of classes in the same superclass. A new class assigned to the superclass is easier to classify by leveraging these relevant features. Besides, in order to effectively train the hierarchy-based-detector in FSIS, we apply the label refinement to further describe the associations between fine-grained classes. The extensive experiments demonstrate the effectiveness of our method on FSIS benchmarks. The source code is available here: https://github.com/nvakhoa/superclass-FSIS.
Anh-Khoa Nguyen Vu, Thanh-Toan Do, Nhat-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Tam V. Nguyen 0002
IEEE Trans. Image Process.5
2022 UIT at VBS 2022: An Unified and Interactive Video Retrieval System with Temporal Search
Khanh Ho, Vu Xuan Dinh, Khiem Le 0001, Khang Dinh Tran, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le
MMM (2)8
2022 Few-shot object detection via baby learning
Anh-Khoa Nguyen Vu, Nhat-Duy Nguyen, Khanh-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002
Image Vis. Comput.5
2021 Dictionary-Guided Scene Text Recognition
abstract
Language prior plays an important role in the way humans detect and recognize text in the wild. Current scene text recognition methods do use lexicons to improve recognition performance, but their naive approach of casting the output into a dictionary word based purely on the edit distance has many limitations. In this paper, we present a novel approach to incorporate a dictionary in both the training and inference stage of a scene text recognition system. We use the dictionary to generate a list of possible outcomes and find the one that is most compatible with the visual appearance of the text. The proposed method leads to a robust scene text recognition model, which is better at handling ambiguous cases encountered in the wild, and improves the overall performance of state-of-the-art scene text spotting frameworks. Our work suggests that incorporating language prior is a potential approach to advance scene text detection and recognition methods. Besides, we contribute VinText, a challenging scene text dataset for Vietnamese, where some characters are equivocal in the visual form due to accent symbols. This dataset will serve as a challenging benchmark for measuring the applicability and robustness of scene text detection and recognition algorithms. Code and dataset are available at https://github.com/VinAIResearch/dict-guided.
Thu Nguyen 0003, Vinh Tran 0005, Minh-Triet Tran, Thanh Duc Ngo, Thien Huu Nguyen, Minh Hoai
CVPR5
2018 Video Search Based on Semantic Extraction and Locally Regional Object Proposal
Thanh-Dat Truong, Vinh-Tiep Nguyen, Minh-Triet Tran, Trang-Vinh Trieu, Tien Do, Thanh Duc Ngo, Duy-Dinh Le
MMM (2)6
2017 Evaluation of Deep Models for Real-Time Small Object Detection
Phuoc Pham, Tien Do, Thanh Duc Ngo, Duy-Dinh Le
ICONIP (3)4
2017 Semantic Extraction and Object Proposal for Video Search
Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Minh-Triet Tran, Duc Anh Duong, Shin'ichi Satoh 0001
MMM (2)2
2017 Efficient large-scale multi-class image classification by learning balanced trees
Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001
Comput. Vis. Image Underst.2
2017 Scalable Face Track Retrieval in Video Archives Using Bag-of-Faces Sparse Representation
abstract
Huge video archives consisting of news programs, dramas, movies, and Web videos (e.g., YouTube) are available in our daily life. In all these videos, human is usually one of the most important subjects. Using state-of-the-art techniques, we can efficiently detect and track faces in the videos. In order to organize large-scale face tracks, containing sequences of (detected) consecutive faces in the videos, we propose an efficient method to retrieve human face tracks using bag-of-faces sparse representation (BoF-SR). Using the proposed method, a face track is encoded as a single BoF-SR, therefore allowing an efficient indexing method to handle large-scale data. To further consider the possible variations in face tracks, we generalize our method to find multiple SRs, in an unsupervised manner, to represent a bag of faces and balance the tradeoff between performance and retrieval time. The experimental results on two real-world (million-scale) data sets confirm that the proposed methods achieve significant performance gains compared with different state-of-the-art methods.
Bor-Chun Chen, Yan-Ying Chen, Yin-Hsi Kuo, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Winston H. Hsu
IEEE Trans. Circuits Syst. Video Technol.4
2016 Efficient Large Scale Image Classification via Prediction Score Decomposition
Duy-Dinh Le, Tien-Dung Mai, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong
ECCV (6)4
2016 Using node relationships for hierarchical classification
abstract
Hierarchical classification is a computational efficient approach for large-scale image classification. The main challenging issue of this approach is to deal with error propagation. Irrelevant branching decision made at a parent node cannot be corrected at its child nodes in traversing the tree for classification. This paper presents a novel approach to reduce branching error at a node by taking its relative relationship into account. Given a node on the tree, we model each candidate branch by considering classification response of its child nodes, grandchild nodes and their differences with siblings. A maximum margin classifier is then applied to select the most discriminating candidate. Our proposed approach outperforms related approaches on Caltech-256, SUN-397 and ILSVRC2010-1K.
Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001
ICIP2
2016 News Archive Exploration Combining Face Detection and Tracking with Network Visual Analytics
abstract
Visual analytics helps analytical reasoning and exploration of complex systems, for which it combines the means of interactive visualization, with the power of data analytics. The recent progress in computer vision techniques opens wide applications in real world video archives. Particularly, recent advances in face detection and recognition have been put under the spotlight. The applications of such techniques are often concern intelligence, or peer recognition in photo posted in social networks. We propose to combine those two domains by demonstrating a visual exploration of over a decade of the news program from the Japanese broadcaster NHK News 7. We derive social networks from face detection and tracking of this large dataset. With the help of a little domain knowledge, we monitor the activity of political public figures and explore the archive. This allows understanding and comparison of the politico-media scene presented by NHK under different Prime Minister's governance. The social networks are interactive, and also allow to explore the multimedia database and explore its video content.
Benjamin Renoust, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001
ACM Multimedia2
2015 Transfer AdaBoost SVM for Link Prediction in Newly Signed Social Networks using Explicit and PNR Features
abstract
In signed social network, the user-generated content and interactions have overtaken the web. Questions of whom and what to trust has become increasingly important. We must have methods which predict the signs of links in the social network to solve this problem. We study signed social networks with positive links (friendship, fan, like, etc) and negative links (opposition, anti-fan, dislike, etc). Specifically, we focus how to effectively predict positive and negative links in newly signed social networks. With SVM model, the small amount of edge sign information in newly signed network is not adequate to train a good classifier. In this paper, we introduce an effective solution to this problem. We present a novel transfer learning framework is called Transfer AdaBoost with SVM (TAS) which extends boosting-based learning algorithms and incorporates properly designed RBFSVM (SVM with the RBF kernel) component classifiers. With our framework, we use explicit topological features and Positive Negative Ratio (PNR) features which are based on decision-making theory. Experimental results on three networks (Epinions, Slashdot and Wiki) demonstrate our method that can improve the prediction accuracy by 40% over baseline methods. Additionally, our method has faster performance time.
Thu Nguyen 0003, Phuc Quang Nguyen, Thanh Duc Ngo, Tu-Anh Nguyen-Hoang
KES3
2015 NII-UIT Browser: A Multimodal Video Search System
Thanh Duc Ngo, Vinh-Tiep Nguyen, Vu Hoang Nguyen, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001
MMM (2)1
2015 AttRel: An Approach to Person Re-Identification by Exploiting Attribute Relationships
Ngoc-Bao Nguyen, Vu Hoang Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong
MMM (2)3
2015 Human Action recognition from depth videos using multi-projection based representation
abstract
In this paper, a novel method for human action recognition from depth videos is proposed. We project 3D data on to multiple 2D-planes from which dense trajectories features are extracted. In the training stage, for each projection, a classifier is trained using the training data. In the testing stage, for each test video, the multiple trained classifiers are applied and the predicted scores are combined for final decision. We propose a greedy-based method to select a subset of the trained classifiers for optimal combination. Experiments on the MSR Action 3D dataset show that the proposed method outperforms the baseline method that does not use multi-projection-based features.
Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong
MMSP2
2015 Large scale multi-class classification using latent classifiers
abstract
We study the problem of multi-class image classification with large number of classes, of which the one-vs-all based approach is prohibitive in practical applications. Recent state-of-the-art approaches rely on label tree to reduce classification complexity. However, building optimal tree structures and learning precise classifiers to optimize tree loss is challenging. In this paper, we introduce a novel approach using latent classifiers that can achieve comparable speed but better performance. The key idea is that instead of using C one-vs-all classifiers (C is the number of classes) to generate the score matrix for label prediction, a much smaller number of classifiers are used. These classifiers, called latent classifiers, are generated by analyzing the correlation among classes and removing redundancy. Experiments on several large datasets including ImageNet-1K, SUN-397, and Caltech-256 show the efficiency of our approach.
Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001
MMSP2
2015 Cross-View Action Recognition by Projection-Based Augmentation
Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong
PSIVT2
2014 Integrating Spatial Information into Inverted Index for Large-Scale Image Retrieval
abstract
In recent years, large-scale image retrieval has been shown remarkable potential in real-life applications. To reduce retrieval time as searched database may contain thousands of images, Inverted Indexing is the basic technique, given images are represented by Bag-of-Words model. However, one major limitation of both standard Inverted Index and Bag-of-Words model is that they ignore spatial information of the visual words in images. This might reduce retrieval accuracy. In this paper, we introduce an approach to integrate spatial information into inverted index to improve accuracy while maintaining short retrieval time. Experiments conducted on several benchmark datasets (Oxford Building 5K, Paris 6K and Oxford Building 5K+100K) demonstrate the effectiveness of our proposed approach.
Bien-Van Nguyen, Duy Pham, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong
ISM3
2014 NII-UIT: A Tool for Known Item Search by Sequential Pattern Filtering
Thanh Duc Ngo, Vu Hoang Nguyen, Vu Lam, Sang Phan Le, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001
MMM (2)1
2013 NII-UIT-VBS: A Video Browsing Tool for Known Item Search
Duy-Dinh Le, Vu Lam, Thanh Duc Ngo, Vinh Quang Tran, Vu Hoang Nguyen, Duc Anh Duong, Shin'ichi Satoh 0001
MMM (2)3
2012 Robust eye localization in video by combining eye detector and eye tracker
Chi Nhan Duong, Thang Cap Pham Dinh, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Bac Le, Shin'ichi Satoh 0001
ICPR3
2011 Boosting global scene classification accuracy by discriminative region localization
abstract
Combining global scene classification with object detection has helped in improving the classification accuracy. However, training an object detector requires a large amount of manual annotation. The object detector may also fail when the object is occluded. Meanwhile, the presence of the object is not only indicated by the entire object region but any of its parts or its correlations with other regions in the image. To overcome these limitations, we propose using discriminative region localization instead of object detection in the combination. Our contribution is two-fold, a) a complete framework that combines global scene classification with discriminative region localization for image classification and b) a weakly supervised discriminative region localization approach that utilizes spatial context to improve the learning accuracy. Our experimental results on benchmark datasets demonstrated that the proposed discriminative region localization approach outperforms the state-of-the-art approach. In addition, the combination significantly increases the classification performance.
Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001
ICIP1
2011 Fast face sequence matching in large-scale video databases
abstract
There have recently been many methods proposed for matching face sequences in the field of face retrieval. However, most of them have proven to be inefficient in large-scale video databases because they frequently require a huge amount of computational cost to obtain a high degree of accuracy. We present an efficient matching method that is based on the face sequences (called face tracks) in large-scale video databases. The key idea is how to capture the distribution of a face track in the fewest number of low-computational steps. In order to do that, each face track is represented by a vector that approximates the first principal component of the face track distribution and the similarity of face tracks bases on the similarity of these vectors. Our experimental results from a large-scale database of 457,320 human faces extracted from 370 hours of TRECVID videos from 2004-2006 show that the proposed method easily handles the scalability by maintaining a good balance between the speed and the accuracy.
Hung Thanh Vu, Thanh Duc Ngo, Thao Ngoc Nguyen, Duy-Dinh Le, Shin'ichi Satoh 0001, Bac Le, Duc Anh Duong
ICIP2
2008 A text segmentation based approach to video shot boundary detection
abstract
Video shot boundary detection is one of the fundamental tasks of video indexing and retrieval applications. Although many methods have been proposed for this task, finding a general and robust shot boundary method that is able to handle the various transition types caused by photo flashes, rapid camera movement and object movement is still challenging. We present a novel approach for detecting video shot boundaries in which we cast the problem of shot boundary detection into the problem of text segmentation in natural language processing. This is possible by assuming that each frame is a word and then the shot boundaries are treated as text segment boundaries (e.g. topics). The text segmentation based approaches in natural language processing can be used. The experimental results from various long video sequences have proved the effectiveness of our approach.
Duy-Dinh Le, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong
MMSP3