EDBT 2026 Demo / reviewers in the wild / expert
Thanh Duc Ngo
dblp:65/3565 · also Thanh-Duc Ngo
· DBLP profile ↗
36ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-6882-0070ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NII-UIT at VBS2026: Towards Effective Visual Question Answering for Interactive and Multimodal Video Retrieval
Bao Tran, Tien Do, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
MMM (4) | 3 |
| 2026 | ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language RetrievalabstractVision Language Models (VLMs) have rapidly advanced and show strong promise for text-based person search (TBPS), a task that requires capturing fine-grained relationships between images and text to distinguish individuals. Previous methods address these challenges through local alignment, yet they are often prone to shortcut learning and spurious correlations, yielding misalignment. Moreover, injecting prior knowledge can distort intra-modality structure. Motivated by our finding that encoder attention surfaces spatially precise evidence from the earliest training epochs, and to alleviate these issues, we introduce ITSELF, an attention-guided framework for implicit local alignment. At its core, Guided Representation with Attentive Bank (GRAB) converts the model’s own attention into an Attentive Bank of high-saliency tokens and applies local objectives on this bank, learning fine-grained correspondences without extra supervision. To make the selection reliable and non-redundant, we introduce Multi-Layer Attention for Robust Selection (MARS), which aggregates attention across layers and performs diversity-aware top-k selection; and Adaptive Token Scheduler (ATS), which schedules the retention budget from coarse to fine over training, preserving context early while progressively focusing on discriminative details. Extensive experiments on three widely used TBPS benchmarks show state-of-the-art performance and strong cross-dataset generalization, confirming the effectiveness and robustness of our approach without additional prior supervision. Our project is publicly available at https://trhuuloc.github.io/itself Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo |
WACV | 3 |
| 2026 | Skeleton-guided artistic text recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo |
Int. J. Document Anal. Recognit. | 5 |
| 2026 | Towards scalable and context-aware multimodal interactive video retrieval
Bao Tran, Khiem Le 0001, Thanh Duc Ngo |
Multim. Syst. | 3 |
| 2025 | Towards Understanding the Logical Layout of Scene Text in Signboard Images
Giang Tran Thi Cam, Cam-Nguyen Tran-Nhu, Thuyen Tran 0001, Thanh Duc Ngo |
ICDAR (3) | 4 |
| 2025 | Skeleton-Guided Artistic Text Recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo |
ICDAR (5) | 5 |
| 2025 | Multi-Perspective Data Augmentation for Few-shot Object DetectionabstractRecent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of $17.5\%$ in nAP50 over the baseline on PASCAL VOC. Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
ICLR | 4 |
| 2025 | NII-UIT at VBS2025: Multimodal Video Retrieval with LLM Integration and Dynamic Temporal Search
Bao Tran Gia, Tuong Bui Cong Khanh, Tam Le Thi Thanh, Thuyen Tran 0001, Khiem Le 0001, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
MMM (5) | 8 |
| 2025 | Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text RecognitionabstractUnsupervised domain adaptation (UDA) has become increasingly prevalent in scene text recognition (STR), especially where training and testing data reside in different domains. The efficacy of existing UDA approaches tends to degrade when there is a large gap between the source and target domains. To deal with this problem, gradually shifting or progressively learning to shift from domain to domain is the key issue. In this paper, we introduce the Stratified Domain Adaptation (StrDA) approach, which examines the gradual escalation of the domain gap for the learning process. The objective is to partition the target data into subsets so that the progressively self-trained model can adapt to gradual changes. We stratify the target data by evaluating the proximity of each data sample to both the source and target domains. We propose a novel method for employing domain discriminators to estimate the out-of-distribution and domain discriminative levels of data samples. Extensive experiments on benchmark scene-text datasets show that our approach significantly improves the performance of baseline (source-trained) STR models. The source code is available at https://github.com/KhaLee2307/StrDA. Kha Nhat Le, Hoang-Tuan Nguyen, Hung Tien Tran, Thanh Duc Ngo |
WACV | 4 |
| 2023 | Unsupervised Domain Adaptation with Imbalanced Character Distribution for Scene Text RecognitionabstractRecent deep learning based methods have demonstrated promising results in scene text recognition. One of the major difficulty is the lack of manually annotated data. Synthetic data are then used to eliminate the requirement for human annotation. However, the domain gap between synthetic and real-world data remains a challenging issue. To bridge the gap, unsupervised domain adaptation (UDA) was introduced to transfer knowledge from a labeled source domain to a target domain. In this work, we introduce an unsupervised domain adaptation method based on a sequence-to-sequence attention model. We take into account imbalanced distribution of characters to optimize the adaptation process. We propose to use focal loss as the classification loss for the labeled source domain and focal entropy as the entropy loss for the unlabeled target domain. Our proposed method, named ICD-DA, outperforms other UDA methods on official benchmarks. Hung Tran Tien, Thanh Duc Ngo |
ICIP | 2 |
| 2023 | Abstraction-perception preserving cartoon face synthesis
Sy-Tuyen Ho, Manh-Khanh Ngo Huu, Thanh-Danh Nguyen, Nguyen Phan, Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 6 |
| 2023 | Instance-Level Few-Shot Learning With Class Hierarchy MiningabstractFew-shot learning is proposed to tackle the problem of scarce training data in novel classes. However, prior works in instance-level few-shot learning have paid less attention to effectively utilizing the relationship between categories. In this paper, we exploit the hierarchical information to leverage discriminative and relevant features of base classes to effectively classify novel objects. These features are extracted from abundant data of base classes, which could be utilized to reasonably describe classes with scarce data. Specifically, we propose a novel superclass approach that automatically creates a hierarchy considering base and novel classes as fine-grained classes for few-shot instance segmentation (FSIS). Based on the hierarchical information, we design a novel framework called Soft Multiple Superclass (SMS) to extract relevant features or characteristics of classes in the same superclass. A new class assigned to the superclass is easier to classify by leveraging these relevant features. Besides, in order to effectively train the hierarchy-based-detector in FSIS, we apply the label refinement to further describe the associations between fine-grained classes. The extensive experiments demonstrate the effectiveness of our method on FSIS benchmarks. The source code is available here: https://github.com/nvakhoa/superclass-FSIS. Anh-Khoa Nguyen Vu, Thanh-Toan Do, Nhat-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Tam V. Nguyen 0002 |
IEEE Trans. Image Process. | 5 |
| 2022 | UIT at VBS 2022: An Unified and Interactive Video Retrieval System with Temporal Search
Khanh Ho, Vu Xuan Dinh, Khiem Le 0001, Khang Dinh Tran, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le |
MMM (2) | 8 |
| 2022 | Few-shot object detection via baby learning
Anh-Khoa Nguyen Vu, Nhat-Duy Nguyen, Khanh-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
Image Vis. Comput. | 5 |
| 2021 | Dictionary-Guided Scene Text RecognitionabstractLanguage prior plays an important role in the way humans detect and recognize text in the wild. Current scene text recognition methods do use lexicons to improve recognition performance, but their naive approach of casting the output into a dictionary word based purely on the edit distance has many limitations. In this paper, we present a novel approach to incorporate a dictionary in both the training and inference stage of a scene text recognition system. We use the dictionary to generate a list of possible outcomes and find the one that is most compatible with the visual appearance of the text. The proposed method leads to a robust scene text recognition model, which is better at handling ambiguous cases encountered in the wild, and improves the overall performance of state-of-the-art scene text spotting frameworks. Our work suggests that incorporating language prior is a potential approach to advance scene text detection and recognition methods. Besides, we contribute VinText, a challenging scene text dataset for Vietnamese, where some characters are equivocal in the visual form due to accent symbols. This dataset will serve as a challenging benchmark for measuring the applicability and robustness of scene text detection and recognition algorithms. Code and dataset are available at https://github.com/VinAIResearch/dict-guided. Thu Nguyen 0003, Vinh Tran 0005, Minh-Triet Tran, Thanh Duc Ngo, Thien Huu Nguyen, Minh Hoai |
CVPR | 5 |
| 2018 | Video Search Based on Semantic Extraction and Locally Regional Object Proposal
Thanh-Dat Truong, Vinh-Tiep Nguyen, Minh-Triet Tran, Trang-Vinh Trieu, Tien Do, Thanh Duc Ngo, Duy-Dinh Le |
MMM (2) | 6 |
| 2017 | Evaluation of Deep Models for Real-Time Small Object Detection
Phuoc Pham, Tien Do, Thanh Duc Ngo, Duy-Dinh Le |
ICONIP (3) | 4 |
| 2017 | Semantic Extraction and Object Proposal for Video Search
Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Minh-Triet Tran, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 2 |
| 2017 | Efficient large-scale multi-class image classification by learning balanced trees
Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
Comput. Vis. Image Underst. | 2 |
| 2017 | Scalable Face Track Retrieval in Video Archives Using Bag-of-Faces Sparse RepresentationabstractHuge video archives consisting of news programs, dramas, movies, and Web videos (e.g., YouTube) are available in our daily life. In all these videos, human is usually one of the most important subjects. Using state-of-the-art techniques, we can efficiently detect and track faces in the videos. In order to organize large-scale face tracks, containing sequences of (detected) consecutive faces in the videos, we propose an efficient method to retrieve human face tracks using bag-of-faces sparse representation (BoF-SR). Using the proposed method, a face track is encoded as a single BoF-SR, therefore allowing an efficient indexing method to handle large-scale data. To further consider the possible variations in face tracks, we generalize our method to find multiple SRs, in an unsupervised manner, to represent a bag of faces and balance the tradeoff between performance and retrieval time. The experimental results on two real-world (million-scale) data sets confirm that the proposed methods achieve significant performance gains compared with different state-of-the-art methods. Bor-Chun Chen, Yan-Ying Chen, Yin-Hsi Kuo, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Winston H. Hsu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Efficient Large Scale Image Classification via Prediction Score Decomposition
Duy-Dinh Le, Tien-Dung Mai, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong |
ECCV (6) | 4 |
| 2016 | Using node relationships for hierarchical classificationabstractHierarchical classification is a computational efficient approach for large-scale image classification. The main challenging issue of this approach is to deal with error propagation. Irrelevant branching decision made at a parent node cannot be corrected at its child nodes in traversing the tree for classification. This paper presents a novel approach to reduce branching error at a node by taking its relative relationship into account. Given a node on the tree, we model each candidate branch by considering classification response of its child nodes, grandchild nodes and their differences with siblings. A maximum margin classifier is then applied to select the most discriminating candidate. Our proposed approach outperforms related approaches on Caltech-256, SUN-397 and ILSVRC2010-1K. Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
ICIP | 2 |
| 2016 | News Archive Exploration Combining Face Detection and Tracking with Network Visual AnalyticsabstractVisual analytics helps analytical reasoning and exploration of complex systems, for which it combines the means of interactive visualization, with the power of data analytics. The recent progress in computer vision techniques opens wide applications in real world video archives. Particularly, recent advances in face detection and recognition have been put under the spotlight. The applications of such techniques are often concern intelligence, or peer recognition in photo posted in social networks. We propose to combine those two domains by demonstrating a visual exploration of over a decade of the news program from the Japanese broadcaster NHK News 7. We derive social networks from face detection and tracking of this large dataset. With the help of a little domain knowledge, we monitor the activity of political public figures and explore the archive. This allows understanding and comparison of the politico-media scene presented by NHK under different Prime Minister's governance. The social networks are interactive, and also allow to explore the multimedia database and explore its video content. Benjamin Renoust, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ACM Multimedia | 2 |
| 2015 | Transfer AdaBoost SVM for Link Prediction in Newly Signed Social Networks using Explicit and PNR FeaturesabstractIn signed social network, the user-generated content and interactions have overtaken the web. Questions of whom and what to trust has become increasingly important. We must have methods which predict the signs of links in the social network to solve this problem. We study signed social networks with positive links (friendship, fan, like, etc) and negative links (opposition, anti-fan, dislike, etc). Specifically, we focus how to effectively predict positive and negative links in newly signed social networks. With SVM model, the small amount of edge sign information in newly signed network is not adequate to train a good classifier. In this paper, we introduce an effective solution to this problem. We present a novel transfer learning framework is called Transfer AdaBoost with SVM (TAS) which extends boosting-based learning algorithms and incorporates properly designed RBFSVM (SVM with the RBF kernel) component classifiers. With our framework, we use explicit topological features and Positive Negative Ratio (PNR) features which are based on decision-making theory. Experimental results on three networks (Epinions, Slashdot and Wiki) demonstrate our method that can improve the prediction accuracy by 40% over baseline methods. Additionally, our method has faster performance time. Thu Nguyen 0003, Phuc Quang Nguyen, Thanh Duc Ngo, Tu-Anh Nguyen-Hoang |
KES | 3 |
| 2015 | NII-UIT Browser: A Multimodal Video Search System
Thanh Duc Ngo, Vinh-Tiep Nguyen, Vu Hoang Nguyen, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 1 |
| 2015 | AttRel: An Approach to Person Re-Identification by Exploiting Attribute Relationships
Ngoc-Bao Nguyen, Vu Hoang Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong |
MMM (2) | 3 |
| 2015 | Human Action recognition from depth videos using multi-projection based representationabstractIn this paper, a novel method for human action recognition from depth videos is proposed. We project 3D data on to multiple 2D-planes from which dense trajectories features are extracted. In the training stage, for each projection, a classifier is trained using the training data. In the testing stage, for each test video, the multiple trained classifiers are applied and the predicted scores are combined for final decision. We propose a greedy-based method to select a subset of the trained classifiers for optimal combination. Experiments on the MSR Action 3D dataset show that the proposed method outperforms the baseline method that does not use multi-projection-based features. Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong |
MMSP | 2 |
| 2015 | Large scale multi-class classification using latent classifiersabstractWe study the problem of multi-class image classification with large number of classes, of which the one-vs-all based approach is prohibitive in practical applications. Recent state-of-the-art approaches rely on label tree to reduce classification complexity. However, building optimal tree structures and learning precise classifiers to optimize tree loss is challenging. In this paper, we introduce a novel approach using latent classifiers that can achieve comparable speed but better performance. The key idea is that instead of using C one-vs-all classifiers (C is the number of classes) to generate the score matrix for label prediction, a much smaller number of classifiers are used. These classifiers, called latent classifiers, are generated by analyzing the correlation among classes and removing redundancy. Experiments on several large datasets including ImageNet-1K, SUN-397, and Caltech-256 show the efficiency of our approach. Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
MMSP | 2 |
| 2015 | Cross-View Action Recognition by Projection-Based Augmentation
Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong |
PSIVT | 2 |
| 2014 | Integrating Spatial Information into Inverted Index for Large-Scale Image RetrievalabstractIn recent years, large-scale image retrieval has been shown remarkable potential in real-life applications. To reduce retrieval time as searched database may contain thousands of images, Inverted Indexing is the basic technique, given images are represented by Bag-of-Words model. However, one major limitation of both standard Inverted Index and Bag-of-Words model is that they ignore spatial information of the visual words in images. This might reduce retrieval accuracy. In this paper, we introduce an approach to integrate spatial information into inverted index to improve accuracy while maintaining short retrieval time. Experiments conducted on several benchmark datasets (Oxford Building 5K, Paris 6K and Oxford Building 5K+100K) demonstrate the effectiveness of our proposed approach. Bien-Van Nguyen, Duy Pham, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong |
ISM | 3 |
| 2014 | NII-UIT: A Tool for Known Item Search by Sequential Pattern Filtering
Thanh Duc Ngo, Vu Hoang Nguyen, Vu Lam, Sang Phan Le, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 1 |
| 2013 | NII-UIT-VBS: A Video Browsing Tool for Known Item Search
Duy-Dinh Le, Vu Lam, Thanh Duc Ngo, Vinh Quang Tran, Vu Hoang Nguyen, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 3 |
| 2012 | Robust eye localization in video by combining eye detector and eye tracker
Chi Nhan Duong, Thang Cap Pham Dinh, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Bac Le, Shin'ichi Satoh 0001 |
ICPR | 3 |
| 2011 | Boosting global scene classification accuracy by discriminative region localizationabstractCombining global scene classification with object detection has helped in improving the classification accuracy. However, training an object detector requires a large amount of manual annotation. The object detector may also fail when the object is occluded. Meanwhile, the presence of the object is not only indicated by the entire object region but any of its parts or its correlations with other regions in the image. To overcome these limitations, we propose using discriminative region localization instead of object detection in the combination. Our contribution is two-fold, a) a complete framework that combines global scene classification with discriminative region localization for image classification and b) a weakly supervised discriminative region localization approach that utilizes spatial context to improve the learning accuracy. Our experimental results on benchmark datasets demonstrated that the proposed discriminative region localization approach outperforms the state-of-the-art approach. In addition, the combination significantly increases the classification performance. Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ICIP | 1 |
| 2011 | Fast face sequence matching in large-scale video databasesabstractThere have recently been many methods proposed for matching face sequences in the field of face retrieval. However, most of them have proven to be inefficient in large-scale video databases because they frequently require a huge amount of computational cost to obtain a high degree of accuracy. We present an efficient matching method that is based on the face sequences (called face tracks) in large-scale video databases. The key idea is how to capture the distribution of a face track in the fewest number of low-computational steps. In order to do that, each face track is represented by a vector that approximates the first principal component of the face track distribution and the similarity of face tracks bases on the similarity of these vectors. Our experimental results from a large-scale database of 457,320 human faces extracted from 370 hours of TRECVID videos from 2004-2006 show that the proposed method easily handles the scalability by maintaining a good balance between the speed and the accuracy. Hung Thanh Vu, Thanh Duc Ngo, Thao Ngoc Nguyen, Duy-Dinh Le, Shin'ichi Satoh 0001, Bac Le, Duc Anh Duong |
ICIP | 2 |
| 2008 | A text segmentation based approach to video shot boundary detectionabstractVideo shot boundary detection is one of the fundamental tasks of video indexing and retrieval applications. Although many methods have been proposed for this task, finding a general and robust shot boundary method that is able to handle the various transition types caused by photo flashes, rapid camera movement and object movement is still challenging. We present a novel approach for detecting video shot boundaries in which we cast the problem of shot boundary detection into the problem of text segmentation in natural language processing. This is possible by assuming that each frame is a word and then the shot boundaries are treated as text segment boundaries (e.g. topics). The text segmentation based approaches in natural language processing can be used. The experimental results from various long video sequences have proved the effectiveness of our approach. Duy-Dinh Le, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong |
MMSP | 3 |