EDBT 2026 Demo / reviewers in the wild / expert
Vinh-Tiep Nguyen
dblp:60/11111
· DBLP profile ↗
33ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0003-4260-7874ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FA-Seg: A fast and accurate diffusion-based method for open-vocabulary segmentation
Quang Huy Che, Vinh-Tiep Nguyen |
Neurocomputing | 2 |
| 2026 | DisenID: Identity-preserving disentangled personalization for multi-subject generation
Gia-Nghia Tran, Quang Huy Che, Trong-Tai Dam Vu, Bich-Nga Pham, Vinh-Tiep Nguyen, Trung-Nghia Le, Minh-Triet Tran |
Neurocomputing | 5 |
| 2025 | A Generative Approach at the Instance-Level for Image Segmentation Under Limited Training Data Conditions (Student Abstract)abstractHigh-accuracy image segmentation models require abundant training annotated data which is costly for pixel-level annotations. Our work addresses a high-cost manual annotating process or the lack of detailed annotations via a generative approach. In particular, our approach (1) proposes the conditional instance-level synthesis to enrich the limited data to enhance the segmentation performance, and (2) employs the generative architectures to complete the segmentation task under few-shot learning concepts. The initial results on the Cityscapes benchmark emphasize our potential generative solution on the instance segmentation task given limited data. Thanh-Danh Nguyen, Vinh-Tiep Nguyen, Tam V. Nguyen 0002 |
AAAI | 2 |
| 2025 | FaR: Enhancing Multi-concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
Gia-Nghia Tran, Quang Huy Che, Trong-Tai Dam Vu, Bich-Nga Pham, Vinh-Tiep Nguyen, Trung-Nghia Le, Minh-Triet Tran |
ICCCI (2) | 5 |
| 2025 | Multi-Perspective Data Augmentation for Few-shot Object DetectionabstractRecent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of $17.5\%$ in nAP50 over the baseline on PASCAL VOC. Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
ICLR | 3 |
| 2025 | Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance
Quang Huy Che, Duc-Tri Le, Bich-Nga Pham, Duc Khai Lam, Vinh-Tiep Nguyen |
ICPRAM | 5 |
| 2025 | A Re-Ranking Method Using K-Nearest Weighted Fusion for Person Re-IdentificationabstractIn person re-identification, re-ranking is a crucial step to enhance the overall accuracy by refining the initial ranking of retrieved results. Previous studies have mainly focused on features from single-view images, which can cause view bias and issues like pose variation, viewpoint changes, and occlusions. Using multi-view features to present a person can help reduce view bias. In this work, we present an efficient re-ranking method that generates multi-view features by aggregating neighbors' features using K-nearest Weighted Fusion (KWF) method. Specifically, we hypothesize that features extracted from re-identification models are highly similar when representing the same identity. Thus, we select K neighboring features in an unsupervised manner to generate multi-view features. Additionally, this study explores the weight selection strategies during feature aggregation, allowing us to identify an effective strategy. Our re-ranking approach does not require model fine-tuning or extra annotations, making it applicable to large-scale datasets. We evaluate our method on the person re-identification datasets Market1501, MSMT17, and Occluded-DukeMTMC. The results show that our method significantly improves Rank@1 and mAP when re-ranking the top M candidates from the initial ranking results. Specifically, compared to the initial results, our re-ranking method achieves improvements of 9.8%/22.0% in Rank@1 on the challenging datasets: MSMT17 and Occluded-DukeMTMC, respectively. Furthermore, our approach demonstrates substantial enhancements in computational efficiency compared to other re-ranking methods. Code is available at https://github.com/chequanghuy/Enhancing-Person-Re-Identification-via-UFFM-and-AMC. Quang Huy Che, Le-Chuong Nguyen, Gia-Nghia Tran, Dinh-Duy Phan, Vinh-Tiep Nguyen |
ICPRAM | 5 |
| 2025 | IMSearch 2.0: Toward User-Centric and Efficient Interactive Multimedia Retrieval System
Duc-Tuan Luu, Khanh-An C. Quan, Duy-Ngoc Nguyen, Khanh-Linh Bui-Le, Nhat-Sang Doan, Minh-Duc Le-Ngo, Vinh-Tiep Nguyen, Minh-Triet Tran |
MMM (5) | 7 |
| 2025 | SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Viet-Tham Huynh, Hoang-Phuc Nguyen, Long Bao Le, Thai Hoang Minh, Minh Nguyen Anh, Thang Nguyen Tien, Phat Nguyen Thuan, Huy Nguyen Phong, Bao Huynh Thai, Vinh-Tiep Nguyen, Duc-Vu Nguyen, Phu-Hoa Pham, Minh-Huy Le-Hoang, Nguyen-Khang Le, Minh-Chinh Nguyen, Minh-Quan Ho, Ngoc-Long Tran, Hien-Long Le-Hoang, Man-Khoi Tran, Anh-Duong Tran, Quan Nguyen Hung, Dat Phan Thanh, Hoang Tran Van, Tien Huynh Viet, Nhan Nguyen Viet Thien, Dinh-Khoi Vo, Van-Loc Nguyen, Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Triet Tran |
Comput. Graph. | 12 |
| 2025 | Few-shot object detection via synthetic features with optimal transport
Anh-Khoa Nguyen Vu, Thanh-Toan Do, Vinh-Tiep Nguyen, Tam Le, Minh-Triet Tran, Tam V. Nguyen 0002 |
Comput. Vis. Image Underst. | 3 |
| 2025 | Enhancing person re-identification via Uncertainty Feature Fusion Method and Auto-weighted Measure Combination
Quang Huy Che, Le-Chuong Nguyen, Duc-Tuan Luu, Vinh-Tiep Nguyen |
Knowl. Based Syst. | 4 |
| 2024 | Adaptive Scheme of Clustering-Based Unsupervised Learning for Person Re-identification
Anh-Vu Vo Duy, Quang Huy Che, Vinh-Tiep Nguyen |
ACIIDS (2) | 3 |
| 2024 | IMSearch: An Interactive Multimedia Video-Moment Search SystemabstractIn recent years, video content has become increasingly popular due to the development of technologies, especially for recording devices. This explosion has driven the need for an advanced video-moment retrieval system that can accurately search for specific video segments matching the intentions of the user's queries. A prominent challenge in this field lies in the multimedia nature of the video data, which includes visual, auditory and even textual information. The accuracy and relevance of search results depend on how efficiently a retrieval system processes these types of metadata. In this paper, we introduce IMSearch, an interactive multimedia search system that can retrieve precise video-moment content by executing queries across various information types, such as text, audio, location of the objects, sketch, and human pose. Additionally, we implement a re-ranking mechanism based on user feedback to simultaneously improve and optimize the performance of our IMSearch system. We also explore two different vector libraries (FAISS and HNSW) for vector searching and report the retrieval accuracy and executed time to demonstrate our proposed paradigm. Duc-Tuan Luu, Duy-Ngoc Nguyen, Khanh-Linh Bui-Le, Vinh-Tiep Nguyen, Minh-Triet Tran |
CBMI | 4 |
| 2024 | Nighttime scene understanding with label transfer scene parser
Thanh-Danh Nguyen, Nguyen Phan, Tam V. Nguyen 0002, Vinh-Tiep Nguyen, Minh-Triet Tran |
Image Vis. Comput. | 4 |
| 2024 | Local part attention for image stylization with text prompt
Quoc-Truong Truong, Vinh-Tiep Nguyen, Lan-Phuong Nguyen, Hung-Phu Cao, Duc-Tuan Luu |
Neural Comput. Appl. | 2 |
| 2023 | TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 11 |
| 2023 | SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 11 |
| 2023 | Abstraction-perception preserving cartoon face synthesis
Sy-Tuyen Ho, Manh-Khanh Ngo Huu, Thanh-Danh Nguyen, Nguyen Phan, Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 5 |
| 2023 | Instance-Level Few-Shot Learning With Class Hierarchy MiningabstractFew-shot learning is proposed to tackle the problem of scarce training data in novel classes. However, prior works in instance-level few-shot learning have paid less attention to effectively utilizing the relationship between categories. In this paper, we exploit the hierarchical information to leverage discriminative and relevant features of base classes to effectively classify novel objects. These features are extracted from abundant data of base classes, which could be utilized to reasonably describe classes with scarce data. Specifically, we propose a novel superclass approach that automatically creates a hierarchy considering base and novel classes as fine-grained classes for few-shot instance segmentation (FSIS). Based on the hierarchical information, we design a novel framework called Soft Multiple Superclass (SMS) to extract relevant features or characteristics of classes in the same superclass. A new class assigned to the superclass is easier to classify by leveraging these relevant features. Besides, in order to effectively train the hierarchy-based-detector in FSIS, we apply the label refinement to further describe the associations between fine-grained classes. The extensive experiments demonstrate the effectiveness of our method on FSIS benchmarks. The source code is available here: https://github.com/nvakhoa/superclass-FSIS. Anh-Khoa Nguyen Vu, Thanh-Toan Do, Nhat-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Tam V. Nguyen 0002 |
IEEE Trans. Image Process. | 4 |
| 2022 | CDC: Color-Based Diffusion Model with Caption Embedding in VBS 2022
Duc-Tuan Luu, Khanh-An C. Quan, Thinh-Quyen Nguyen, Van-Son Hua, Minh-Chau Nguyen, Minh-Triet Tran, Vinh-Tiep Nguyen |
MMM (2) | 7 |
| 2022 | Few-shot object detection via baby learning
Anh-Khoa Nguyen Vu, Nhat-Duy Nguyen, Khanh-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
Image Vis. Comput. | 4 |
| 2020 | Flood Level Prediction via Human Pose Estimation from Social Media ImagesabstractFloods are the most common natural and among the most dangerous disasters in the world. It is important to get up-to-date information about flooding and the flood level for flood preparation and prevention. In this paper, we propose an efficient method to determine the flood level from daily activity photos on social media. Our method is based on the idea of matching the water level with human pose to determine the level of severity of flooding. Extensive experiments conducted on the dataset of Multimodal Flood Level Estimation show the superiority of our proposed method. We achieve the first rank in MediaEval 2019 and this demonstrates the potential applications of our method to analyze flood information. Khanh-An C. Quan, Vinh-Tiep Nguyen, Tan-Cong Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran |
ICMR | 2 |
| 2019 | Enhancing Endoscopic Image Classification with Symptom Localization and Data AugmentationabstractInspired by recent advances in computer vision and deep learning, we propose new enhancements to tackle problems appearing in endoscopic image analysis, especially abnormality finding and anatomical landmark detection. In details, a combination of Residual Neural Network and Faster R-CNN are jointly applied in order to take all of their advantages and improve the overall performance. Nevertheless, novel data augmentation is designed and adapted to corresponding domains. Our approaches prove their competitive results in term of not only the accuracy but also the inference time in Medico: The 2018 Multimedia for Medicine Task and The Biomedia ACM MM Grand Challenge 2019. These results show the great potential of the collaborating between deep learning models and data augmentation in medical image analysis applications. Especially, more than 4900 bounding boxes localizing the symptom of some classes from KVASIR dataset that we annotated and used in this project are shared online for future research. Trung-Hieu Hoang, Hai-Dang Nguyen, Thanh-An Nguyen, Vinh-Tiep Nguyen, Minh-Triet Tran |
ACM Multimedia | 5 |
| 2018 | Lightweight Deep Convolutional Network for Tiny Object Recognition
Thanh-Dat Truong, Vinh-Tiep Nguyen, Minh-Triet Tran |
ICPRAM | 2 |
| 2018 | Video Search Based on Semantic Extraction and Locally Regional Object Proposal
Thanh-Dat Truong, Vinh-Tiep Nguyen, Minh-Triet Tran, Trang-Vinh Trieu, Tien Do, Thanh Duc Ngo, Duy-Dinh Le |
MMM (2) | 2 |
| 2017 | Video Indexing, Search, Detection, and Description with Focus on TRECVIDabstractThere has been a tremendous growth in video data the last decade. People are using mobile phones and tablets to take, share or watch videos more than ever before. Video cameras are around us almost everywhere in the public domain (e.g. stores, streets, public facilities, ...etc). Efficient and effective retrieval methods are critically needed in different applications. The goal of TRECVID is to encourage research in content-based video retrieval by providing large test collections, uniform scoring procedures, and a forum for organizations interested in comparing their results. In this tutorial, we present and discuss some of the most important and fundamental content-based video retrieval problems such as recognizing predefined visual concepts, searching in videos for complex ad-hoc user queries, searching by image/video examples in a video dataset to retrieve specific objects, persons, or locations, detecting events, and finally bridging the gap between vision and language by looking into how can systems automatically describe videos in a natural language. A review of the state of the art, current challenges, and future directions along with pointers to useful resources will be presented by different regular TRECVID participating teams. Each team will present one of the following tasks: George Awad, Duy-Dinh Le, Chong-Wah Ngo, Vinh-Tiep Nguyen, Georges Quénot, Cees Snoek, Shin'ichi Satoh 0001 |
ICMR | 4 |
| 2017 | Semantic Extraction and Object Proposal for Video Search
Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Minh-Triet Tran, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 1 |
| 2016 | NowAndThen: a social network-based photo recommendation tool supporting reminiscenceabstractPeople frequently post their photos on social network sites (e.g. Facebook, Instagram) as a way to share memorable moments, emotions, or locations visited. While sharing photos with the same subjects as past photos could lead to user reminiscence and potentially create valuable benefits, existing products still cannot support this adequately. Based on a survey on user habits in sharing and revisiting photos on social network sites, we propose NowAndThen, a photo recommendation concept and tool that assists reminiscence when sharing photos on social network sites. By combining visual features of the photos and associated tags, a prototype we developed can recommend old photos that have common subjects with the user's current photos of interest. Our study with the prototype shows that this approach can help users positively revive past memories and connections with their friends. In addition, our results include various design insights and implications for future reminiscence-support systems. Vinh-Tiep Nguyen, Khanh-Duy Le, Minh-Triet Tran, Morten Fjeld |
MUM | 1 |
| 2015 | Adaptive WildNet Face network for detecting face in the wildabstractCombining Convolutional Neural Network and Deformable Part Models is a new trend in object detection area. Following this trend, we propose Adaptive WildNet Face network using Deformable Part Models structure to exploit advantages of two methods in face detection area. We evaluate the merit of our method on Face Detection Data Set and Benchmark. Experimental results show that our method achieves up to 86.22% true positive images in 1000 false positive images in FDDB. Our method becomes one of state-of-the-art methods in FDDB dataset and it opens a new way to detect faces of images in the wild. Dinh-Luan Nguyen, Vinh-Tiep Nguyen, Minh-Triet Tran, Atsuo Yoshitaka |
ICMV | 2 |
| 2015 | NII-UIT Browser: A Multimodal Video Search System
Thanh Duc Ngo, Vinh-Tiep Nguyen, Vu Hoang Nguyen, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 2 |
| 2015 | Query-adaptive late fusion with neural network for instance searchabstractBag-of-Word based model is one of the state-of-the-art approaches for object retrieval or also known as instance search problem. Although this model and its extensions are good for rich-textured objects, it is still unsolved for searching on textureless ones. In this paper, we propose to combine this model with Deformable Part Models object detector using late fusion technique to improve final result. To find the optimal weights for each type of query objects, we further propose to use a neural network to learn query features including object area, number of shared visual words to get optimal weights for each model. Experimental results on TRECVID Instance Search (INS) dataset with queries in INS2013 and INS2014 show that our proposed method significantly improves 18.48% and 14.63% in mAP respectively comparing to standard BOW model and outperform other state-of-the-art methods. This method opens a new way of adaptively combining DPM, an object detector, in a hybrid model for visual instance search. Vinh-Tiep Nguyen, Dinh-Luan Nguyen, Minh-Triet Tran, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMSP | 1 |
| 2015 | Deep Convolutional Neural Network in Deformable Part Models for Face Detection
Dinh-Luan Nguyen, Vinh-Tiep Nguyen, Minh-Triet Tran, Atsuo Yoshitaka |
PSIVT | 2 |
| 2012 | Realtime arbitrary-shaped template matching processabstractTemplate matching is one of the significant methods used in most computer vision applications. However, one disadvantage of existing template matching methods is that a template must be a rectangular image. Several solutions can be used to overcome this problem, such as feature-based, edge-based, etc., but most of them do not meet the requirements of practical applications because of their slow speed or low accuracy. In this paper, the authors propose a process for matching a template with an arbitrary shape using the idea of coarse-fine and refining interesting regions. Variety of measures can be used to provide different levels of performance for each practical situation. Experiments show that the process performs an impressive performance in average. Using the same measures, the proposed process works 5 times faster than the FFT template matching method (implemented in OpenCV), while achieving the similar accuracy. Moreover, most parts of the proposed process can be executed in parallel on multi-core CPUs or GPUs. Dai-Duong Truong, Vinh-Tiep Nguyen, Anh Duc Duong, Chau-Sang Nguyen Ngoc, Minh-Triet Tran |
ICARCV | 2 |