VLDB 2026 Research / reviewers in the wild / expert
Trung-Nghia Le
dblp:00/11111
· DBLP profile ↗
45ranked-venue papers
15as first author
35since 2021 · last 2026
0000-0002-7363-2610ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 12 first-author · 27 since 2021Artificial intelligence and machine learning · 21 · 8 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VietFashion: Benchmarking Sketch-Text Composed Image Retrieval for Cultural OutfitsabstractCultural garments pose a unique challenge for visual retrieval systems, as their identity often depends on subtle structural and symbolic details that are poorly captured by standard AI models. We introduce VietFashion, a new benchmark for sketch–text composed image retrieval centered on the Ao Dai, a traditional Vietnamese garment. VietFashion enables designers and researchers to retrieve culturally meaningful outfits using a combination of hand-drawn sketches, which convey garment structure, and textual descriptions, which encode cultural semantics. The dataset is initialized with 650 sketches and expanded using generative models to produce over 21,000 photorealistic images with aligned captions. Textual prompts that describe detailed outfit attributes, which are extracted from fashion magazines to ensure authenticity and diversity. To better reflect the inherent ambiguity of design intent, VietFashion adopts a multi-target retrieval setting, where a single query may correspond to multiple valid results. We establish standardized evaluation protocols and benchmark state-of-the-art composed image retrieval methods. Experimental results reveal significant performance gaps in modeling fine-grained cultural semantics and multi-modal composition, positioning VietFashion as a challenging benchmark for fine-grained fashion retrieval. The dataset is publicly available at: https://hng0303.github.io/VietFashion. Hoang-Nguyen Cao, Le-Hoang Bui, Dinh-Khoi Vo, Minh-Triet Tran, Trung-Nghia Le |
ICMR | 5 |
| 2026 | ShowFlow: From robust single concept to condition-free multi-concept generation
Trong-Vu Hoang, Quang-Binh Nguyen, Thanh-Toan Do, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
Neurocomputing | 6 |
| 2026 | DisenID: Identity-preserving disentangled personalization for multi-subject generation
Gia-Nghia Tran, Quang Huy Che, Trong-Tai Dam Vu, Bich-Nga Pham, Vinh-Tiep Nguyen, Trung-Nghia Le, Minh-Triet Tran |
Neurocomputing | 6 |
| 2026 | You can see me: Retrieval-augmented framework for divergent perception in animal art generation
Quoc-Duy Tran, Anh-Tuan Vo, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
Pattern Recognit. Lett. | 5 |
| 2025 | Toward Content-Based Indexing and Retrieval of Head and Neck CT With Abscess SegmentationabstractAbscesses in the head and neck represent an acute infectious process that can potentially lead to sepsis or mortality if not diagnosed and managed promptly. Accurate detection and delineation of these lesions on imaging are essential for diagnosis, treatment planning, and surgical intervention. In this study, we introduce AbscessHeNe, a curated and comprehensively annotated dataset comprising 4,926 contrastenhanced CT slices with clinically confirmed head and neck abscesses. The dataset is designed to facilitate the development of robust semantic segmentation models that can accurately delineate abscess boundaries and evaluate deep neck space involvement, thereby supporting informed clinical decision-making. To establish performance baselines, we evaluate several state-of-the-art segmentation architectures, including CNN, Transformer, and Mamba-based models. The highestperforming model achieved a Dice Similarity Coefficient of 0.39, Intersection-over-Union of 0.27, and Normalized Surface Distance of 0.67, indicating the challenges of this task and the need for further research. Beyond segmentation, AbscessHeNe is structured for future applications in content-based multimedia indexing and case-based retrieval. Each CT scan is linked with pixel-level annotations and clinical metadata, providing a foundation for building intelligent retrieval systems and supporting knowledge-driven clinical workflows. The dataset will be made publicly available at https://github.com/drthaodao3101/AbscessHeNe.git. Thao Thi Phuong Dao, Tan-Cong Nguyen, Trong-Le Do, Truong Hoang Viet, Nguyen Chi Thanh, Huynh Nguyen Thuan, Do Vo Cong Nguyen, Minh-Khoi Pham, Mai-Khiem Tran, Viet-Tham Huynh, Trung-Nghia Le, Thanh-Nhan Vo, Tam V. Nguyen 0002, Minh-Triet Tran, Thanh Dinh Le |
CBMI | 12 |
| 2025 | GenFlow: Interactive Modular System for Image GenerationabstractGenerative art unlocks boundless creative possibilities, yet its full potential remains untapped due to the technical expertise required for advanced architectural concepts and computational workflows. To bridge this gap, we present GenFlow, a novel modular framework that empowers users of all skill levels to generate images with precision and ease. Featuring a node-based editor for seamless customization and an intelligent assistant powered by natural language processing, GenFlow transforms the complexity of workflow creation into an intuitive and accessible experience. By automating deployment processes and minimizing technical barriers, our framework makes cutting-edge generative art tools available to everyone. A user study demonstrated GenFlow's ability to optimize workflows, reduce task completion times, and enhance user understanding through its intuitive interface and adaptive features. These results position GenFlow as a groundbreaking solution that redefines accessibility and efficiency in the realm of generative art. Duc-Hung Nguyen, Huu-Phuc Huynh, Minh-Triet Tran, Trung-Nghia Le |
CBMI | 4 |
| 2025 | Automated Image Recognition Framework
Quang-Binh Nguyen, Trong-Vu Hoang, Ngoc-Do Tran, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ICCCI (2) | 6 |
| 2025 | FaR: Enhancing Multi-concept Text-to-Image Diffusion via Concept Fusion and Localized Refinement
Gia-Nghia Tran, Quang Huy Che, Trong-Tai Dam Vu, Bich-Nga Pham, Vinh-Tiep Nguyen, Trung-Nghia Le, Minh-Triet Tran |
ICCCI (2) | 6 |
| 2025 | Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image ClassificationabstractWe present a unified vision-language framework tailored for ENT endoscopy image analysis that simultaneously tackles three clinically-relevant tasks: image classification, image-to-image retrieval, and text-to-image retrieval. Unlike conventional CNN-based pipelines that struggle to capture cross-modal semantics, our approach leverages the CLIP ViT-B/16 backbone and enhances it through Low-Rank Adaptation, multi-level CLS token aggregation, and spherical feature interpolation. These components collectively enable efficient fine-tuning on limited medical data while improving representation diversity and semantic alignment across modalities. To bridge the gap between visual inputs and textual diagnostic context, we introduce class-specific natural language prompts that guide the image encoder through a joint training objective combining supervised classification with contrastive learning. We validated our framework through participation in the ACM MM'25 ENTRep Grand Challenge, achieving 95% accuracy and F1-score in classification, Recall@1 of 0.93 and 0.92 for image-to-image and text-to-image retrieval respectively, and MRR scores of 0.97 and 0.96. Ablation studies demonstrated the incremental benefits of each architectural component, validating the effectiveness of our design for robust multimodal medical understanding in low-resource clinical settings. Y. Hop Nguyen, Doan Anh Phan Huu, Trung Thai Tran, Nhat Nam Mai, Van Toi Giap, Thao Thi Phuong Dao, Trung-Nghia Le |
ACM Multimedia | 7 |
| 2025 | ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian NormalizationabstractImage captioning systems often produce generic descriptions that fail to capture event-level semantics which are crucial for applications like news reporting and digital archiving. We present ReCap, a novel pipeline for event-enriched image retrieval and captioning that incorporates broader contextual information from relevant articles to generate narrative-rich, factually grounded captions. Our approach addresses the limitations of standard vision-language models that typically focus on visible content while missing temporal, social, and historical contexts. ReCap comprises three integrated components: (1) a robust two-stage article retrieval system using DINOv2 embeddings with global feature similarity for initial candidate selection followed by patch-level mutual nearest neighbor similarity re-ranking; (2) a context extraction framework that synthesizes information from article summaries, generic captions, and original source metadata; and (3) a large language model-based caption generation system with Semantic Gaussian Normalization to enhance fluency and relevance. Evaluated on the OpenEvents V1 dataset as part of Track 1 in the EVENTA 2025 Grand Challenge, ReCap achieved a strong overall score of 0.54666, ranking 2nd on the private test set. These results highlight ReCap's effectiveness in bridging visual perception with real-world knowledge, offering a practical solution for context-aware image understanding in high-stakes domains. The code is available at https://github.com/Noridom1/EVENTA2025-Event-Enriched-Image-Captioning. Thinh-Phuc Nguyen, Gia-Huy Dinh, Lam-Huy Nguyen, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 6 |
| 2025 | OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event GroundingabstractWe introduce OpenEvents V1, a large-scale benchmark dataset designed to advance event-centric vision-language understanding. Unlike conventional image captioning and retrieval datasets that focus on surface-level descriptions, OpenEvents V1 dataset emphasizes contextual and temporal grounding through three primary tasks: (1) generating rich, event-aware image captions, (2) retrieving event-relevant news articles from image queries, and (3) retrieving event-relevant images from narrative-style textual queries. The dataset comprises over 200,000 news articles and 400,000 associated images sourced from CNN and The Guardian, spanning diverse domains and time periods. We provide extensive baseline results and standardized evaluation protocols for all tasks. OpenEvents V1 establishes a robust foundation for developing multimodal AI systems capable of deep reasoning over complex real-world events. The dataset is publicly available at https://ltnghia.github.io/eventa/openevents-v1. Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 7 |
| 2025 | Streamlining Virtual KOL Generation Through Modular Generative AI ArchitectureabstractKey Opinion Leaders (KOLs) play a crucial role in modern marketing by shaping consumer perceptions and enhancing brand credibility. However, collaborating with human KOLs often involves high costs and logistical challenges. To address this, we present GenKOL, an interactive system that empowers marketing professionals to efficiently generate high-quality virtual KOL images using generative AI. GenKOL enables users to dynamically compose promotional visuals through an intuitive interface that integrates multiple AI capabilities, including garment generation, makeup transfer, background synthesis, and hair editing. These capabilities are implemented as modular, interchangeable services that can be deployed flexibly on local machines or in the cloud. This modular architecture ensures adaptability across diverse use cases and computational environments. Our system can significantly streamline the production of branded content, lowering costs and accelerating marketing workflows through scalable virtual KOL creation. Video demo is available at https://youtu.be/uXpXmEbjg3M. Tan-Hiep To, Duy-Khang Nguyen, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 4 |
| 2025 | Event-Enriched Image Analysis Grand Challenge At ACM Multimedia 2025abstractThe Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses this gap by integrating contextual, temporal, and semantic information to capture the who, when, where, what, and why behind an image. Built upon the OpenEvents V1 dataset, the challenge features two tracks: Event-Enriched Image Retrieval and Captioning, and Event-Based Image Retrieval. A total of 45 teams from six countries participated, with evaluation conducted through Public and Private Test phases to ensure fairness and reproducibility. The top three teams were invited to present their solutions at ACM Multimedia 2025. EVENTA establishes a foundation for context-aware, narrative-driven multimedia AI, with applications in journalism, media analysis, cultural archiving, and accessibility. Further details about the challenge are available at the official homepage: https://ltnghia.github.io/eventa/eventa-2025. Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le |
ACM Multimedia | 10 |
| 2025 | Advancing Fashion Design Through Intelligent Sketchpad StudioabstractLine sketches serve as the visual DNA of fashion design, forming the essential foundation where concepts take shape, yet today's digital tools often lack the fluidity, personalization, and intelligence needed to truly support this creative process. We present FashSketch, an interactive, multimedia-driven system that reimagines fashion sketching through the lens of generative AI. Designed with a layer-based creative interface, FashSketch empowers designers to ideate, customize, and iterate on sketches seamlessly. By integrating state-of-the-art generative models, sketch-based retrieval, and large language models, the system supports advanced functionalities such as text-to-sketch generation and context-aware sketch recommendation. FashSketch not only enhances the sketching experience but also opens new multimodal pathways for creative expression, making it a powerful co-creative partner in the early stages of fashion design. Video demo is available at https://youtu.be/BX-Edz7Z7ZY. Nhu-Binh Nguyen Truc, Nhu-Vinh Hoang, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 5 |
| 2025 | EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic CaptionsabstractEvent-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and real-world knowledge. Conventional vision-language retrieval approaches often fall short when captions describe abstract events, implicit causality, temporal context, or contain long, complex narratives. To tackle these issues, we introduce a multi-stage retrieval framework combining dense article retrieval, event-aware language model reranking, and efficient image collection, followed by caption-guided semantic matching and rank-aware selection. We leverage Qwen3 for article search, Qwen3-Reranker for contextual alignment, and Qwen2-VL for precise image scoring. To further enhance performance and robustness, we fuse outputs from multiple configurations using Reciprocal Rank Fusion (RRF). Our system achieves the top-1 score on the private test set of Track 2 in the EVENTA 2025 Grand Challenge, demonstrating the effectiveness of combining language-based reasoning and multimodal retrieval for complex, real-world image understanding. The code is available at https://github.com/vdkhoi20/EVENT-Retriever. Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 4 |
| 2025 | CamoFA: A Learnable Fourier-Based Augmentation for Camouflage SegmentationabstractCamouflaged object detection (COD) and camouflaged instance segmentation (CIS) aim to recognize and segment objects that are blended into their surroundings, respectively. While several deep neural network models have been proposed to tackle those tasks, augmentation methods for COD and CIS have not been thoroughly explored. strategies can help improve models' performance by increasing the size and diversity of the training data and exposing the model to a wider range of variations in the data. Besides, we aim to automatically learn transformations that help to reveal the underlying structure of camou-flaged objects and allow the model to learn to identify better and segment camouflaged objects. To achieve this, we pro-pose a learnable augmentation method in the frequency domain for COD and CIS via the Fourier transform approach, dubbed CamoFA. Our method leverages a conditional generative adversarial network and cross-attention mechanism to generate a reference image and an adaptive hybrid swapping with parameters to mix the low-frequency component of the reference image and the high-frequency component of the input image. This approach aims to make camouflaged objects more visible for detection and segmentation models. Without bells and whistles, our proposed augmentation method boosts the performance of camouflaged object detectors and instance segmenters by large margins. Minh-Quan Le, Minh-Triet Tran, Trung-Nghia Le, Tam V. Nguyen 0002, Thanh-Toan Do |
WACV | 3 |
| 2025 | SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Viet-Tham Huynh, Hoang-Phuc Nguyen, Long Bao Le, Thai Hoang Minh, Minh Nguyen Anh, Thang Nguyen Tien, Phat Nguyen Thuan, Huy Nguyen Phong, Bao Huynh Thai, Vinh-Tiep Nguyen, Duc-Vu Nguyen, Phu-Hoa Pham, Minh-Huy Le-Hoang, Nguyen-Khang Le, Minh-Chinh Nguyen, Minh-Quan Ho, Ngoc-Long Tran, Hien-Long Le-Hoang, Man-Khoi Tran, Anh-Duong Tran, Quan Nguyen Hung, Dat Phan Thanh, Hoang Tran Van, Tien Huynh Viet, Nhan Nguyen Viet Thien, Dinh-Khoi Vo, Van-Loc Nguyen, Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Triet Tran |
Comput. Graph. | 31 |
| 2025 | GUNNEL: guided mixup augmentation and multi-model fusion for aquatic animal segmentation
Minh-Quan Le, Trung-Nghia Le, Tam V. Nguyen 0002, Isao Echizen, Minh-Triet Tran |
Neural Comput. Appl. | 2 |
| 2024 | MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance SegmentationabstractFew-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, known as point estimation. However, this mechanism depends on prototypes (e.g. mean of K-shot) for prediction, leading to performance instability. To overcome the disadvantage of the point estimation mechanism, we propose a novel approach, dubbed MaskDiff, which models the underlying conditional distribution of a binary mask, which is conditioned on an object region and K-shot information. Inspired by augmentation approaches that perturb data with Gaussian noise for populating low data density regions, we model the mask distribution with a diffusion probabilistic model. We also propose to utilize classifier-free guided mask sampling to integrate category information into the binary mask generation process. Without bells and whistles, our proposed method consistently outperforms state-of-the-art methods on both base and novel classes of the COCO dataset while simultaneously being more stable than existing methods. The source code is available at: https://github.com/minhquanlecs/MaskDiff. Minh-Quan Le, Tam V. Nguyen 0002, Trung-Nghia Le, Thanh-Toan Do, Minh N. Do, Minh-Triet Tran |
AAAI | 3 |
| 2024 | CrossPAR: Enhancing Pedestrian Attribute Recognition with Vision-Language Fusion and Human-Centric Pre-training
Bach-Hoang Ngo, Si-Tri Ngo, Phu-Duc Le, Quang-Minh Phan, Minh-Triet Tran, Trung-Nghia Le |
ACCV (6) | 6 |
| 2024 | Rethinking Sampling for Music-Driven Long-Term Dance Generation
Tuong-Vy Truong-Thuy, Gia-Cat Bui-Le, Hai-Dang Nguyen, Trung-Nghia Le |
ACCV (5) | 4 |
| 2024 | NearbyPatchCL: Leveraging Nearby Patches for Self-supervised Patch-Level Multi-class Classification in Whole-Slide Images
Gia-Bao Le, Van-Tien Nguyen, Trung-Nghia Le, Minh-Triet Tran |
MMM (2) | 3 |
| 2024 | Artificial intelligence for laryngoscopy in vocal fold diseases: a review of dataset, technology, and ethics
Thao Thi Phuong Dao, Tan-Cong Nguyen, Viet-Tham Huynh, Xuan-Hai Bui, Trung-Nghia Le, Minh-Triet Tran |
Mach. Learn. | 5 |
| 2023 | Efficient 3D Brain Tumor Segmentation with Axial-Coronal-Sagittal Embedding
Tuan-Luc Huynh, Thanh-Danh Le, Tam V. Nguyen 0002, Trung-Nghia Le, Minh-Triet Tran |
PSIVT | 4 |
| 2023 | Cluster-Based Video Summarization with Temporal Context Awareness
Hai-Dang Huynh-Lam, Ngoc-Phuong Ho-Thi, Minh-Triet Tran, Trung-Nghia Le |
PSIVT | 4 |
| 2023 | Analysis of Master Vein Attacks on Finger Vein Recognition SystemsabstractFinger vein recognition (FVR) systems have been commercially used, especially in ATMs, for customer verification. Thus, it is essential to measure their robustness against various attack methods, especially when a handcrafted FVR system is used without any countermeasure methods. In this paper, we are the first in the literature to introduce master vein attacks in which we craft a vein-looking image so that it can falsely match with as many identities as possible by the FVR systems. We present two methods for generating master veins for use in attacking these systems. The first uses an adaptation of the latent variable evolution algorithm with a proposed generative model (a multi-stage combination of β-VAE and WGAN-GP models). The second uses an adversarial machine learning attack method to attack a strong surrogate CNN-based recognition system. The two methods can be easily combined to boost their attack ability. Experimental results demonstrated that the proposed methods alone and together achieved false acceptance rates up to 73.29% and 88.79%, respectively, against Miura’s hand-crafted FVR system. We also point out that Miura’s system is easily compromised by non-vein-looking samples generated by a WGAN-GP model with false acceptance rates up to 94.21%. The results raise the alarm about the robustness of such systems and suggest that master vein attacks should be considered an important security measure. Huy H. Nguyen, Trung-Nghia Le, Junichi Yamagishi, Isao Echizen |
WACV | 2 |
| 2023 | Closer Look at the Transferability of Adversarial Examples: How They Fool Different Models DifferentlyabstractDeep neural networks are vulnerable to adversarial examples (AEs), which have adversarial transferability: AEs generated for the source model can mislead another (target) model’s predictions. However, the transferability has not been understood in terms of to which class target model’s predictions were misled (i.e., class-aware transferability). In this paper, we differentiate the cases in which a target model predicts the same wrong class as the source model ("same mistake") or a different wrong class ("different mistake") to analyze and provide an explanation of the mechanism. We find that (1) AEs tend to cause same mistakes, which correlates with "non-targeted transferability"; how-ever, (2) different mistakes occur even between similar models, regardless of the perturbation size. Furthermore, we present evidence that the difference between same mistakes and different mistakes can be explained by non-robust features, predictive but human-uninterpretable patterns: different mistakes occur when non-robust features in AEs are used differently by models. Non-robust features can thus provide consistent explanations for the class-aware transferability of AEs. Futa Waseda, Sosuke Nishikawa, Trung-Nghia Le, Huy H. Nguyen, Isao Echizen |
WACV | 3 |
| 2023 | TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 1 |
| 2023 | SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 1 |
| 2022 | Contextual Guided Segmentation Framework for Semi-supervised Video Instance Segmentation
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Triet Tran |
Mach. Vis. Appl. | 1 |
| 2022 | Camouflaged Instance Segmentation In-the-Wild: Dataset, Method, and Benchmark SuiteabstractThis paper pushes the envelope on decomposing camouflaged regions in an image into meaningful components, namely, camouflaged instances. To promote the new task of camouflaged instance segmentation of in-the-wild images, we introduce a dataset, dubbed CAMO++, that extends our preliminary CAMO dataset (camouflaged object segmentation) in terms of quantity and diversity. The new dataset substantially increases the number of images with hierarchical pixel-wise ground truths. We also provide a benchmark suite for the task of camouflaged instance segmentation. In particular, we present an extensive evaluation of state-of-the-art instance segmentation methods on our newly constructed CAMO++ dataset in various scenarios. We also present a camouflage fusion learning (CFL) framework for camouflaged instance segmentation to further improve the performance of state-of-the-art methods. The dataset, model, evaluation suite, and benchmark will be made publicly available on our project page. Trung-Nghia Le, Yubo Cao, Tan-Cong Nguyen, Minh-Quan Le, Khanh-Duy Nguyen, Thanh-Toan Do, Minh-Triet Tran, Tam V. Nguyen 0002 |
IEEE Trans. Image Process. | 1 |
| 2021 | CamouFinder: Finding Camouflaged Instances in ImagesabstractIn this paper, we investigate the interesting yet challenging problem of camouflaged instance segmentation. To this end, we first annotate the available CAMO dataset at the instance level. We also embed the data augmentation in order to increase the number of training samples. Then, we train different state-of-the-art instance segmentation on the CAMO-instance data. Last but not least, we develop an interactive user interface which demonstrates the performance of different state-of-the-art instance segmentation methods on the task of camouflaged instance segmentation. The users are able to compare the results of different methods on the given input images. Our work is expected to push the envelope of the camouflage analysis problem. Trung-Nghia Le, Vuong Nguyen, Cong Le, Tan-Cong Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002 |
AAAI | 1 |
| 2021 | Interactive Video Object Mask AnnotationabstractIn this paper, we introduce a practical system for interactive video object mask annotation, which can support multiple back-end methods. To demonstrate the generalization of our system, we introduce a novel approach for video object annotation. Our proposed system takes scribbles at a chosen key-frame from the end-users via a user-friendly interface and produces masks of corresponding objects at the key-frame via the Control-Point-based Scribbles-to-Mask (CPSM) module. The object masks at the key-frame are then propagated to other frames and refined through the Multi-Referenced Guided Segmentation (MRGS) module. Last but not least, the user can correct wrong segmentation at some frames, and the corrected mask is continuously propagated to other frames in the video via the MRGS to produce the object masks at all video frames. Trung-Nghia Le, Tam V. Nguyen 0002, Quoc-Cuong Tran, Trung-Hieu Hoang, Minh-Quan Le, Minh-Triet Tran |
AAAI | 1 |
| 2021 | Effectiveness of Detection-based and Regression-based Approaches for Estimating Mask-Wearing RatioabstractEstimating the mask-wearing ratio in public places is important as it enables health authorities to promptly analyze and implement policies. Methods for estimating the mask-wearing ratio on the basis of image analysis have been reported. However, there is still a lack of comprehensive research on both methodologies and datasets. Most recent reports straightforwardly propose estimating the ratio by applying conventional object detection and classification methods. It is feasible to use regression-based approaches to estimate the number of people wearing masks, especially for congested scenes with tiny and occluded faces, but this has not been well studied. A large-scale and well-annotated dataset is still in demand. In this paper, we proposed two different methods for ratio estimation that are leveraged by either detection-based or regression-based approaches. For the detection-based approach, we improved a state-of-the-art face detector, RetinaFace, for the ratio estimation. For the regression-based approach, we utilized a baseline network, CSRNet, and finetuned it to estimate the density maps for masked and unmasked faces. We also proposed the first large-scale dataset, the “NFM,” which contains 581,108 face annotations extracted from 18,088 video frames in 17 street-view videos11The annotations (bounding boxes and labels), and pre-trained models will be released with the publication of our paper. , Through experiments, the RetinaFace-based method achieves better accuracy under different situations, while the CSRNet-based method is superior in terms of operation time thanks to its compactness. Khanh-Duy Nguyen, Hai-Dang Nguyen, Trung-Nghia Le, Junichi Yamagishi, Isao Echizen |
FG | 3 |
| 2021 | OpenForensics: Large-Scale Challenging Dataset For Multi-Face Forgery Detection And Segmentation In-The-WildabstractThe proliferation of deepfake media is raising concerns among the public and relevant authorities. It has become essential to develop countermeasures against forged faces in social media. This paper presents a comprehensive study on two new countermeasure tasks: multi-face forgery detection and segmentation in-the-wild. Localizing forged faces among multiple human faces in unrestricted natural scenes is far more challenging than the traditional deepfake recognition task. To promote these new tasks, we have created the first large-scale dataset posing a high level of challenges that is designed with face-wise rich annotations explicitly for face forgery detection and segmentation, namely Open-Forensics. With its rich annotations, our OpenForensics dataset has great potentials for research in both deepfake prevention and general human face detection. We have also developed a suite of benchmarks for these tasks by conducting an extensive evaluation of state-of-the-art instance detection and segmentation methods on our newly constructed dataset in various scenarios. Trung-Nghia Le, Huy H. Nguyen, Junichi Yamagishi, Isao Echizen |
ICCV | 1 |
| 2020 | Attention R-CNN for Accident DetectionabstractThis paper addresses accident detection where we not only detect objects with classes, but also recognize their characteristic properties. More specifically, we aim at simultaneously detecting object class bounding boxes on roads and recognizing their status such as safe, dangerous, or crashed. To achieve this goal, we construct a new dataset and propose a baseline method for benchmarking the task of accident detection. We design an accident detection network, called Attention R-CNN, which consists of two streams: one is for object detection with classes and one for characteristic property computation. As an attention mechanism capturing contextual information in the scene, we integrate global contexts exploited from the scene into the stream for object detection. This introduced attention mechanism enables us to recognize object characteristic properties. Extensive experiments on the newly constructed dataset demonstrate the effectiveness of our proposed network. The dataset and source code are publicly available on our project page. Trung-Nghia Le, Shintaro Ono, Akihiro Sugimoto, Hiroshi Kawasaki |
IV | 1 |
| 2020 | Text-to-Image Synthesis via Aesthetic LayoutabstractIn this work, we introduce a practical system which synthesizes an appealing image from natural language descriptions such that the generated image should maintain the aesthetic level of photographs. Our proposed method takes the text from the end-users via a user-friendly interface and produces a set of different label maps via the primary generator PG. Then, choosing a subset from the label maps set is performed through the primary aesthetic appreciation PAA. Next, our subset of label maps is fed into the accessory generator AG, which is the state-of-the-art image-to-image translation. Last but not least, our subset of generated images is ranked via the accessory aesthetic appreciation AAA, and the most appealing image is produced. Samah Saeed Baraheem, Trung-Nghia Le, Tam V. Nguyen 0002 |
ACM Multimedia | 2 |
| 2020 | Toward Interactive Self-Annotation For Video Object Bounding Box: Recurrent Self-Learning And Hierarchical Annotation Based FrameworkabstractAmount and variety of training data drastically affect the performance of CNNs. Thus, annotation methods are becoming more and more critical to collect data efficiently. In this paper, we propose a simple yet efficient Interactive Self-Annotation framework to cut down both time and human labor cost for video object bounding box annotation. Our method is based on recurrent self-supervised learning and consists of two processes: automatic process and interactive process, where the automatic process aims to build a supported detector to speed up the interactive process. In the Automatic Recurrent Annotation, we let an off-the-shelf detector watch unlabeled videos repeatedly to reinforce itself automatically. At each iteration, we utilize the trained model from the previous iteration to generate better pseudo ground-truth bounding boxes than those at the previous iteration, recurrently improving self-supervised training the detector. In the Interactive Recurrent Annotation, we tackle the human-in-the-loop annotation scenario where the detector receives feedback from the human annotator. To this end, we propose a novel Hierarchical Correction module, where the annotated frame-distance binarizedly decreases at each time step, to utilize the strength of CNN for neighbor frames. Experimental results on various video datasets demonstrate the advantages of the proposed framework in generating high-quality annotations while reducing annotation time and human labor costs. Trung-Nghia Le, Akihiro Sugimoto, Shintaro Ono, Hiroshi Kawasaki |
WACV | 1 |
| 2019 | Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance SegmentationabstractFocusing on only semantic instances that only salient in a scene gains more benefits for robot navigation and self-driving cars than looking at all objects in the whole scene. This paper pushes the envelope on salient regions in a video to decompose them into semantically meaningful components, namely, semantic salient instances. We provide the baseline for the new task of video semantic salient instance segmentation (VSSIS), that is, Semantic Instance - Salient Object (SISO) framework. The SISO framework is simple yet efficient, leveraging advantages of two different segmentation tasks, i.e. semantic instance segmentation and salient object segmentation to eventually fuse them for the final result. In SISO, we introduce a sequential fusion by looking at overlapping pixels between semantic instances and salient regions to have non-overlapping instances one by one. We also introduce a recurrent instance propagation to refine the shapes and semantic meanings of instances, and an identity tracking to maintain both the identity and the semantic meaning of instances over the entire video. Experimental results demonstrated the effectiveness of our SISO baseline, which can handle occlusions in videos. In addition, to tackle the task of VSSIS, we augment the DAVIS-2017 benchmark dataset by assigning semantic ground-truth for salient instance labels, obtaining SEmantic Salient Instance Video (SESIV) dataset. Our SESIV dataset consists of 84 high-quality video sequences with pixel-wisely per-frame ground-truth labels. Trung-Nghia Le, Akihiro Sugimoto |
WACV | 1 |
| 2019 | Anabranch network for camouflaged object segmentation
Trung-Nghia Le, Tam V. Nguyen 0002, Zhongliang Nie, Minh-Triet Tran, Akihiro Sugimoto |
Comput. Vis. Image Underst. | 1 |
| 2018 | Balancing Content and Style with Two-Stream FCNs for Style TransferabstractStyle transfer is to render given image contents in given styles, and it has an important role in both computer vision fundamental research and industrial applications. Following the success ofdeep learning based approaches, this problem has been re-launched very recently, but still remains a difficult task because of trade-of between preserving contents and faithful rendering of styles. In this paper, we propose an end-to-end two-stream Fully Convolutional Networks (FCNs) aiming at balancing the contributions of the content and the style in rendered images. Our proposed network consists ofthe encoder and decoder parts. The encoder part utilizes a FCN for content and a FCN for style where the two FCNs are independently trained to preserve the semantic content and to learn the faithful style representation in each. The semantic content feature and the style representationfeature are then concatenated adaptively and fed into the decoder to generate style-transferred (stylized) images. In order to train our proposed network, we employ a loss network, the pre-trained VGG-I6, to compute content loss and style loss, both of which are efficiently used for the feature concatenation. Our intensive experiments show that our proposed model generates more balanced stylized images in content and style than state-of-theart methods. Moreover, our proposed network achieves efficiency in speed. Duc Minh Vo, Trung-Nghia Le, Akihiro Sugimoto |
WACV | 2 |
| 2018 | Video Salient Object Detection Using Spatiotemporal Deep FeaturesabstractThis paper presents a method for detecting salient objects in videos, where temporal information in addition to spatial information is fully taken into account. Following recent reports on the advantage of deep features over conventional handcrafted features, we propose a new set of spatiotemporal deep (STD) features that utilize local and global contexts over frames. We also propose new spatiotemporal conditional random field (STCRF) to compute saliency from STD features. STCRF is our extension of CRF to the temporal domain and describes the relationships among neighboring regions both in a frame and over frames. STCRF leads to temporally consistent saliency maps over frames, contributing to accurate detection of salient objects' boundaries and noise reduction during detection. Our proposed method first segments an input video into multiple scales and then computes a saliency map at each scale level using STD features with STCRF. The final saliency map is computed by fusing saliency maps at different scale levels. Our experiments, using publicly available benchmark datasets, confirm that the proposed method significantly outperforms the state-of-the-art methods. We also applied our saliency computation to the video object segmentation task, showing that our method outperforms existing video object segmentation methods. Trung-Nghia Le, Akihiro Sugimoto |
IEEE Trans. Image Process. | 1 |
| 2017 | Deeply Supervised 3D Recurrent FCN for Salient Object Detection in Videos
Trung-Nghia Le, Akihiro Sugimoto |
BMVC | 1 |
| 2015 | Contrast Based Hierarchical Spatial-Temporal Saliency for Video
Trung-Nghia Le, Akihiro Sugimoto |
PSIVT | 1 |
| 2014 | Essential keypoints to enhance visual object recognition with saliency-based metricsabstractThe authors propose a novel pre-processing phase that can be integrated into conventional methods to detect and recognize planar visual objects in printed materials with low computational cost and higher accuracy. A simple yet efficient visual saliency estimation technique based on regional contrast is developed to quickly filter out low informative regions in printed materials. By eliminating noisy or unimportant keypoint candidates, our proposed method not only reduces unnecessary computational cost of keypoint descriptors but also increases robustness and accuracy of visual object recognition. Our experimental results show that the whole visual object recognition process can be speeded up 46 times and the accuracy can increase up to 23%. These are desirable advantages for an augmented reality system, especially on mobile devices. Furthermore, this pre-processing stage is independent of the choice of features and matching model in a general process. Therefore it can be used to boost the performance of existing systems into real-time manner. Trung-Nghia Le, Yen-Thanh Le, Minh-Triet Tran, Anh Duc Duong |
ICARCV | 1 |