VLDB 2026 Research / reviewers in the wild / expert
Tam V. Nguyen 0002
dblp:119/1364-2 · also Tam Van Nguyen 0002
· DBLP profile ↗
86ranked-venue papers
21as first author
44since 2021 · last 2026
0000-0003-0236-7992ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 17 first-author · 30 since 2021Artificial intelligence and machine learning · 32 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TapesVRy: Immersive Panoramic Exploration in Large-Scale Video Retrieval
Viet-Tham Huynh, Nhut-Thanh Le-Hinh, Thang-Long Nguyen-Ho, Cathal Gurrin, Tam V. Nguyen 0002, Minh-Triet Tran |
MMM (4) | 6 |
| 2026 | VENUS: Visual Editing with Noise Inversion Using Scene Graphs
Thanh-Nhan Vo, Tam V. Nguyen 0002, Minh-Triet Tran |
MMM (2) | 3 |
| 2026 | PoseGaussian: Pose-Driven Novel View Synthesis for Robust 3D Human ReconstructionabstractWe propose PoseGaussian, a pose-guided Gaussian Splatting framework for high-fidelity human novel view synthesis. Human body pose serves a dual purpose in our design: as a structural prior, it is fused with a color encoder to refine depth estimation; as a temporal cue, it is processed by a dedicated pose encoder to enhance temporal consistency across frames. These components are integrated into a fully differentiable, end-to-end trainable pipeline. Unlike prior works that use pose only as a condition or for warping, PoseGaussian embeds pose signals into both geometric and temporal stages to improve robustness and generalization. It is specifically designed to address challenges inherent in dynamic human scenes, such as articulated motion and severe self-occlusion. Notably, our framework achieves real-time rendering at 100 FPS, maintaining the efficiency of standard Gaussian Splatting pipelines. We validate our approach on ZJU-MoCap, THuman2.0, and in-house datasets, demonstrating state-of-the-art performance in perceptual quality and structural accuracy (PSNR 30.86, SSIM 0.979, LPIPS 0.028). Ju Shen, Chen Chen 0001, Tam V. Nguyen 0002, Vijayan K. Asari |
WACV | 3 |
| 2026 | ShowFlow: From robust single concept to condition-free multi-concept generation
Trong-Vu Hoang, Quang-Binh Nguyen, Thanh-Toan Do, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
Neurocomputing | 4 |
| 2026 | You can see me: Retrieval-augmented framework for divergent perception in animal art generation
Quoc-Duy Tran, Anh-Tuan Vo, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
Pattern Recognit. Lett. | 3 |
| 2025 | A Generative Approach at the Instance-Level for Image Segmentation Under Limited Training Data Conditions (Student Abstract)abstractHigh-accuracy image segmentation models require abundant training annotated data which is costly for pixel-level annotations. Our work addresses a high-cost manual annotating process or the lack of detailed annotations via a generative approach. In particular, our approach (1) proposes the conditional instance-level synthesis to enrich the limited data to enhance the segmentation performance, and (2) employs the generative architectures to complete the segmentation task under few-shot learning concepts. The initial results on the Cityscapes benchmark emphasize our potential generative solution on the instance segmentation task given limited data. Thanh-Danh Nguyen, Vinh-Tiep Nguyen, Tam V. Nguyen 0002 |
AAAI | 3 |
| 2025 | Toward Content-Based Indexing and Retrieval of Head and Neck CT With Abscess SegmentationabstractAbscesses in the head and neck represent an acute infectious process that can potentially lead to sepsis or mortality if not diagnosed and managed promptly. Accurate detection and delineation of these lesions on imaging are essential for diagnosis, treatment planning, and surgical intervention. In this study, we introduce AbscessHeNe, a curated and comprehensively annotated dataset comprising 4,926 contrastenhanced CT slices with clinically confirmed head and neck abscesses. The dataset is designed to facilitate the development of robust semantic segmentation models that can accurately delineate abscess boundaries and evaluate deep neck space involvement, thereby supporting informed clinical decision-making. To establish performance baselines, we evaluate several state-of-the-art segmentation architectures, including CNN, Transformer, and Mamba-based models. The highestperforming model achieved a Dice Similarity Coefficient of 0.39, Intersection-over-Union of 0.27, and Normalized Surface Distance of 0.67, indicating the challenges of this task and the need for further research. Beyond segmentation, AbscessHeNe is structured for future applications in content-based multimedia indexing and case-based retrieval. Each CT scan is linked with pixel-level annotations and clinical metadata, providing a foundation for building intelligent retrieval systems and supporting knowledge-driven clinical workflows. The dataset will be made publicly available at https://github.com/drthaodao3101/AbscessHeNe.git. Thao Thi Phuong Dao, Tan-Cong Nguyen, Trong-Le Do, Truong Hoang Viet, Nguyen Chi Thanh, Huynh Nguyen Thuan, Do Vo Cong Nguyen, Minh-Khoi Pham, Mai-Khiem Tran, Viet-Tham Huynh, Trung-Nghia Le, Thanh-Nhan Vo, Tam V. Nguyen 0002, Minh-Triet Tran, Thanh Dinh Le |
CBMI | 14 |
| 2025 | Automated Image Recognition Framework
Quang-Binh Nguyen, Trong-Vu Hoang, Ngoc-Do Tran, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ICCCI (2) | 4 |
| 2025 | Multi-Perspective Data Augmentation for Few-shot Object DetectionabstractRecent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of $17.5\%$ in nAP50 over the baseline on PASCAL VOC. Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
ICLR | 6 |
| 2025 | ACM Multimedia Grand Challenge on ENT Endoscopy AnalysisabstractAutomated analysis of endoscopic imagery is a critical yet underdeveloped component of ENT (ear, nose, and throat) care, hindered by variability in devices and operators, subtle and localized findings, and fine-grained distinctions such as laterality and vocal-fold state. In addition to classification, clinicians require reliable retrieval of similar cases, both visually and through concise textual descriptions. These capabilities are rarely supported by existing public benchmarks. To this end, we introduce ENTRep, the ACM Multimedia 2025 Grand Challenge on ENT endoscopy analysis, which integrates fine-grained anatomical classification with image-to-image and text-to-image retrieval under bilingual (Vietnamese and English) clinical supervision. Specifically, the dataset comprises expert-annotated images, labeled for anatomical region and normal or abnormal status, and accompanied by dual-language narrative descriptions. In addition, we define three benchmark tasks, standardize the submission protocol, and evaluate performance on public and private test splits using server-side scoring. Moreover, we report results from the top-performing teams and provide an insightful discussion. Viet-Tham Huynh, Thao Thi Phuong Dao, Mai-Khiem Tran, Ha Nguyen Thi, Tien To Vu Thuy, Uyen Hanh Tran, Tam V. Nguyen 0002, Minh-Triet Tran, Thanh Dinh Le |
ACM Multimedia | 8 |
| 2025 | OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event GroundingabstractWe introduce OpenEvents V1, a large-scale benchmark dataset designed to advance event-centric vision-language understanding. Unlike conventional image captioning and retrieval datasets that focus on surface-level descriptions, OpenEvents V1 dataset emphasizes contextual and temporal grounding through three primary tasks: (1) generating rich, event-aware image captions, (2) retrieving event-relevant news articles from image queries, and (3) retrieving event-relevant images from narrative-style textual queries. The dataset comprises over 200,000 news articles and 400,000 associated images sourced from CNN and The Guardian, spanning diverse domains and time periods. We provide extensive baseline results and standardized evaluation protocols for all tasks. OpenEvents V1 establishes a robust foundation for developing multimodal AI systems capable of deep reasoning over complex real-world events. The dataset is publicly available at https://ltnghia.github.io/eventa/openevents-v1. Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 5 |
| 2025 | Event-Enriched Image Analysis Grand Challenge At ACM Multimedia 2025abstractThe Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses this gap by integrating contextual, temporal, and semantic information to capture the who, when, where, what, and why behind an image. Built upon the OpenEvents V1 dataset, the challenge features two tracks: Event-Enriched Image Retrieval and Captioning, and Event-Based Image Retrieval. A total of 45 teams from six countries participated, with evaluation conducted through Public and Private Test phases to ensure fairness and reproducibility. The top three teams were invited to present their solutions at ACM Multimedia 2025. EVENTA establishes a foundation for context-aware, narrative-driven multimedia AI, with applications in journalism, media analysis, cultural archiving, and accessibility. Further details about the challenge are available at the official homepage: https://ltnghia.github.io/eventa/eventa-2025. Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le |
ACM Multimedia | 4 |
| 2025 | Advancing Fashion Design Through Intelligent Sketchpad StudioabstractLine sketches serve as the visual DNA of fashion design, forming the essential foundation where concepts take shape, yet today's digital tools often lack the fluidity, personalization, and intelligence needed to truly support this creative process. We present FashSketch, an interactive, multimedia-driven system that reimagines fashion sketching through the lens of generative AI. Designed with a layer-based creative interface, FashSketch empowers designers to ideate, customize, and iterate on sketches seamlessly. By integrating state-of-the-art generative models, sketch-based retrieval, and large language models, the system supports advanced functionalities such as text-to-sketch generation and context-aware sketch recommendation. FashSketch not only enhances the sketching experience but also opens new multimodal pathways for creative expression, making it a powerful co-creative partner in the early stages of fashion design. Video demo is available at https://youtu.be/BX-Edz7Z7ZY. Nhu-Binh Nguyen Truc, Nhu-Vinh Hoang, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le |
ACM Multimedia | 3 |
| 2025 | ChemersiveLLM: Prompt-to-VR Simulation of Chemistry Experiments Using Generative AIabstractLarge Language Models (LLMs) offer significant potential for integration with Virtual Reality (VR), but current AI systems struggle to generate accurate 3D environments and support semantic interaction. We present ChemersiveLLM, a VR-based chemistry learning platform that leverages LLMs for instruction sequencing, natural language grounding, and real-time guidance. Using a semantic action-mapping framework, the system translates AI-generated content into structured lab actions, enabling multimodal interaction, embodied experimentation, and intelligent feedback. Comparative evaluation across textbook, chatbot-based, and VR learning shows that our system improves engagement, comprehension, and satisfaction, underscoring its promise as a next-generation tool for science education. Thanh Ngoc-Dat Tran, Viet-Tham Huynh, G. Michael Poor, Minh-Triet Tran, Tam V. Nguyen 0002 |
VRST | 5 |
| 2025 | CamoFA: A Learnable Fourier-Based Augmentation for Camouflage SegmentationabstractCamouflaged object detection (COD) and camouflaged instance segmentation (CIS) aim to recognize and segment objects that are blended into their surroundings, respectively. While several deep neural network models have been proposed to tackle those tasks, augmentation methods for COD and CIS have not been thoroughly explored. strategies can help improve models' performance by increasing the size and diversity of the training data and exposing the model to a wider range of variations in the data. Besides, we aim to automatically learn transformations that help to reveal the underlying structure of camou-flaged objects and allow the model to learn to identify better and segment camouflaged objects. To achieve this, we pro-pose a learnable augmentation method in the frequency domain for COD and CIS via the Fourier transform approach, dubbed CamoFA. Our method leverages a conditional generative adversarial network and cross-attention mechanism to generate a reference image and an adaptive hybrid swapping with parameters to mix the low-frequency component of the reference image and the high-frequency component of the input image. This approach aims to make camouflaged objects more visible for detection and segmentation models. Without bells and whistles, our proposed augmentation method boosts the performance of camouflaged object detectors and instance segmenters by large margins. Minh-Quan Le, Minh-Triet Tran, Trung-Nghia Le, Tam V. Nguyen 0002, Thanh-Toan Do |
WACV | 4 |
| 2025 | SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
Viet-Tham Huynh, Hoang-Phuc Nguyen, Long Bao Le, Thai Hoang Minh, Minh Nguyen Anh, Thang Nguyen Tien, Phat Nguyen Thuan, Huy Nguyen Phong, Bao Huynh Thai, Vinh-Tiep Nguyen, Duc-Vu Nguyen, Phu-Hoa Pham, Minh-Huy Le-Hoang, Nguyen-Khang Le, Minh-Chinh Nguyen, Minh-Quan Ho, Ngoc-Long Tran, Hien-Long Le-Hoang, Man-Khoi Tran, Anh-Duong Tran, Quan Nguyen Hung, Dat Phan Thanh, Hoang Tran Van, Tien Huynh Viet, Nhan Nguyen Viet Thien, Dinh-Khoi Vo, Van-Loc Nguyen, Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Triet Tran |
Comput. Graph. | 32 |
| 2025 | SHREC GS-3DORC: Towards the advancements of 3D Gaussian Splatting object part retrieval
Thien-Phuc Tran, Minh-Quang Nguyen, Thanh-Khoi Nguyen, Nam-Quan Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran |
Comput. Graph. | 5 |
| 2025 | Few-shot object detection via synthetic features with optimal transport
Anh-Khoa Nguyen Vu, Thanh-Toan Do, Vinh-Tiep Nguyen, Tam Le, Minh-Triet Tran, Tam V. Nguyen 0002 |
Comput. Vis. Image Underst. | 6 |
| 2025 | Small object detection in aerial traffic imagery: A benchmark for motorbike-dominated road scenes
Dung Truong, Khanh-Duy Nguyen, Tam V. Nguyen 0002, Khang Nguyen 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Towards safer roads: benchmarking object detection models in complex weather scenarios
Ba-Thinh Tran-Le, Vatsa S. Patel, Viet-Tham Huynh, Mai-Khiem Tran, Kunal Agrawal 0004, Minh-Triet Tran, Tam V. Nguyen 0002 |
Mach. Vis. Appl. | 7 |
| 2025 | Correction: Real estate pricing prediction via textual and visual features
Amira Yousif, Samah Saeed Baraheem, Sai Surya Vaddi, Vatsa S. Patel, Ju Shen, Tam V. Nguyen 0002 |
Mach. Vis. Appl. | 6 |
| 2025 | GUNNEL: guided mixup augmentation and multi-model fusion for aquatic animal segmentation
Minh-Quan Le, Trung-Nghia Le, Tam V. Nguyen 0002, Isao Echizen, Minh-Triet Tran |
Neural Comput. Appl. | 3 |
| 2024 | MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance SegmentationabstractFew-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, known as point estimation. However, this mechanism depends on prototypes (e.g. mean of K-shot) for prediction, leading to performance instability. To overcome the disadvantage of the point estimation mechanism, we propose a novel approach, dubbed MaskDiff, which models the underlying conditional distribution of a binary mask, which is conditioned on an object region and K-shot information. Inspired by augmentation approaches that perturb data with Gaussian noise for populating low data density regions, we model the mask distribution with a diffusion probabilistic model. We also propose to utilize classifier-free guided mask sampling to integrate category information into the binary mask generation process. Without bells and whistles, our proposed method consistently outperforms state-of-the-art methods on both base and novel classes of the COCO dataset while simultaneously being more stable than existing methods. The source code is available at: https://github.com/minhquanlecs/MaskDiff. Minh-Quan Le, Tam V. Nguyen 0002, Trung-Nghia Le, Thanh-Toan Do, Minh N. Do, Minh-Triet Tran |
AAAI | 2 |
| 2024 | LUMOS-DM: Landscape-Based Multimodal Scene Retrieval Enhanced by Diffusion Model
Viet-Tham Huynh, Mai-Khiem Tran, Tam V. Nguyen 0002, Minh-Triet Tran |
MMM (4) | 5 |
| 2024 | Nighttime scene understanding with label transfer scene parser
Thanh-Danh Nguyen, Nguyen Phan, Tam V. Nguyen 0002, Vinh-Tiep Nguyen, Minh-Triet Tran |
Image Vis. Comput. | 3 |
| 2024 | Sketch-to-image synthesis via semantic masks
Samah Saeed Baraheem, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 2 |
| 2024 | Image de-photobombing benchmarkabstractAbstract Removing photobombing elements from images is a challenging task that requires sophisticated image inpainting techniques. Despite the availability of various methods, their effectiveness depends on the complexity of the image and the nature of the distracting element. To address this issue, we conducted a benchmark study to evaluate 10 state-of-the-art photobombing removal methods on a dataset of over 300 images. Our study focused on identifying the most effective image inpainting techniques for removing unwanted regions from images. We annotated the photobombed regions that require removal and evaluated the performance of each method using peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and Fréchet inception distance (FID). The results show that image inpainting techniques can effectively remove photobombing elements, but more robust and accurate methods are needed to handle various image complexities. Our benchmarking study provides a valuable resource for researchers and practitioners to select the most suitable method for their specific photobombing removal task. Vatsa S. Patel, Kunal Agrawal 0004, Samah Saeed Baraheem, Amira Yousif, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 5 |
| 2024 | Transformer with multi-level grid features and depth pooling for image captioning
Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001 |
Mach. Vis. Appl. | 2 |
| 2024 | Transformer-Based Spatio-Temporal Unsupervised Traffic Anomaly Detection in Aerial VideosabstractAnomaly detection is an area of video analysis and plays an increasing role in ensuring safety, preventing risks, and guaranteeing quick response in intelligent surveillance systems. It has become a popular research topic and has piqued the interest of researchers in different communities, such as computer vision, machine learning, remote sensing, and data mining, in recent years. This promotes novel mobile systems where drones are equipped with cameras to help people find better and more efficient solutions to automatically detect anomalies (e.g., car accidents, traffic congestion, street fighting) in traffic surveillance videos. However, anomaly detection methods are still rarely studied and developed in the remote sensing community due to anomalous events rarely occurring in real life, along with the high similarities between the objects of interest with small sizes, multi-scale objects, complex backgrounds of great variations, and high overlap between objects. Therefore, in order to fully exploit the spatio-temporal information for anomaly detection in traffic surveillance circumstances, we propose a future frame prediction network based on transformer architectures to detect abnormal events from drone videography in an unsupervised way. Our model treats consecutive video frames from an input clip and feeds features to a transformer encoder to capture spatial and temporal representations from the sequence. Then, it leverages a decoder to predict the next frame. Furthermore, an event with high reconstruction error is identified as an anomaly in the test phase. Thoroughly empirical studies demonstrate that our method achieves superior performance on the UIT-ADrone dataset and largely outperforms the state-of-the-art anomaly methods on the Drone-Anomaly dataset in aerial surveillance. The source code is available online at https://github.com/Tungufm/ASTT. Tung Minh Tran, Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Efficient 3D Brain Tumor Segmentation with Axial-Coronal-Sagittal Embedding
Tuan-Luc Huynh, Thanh-Danh Le, Tam V. Nguyen 0002, Trung-Nghia Le, Minh-Triet Tran |
PSIVT | 3 |
| 2023 | MobileNet-SA: Lightweight CNN with Self Attention for Sketch Classification
Viet-Tham Huynh, Tam V. Nguyen 0002, Minh-Triet Tran |
PSIVT | 3 |
| 2023 | TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 2 |
| 2023 | SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran |
Comput. Graph. | 2 |
| 2023 | Abstraction-perception preserving cartoon face synthesis
Sy-Tuyen Ho, Manh-Khanh Ngo Huu, Thanh-Danh Nguyen, Nguyen Phan, Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 8 |
| 2023 | Real estate pricing prediction via textual and visual features
Amira Yousif, Samah Saeed Baraheem, Sai Surya Vaddi, Vatsa S. Patel, Ju Shen, Tam V. Nguyen 0002 |
Mach. Vis. Appl. | 6 |
| 2023 | Instance-Level Few-Shot Learning With Class Hierarchy MiningabstractFew-shot learning is proposed to tackle the problem of scarce training data in novel classes. However, prior works in instance-level few-shot learning have paid less attention to effectively utilizing the relationship between categories. In this paper, we exploit the hierarchical information to leverage discriminative and relevant features of base classes to effectively classify novel objects. These features are extracted from abundant data of base classes, which could be utilized to reasonably describe classes with scarce data. Specifically, we propose a novel superclass approach that automatically creates a hierarchy considering base and novel classes as fine-grained classes for few-shot instance segmentation (FSIS). Based on the hierarchical information, we design a novel framework called Soft Multiple Superclass (SMS) to extract relevant features or characteristics of classes in the same superclass. A new class assigned to the superclass is easier to classify by leveraging these relevant features. Besides, in order to effectively train the hierarchy-based-detector in FSIS, we apply the label refinement to further describe the associations between fine-grained classes. The extensive experiments demonstrate the effectiveness of our method on FSIS benchmarks. The source code is available here: https://github.com/nvakhoa/superclass-FSIS. Anh-Khoa Nguyen Vu, Thanh-Toan Do, Nhat-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Tam V. Nguyen 0002 |
IEEE Trans. Image Process. | 6 |
| 2023 | Multimodal Mutual Information Maximization: A Novel Approach for Unsupervised Deep Cross-Modal HashingabstractIn this article, we adopt the maximizing mutual information (MI) approach to tackle the problem of unsupervised learning of binary hash codes for efficient cross-modal retrieval. We proposed a novel method, dubbed cross-modal info-max hashing (CMIMH). First, to learn informative representations that can preserve both intramodal and intermodal similarities, we leverage the recent advances in estimating variational lower bound of MI to maximizing the MI between the binary representations and input features and between binary representations of different modalities. By jointly maximizing these MIs under the assumption that the binary representations are modeled by multivariate Bernoulli distributions, we can learn binary representations, which can preserve both intramodal and intermodal similarities, effectively in a mini-batch manner with gradient descent. Furthermore, we find out that trying to minimize the modality gap by learning similar binary representations for the same instance from different modalities could result in less informative representations. Hence, balancing between reducing the modality gap and losing modality-private information is important for the cross-modal retrieval tasks. Quantitative evaluations on standard benchmark datasets demonstrate that the proposed method consistently outperforms other state-of-the-art cross-modal retrieval methods. Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Few-shot object detection via baby learning
Anh-Khoa Nguyen Vu, Nhat-Duy Nguyen, Khanh-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002 |
Image Vis. Comput. | 7 |
| 2022 | Contextual Guided Segmentation Framework for Semi-supervised Video Instance Segmentation
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Triet Tran |
Mach. Vis. Appl. | 2 |
| 2022 | Camouflaged Instance Segmentation In-the-Wild: Dataset, Method, and Benchmark SuiteabstractThis paper pushes the envelope on decomposing camouflaged regions in an image into meaningful components, namely, camouflaged instances. To promote the new task of camouflaged instance segmentation of in-the-wild images, we introduce a dataset, dubbed CAMO++, that extends our preliminary CAMO dataset (camouflaged object segmentation) in terms of quantity and diversity. The new dataset substantially increases the number of images with hierarchical pixel-wise ground truths. We also provide a benchmark suite for the task of camouflaged instance segmentation. In particular, we present an extensive evaluation of state-of-the-art instance segmentation methods on our newly constructed CAMO++ dataset in various scenarios. We also present a camouflage fusion learning (CFL) framework for camouflaged instance segmentation to further improve the performance of state-of-the-art methods. The dataset, model, evaluation suite, and benchmark will be made publicly available on our project page. Trung-Nghia Le, Yubo Cao, Tan-Cong Nguyen, Minh-Quan Le, Khanh-Duy Nguyen, Thanh-Toan Do, Minh-Triet Tran, Tam V. Nguyen 0002 |
IEEE Trans. Image Process. | 8 |
| 2021 | CamouFinder: Finding Camouflaged Instances in ImagesabstractIn this paper, we investigate the interesting yet challenging problem of camouflaged instance segmentation. To this end, we first annotate the available CAMO dataset at the instance level. We also embed the data augmentation in order to increase the number of training samples. Then, we train different state-of-the-art instance segmentation on the CAMO-instance data. Last but not least, we develop an interactive user interface which demonstrates the performance of different state-of-the-art instance segmentation methods on the task of camouflaged instance segmentation. The users are able to compare the results of different methods on the given input images. Our work is expected to push the envelope of the camouflage analysis problem. Trung-Nghia Le, Vuong Nguyen, Cong Le, Tan-Cong Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002 |
AAAI | 6 |
| 2021 | Interactive Video Object Mask AnnotationabstractIn this paper, we introduce a practical system for interactive video object mask annotation, which can support multiple back-end methods. To demonstrate the generalization of our system, we introduce a novel approach for video object annotation. Our proposed system takes scribbles at a chosen key-frame from the end-users via a user-friendly interface and produces masks of corresponding objects at the key-frame via the Control-Point-based Scribbles-to-Mask (CPSM) module. The object masks at the key-frame are then propagated to other frames and refined through the Multi-Referenced Guided Segmentation (MRGS) module. Last but not least, the user can correct wrong segmentation at some frames, and the corrected mask is continuously propagated to other frames in the video via the MRGS to produce the object masks at all video frames. Trung-Nghia Le, Tam V. Nguyen 0002, Quoc-Cuong Tran, Trung-Hieu Hoang, Minh-Quan Le, Minh-Triet Tran |
AAAI | 2 |
| 2021 | Parsing Digitized Vietnamese Paper Documents
Linh Truong Dieu, Nguyen D. Vo, Tam V. Nguyen 0002, Khang Nguyen 0001 |
CAIP (1) | 4 |
| 2021 | CCBANet: Cascading Context and Balancing Attention for Polyp Segmentation
Tan-Cong Nguyen, Tien-Phat Nguyen, Gia-Han Diep, Anh-Huy Tran-Dinh, Tam V. Nguyen 0002, Minh-Triet Tran |
MICCAI (1) | 5 |
| 2020 | Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural NetworksabstractThis paper proposes two novel techniques to train deep convolutional neural networks with low bit-width weights and activations. First, to obtain low bit-width weights, most existing methods obtain the quantized weights by performing quantization on the full-precision network weights. However, this approach would result in some mismatch: the gradient descent updates full-precision weights, but it does not update the quantized weights. To address this issue, we propose a novel method that enables direct updating of quantized weights with learnable quantization levels to minimize the cost function using gradient descent. Second, to obtain low bit-width activations, existing works consider all channels equally. However, the activation quantizers could be biased toward a few channels with high-variance. To address this issue, we propose a method to take into account the quantization errors of individual channels. With this approach, we can learn activation quantizers that minimize the quantization errors in the majority of channels. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the image classification task, using AlexNet, ResNet and MobileNetV2 architectures on CIFAR-100 and ImageNet datasets. Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung |
IJCAI | 3 |
| 2020 | Flood Level Prediction via Human Pose Estimation from Social Media ImagesabstractFloods are the most common natural and among the most dangerous disasters in the world. It is important to get up-to-date information about flooding and the flood level for flood preparation and prevention. In this paper, we propose an efficient method to determine the flood level from daily activity photos on social media. Our method is based on the idea of matching the water level with human pose to determine the level of severity of flooding. Extensive experiments conducted on the dataset of Multimodal Flood Level Estimation show the superiority of our proposed method. We achieve the first rank in MediaEval 2019 and this demonstrates the potential applications of our method to analyze flood information. Khanh-An C. Quan, Vinh-Tiep Nguyen, Tan-Cong Nguyen, Tam V. Nguyen 0002, Minh-Triet Tran |
ICMR | 4 |
| 2020 | Text-to-Image Synthesis via Aesthetic LayoutabstractIn this work, we introduce a practical system which synthesizes an appealing image from natural language descriptions such that the generated image should maintain the aesthetic level of photographs. Our proposed method takes the text from the end-users via a user-friendly interface and produces a set of different label maps via the primary generator PG. Then, choosing a subset from the label maps set is performed through the primary aesthetic appreciation PAA. Next, our subset of label maps is fed into the accessory generator AG, which is the state-of-the-art image-to-image translation. Last but not least, our subset of generated images is ranked via the accessory aesthetic appreciation AAA, and the most appealing image is produced. Samah Saeed Baraheem, Trung-Nghia Le, Tam V. Nguyen 0002 |
ACM Multimedia | 3 |
| 2020 | Text-to-image via mask anchor points
Samah Saeed Baraheem, Tam V. Nguyen 0002 |
Pattern Recognit. Lett. | 2 |
| 2020 | Unsupervised Deep Cross-modality Spectral HashingabstractThis paper presents a novel framework, namely Deep Cross-modality Spectral Hashing (DCSH), to tackle the unsupervised learning problem of binary hash codes for efficient cross-modal retrieval. The framework is a two-step hashing approach which decouples the optimization into (1) binary optimization and (2) hashing function learning. In the first step, we propose a novel spectral embedding-based algorithm to simultaneously learn single-modality and binary cross-modality representations. While the former is capable of well preserving the local structure of each modality, the latter reveals the hidden patterns from all modalities. In the second step, to learn mapping functions from informative data inputs (images and word embeddings) to binary codes obtained from the first step, we leverage the powerful CNN for images and propose a CNN-based deep architecture to learn text modality. Quantitative evaluations on three standard benchmark datasets demonstrate that the proposed DCSH method consistently outperforms other state-of-the-art methods. Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung |
IEEE Trans. Image Process. | 3 |
| 2019 | LiveSense: Contextual Advertising in Live Streaming VideosabstractLive streaming has become a new form of entertainment, which attracts hundreds of millions of users worldwide. The huge amount of multimedia data in live streaming platforms creates tremendous opportunities for online advertising. However, existing state-of-the-art video advertising strategies (e.g., pre-roll and contextual mid-roll advertising) that rely on analyzing the whole video, are not applicable to live streaming videos. This paper describes a novel monetization framework, named LiveSense, for live streaming videos, which is able to display a contextually relevant ad at a suitable timestamp in a non-intrusive way. Specifically, given a live streaming video, we first employ a deep neural network to determine whether the current moment is appropriate for displaying an ad using the historical streaming data. Then, we detect a set of candidate ad insertion areas by incorporating image saliency, background map, and location priorities, so that the ad is displayed over the non-important area. We introduce three types of relevance metrics including textual relevance, global visual relevance and local visual relevance to select the contextually relevant ad. To minimize user intrusiveness, we initially display the ad at a non-important area. If the user is interested in the ad, we will show the ad in an overlaid window with a translucent background. Empirical evaluation on a real-world dataset demonstrates that our proposed framework is able to effectively display ads in live streaming videos while maintaining users' online experience. Xiang Chen 0010, Tam V. Nguyen 0002, Zhiqi Shen 0002, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2019 | Anabranch network for camouflaged object segmentation
Trung-Nghia Le, Tam V. Nguyen 0002, Zhongliang Nie, Minh-Triet Tran, Akihiro Sugimoto |
Comput. Vis. Image Underst. | 2 |
| 2019 | You always look again: Learning to detect the unseen objects
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | YADA: you always dream again for better object detection
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 5 |
| 2019 | Simultaneous Feature Aggregating and Hashing for Compact Binary Code LearningabstractRepresenting images by compact hash codes is an attractive approach for large-scale content-based image retrieval. In most state-of-the-art hashing-based image retrieval systems, for each image, local descriptors are first aggregated as a global representation vector. This global vector is then subjected to a hashing function to generate a binary hash code. In previous works, the aggregating and the hashing processes are designed independently. Hence, these frameworks may generate suboptimal hash codes. In this paper, we first propose a novel unsupervised hashing framework in which feature aggregating and hashing are designed simultaneously and optimized jointly. Specifically, our joint optimization generates aggregated representations that can be better reconstructed by some binary codes. This leads to more discriminative binary hash codes and improved retrieval accuracy. In addition, the proposed method is flexible. It can be extended for supervised hashing. When the data label is available, the framework can be adapted to learn binary codes which minimize the reconstruction loss with respect to label vectors. Furthermore, we also propose a fast version of the state-of-the-art hashing method Binary Autoencoder to be used in our proposed frameworks. Extensive experiments on benchmark datasets under various settings show that the proposed methods outperform the state-of-the-art unsupervised and supervised hashing methods. Thanh-Toan Do, Khoa Le, Tuan Hoang, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung |
IEEE Trans. Image Process. | 5 |
| 2019 | Semantic Prior Analysis for Salient Object DetectionabstractSalient object detection aims to detect the main objects in the given image. In this paper, we proposed an approach that integrates semantic priors into the salient object detection process. The method first obtains an explicit saliency map that is refined by the explicit semantic priors learned from data. Then an implicit saliency map is constructed using a trained model that maps the implicit semantic priors embedded into superpixel features with the saliency values. Next, the fusion saliency map is computed by adaptively fusing both the explicit and implicit semantic maps. The final saliency map is eventually computed via the post-processing refinement step. Experimental results have demonstrated the effectiveness of the proposed method, particularly, it achieves competitive performance with the state-of-the-art baselines on three challenging datasets, namely, ECSSD, HKUIS, and iCoSeg. Tam V. Nguyen 0002, Khanh-Duy Nguyen, Thanh-Toan Do |
IEEE Trans. Image Process. | 1 |
| 2019 | From Selective Deep Convolutional Features to Compact Binary Representations for Image RetrievalabstractIn the large-scale image retrieval task, the two most important requirements are the discriminability of image representations and the efficiency in computation and storage of representations. Regarding the former requirement, Convolutional Neural Network is proven to be a very powerful tool to extract highly discriminative local descriptors for effective image search. Additionally, to further improve the discriminative power of the descriptors, recent works adopt fine-tuned strategies. In this article, taking a different approach, we propose a novel, computationally efficient, and competitive framework. Specifically, we first propose various strategies to compute masks, namely, SIFT-masks , SUM-mask , and MAX-mask , to select a representative subset of local convolutional features and eliminate redundant features. Our in-depth analyses demonstrate that proposed masking schemes are effective to address the burstiness drawback and improve retrieval accuracy. Second, we propose to employ recent embedding and aggregating methods that can significantly boost the feature discriminability. Regarding the computation and storage efficiency, we include a hashing module to produce very compact binary image representations. Extensive experiments on six image retrieval benchmarks demonstrate that our proposed framework achieves the state-of-the-art retrieval performances. Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2018 | Revisit of Region-Feature Combinations in Facial AnalysisabstractFacial analysis has significant applications in human-computer interaction and multimedia communication research. In this paper, we aim to investigate the impact of facial regions and features in the facial analysis tasks. In particular, we compare the performances of various facial features extracted from different regions in two specific tasks, namely, age group prediction and twin recognition. Furthermore, we also evaluate the fusion of handcrafted features and deep learning features in both tasks. Through the experiments, we highlight different impacts of facial regions and facial features combination in facial analysis tasks. Zhongliang Nie, Akhil Mattey, Zhe Huang 0008, Tam V. Nguyen 0002 |
SMC | 4 |
| 2018 | Attentive Systems: A Survey
Tam V. Nguyen 0002, Shuicheng Yan |
Int. J. Comput. Vis. | 1 |
| 2018 | Kernel based online learning for imbalance multiclass classification
Shuya Ding, Bilal Mirza, Zhiping Lin 0001, Jiuwen Cao, Xiaoping Lai, Tam V. Nguyen 0002, Jose Sepulveda |
Neurocomputing | 6 |
| 2017 | Novel evaluation metrics for seam carving based image retargetingabstractImage retargeting effectively resizes images by preserving the recognizability of important image regions. Most of retargeting methods rely on good importance maps as a cue to retain or remove certain regions in the input image. In addition, the traditional evaluation exhaustively depends on user ratings. There is a legitimate need for a methodological approach for evaluating retargeted results. Therefore, in this paper, we conduct a study and analysis on the prominent method in image retargeting, Seam Carving. First, we introduce two novel evaluation metrics which can be considered as the proxy of user ratings. Second, we exploit salient object dataset as a benchmark for this task. We then investigate different types of importance maps for this particular problem. The experiments show that humans in general agree with the evaluation metrics on the retargeted results and some importance map methods are consistently more favorable than others. Tam V. Nguyen 0002, Guangyu Gao |
ICIP | 1 |
| 2017 | Salient Object Detection with Semantic PriorsabstractSalient object detection has increasingly become a popular topic in cognitive and computational sciences, including computer vision and artificial intelligence research. In this paper, we propose integrating semantic priors into the salient object detection process. Our algorithm consists of three basic steps. Firstly, the explicit saliency map is obtained based on the semantic segmentation refined by the explicit saliency priors learned from the data. Next, the implicit saliency map is computed based on a trained model which maps the implicit saliency priors embedded into regional features with the saliency values. Finally, the explicit semantic map and the implicit map are adaptively fused to form a pixel-accurate saliency map which uniformly covers the objects of interest. We further evaluate the proposed framework on two challenging datasets, namely, ECSSD and HKUIS. The extensive experimental results demonstrate that our method outperforms other state-of-the-art methods. Tam V. Nguyen 0002, Luoqi Liu |
IJCAI | 1 |
| 2017 | Shadow Puppetry with Robotic ArmsabstractIn this work, we aim to introduce our framework that performs the hand shadow puppetry using robotic arms. Human arms are not flexible and dynamic enough to produce many complicated shadow poses so we aim to utilize robotic arms for shadow puppetry. Firstly, we construct a shadow library to cover all the possible shadow images for a single robotic hand. For the input image, we first extract the shadow image by using a salient object detector. Then we match the shape correspondences between the input shadow image with the ones in the shadow library. Finally, we transfer the corresponding parameters of the best matching shadow in the library into proper utility to control the physical robotics arms. Zhe Huang 0008, Vamshi Krishna Madaram, Saad Albadrani, Tam V. Nguyen 0002 |
ACM Multimedia | 4 |
| 2017 | Smart Mirror: Intelligent Makeup Recommendation and SynthesisabstractThe female facial image beautification usually requires professional editing softwares, which are relatively difficult for common users. In this demo, we introduce a practical system for automatic and personalized facial makeup recommendation and synthesis. First, a model describing the relations among facial features, facial attributes and makeup attributes is learned as the makeup recommendation model for suggesting the most suitable makeup attributes. Then the recommended makeup attributes are seamlessly synthesized onto the input facial image. Tam V. Nguyen 0002, Luoqi Liu |
ACM Multimedia | 1 |
| 2017 | Dual-layer kernel extreme learning machine for action recognition
Tam V. Nguyen 0002, Bilal Mirza |
Neurocomputing | 1 |
| 2017 | As-similar-as-possible saliency fusion
Tam V. Nguyen 0002, Mohan Kankanhalli |
Multim. Tools Appl. | 1 |
| 2016 | Exploiting generic multi-level convolutional neural networks for scene understandingabstractIn this paper, we introduce the application of generic multi-level Convolutional Neural Networks (CNN) approach into the scene understanding or image parsing task. Given an input image, first, a set of similar images from the training set are retrieved based on global-level CNN feature matching similarities. Then, the input test image and the similar images are oversegmented into superpixels. Next, the class of each test image's superpixel is initialized by the majority vote of the k-nearest-neighbor superpixels based on regional-level CNN features and hand-crafted features matching. The initial superpixel parsing is later combined with per-exemplar sliding windows to roughly form the pixel labels. Eventually, the final labels are further refined by the contextual smoothing. Extensive experiments on different challenging datasets demonstrate the potentials of the proposed method. Tam V. Nguyen 0002, Luoqi Liu, Khang Nguyen 0001 |
ICARCV | 1 |
| 2016 | MARIM: Mobile Augmented Reality for Interactive ManualsabstractIn this work, we present a practical system which uses mobile devices for interactive manuals. In particular, there are two modes provided in the system, namely, expert/trainer and trainee modes. Given the expert/trainer editor, experts design the step-by-step interactive manuals. For each step, the experts capture the images by using phones/tablets and provide visual instructions such as interest regions, text, and action animations. In the trainee mode, the system utilizes the existing object detection and tracking algorithms to identify the step scene and retrieve the respective instruction to be displayed on the mobile device. The trainee then follows the displayed instruction. Once each step is performed, the trainee commands the devices to proceed to the next step. Tam V. Nguyen 0002, Dorothy Tan, Bilal Mirza, Jose Sepulveda |
ACM Multimedia | 1 |
| 2016 | Augmented immersion: video cutout and gesture-guided embedding for gaming applications
Tam V. Nguyen 0002, Jose Sepulveda |
Multim. Tools Appl. | 1 |
| 2015 | Salient Object Detection via Objectness ProposalsabstractSalient object detection has gradually become a popular topic in robotics and computer vision research. This paper presents a real-time system that detects salient object by integrating objectness, foreground and compactness measures. Our algorithm consists of four basic steps. First, our method generates the objectness map via object proposals. Based on the objectness map, we estimate the background margin and compute the corresponding foreground map which prefers the foreground objects. From the objectness map and the foreground map, the compactness map is formed to favor the compact objects. We then integrate those cues to form a pixel-accurate saliency map which covers the salient objects and consistently separates fore- and background. Tam V. Nguyen 0002 |
AAAI | 1 |
| 2015 | Salient Object Detection via Augmented Hypotheses
Tam V. Nguyen 0002, Jose Sepulveda |
IJCAI | 1 |
| 2015 | SalAd: A Multimodal Approach for Contextual Video AdvertisingabstractThe explosive growth of multimedia data on Internet has created huge opportunities for online video advertising. In this paper, we propose a novel advertising technique called SalAd, which utilizes textual information, visual content and the webpage saliency, to automatically associate the most suitable companion ads with online videos. Unlike most existing approaches that only focus on selecting the most relevant ads, SalAd further considers the saliency of selected ads to reduce intentional ignorance. SalAd consists of three basic steps. Given an online video and a set of advertisements, we first roughly identify a set of relevant ads based on the textual information matching. We then carefully select a sub-set of candidates based on visual content matching. In this regard, our selected ads are contextually relevant to online video content in terms of both textual information and visual content. We finally select the most salient ad among the relevant ads as the most appropriate one. To demonstrate the effectiveness of our method, we have conducted a rigorous eye-tracking experiment on two ad-datasets. The experimental results show that our method enhances the user engagement with the ad content while maintaining users' quality of video viewing experience. Chen Xiang, Tam V. Nguyen 0002, Mohan Kankanhalli |
ISM | 2 |
| 2015 | Sense Beyond Expressions: CutenessabstractWith the development of Internet culture, cute has become a popular concept. Many people are curious about what factors making a person look cute. However, there is rare research to answer this interesting question. In this work, we construct a dataset of personal images with comprehensively annotated cuteness scores and facial attributes to investigate this high-level concept in depth. Based on this dataset, through an automatic attributes mining process, we find several critical attributes determining the cuteness of a person. We also develop a novel Continuous Latent Support Vector Machine (C-LSVM) method to predict the cuteness score of one person given only his image. Extensive evaluations validate the effectiveness of the proposed method for cuteness prediction. Kang Wang 0002, Tam V. Nguyen 0002, Jiashi Feng, Jose Sepulveda |
ACM Multimedia | 2 |
| 2015 | Adaptive Nonparametric Image ParsingabstractIn this paper, we present an adaptive nonparametric solution to the image parsing task, namely, annotating each image pixel with its corresponding category label. For a given test image, first, a locality-aware retrieval set is extracted from the training data based on superpixel matching similarities, which are augmented with feature extraction for better differentiation of local superpixels. Then, the category of each superpixel is initialized by the majority vote of the k -nearest-neighbor superpixels in the retrieval set. Instead of fixing k as in traditional nonparametric approaches, here, we propose a novel adaptive nonparametric approach that determines the sample-specific k for each test image. In particular, k is adaptively set to be the number of the fewest nearest superpixels that the images in the retrieval set can use to get the best category prediction. Finally, the initial superpixel labels are further refined by contextual smoothing. Extensive experiments on challenging data sets demonstrate the superiority of the new solution over other state-of-the-art nonparametric solutions. Tam V. Nguyen 0002, Canyi Lu, Jose Sepulveda, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | STAP: Spatial-Temporal Attention-Aware Pooling for Action RecognitionabstractHuman action recognition is valuable for numerous practical applications, e.g., gaming, video surveillance, and video search. In this paper we hypothesize that the classification of actions can be boosted by designing a smart feature pooling strategy under the prevalently used bag-of-words-based representation. Founded on automatic video saliency analysis, we propose the spatial-temporal attention-aware pooling scheme for feature pooling. First, the video saliencies are predicted using the video saliency model, and the localized spatial-temporal features are pooled at different saliency levels and video-saliency-guided channels are formed. Saliency-aware matching kernels are thus derived as the similarity measurement of these channels. Intuitively, the proposed kernels calculate the similarities of the video foreground (salient areas) or background (nonsalient areas) at different levels. Finally, the kernels are fed into popular support vector machines for action classification. Extensive experiments on three popular data sets for action classification validate the effectiveness of our proposed method, which outperforms the state-of-the-art methods, namely 95.3% on UCF Sports (better by 4.0%), 87.9% on YouTube data set (better by 2.5%), and achieves comparable results on Hollywood2 dataset. Tam V. Nguyen 0002, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Seeing Human Weight from a Single RGB-D Image
Tam V. Nguyen 0002, Jiashi Feng, Shuicheng Yan |
J. Comput. Sci. Technol. | 1 |
| 2014 | Audio Matters in Visual AttentionabstractThere is a dearth of information on how perceived auditory information guides image-viewing behavior. To investigate auditory-driven visual attention, we first generated a human eye-fixation database from a pool of 200 static images and 400 image-audio pairs viewed by 48 subjects. The eye tracking data for the image-audio pairs were captured while participants viewed images, which took place immediately after exposure to coherent/incoherent audio samples. The database was analyzed in terms of time to first fixation, fixation durations on the target object, entropy, AUC, and saliency ratio. It was found that coherent audio information is an important cue for enhancing the feature-specific response to the target object. Conversely, incoherent audio information attenuates this response. Finally, a system predicting the image-viewing with the influence of different audio sources was developed. The detailedly discussed top-down module in the system is composed of auditory estimation based on Gaussian mixture model-maximum a posteriori algorithm-universal background model structure, as well as visual estimation based on the conditional random field model and sparse latent variables. The evaluation experiments show that the proposed models in the system exhibit strong consistency with eye fixations. Tam V. Nguyen 0002, Mohan Kankanhalli, Shuicheng Yan, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Touch Saliency: Characteristics and PredictionabstractIn this work, we propose an alternative ground truth to the eye fixation map in visual attention study, called touch saliency. As it can be directly collected from the recorded data of users' daily browsing behavior on widely used smart phone devices with touch screens, the touch saliency data is easy to obtain. Due to the limited screen size, smart phone users usually move and zoom in the images, and fix the region of interest on the screen when browsing images. Our studies are two-fold. First, we collect and study the characteristics of these touch screen fixation maps (named touch saliency) by comprehensive comparisons with their counterpart, the eye-fixation maps (namely, visual saliency). The comparisons show that the touch saliency is highly correlated with the eye fixations for the same stimuli, which indicates its utility in data collection for visual attention study. Based on the consistency between both touch saliency and visual saliency, our second task is to propose a unified saliency prediction model for both visual and touch saliency detection. This model utilizes middle-level object category features extracted from pre-segmented image superpixels as input to the recently proposed multitask sparsity pursuit (MTSP) framework for saliency prediction. Extensive evaluations show that the proposed middle-level category features can considerably improve the saliency prediction performance when taking both touch saliency and visual saliency as ground truth. Bingbing Ni, Mengdi Xu, Tam V. Nguyen 0002, Meng Wang 0001, Congyan Lang, ZhongYang Huang, Shuicheng Yan |
IEEE Trans. Multim. | 3 |
| 2013 | Static saliency vs. dynamic saliency: a comparative studyabstractRecently visual saliency has attracted wide attention of researchers in the computer vision and multimedia field. However, most of the visual saliency-related research was conducted on still images for studying static saliency. In this paper, we give a comprehensive comparative study for the first time of dynamic saliency (video shots) and static saliency (key frames of the corresponding video shots), and two key observations are obtained: 1) video saliency is often different from, yet quite related with, image saliency, and 2) camera motions, such as tilting, panning or zooming, affect dynamic saliency significantly. Motivated by these observations, we propose a novel camera motion and image saliency aware model for dynamic saliency prediction. The extensive experiments on two static-vs-dynamic saliency datasets collected by us show that our proposed method outperforms the state-of-the-art methods for dynamic saliency prediction. Finally, we also introduce the application of dynamic saliency prediction for dynamic video captioning, assisting people with hearing impairments to better entertain videos with only off-screen voices, e.g., documentary films, news videos and sports videos. Tam V. Nguyen 0002, Mengdi Xu, Guangyu Gao, Mohan Kankanhalli, Qi Tian 0001, Shuicheng Yan |
ACM Multimedia | 1 |
| 2013 | Image Re-AttentionizingabstractIn this paper, we propose a computational framework, called Image Re-Attentionizing, to endow the target region in an image with the ability of attracting human visual attention. In particular, the objective is to recolor the target patches by color transfer with naturalness and smoothness preserved yet visual attention augmented. We propose to approach this objective within the Markov Random Field (MRF) framework and an extended graph cuts method is developed to pursue the solution. The input image is first over-segmented into patches, and the patches within the target region as well as their neighbors are used to construct the consistency graphs. Within the MRF framework, the unitary potentials are defined to encourage each target patch to match the patches with similar shapes and textures from a large salient patch database, each of which corresponds to a high-saliency region in one image, while the spatial and color coherence is reinforced as pairwise potentials. We evaluate the proposed method on the direct human fixation data. The results demonstrate that the target region(s) successfully attract human attention and in the meantime both spatial and color coherence is well preserved. Tam V. Nguyen 0002, Bingbing Ni, Hairong Liu, Jiebo Luo 0001, Mohan Kankanhalli, Shuicheng Yan |
IEEE Trans. Multim. | 1 |
| 2013 | Towards decrypting attractiveness via multi-modality cuesabstractDecrypting the secret of beauty or attractiveness has been the pursuit of artists and philosophers for centuries. To date, the computational model for attractiveness estimation has been actively explored in computer vision and multimedia community, yet with the focus mainly on facial features. In this article, we conduct a comprehensive study on female attractiveness conveyed by single/multiple modalities of cues, that is, face, dressing and/or voice, and aim to discover how different modalities individually and collectively affect the human sense of beauty. To extensively investigate the problem, we collect the Multi-Modality Beauty (M2B) dataset, which is annotated with attractiveness levels converted from manualk-wise ratings and semantic attributes of different modalities. Inspired by the common consensus that middle-level attribute prediction can assist higher-level computer vision tasks, we manually labeled many attributes for each modality. Next, a tri-layer Dual-supervised Feature-Attribute-Task (DFAT) network is proposed to jointly learn the attribute model and attractiveness model of single/multiple modalities. To remedy possible loss of information caused by incomplete manual attributes, we also propose a novel Latent Dual-supervised Feature-Attribute-Task (LDFAT) network, where latent attributes are combined with manual attributes to contribute to the final attractiveness estimation. The extensive experimental evaluations on the collected M2B dataset well demonstrate the effectiveness of the proposed DFAT and LDFAT networks for female attractiveness prediction. Tam V. Nguyen 0002, Si Liu 0001, Bingbing Ni, Yong Rui, Shuicheng Yan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | Depth Matters: Influence of Depth Cues on Visual Saliency
Congyan Lang, Tam V. Nguyen 0002, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Shuicheng Yan |
ECCV (2) | 2 |
| 2012 | Hi, magic closet, tell me what to wear!abstractIn this demo, we present a practical system, magic closet, for automatic occasion-oriented clothing pairing. Given a user-input occasion, e.g., wedding or shopping, the magic closet intelligently and automatically pairs the user-specified reference clothing (upper-body or lower-body) with the most suitable one from online shops. Two key criteria are explicitly considered for the magic closet system. One criterion is to wear properly, e.g., compared to suit pants, it is more decent to wear a cocktail dress for a banquet occasion. The other criterion is to wear aesthetically, e.g., a red T-shirt matches better white pants than green pants. To narrow the semantic gap between the low-level visual features and the high-level occasion categories, we propose to adopt middle-level clothing attributes (e.g., clothing category, color, pattern) as a bridge. More specifically, the clothing attributes are treated as latent variables in our proposed latent Support Vector Machine (SVM) based recommendation model. The wearing properly criterion is described through a feature-occasion potential and an attribute-occasion potential, while the wearing aesthetically criterion is expressed by an attribute-attribute potential. Si Liu 0001, Tam V. Nguyen 0002, Jiashi Feng, Meng Wang 0001, Shuicheng Yan |
ACM Multimedia | 2 |
| 2012 | Sense beauty via face, dressing, and/or voiceabstractDiscovering the secret of beauty has been the pursuit of artists and philosophers for centuries. Nowadays, the computational model for beauty estimation has been actively explored in computer science community, yet with the focus mainly on facial features. In this work, we perform a comprehensive study of female attractiveness conveyed by single/multiple modalities of cues, i.e., face, dressing and/or voice, and aim to uncover how different modalities individually and collectively affect the human sense of beauty. To this end, we collect the first Multi-Modality Beauty (M2B) dataset in the world for female attractiveness study, which is thoroughly annotated with attractiveness levels converted from manual k-wise ratings and semantic attributes of different modalities. A novel Dual-supervised Feature-Attribute-Task (DFAT) network is proposed to jointly learn the beauty estimation models of single/multiple modalities as well as the attribute estimation models. The DFAT network differentiates itself by its supervision in both attribute and task layers. Several interesting beauty-sense observations over single/multiple modalities are reported, and the extensive experimental evaluations on the collected M2B dataset well demonstrate the effectiveness of the proposed DFAT network for female attractiveness estimation. Tam V. Nguyen 0002, Si Liu 0001, Bingbing Ni, Yong Rui, Shuicheng Yan |
ACM Multimedia | 1 |
| 2012 | 3DME: 3D media express from RGB-D imagesabstractConsidering the continuously increasing availability and accessibility of 3D media and the depth camera such as Kinect, we demonstrate an innovative 3D media system called 3DME. The objective of this demo is three-fold. First, the demo exhibits the creation of 3D images from RGB-D images. Second, 3DME allows a user to insert impressive effects to the produced 3D content. Last but not least, our demo is one of the first attempts towards advertising for 3D content which enables both the advertisers and content providers deliver more effective ads carried through 3D media. Tam V. Nguyen 0002, Lusong Li, Shuicheng Yan |
ACM Multimedia | 1 |
| 2009 | How to Maximize User Satisfaction Degree in Multi-service IP NetworksabstractBandwidth allocation is a fundamental problem in communication networks. With current network moving towards the future Internet model, the problem is further intensified as network traffic demanding far from exceeds network bandwidth capability. Maintaining a certain user satisfaction degree therefore becomes a challenge research topic. In this paper, we deal with the problem by proposing BASMIN, a novel bandwidth allocation scheme that aims to maximize network userpsilas happiness. We also defined a new metric for evaluating network user satisfaction degree: network worth. A three-step evaluation process is then conducted to compare BASMIN efficiency with other three popular bandwidth allocation schemes. Throughout the tests, we experienced BASMIN's advantages over the others; we even found out that one of the most widely used bandwidth allocation schemes, in fact, is not effective at all. Huy Anh Nguyen, Tam V. Nguyen 0002, Deokjai Choi 0001 |
ACIIDS | 2 |
| 2009 | CCBR: Chaining Case Based Reasoning in Context-Aware Smart HomeabstractIn ubiquitous computing environments like smart home, context awareness is a very important component which aims at providing automatic services. This paper proposes chaining case based reasoning (CCBR) as the reasoning method which solves the vagueness of traditional case based reasoning (CBR) approach "In the certain case with more than one solution, we donpsilat know which solution or activity will be chosen to satisfy user's needs in smart homepsilas context". The context's contents in smart home are described in this paper. Also, we introduce the framework of context awareness based on CCBR, and discuss the case representation, case adaptation, and similarity computation in detail. Our proposed CCBR integrated into the virtual smart home environment acquires knowledge about user actions that are recorded to determine their preferences and then simultaneously activates the devices with predefined settings. Tam V. Nguyen 0002, Yi Chang Woo, Deokjai Choi 0001 |
ACIIDS | 1 |