Khoa Luu

dblp:43/8092 · DBLP profile ↗
← Back
73ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0003-2104-0901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-author · 17 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 COBRA: A Continual Learning Approach to Vision-Brain Understanding
Xuan-Bac Nguyen, Manuel Serna-Aguilera, Arabinda Kumar Choudhary, Pawan Sinha, Xin Li 0005, Khoa Luu
Int. J. Comput. Vis.6
2026 BRACTIVE: A Brain Activation Approach to Human Visual Brain Learning
abstract
The human brain is a highly efficient processing unit, and understanding how it works can inspire new algorithms and architectures in machine learning. In this work, we introduce a novel framework named Brain Activation Network (BRACTIVE), a transformer-based approach to studying the human visual brain. The primary objective of BRACTIVE is to align the visual features of subjects with their corresponding brain representations using functional Magnetic Resonance Imaging (fMRI) signals. It enables us to identify the brain's Regions of Interest (ROIs) in the subjects. Unlike previous brain research methods, which can only identify ROIs for one subject at a time and are limited by the number of subjects, BRACTIVE automatically extends this identification to multiple subjects and ROIs. Our experiments demonstrate that BRACTIVE effectively identifies person-specific regions of interest, such as face and body-selective areas, aligning with neuroscience findings and indicating potential applicability to various object categories. More importantly, we found that leveraging human visual brain activity to guide deep neural networks enhances performance across various benchmarks. It encourages the potential of BRACTIVE in both neuroscience and machine intelligence studies.
Xuan-Bac Nguyen, Hojin Jang, Xin Li 0005, Samee Ullah Khan, Pawan Sinha, Khoa Luu
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
abstract
Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their ability to handle complex multi-object interactions and reasoning. To this end, we propose the Multimodal Large Language Models (LLMs) on a Scene HyperGraph (HyperGLM), promoting reasoning about multi-way interactions and higher-order relationships. Our approach uniquely integrates entity scene graphs, which capture spatial relationships between objects, with a procedural graph that models their causal transitions, forming a unified HyperGraph. Significantly, HyperGLM enables reasoning by injecting this unified HyperGraph into LLMs. Additionally, we introduce a new Video Scene Graph Reasoning (VSGR) dataset featuring 1.9M frames from third-person, egocentric, and drone views and support five tasks. Empirically, HyperGLM consistently outperforms state-of-the-art methods, effectively modeling and reasoning complex relationships in diverse scenes.
Pha Nguyen, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu
CVPR5
2025 FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding
abstract
Continual Learning in semantic scene segmentation aims to continually learn new unseen classes in dynamic environments while maintaining previously learned knowledge. Prior studies focused on modeling the catastrophic forgetting and background shift challenges in continual learning. However, fairness, another major challenge that causes unfair predictions leading to low performance among major and minor classes, still needs to be well addressed. In addition, prior methods have yet to model the unknown classes well, thus resulting in producing non-discriminative features among unknown classes. This work presents a novel Fairness Learning via Contrastive Attention Approach to continual learning in semantic scene understanding. In particular, we first introduce a new Fairness Contrastive Clustering loss to address the problems of catastrophic forgetting and fairness. Then, we propose an attention-based visual grammar approach to effectively model the background shift problem and unknown classes, producing better feature representations for different unknown classes. Through our experiments, our proposed approach achieves State-of-the-Art (SoTA) performance on different continual learning benchmarks, i.e., ADE20K, Cityscapes, and Pascal VOC. It promotes the fairness of the continual semantic segmentation model.
Thanh-Dat Truong, Utsav Prabhu, Bhiksha Raj, Jackson David Cothren, Khoa Luu
CVPR5
2025 MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
abstract
Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the multimodal model cannot capture the essential features of each modality, making it difficult to comprehend complex structures and correlations of multimodal inputs. This paper introduces a novel Multimodal Attention-based Normalizing Flow (MANGO) approach to developing explicit, interpretable, and tractable multimodal fusion learning. In particular, we propose a new Invertible Cross-Attention (ICA) layer to develop the Normalizing Flow-based Model for multimodal data. To efficiently capture the complex, underlying correlations in multimodal data in our proposed invertible cross-attention layer, we propose three new cross-attention mechanisms: Modality-to-Modality Cross-Attention (MMCA), Inter-Modality Cross-Attention (IMCA), and Learnable Inter-Modality Cross-Attention (LICA). Finally, we introduce a new Multimodal Attention-based Normalizing Flow to enable the scalability of our proposed method to high-dimensional multimodal data. Our experimental results on three different multimodal learning tasks, i.e., semantic segmentation, image-to-image translation, and movie genre classification, have illustrated the state-of-the-art (SoTA) performance of the proposed approach.
Thanh-Dat Truong, Christophe Bobda, Khoa Luu
NeurIPS4
2025 Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
abstract
Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and textual features. In this paper, we introduce a simple but efficient learning mechanism for improving the robust alignment between visual and textual modalities by solving shuffling problems. In particular, the proposed approach can improve reasoning capability, visual understanding, and cross-modality alignment by introducing two new tasks: reconstructing the image order and the text order into the LMM's pre-training and fine-tuning phases. In addition, we propose a new directed-token approach to capture visual and textual knowledge, enabling the capability to reconstruct the correct order of visual inputs. Then, we introduce a new Image-to-Response Guided loss to further improve the visual understanding of the LMM in its responses. The proposed approach consistently achieves state-of-the-art (SoTA) performance compared with prior LMMs on academic task-oriented and instruction-following LMM benchmarks.
Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj, Khoa Luu
NeurIPS5
2025 LiGAR: LiDAR-Guided Hierarchical Transformer for Multi-Modal Group Activity Recognition
abstract
Group Activity Recognition (GAR) remains challenging in computer vision due to the complex nature of multi-agent interactions. This paper introduces LiGAR, a LIDAR-Guided Hierarchical Transformer for Multi-Modal Group Activity Recognition. LiGAR leverages LiDAR data as a structural backbone to guide the processing of visual and textual information, enabling robust handling of occlusions and complex spatial arrangements. Our framework incorporates a Multi-Scale LIDAR Transformer, Cross-Modal Guided Attention, and an Adaptive Fusion Module to integrate multi-modal data at different semantic levels effectively. LiGAR's hierarchical architecture captures group activities at various granularities, from individual actions to scene-level dynamics. Extensive experiments on the JRDB-PAR, Volleyball, and NBA datasets demonstrate LiGAR's superior performance, achieving state-of-the-art results with improvements of up to 10.6% in F1-score on JRDB-PAR and 5.9% in Mean Per Class Accuracy on the NBA dataset. Notably, LiGAR maintains high performance even when LiDAR data is unavailable during inference, showcasing its adaptability. Our ablation studies highlight the significant contributions of each component and the effectiveness of our multi-modal, multi-scale approach in advancing the field of group activity recognition.
Naga Venkata Sai Raviteja Chappa, Khoa Luu
WACV2
2025 Autoregressive Temporal Modeling for Advanced Tracking-by-Diffusion
abstract
Object tracking is a widely studied computer vision task with video and instance analysis applications. While paradigms such as tracking-by-regression , -detection , -attention have advanced the field, generative modeling offers new potential. Although some studies explore the generative process in instance-based understanding tasks, they rely on prediction refinement in the coordinate space rather than the visual domain. Instead, this paper presents Tracking-by-Diffusion , a novel paradigm for object tracking in video, leveraging visual generative models via the perspective of autoregressive models. This paradigm demonstrates broad applicability across point, box, and mask modalities while uniquely enabling textual guidance. We present DIFTracker, a framework that utilizes iterative latent variable diffusion models to redefine tracking as a next-frame reconstruction task. Our approach uniquely combines spatial and temporal dependencies in video data, offering a unified solution that encompasses existing tracking paradigms within a single Inversion-Reconstruction process. DIFTracker operates online and auto-regressively, enabling flexible instance-based video understanding. It allows us to overcome difficulties in variable-length video understanding encountered by video-inflated models and perform superior performance on seven benchmarks across five modalities. This paper not only introduces a new perspective on visual autoregressive modeling in understanding sequential visual data, specifically videos, but also provides robust theoretical validations and demonstrates broader applications in visual tracking and computer vision.
Pha A. Nguyen, Rishi Madhok, Bhiksha Raj, Khoa Luu
Int. J. Comput. Vis.4
2025 Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
abstract
Abstract Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data. However, current conversational models still lack knowledge about visual insects since they are often trained on the general knowledge of vision-language data. Meanwhile, understanding insects is a fundamental problem in precision agriculture, helping to promote sustainable development in agriculture. Therefore, this paper proposes a novel multimodal conversational model, Insect-LLaVA, to promote visual understanding in insect-domain knowledge. In particular, we first introduce a new large-scale Multimodal Insect Dataset with Visual Insect Instruction Data that enables the capability of learning the multimodal foundation models. Our proposed dataset enables conversational models to comprehend the visual and semantic features of the insects. Second, we propose a new Insect-LLaVA model, a new general Large Language and Vision Assistant in Visual Insect Understanding. Then, to enhance the capability of learning insect features, we develop an Insect Foundation Model by introducing a new micro-feature self-supervised learning with a Patch-wise Relevant Attention mechanism to capture the subtle differences among insect images. We also present Description Consistency loss to improve micro-feature learning via text descriptions. The experimental results evaluated on our new Visual Insect Question Answering benchmarks illustrate the effective performance of our proposed approach in visual insect understanding and achieve State-of-the-Art performance on standard benchmarks of insect-related tasks Project Page: https://uarkcviu.github.io/projects/insectfoundation .
Thanh-Dat Truong, Hoang-Quan Nguyen, Xuan-Bac Nguyen, Ashley Dowling, Xin Li 0005, Khoa Luu
Int. J. Comput. Vis.6
2025 Brainformer: Mimic human visual brain functions to machine vision models via fMRI
Xuan-Bac Nguyen, Xin Li 0005, Pawan Sinha, Samee Ullah Khan, Khoa Luu
Neurocomputing5
2025 Cross-view action recognition understanding from exocentric to egocentric perspective
Thanh-Dat Truong, Khoa Luu
Neurocomputing2
2024 HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
abstract
Visual interactivity understanding within visual scenes presents a significant challenge in computer vision. Existing methods focus on complex interactivities while leveraging a simple relationship model. These methods, however, struggle with a diversity of appearance, situation, position, interaction, and relation in videos. This limitation hinders the ability to fully comprehend the interplay within the complex visual dynamics of subjects. In this paper, we delve into interactivities understanding within visual content by deriving scene graph representations from dense interactivities among humans and objects. To achieve this goal, we first present a new dataset containing Appearance-Situation-Position-Interaction-Relation predicates, named ASPIRe, offering an extensive collection of videos marked by a wide range of interactivities. Then, we propose a new approach named Hierarchical Interlacement Graph (HIG), which leverages a unified layer and graph within a hierarchical structure to provide deep insights into scene changes across five distinct tasks. Our approach demonstrates superior performance to other methods through extensive experiments conducted in various scenarios.
Pha A. Nguyen, Khoa Luu
CVPR3
2024 Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect Understanding
abstract
In precision agriculture, the detection and recognition of insects play an essential role in the ability of crops to grow healthy and produce a high-quality yield. The current machine vision model requires a large volume of data to achieve high performance. However, there are approximately 5.5 million different insect species in the world. None of the existing insect datasets can cover even a fraction of them due to varying geographic locations and acquisition costs. In this paper, we introduce a novel “Insect-1M” dataset, a game-changing resource poised to revolutionize insect-related foundation model training. Covering a vast spectrum of insect species, our dataset, including 1 million images with dense identification labels of taxonomy hierarchy and insect descriptions, offers a panoramic view of entomology, enabling foundation models to comprehend visual and semantic information about insects like never before. Then, to efficiently establish an Insect Foundation Model, we develop a micro-feature self-supervised learning method with a Patch-wise Relevant Attention mechanism capable of discerning the subtle differences among insect images. In addition, we introduce Description Consistency loss to improve micro-feature modeling via insect descriptions. Through our experiments, we illustrate the effectiveness of our proposed approach in insect modeling and achieve State-of-the-Art performance on standard benchmarks of insect-related tasks. Our Insect Foundation Model and Dataset promise to empower the next generation of insect-related vision models, bringing them closer to the ultimate goal of precision agriculture.
Hoang-Quan Nguyen, Thanh-Dat Truong, Xuan-Bac Nguyen, Ashley Dowling, Xin Li 0005, Khoa Luu
CVPR6
2024 DINTR: Tracking via Diffusion-based Interpolation
abstract
Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work proposes a novel diffusion-based methodology to formulate the tracking task. Firstly, their conditional process allows for injecting indications of the target object into the generation process. Secondly, diffusion mechanics can be developed to inherently model temporal correspondences, enabling the reconstruction of actual frames in video. However, existing diffusion models rely on extensive and unnecessary mapping to a Gaussian noise domain, which can be replaced by a more efficient and stable interpolation process. Our proposed interpolation mechanism draws inspiration from classic image-processing techniques, offering a more interpretable, stable, and faster approach tailored specifically for the object tracking task. By leveraging the strengths of diffusion models while circumventing their limitations, our Diffusion-based INterpolation TrackeR (DINTR) presents a promising new paradigm and achieves a superior multiplicity on seven benchmarks across five indicator representations.
Pha A. Nguyen, T. Hoang Ngan Le, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu
NeurIPS5
2024 CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos
abstract
Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics in video sequences. In this paper, we introduce the new AeroEye dataset that focuses on multi-object relationship modeling in aerial videos. Our AeroEye dataset features various drone scenes and includes a visually comprehensive and precise collection of predicates that capture the intricate relationships and spatial arrangements among objects. To this end, we propose the novel Cyclic Graph Transformer (CYCLO) approach that allows the model to capture both direct and long-range temporal dependencies by continuously updating the history of interactions in a circular manner. The proposed approach also allows one to handle sequences with inherent cyclical patterns and process object relationships in the correct sequential order. Therefore, it can effectively capture periodic and overlapping relationships while minimizing information loss. The extensive experiments on the AeroEye dataset demonstrate the effectiveness of the proposed CYCLO model, demonstrating its potential to perform scene understanding on drone videos. Finally, the CYCLO method consistently achieves State-of-the-Art (SOTA) results on two in-the-wild scene graph generation benchmarks, i.e., PVSG and ASPIRe.
Pha A. Nguyen, Xin Li 0005, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu
NeurIPS6
2024 EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding
abstract
Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language models is effective in open-set settings because it can learn diverse concepts and categories. However, these prior methods fail to generalize across different camera views due to the lack of cross-view geometric modeling. At present, there are limited studies analyzing cross-view learning. To address this problem, we introduce a novel Unsupervised Cross-view Adaptation Learning approach to modeling the geometric structural change across views in Semantic Scene Understanding. First, we introduce a novel Cross-view Geometric Constraint on Unpaired Data to model structural changes in images and segmentation masks across cameras. Second, we present a new Geodesic Flow-based Correlation Metric to efficiently measure the geometric structural changes across camera views. Third, we introduce a novel view-condition prompting mechanism to enhance the view-information modeling of the open-vocabulary segmentation network in cross-view adaptation learning. The experiments on different cross-view adaptation benchmarks have shown the effectiveness of our approach in cross-view modeling, demonstrating that we achieve State-of-the-Art (SOTA) performance compared to prior unsupervised domain adaptation and open-vocabulary semantic segmentation methods.
Thanh-Dat Truong, Utsav Prabhu, Dongyi Wang, Bhiksha Raj, Susan Gauch, Jeyamkondan Subbiah, Khoa Luu
NeurIPS7
2024 React: recognize every action everywhere all at once
abstract
Abstract In the realm of computer vision, Group Activity Recognition (GAR) plays a vital role, finding applications in sports video analysis, surveillance, and social scene understanding. This paper introduces R ecognize E very Act ion Everywhere All At Once (REACT), a novel architecture designed to model complex contextual relationships within videos. REACT leverages advanced transformer-based models for encoding intricate contextual relationships, enhancing understanding of group dynamics. Integrated Vision-Language Encoding facilitates efficient capture of spatiotemporal interactions and multi-modal information, enabling comprehensive scene understanding. The model’s precise action localization refines joint understanding of text and video data, enabling precise bounding box retrieval and enhancing semantic links between textual descriptions and visual reality. Actor-Specific Fusion strikes a balance between actor-specific details and contextual information, improving model specificity and robustness in recognizing group activities. Experimental results demonstrate REACT’s superiority over state-of-the-art GAR approaches, achieving higher accuracy in recognizing and understanding group activities across diverse datasets. This work significantly advances group activity recognition, offering a robust framework for nuanced scene comprehension.
Naga Venkata Sai Raviteja Chappa, Pha A. Nguyen, Page Daniel Dobbs, Khoa Luu
Mach. Vis. Appl.4
2024 Multi-camera multi-object tracking on the move via single-stage global association approach
Pha A. Nguyen, Kha Gia Quach, Chi Nhan Duong, Son Lam Phung, T. Hoang Ngan Le, Khoa Luu
Pattern Recognit.6
2023 Micron-BERT: BERT-Based Facial Micro-Expression Recognition
abstract
Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT ($(\mu$-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed$\mu$-BERT significantly outperforms all previous work in various micro-expression tasks.$\mu$-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show$\mu$-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT
Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li 0005, Susan Gauch, Han-Seok Seo, Khoa Luu
CVPR6
2023 FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding
abstract
Although Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into human-related real-world applications, e.g., autonomous driving, as any unfair predictions could influence human safety. In this paper, we propose a novel Fairness Domain Adaptation (FREDOM) approach to semantic scene segmentation. In particular, from the proposed formulated fairness objective, a new adaptation framework will be introduced based on the fair treatment of class distributions. Moreover, to generally model the context of structural dependency, a new conditional structural constraint is introduced to impose the consistency of predicted segmentation. Thanks to the proposed Conditional Structure Network, the self-attention mechanism has sufficiently modeled the structural information of segmentation. Through the ablation studies, the proposed method has shown the performance improvement of the segmentation models and promoted fairness in the model predictions. The experimental results on the two standard benchmarks, i.e., SYNTHIA$\rightarrow$Cityscapes and GTA5$\rightarrow$Cityscapes, have shown that our method achieved State-of-the-Art (SOTA) performance11The implementation of FREDOM is available at https://github.com/uark-cviu/FREDOM
Thanh-Dat Truong, T. Hoang Ngan Le, Bhiksha Raj, Jackson David Cothren, Khoa Luu
CVPR5
2023 Type-to-Track: Retrieve Any Object via Prompt-based Tracking
abstract
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest. This approach can overcome some limitations of traditional methods that rely on bounding boxes or category annotations. This paper introduces a novel paradigm for Multiple Object Tracking called Type-to-Track, which allows users to track objects in videos by typing natural language descriptions. We present a new dataset for that Grounded Multiple Object Tracking task, called GroOT, that contains videos with various types of objects and their corresponding textual captions describing their appearance and action in detail. Additionally, we introduce two new evaluation protocols and formulate evaluation metrics specifically for this task. We develop a new efficient method that models a transformer-based eMbed-ENcoDE-extRact framework (MENDER) using the third-order tensor decomposition. The experiments in five scenarios show that our MENDER approach outperforms another two-stage design in terms of accuracy and efficiency, up to 14.7\% accuracy and $4\times$ speed faster.
Pha A. Nguyen, Kha Gia Quach, Kris Makoto Kitani, Khoa Luu
NeurIPS4
2023 Fairness Continual Learning Approach to Semantic Scene Understanding in Open-World Environments
abstract
Continual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness is one of the most vital factors in deploying the deep learning model, especially in human-related or safety applications. In this paper, we present a novel Fairness Continual Learning approach to the semantic segmentation problem. In particular, under the fairness objective, a new fairness continual learning framework is proposed based on class distributions. Then, a novel Prototypical Contrastive Clustering loss is proposed to address the significant challenges in continual learning, i.e., catastrophic forgetting and background shift. Our proposed loss has also been proven as a novel, generalized learning paradigm of knowledge distillation commonly used in continual learning. Moreover, the proposed Conditional Structural Consistency loss further regularized the structural constraint of the predicted segmentation. Our proposed approach has achieved State-of-the-Art performance on three standard scene understanding benchmarks, i.e., ADE20K, Cityscapes, and Pascal VOC, and promoted the fairness of the segmentation model.
Thanh-Dat Truong, Hoang-Quan Nguyen, Bhiksha Raj, Khoa Luu
NeurIPS4
2023 POP-HIT: Partially Order-Preserving Hash-Induced Transformation for Privacy Protection in Face Recognition Access Control
Yatish Dubasi, Khoa Luu
SecureComm (2)3
2023 LIAAD: Lightweight attentive angular distillation for large-scale age-invariant face recognition
Thanh-Dat Truong, Chi Nhan Duong, Kha Gia Quach, T. Hoang Ngan Le, Tien D. Bui, Khoa Luu
Neurocomputing6
2022 DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition
abstract
Human action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video action recognition with competitive results. However, these methods have suffered some fundamental limitations such as lack of robustness and generalization, e.g., how does the temporal ordering of video frames affect the recognition results? This work presents a novel end-to-end Transformer-based Directed Attention (Direc-Former) framework11The implementation of DirecFormer is available at https://github.com/uark-cviu/DirecFormer for robust action recognition. The method takes a simple but novel perspective of Transformer-based approach to understand the right order of sequence actions. Therefore, the contributions of this work are three-fold. Firstly, we introduce the problem of ordered temporal learning issues to the action recognition problem. Secondly, a new Directed Attention mechanism is introduced to understand and provide attentions to human actions in the right order. Thirdly, we introduce the conditional dependency in action sequence modeling that includes orders and classes. The proposed approach consistently achieves the state-of-the-art (SOTA) results compared with the recent action recognition methods [4, 18, 72, 74]. on three standard large-scale benchmarks, i.e. Jester, Kinetics-400 and Something-Something-V2.
Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li 0005, Khoa Luu
CVPR7
2022 Self-Supervised Domain Adaptation in Crowd Counting
abstract
Self-training crowd counting has not been attentively explored though it is one of the important challenges in computer vision. In practice, the fully supervised methods usually require an intensive resource of manual annotation. In order to address this challenge, this work introduces a new approach to utilize existing datasets with ground truth to produce more robust predictions on unlabeled datasets, named domain adaptation, in crowd counting. While the network is trained with labeled data, samples without labels from the target domain are also added to the training process. In this process, the entropy map is computed and minimized in addition to the adversarial training process designed in parallel. Experiments on Shanghaitech, UCF_CC_50, and UCF-QNRF datasets prove a more generalized improvement of our method over the other state-of-the-arts in the cross-domain setting.
Pha A. Nguyen, Thanh-Dat Truong, Miaoqing Huang, T. Hoang Ngan Le, Khoa Luu
ICIP6
2022 VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning
abstract
In this paper, we leverage the human perceiving process, that involves vision and language interaction, to generate a coherent paragraph description of untrimmed videos. We propose vision-language (VL) features consisting of two modalities, i.e., (i) vision modality to capture global visual content of the entire scene and (ii) language modality to extract scene elements description of both human and non-human objects (e.g. animals, vehicles, etc), visual and non-visual elements (e.g. relations, activities, etc). Furthermore, we propose to train our proposed VLCap under a contrastive learning VL loss. The experiments and ablation studies on ActivityNet Captions and YouCookII datasets show that our VLCap outperforms existing SOTA methods on both accuracy and diversity metrics. Source code: https://github.com/UARK-AICV/VLCAP
Kashu Yamazaki, Sang Truong, Viet-Khoa Vo-Ho, Michael Kidd, Chase Rainwater, Khoa Luu, T. Hoang Ngan Le
ICIP6
2022 OTAdapt: Optimal Transport-based Approach For Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation is one of the challenging problems in computer vision. This paper presents a novel approach to unsupervised domain adaptations based on the optimal transport-based distance. Our approach allows aligning target and source domains without the requirement of meaningful metrics across domains. In addition, the proposal can associate the correct mapping between source and target domains and guarantee a constraint of topology between source and target domains. The proposed method is evaluated on different datasets in various problems, i.e. (i) digit recognition on MNIST, MNISTM, USPS datasets, (ii) Object recognition on Amazon, Webcam, DSLR, and VisDA datasets, (iii) Insect Recognition on the IP102 dataset. The experimental results show our proposed method consistently improves performance accuracy. Also, our framework can be incorporated with any other CNN frameworks within an end-to-end deep network design for recognition problems to improve their performance.
Thanh-Dat Truong, Naga Venkata Sai Raviteja Chappa, Xuan-Bac Nguyen, T. Hoang Ngan Le, Ashley Dowling, Khoa Luu
ICPR6
2022 Efficient hyperspectral image segmentation for biosecurity scanning using knowledge distillation from multi-head teacher
Minh-Hieu Phan, Son Lam Phung, Khoa Luu, Abdesselam Bouzerdoum
Neurocomputing3
2022 Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
Kha Gia Quach, T. Hoang Ngan Le, Chi Nhan Duong, Ibsa Jalata, Kaushik Roy 0003, Khoa Luu
Pattern Recognit.6
2021 Progressive Semantic Segmentation
abstract
The objective of this work is to segment high-resolution images without overloading GPU memory usage or losing the fine details in the output segmentation map. The memory constraint means that we must either downsample the big image or divide the image into local patches for separate processing. However, the former approach would lose the fine details, while the latter can be ambiguous due to the lack of a global picture. In this work, we present MagNet, a multi-scale framework that resolves local ambiguity by looking at the image at multiple magnification levels. MagNet has multiple processing stages, where each stage corresponds to a magnification level, and the output of one stage is fed into the next stage for coarse-to-fine information propagation. Each stage analyzes the image at a higher resolution than the previous stage, recovering the previously lost details due to the lossy downsampling step, and the segmentation output is progressively refined through the processing stages. Experiments on three high-resolution datasets of urban views, aerial scenes, and medical images show that MagNet consistently outperforms the state-of-the-art methods by a significant margin. Code is available at https://github.com/VinAIResearch/MagNet.
Chuong Huynh, Anh Tuan Tran 0001, Khoa Luu, Minh Hoai
CVPR3
2021 Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark Recognition
abstract
The research in automatic unsupervised visual clustering has received considerable attention over the last couple years. It aims at explaining distributions of unlabeled visual images by clustering them via a parameterized model of appearance. Graph Convolutional Neural Networks (GCN) have recently been one of the most popular clustering methods. However, it has reached some limitations. Firstly, it is quite sensitive to hard or noisy samples. Secondly, it is hard to investigate with various deep network models due to its computational training time. Finally, it is hard to design an end-to-end training model between the deep feature extraction and GCN clustering modeling. This work therefore presents the Clusformer, a simple but new perspective of Transformer based approach, to automatic visual clustering via its unsupervised attention mechanism. The proposed method is able to robustly deal with noisy or hard samples. It is also flexible and effective to collaborate with different deep network models with various model sizes in an end-to-end framework. The proposed method is evaluated on two popular large-scale visual databases, i.e. Google Landmark and MS-Celeb1M face database, and outperforms prior unsupervised clustering methods. Code will be available at https://github.com/VinAIResearch/Clusformer
Xuan-Bac Nguyen, Duc Toan Bui 0002, Chi Nhan Duong, Tien D. Bui, Khoa Luu
CVPR5
2021 DyGLIP: A Dynamic Graph Model With Link Prediction for Accurate Multi-Camera Multiple Object Tracking
abstract
Multi-Camera Multiple Object Tracking (MC-MOT) is a significant computer vision problem due to its emerging applicability in several real-world applications. Despite a large number of existing works, solving the data association problem in any MC-MOT pipeline is arguably one of the most challenging tasks. Developing a robust MC-MOT system, however, is still highly challenging due to many practical issues such as inconsistent lighting conditions, varying object movement patterns, or the trajectory occlusions of the objects between the cameras. To address these problems, this work, therefore, proposes a new Dynamic Graph Model with Link Prediction (DyGLIP) approach1to solve the data association task. Compared to existing methods, our new model offers several advantages, including better feature representations and the ability to recover from lost tracks during camera transitions. Moreover, our model works gracefully regardless of the overlapping ratios between the cameras. Experimental results show that we out-perform existing MC-MOT algorithms by a large margin on several practical datasets. Notably, our model works favor-ably on online settings but can be extended to an incremental approach for large-scale datasets.
Kha Gia Quach, Pha A. Nguyen, Huu Le, Thanh-Dat Truong, Chi Nhan Duong, Minh-Triet Tran, Khoa Luu
CVPR7
2021 BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation
abstract
Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new domain well. In this work, we first introduce a new Unaligned Domain Score to measure the efficiency of a learned model on a new target domain in unsupervised manner. Then, we present the new Bijective Maximum Likelihood1(BiMaL) loss that is a generalized form of the Adversarial Entropy Minimization without any assumption about pixel independence. We have evaluated the proposed BiMaL on two domains. The proposed BiMaL approach consistently outperforms the SOTA methods on empirical experiments on "SYNTHIA to Cityscapes", "GTA5 to Cityscapes", and "SYNTHIA to Vistas".
Thanh-Dat Truong, Chi Nhan Duong, T. Hoang Ngan Le, Son Lam Phung, Chase Rainwater, Khoa Luu
ICCV6
2021 The Right to Talk: An Audio-Visual Transformer Approach
abstract
Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker’s utterances) remains a challenging task. Although some prior methods have partially addressed this task, there still remain some limitations. Firstly, a direct association of Audio and Visual features may limit the correlations to be extracted due to different modalities. Secondly, the relationship across temporal segments helping to maintain the consistency of localization, separation and conversation contexts is not effectively exploited. Finally, the interactions between speakers that usually contain the tracking and anticipatory decisions about transition to a new speaker is usually ignored. Therefore, this work introduces a new Audio-Visual Transformer approach to the problem of localization and highlighting the main speaker in both audio and visual channels of a multi-speaker conversation video in the wild. The proposed method exploits different types of correlations presented in both visual and audio signals. The temporal audio-visual relationships across spatial-temporal space are anticipated and optimized via the self-attention mechanism in a Transformer structure. Moreover, a newly collected dataset is introduced for the main speaker detection. To the best of our knowledge, it is one of the first studies that is able to automatically localize and highlight the main speaker in both visual and audio channels in multi-speaker conversation videos.
Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Bhiksha Raj, T. Hoang Ngan Le, Khoa Luu
ICCV7
2021 PairFlow: Enhancing Portable Chest X-Ray By Flow-Based Deformation For Covid-19 Diagnosing
abstract
This work aims to assist physicians improve their speed and diagnostic accuracy when interpreting portable CXR (p_CXR), which are in especially high demand in the setting of the ongoing COVID-19 pandemic. In this paper, we introduce new deep learning frameworks, named Pair-Flow, to align and enhance the quality of p_CXR to be more consistent, and to more closely match higher quality conventional CXR (c_CXR). The contributions of this work are four folds. Firstly, a new database collection of subject-pair CXR is introduced and available to download. Secondly, a new deep learning-based alignment approach is presented to align subject-pairs dataset to obtain pixel-pairs dataset. Thirdly, a new Pair-Flow approach, an end-to-end invertible transfer deep learning method, to enhance the degraded quality of p_CXR. Finally, the performance of the proposed system is evaluated at both image quality and topological properties.
T. Hoang Ngan Le, James Sorensen, Toan Duc Bui, Arabinda Choudhary, Khoa Luu, Hien Van Nguyen
ICIP5
2021 Deep reinforcement learning in medical imaging: A literature review
Shaohua Kevin Zhou, T. Hoang Ngan Le, Khoa Luu, Hien Van Nguyen, Nicholas Ayache
Medical Image Anal.3
2020 Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition
abstract
Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper presents a novel generative structure with Bijective Metric Learning, namely Bijective Generative Adversarial Networks in a Distillation framework (DiBiGAN), for synthesizing faces of an identity given that person's features. In order to effectively address this problem, this work firstly introduces a bijective metric so that the distance measurement and metric learning process can be directly adopted in image domain for an image reconstruction task. Secondly, a distillation process is introduced to maximize the information exploited from the blackbox face recognition engine. Then a Feature-Conditional Generator Structure with Exponential Weighting Strategy is presented for a more robust generator that can synthesize realistic faces with ID preservation. Results on several benchmarking datasets including CelebA, LFW, AgeDB, CFP-FP against matching engines have demonstrated the effectiveness of DiBiGAN on both image realism and ID preservation properties.
Chi Nhan Duong, Thanh-Dat Truong, Khoa Luu, Kha Gia Quach, Kaushik Roy 0003
CVPR3
2020 Offset Curves Loss for Imbalanced Problem in Medical Segmentation
abstract
Medical image segmentation has played an important role in medical analysis and widely developed for many clinical applications. Deep learning-based approaches have achieved high performance in semantic segmentation but they are limited to pixel-wise setting and imbalanced classes data problem. In this paper, we tackle those limitations by developing a new deep learning-based model which takes into account both higher feature level i.e. region inside contour, intermediate feature level i.e. offset curves around the contour and lower feature level i.e. contour. Our proposed Offset Curves (OsC) loss consists of three main fitting terms. The first fitting term focuses on pixel-wise level segmentation whereas the second fitting term acts as attention model which pays attention to the area around the boundaries (offset curves). The third terms plays a role as regularization term which takes the length of boundaries into account. We evaluate our proposed OsC loss on both 2D network and 3D network. Two common medical datasets, i.e. retina DRIVE and brain tumor BRATS 2018 datasets are used to benchmark our proposed loss performance. The experiments have shown that our proposed OsC loss function outperforms other mainstream loss functions such as Cross-Entropy, Dice, Focal on the most common segmentation networks Unet, FCN.
T. Hoang Ngan Le, Kashu Yamazaki, Toan Duc Bui, Khoa Luu, Marios Savvides
ICPR5
2020 Flow-Based Deformation Guidance for Unpaired Multi-contrast MRI Image-to-Image Translation
Toan Duc Bui, Manh Nguyen 0002, T. Hoang Ngan Le, Khoa Luu
MICCAI (2)4
2019 Automatic Face Aging in Videos via Deep Reinforcement Learning
abstract
This paper presents a novel approach for synthesizing automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of given subjects coherently across video frames. The approach is optimized using a long-term reward, Reinforcement Learning function with deep feature extraction from Deep Convolutional Neural Network. Unlike previous age-progression methods that are only able to synthesize an aged likeness of a face from a single input image, the proposed approach is capable of age-progressing facial likenesses in videos with consistently synthesized facial features across frames. In addition, the deep reinforcement learning method guarantees preservation of the visual identity of input faces after age-progression. Results on videos of our new collected aging face AGFW-v2 database demonstrate the advantages of the proposed solution in terms of both quality of age-progressed faces, temporal smoothness, and cross-age face verification.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Eric Patterson, Tien D. Bui, T. Hoang Ngan Le
CVPR2
2019 Deep Appearance Models: A Deep Boltzmann Machine Approach for Face Modeling
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
Int. J. Comput. Vis.2
2019 Learning from Longitudinal Face Demonstration - Where Tractable Deep Modeling Meets Inverse Reinforcement Learning
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides, Tien D. Bui
Int. J. Comput. Vis.3
2018 Seeing Small Faces From Robust Anchor's Perspective
abstract
This paper introduces a novel anchor design principle to support anchor-based face detection for superior scaleinvariant performance, especially on tiny faces. To achieve this, we explicitly address the problem that anchor-based detectors drop performance drastically on faces with tiny sizes, e.g. less than 16 × 16 pixels. In this paper, we investigate why this is the case. We discover that current anchor design cannot guarantee high overlaps between tiny faces and anchor boxes, which increases the difficulty of training. The new Expected Max Overlapping (EMO) score is proposed which can theoretically explain the low overlapping issue and inspire several effective strategies of new anchor design leading to higher face overlaps, including anchor stride reduction with new network architectures, extra shifted anchors, and stochastic face shifting. Comprehensive experiments show that our proposed method significantly outperforms the baseline anchor-based detector, while consistently achieving state-of-the-art results on challenging face detection datasets with competitive runtime speed.
Chenchen Zhu, Ran Tao 0013, Khoa Luu, Marios Savvides
CVPR3
2018 Enhancing Interior and Exterior Deep Facial Features for Face Detection in the Wild
abstract
Although face detection has been intensely studied for decades, it is still a challenging topic due to numerous conditions, e.g. heavy occlusions, low resolutions, extreme poses, non-face patterns that look like human faces, etc. This paper proposes a novel region-based ConvNet to address these issues. Our approach enhances the interior deep facial features and explicitly incorporates the exterior deep features. The enhanced interior features provide fine details for small faces. The exterior features capture the local information surrounding the face, supporting the detection under challenging conditions. Experiments show that our proposed components improve the baseline method significantly. Additionally, our approach consistently achieves competitive performance in four challenging databases, i.e. Wider Face, AFW, PASCAL Faces, and FDDB. We also introduce a new challenging non-face dataset 1 of 6,000 images to benchmark false positive rates for future research.
Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides
FG3
2018 Deep contextual recurrent residual networks for scene labeling
T. Hoang Ngan Le, Chi Nhan Duong, Ligong Han, Khoa Luu, Kha Gia Quach, Marios Savvides
Pattern Recognit.4
2018 Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation
abstract
Variational Level Set (LS) has been a widely used method in medical segmentation. However, it is limited when dealing with multi-instance objects in the real world. In addition, its segmentation results are quite sensitive to initial settings and highly depend on the number of iterations. To address these issues and boost the classic variational LS methods to a new level of the learnable deep learning approaches, we propose a novel definition of contour evolution named Recurrent Level Set (RLS) 1 to employ Gated Recurrent Unit under the energy minimization of a variational LS functional. The curve deformation process in RLS is formed as a hidden state evolution procedure and updated by minimizing an energy functional composed of fitting forces and contour length. By sharing the convolutional features in a fully end-to-end trainable framework, we extend RLS to Contextual RLS (CRLS) to address semantic segmentation in the wild. The experimental results have shown that our proposed RLS improves both computational time and segmentation accuracy against the classic variational LS-based method whereas the fully end-to-end system CRLS achieves competitive performance compared to the state-of-the-art semantic segmentation approaches.
T. Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, Marios Savvides
IEEE Trans. Image Process.3
2017 Faster than Real-Time Facial Alignment: A 3D Spatial Transformer Network Approach in Unconstrained Poses
Chandrasekhar Bhagavatula, Chenchen Zhu, Khoa Luu, Marios Savvides
ICCV3
2017 Temporal Non-volume Preserving Approach to Facial Age-Progression and Age-Invariant Face Recognition
abstract
Modeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel generative probabilistic model, named Temporal Non-Volume Preserving (TNVP) transformation, is presented to model the facial aging process at each stage. Unlike Generative Adversarial Networks (GANs), which requires an empirical balance threshold, and Restricted Boltzmann Machines (RBM), an intractable model, our proposed TNVP approach guarantees a tractable density function, exact inference and evaluation for embedding the feature transformations between faces in consecutive stages. Our model shows its advantages not only in capturing the non-linear age related variance in each stage but also producing a smooth synthesis in age progression across faces. Our approach can model any face in the wild provided with only four basic landmark points. Moreover, the structure can be transformed into a deep convolutional network while keeping the advantages of probabilistic models with tractable log-likelihood density estimation. Our method is evaluated in both terms of synthesizing age-progressed faces and cross-age face verification and consistently shows the state-of-the-art results in various face aging databases, i.e. FG-NET, MORPH, AginG Faces in the Wild (AGFW), and Cross-Age Celebrity Dataset (CACD). A large-scale face verification on Megaface challenge 1 is also performed to further show the advantages of our proposed approach.
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides
ICCV3
2017 Non-convex online robust PCA: Enhance sparsity via ℓp-norm minimization
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
Comput. Vis. Image Underst.3
2017 Semi self-training beard/moustache detection and segmentation simultaneously
T. Hoang Ngan Le, Khoa Luu, Chenchen Zhu, Marios Savvides
Image Vis. Comput.2
2017 Compressed Submanifold Multifactor Analysis
abstract
Although widely used, Multilinear PCA (MPCA), one of the leading multilinear analysis methods, still suffers from four major drawbacks. First, it is very sensitive to outliers and noise. Second, it is unable to cope with missing values. Third, it is computationally expensive since MPCA deals with large multi-dimensional datasets. Finally, it is unable to maintain the local geometrical structures due to the averaging process. This paper proposes a novel approach named Compressed Submanifold Multifactor Analysis (CSMA) to solve the four problems mentioned above. Our approach can deal with the problem of missing values and outliers via SVD-L1. The Random Projection method is used to obtain the fast low-rank approximation of a given multifactor dataset. In addition, it is able to preserve the geometry of the original data. Our CSMA method can be used efficiently for multiple purposes, e.g. noise and outlier removal, estimation of missing values, biometric applications. We show that CSMA method can achieve good results and is very efficient in the inpainting problem as compared to [1], [2]. Our method also achieves higher face recognition rates compared to LRTC, SPMA, MPCA and some other methods, i.e. PCA, LDA and LPP, on three challenging face databases, i.e. CMU-MPIE, CMU-PIE and Extended YALE-B.
Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 DeepSafeDrive: A grammar-aware driver parsing approach to Driver Behavioral Situational Awareness (DB-SAW)
T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides
Pattern Recognit.4
2016 Longitudinal Face Modeling via Temporal Deep Restricted Boltzmann Machines
abstract
Modeling the face aging process is a challenging task due to large and non-linear variations present in different stages of face development. This paper presents a deep model approach for face age progression that can efficiently capture the non-linear aging process and automatically synthesize a series of age-progressed faces in various age ranges. In this approach, we first decompose the long-term age progress into a sequence of short-term changes and model it as a face sequence. The Temporal Deep Restricted Boltzmann Machines based age progression model together with the prototype faces are then constructed to learn the aging transformation between faces in the sequence. In addition, to enhance the wrinkles of faces in the later age ranges, the wrinkle models are further constructed using Restricted Boltzmann Machines to capture their variations in different facial regions. The geometry constraints are also taken into account in the last step for more consistent age-progressed results. The proposed approach is evaluated using various face aging databases, i.e. FGNET, Cross-Age Celebrity Dataset (CACD) and MORPH, and our collected large-scale aging database named AginG Faces in the Wild (AGFW). In addition, when ground-truth age is not available for input image, our proposed system is able to automatically estimate the age of the input face before aging process is employed.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
CVPR2
2016 Robust hand detection in Vehicles
abstract
The problems of hand detection have been widely addressed in many areas, e.g. human computer interaction environment, driver behaviors monitoring, etc. However, the detection accuracy in recent hand detection systems are still far away from the demands in practice due to a number of challenges, e.g. hand variations, highly occlusions, low-resolution and strong lighting conditions. This paper presents the Multiple Scale Faster Region-based Convolutional Neural Network (MS-FRCNN) to handle the problems of hand detection in given digital images collected under challenging conditions. Our proposed method introduces a multiple scale deep feature extraction approach in order to handle the challenging factors to provide a robust hand detection algorithm. The method is evaluated on the challenging hand database, i.e. the Vision for Intelligent Vehicles and Applications (VIVA) Challenge, and compared against various recent hand detection methods. Our proposed method achieves the state-of-the-art results with 20% of the detection accuracy higher than the second best one in the VIVA challenge.
T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides
ICPR4
2016 Robust Deep Appearance Models
abstract
This paper presents a novel Robust Deep Appearance Models (RDAMs) approach to learn the non-linear correlation between shape and texture of face images. In this approach, two crucial components of face images, i.e. shape and texture, are represented by Deep Boltzmann Machines and Robust Deep Boltzmann Machines (RDBM), respectively. The RDBM, an alternative form of Robust Boltzmann Machines, can separate corrupted/occluded pixels in the texture modeling to achieve better reconstruction results. The two models are connected by Restricted Boltzmann Machines at the top layer to jointly learn and capture the variations of both facial shapes and appearances. This paper also introduces new fitting algorithms with occlusion awareness through the mask obtained from the RDBM reconstruction. The proposed approach is evaluated in various applications by using challenging face datasets, i.e. Labeled Face Parts in the Wild (LFPW), Helen, EURECOM and AR databases, to demonstrate its robustness and capabilities.
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
ICPR3
2016 Depth-based 3D hand pose tracking
abstract
In this paper, we propose two new approaches using the Convolution Neural Network (CNN) and the Recurrent Neural Network (RNN) for tracking 3D hand poses. The first approach is a detection based algorithm while the second is a data driven method. Our first contribution is a new tracking-by-detection strategy extending the CNN based single frame detection method to a multiple frame tracking approach by taking into account prediction history using RNN. Our second contribution is the use of RNN to simulate the fitting of a 3D model to the input data. It helps to relax the need of a carefully designed fitting function and optimization algorithm. With such strategies, we show that our tracking frameworks can automatically correct the fail detection made in previous frames due to occlusions. Our proposed method is evaluated on two public hand datasets, i.e. NYU and ICVL, and compared against other recent hand tracking methods. Experimental results show that our approaches achieve the state-of-the-art accuracy and efficiency in the challenging problem of 3D hand pose estimation.
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
ICPR3
2015 Beyond Principal Components: Deep Boltzmann Machines for face modeling
abstract
The “interpretation through synthesis”, i.e. Active Appearance Models (AAMs) method, has received considerable attention over the past decades. It aims at “explaining” face images by synthesizing them via a parameterized model of appearance. It is quite challenging due to appearance variations of human face images, e.g. facial poses, occlusions, lighting, low resolution, etc. Since these variations are mostly non-linear, it is impossible to represent them in a linear model, such as Principal Component Analysis (PCA). This paper presents a novel Deep Appearance Models (DAMs) approach, an efficient replacement for AAMs, to accurately capture both shape and texture of face images under large variations. In this approach, three crucial components represented in hierarchical layers are modeled using the Deep Boltzmann Machines (DBM) to robustly capture the variations of facial shapes and appearances. DAMs are therefore superior to AAMs in inferring a representation for new face images under various challenging conditions. In addition, DAMs have ability to generate a compact set of parameters in higher level representation that can be used for classification, e.g. face recognition and facial age estimation. The proposed approach is evaluated in facial image reconstruction, facial super-resolution on two databases, i.e. LFPW and Helen. It is also evaluated on FG-NET database for the problem of age estimation.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
CVPR2
2015 IRIS super-resolution via nonparametric over-complete dictionary learning
abstract
This paper presents a novel iris super-resolution approach using a powerful nonparametric Bayesian modeling technique in the framework of sparse representation and over-complete dictionary. Far apart from previous iris super-resolution methods, our proposed approach has ability to automatically discover optimal parameter sets and optimally adapt from a given training data. Particularly, the Beta Process will be employed to build a nonparametric discriminative over-complete dictionary to represent and discriminate input samples simultaneously. Our proposed method will be evaluated on Casia iris database and compared with the linear interpolation super resolution. The result shows that our approach improves the performance of iris recognition.
Raied Aljadaany, Khoa Luu, Shreyas Venugopalan, Marios Savvides
ICIP2
2015 A robust contour sampling and tensor-based approach to facial beard and mustache shape segmentation and matching
abstract
In this paper, we propose a novel system for beard and mustache segmentation and matching in facial images. We first segment out facial hair contours from the image by utilizing a sparse dictionary on self-quotient images to classify regions as either skin or facial hair. We then landmark the shape contour to obtain points around the contour of the image using a combination of two algorithms, a novel non-uniform sampling algorithm, and points obtained from SIFT. We utilize these landmark points to extract inner distance-based shape context features. Finally, these features are used as inputs for a tensor product graph-based matching system. We run experiments on the Multiple Biometric Grand Challenge (MBGC) and the PINELLAS mugshot databases. Our pipeline achieves 90.3% matching accuracy on a subset of the PINELLAS database when divided into four types of facial hair.
Karanhaar Singh, Khoa Luu, T. Hoang Ngan Le, Marios Savvides
ICIP2
2015 Facial aging and asymmetry decomposition based approaches to identification of twins
T. Hoang Ngan Le, Keshav Seshadri, Khoa Luu, Marios Savvides
Pattern Recognit.3
2015 Spartans: Single-Sample Periocular-Based Alignment-Robust Recognition Technique Applied to Non-Frontal Scenarios
abstract
In this paper, we investigate a single-sample periocular-based alignment-robust face recognition technique that is pose-tolerant under unconstrained face matching scenarios. Our Spartans framework starts by utilizing one single sample per subject class, and generate new face images under a wide range of 3D rotations using the 3D generic elastic model which is both accurate and computationally economic. Then, we focus on the periocular region where the most stable and discriminant features on human faces are retained, and marginalize out the regions beyond the periocular region since they are more susceptible to expression variations and occlusions. A novel facial descriptor, high-dimensional Walsh local binary patterns, is uniformly sampled on facial images with robustness toward alignment. During the learning stage, subject-dependent advanced correlation filters are learned for pose-tolerant non-linear subspace modeling in kernel feature space followed by a coupled max-pooling mechanism which further improve the performance. Given any unconstrained unseen face image, the Spartans can produce a highly discriminative matching score, thus achieving high verification rate. We have evaluated our method on the challenging Labeled Faces in the Wild database and solidly outperformed the state-of-the-art algorithms under four evaluation protocols with a high accuracy of 89.69%, a top score among image-restricted and unsupervised protocols. The advancement of Spartans is also proven in the Face Recognition Grand Challenge and Multi-PIE databases. In addition, our learning method based on advanced correlation filters is much more effective, in terms of learning subject-dependent pose-tolerant subspaces, compared with many well-established subspace methods in both linear and non-linear cases.
Felix Juefei-Xu, Khoa Luu, Marios Savvides
IEEE Trans. Image Process.2
2014 Distributed class dependent feature analysis - A big data approach
abstract
Big data has been becoming ubiquitous and applied in numerous fields recently. The challenges to solve a large-scale machine learning problem in big data scenario generally lie in three aspects. Firstly, a proposed machine learning algorithm has to be appropriated for the distributed optimization problem. Secondly, it needs a platform for the distributed implementation. Finally, the communication delays different machines may cause problems in convergence even though the non-distributed algorithm shows a good convergence rate. In order to solve these challenges, we propose a new machine learning approach named Distributed Class-dependent Feature Analysis (DCFA), to combine the advantages of sparse representation in an over-complete dictionary. The classifier is based on the estimation of class-specific optimal filters, by solving an l1-norm optimization problem. We demonstrate how this problem is solved using the Alternating Direction Method of Multipliers and also explore relevant convergency details. More importantly, our proposed framework can be efficiently implemented on a robust distributed framework. Thus, it improves both accuracy and computational time in large-scale databases. Our method achieves very high classification accuracies in face recognition in the presence of occlusions on AR database. It also outperforms the state of the art methods in object recognition on two challenging large-scale object databases, i.e. Caltech101 and Caltech256. It hence shows its applicability to general computer vision and pattern recognition problems. In addition, computational time experiments show our distributed method achieves high speedup of 7.85x on Caltech256 databases with just 10 machine nodes compared to the non-distributed version and can gain even more with more computing resources.
Khoa Luu, Chenchen Zhu, Marios Savvides
IEEE BigData1
2013 SparCLeS: Dynamic 퓁 1 Sparse Classifiers With Level Sets for Robust Beard/Moustache Detection and Segmentation
abstract
Robust facial hair detection and segmentation is a highly valued soft biometric attribute for carrying out forensic facial analysis. In this paper, we propose a novel and fully automatic system, called SparCLeS, for beard/moustache detection and segmentation in challenging facial images. SparCLeS uses the multiscale self-quotient (MSQ) algorithm to preprocess facial images and deal with illumination variation. Histogram of oriented gradients (HOG) features are extracted from the preprocessed images and a dynamic sparse classifier is built using these features to classify a facial region as either containing skin or facial hair. A level set based approach, which makes use of the advantages of both global and local information, is then used to segment the regions of a face containing facial hair. Experimental results demonstrate the effectiveness of our proposed system in detecting and segmenting facial hair regions in images drawn from three databases, i.e., the NIST Multiple Biometric Grand Challenge (MBGC) still face database, the NIST Color Facial Recognition Technology FERET database, and the Labeled Faces in the Wild (LFW) database.
T. Hoang Ngan Le, Khoa Luu, Marios Savvides
IEEE Trans. Image Process.2
2012 A novel energy based filter for cross-blink eye detection
abstract
Based on the fact that eye regions may be considered as a sequence of consecutive low and high spatial frequency regions, we propose a novel and efficient filter based on Isotropic Gaussian Energy which can be used for eye detection. When applied to facial images, the designed filter highlights the eye region with a prominent pattern, which can then be localized by template matching technique such as MACE correlation filter. The proposed filter is proved useful to detect both open and closed eyes. We demonstrate the effectiveness of the proposed filter by conducting experiments on both FERET and MBGC face databases.
T. Hoang Ngan Le, Khoa Luu, Utsav Prabhu, Marios Savvides
ICIP2
2012 Beard and mustache segmentation using sparse classifiers on self-quotient images
abstract
In this paper, we propose a novel system for beard and mustache detection and segmentation in challenging facial images. Our system first eliminates illumination artifacts using the self-quotient algorithm. A sparse classifier is then used on these self-quotient images to classify a region as either containing skin or facial hair. We conduct experiments on the MBGC and color FERET databases to demonstrate the effectiveness of our proposed system.
T. Hoang Ngan Le, Khoa Luu, Keshav Seshadri, Marios Savvides
ICIP2
2012 Facecut - a robust approach for facial feature segmentation
abstract
Segmentation of facial features is a key pre-processing step in enabling facial recognition, building of 3D facial models, expression analysis, and pose estimation. Recently, graph cuts based algorithms have been adapted to carry out this task but many of these methods require manual initialization of points in the foreground and background. In this paper, we propose a novel and fully automatic approach, named Face-Cut, to perform accurate facial feature segmentation. FaceCut combines the positive features of the Modified Active Shape Model (MASM) and GrowCut algorithms to ensure highly accurate and completely automatic segmentation of facial features. We demonstrate the effectiveness of FaceCut on images from two challenging databases.
Khoa Luu, T. Hoang Ngan Le, Keshav Seshadri, Marios Savvides
ICIP1
2012 Compressed Submanifold Multifactor Analysis with adaptive factor structures
Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen
ICPR1
2011 Facial feature fusion and model selection for age estimation
abstract
Automatic face age estimation is challenging due to its complexity owing to genetic difference, behavior and environmental factors, the dynamics of facial aging between different individuals, etc. In this work we propose to fuse the global facial feature extracted from Active Appearance Model (AAM) and the local facial features extracted from Local Binary Pattern (LBP), as the representation of faces. Furthermore, we introduce an advanced age estimation system combining feature fusion and model selection schemes such as Least Angle Regression (LAR) and sequential approaches. Due to the fact that different facial feature representations may come with various types of measurement scales, we compare multiple normalization schemes for both facial features. We demonstrate that the feature fusion with model selection can achieve significant improvement in age estimation over single feature representation alone. Our experiment on multi-ethnicity UIUC-PAL database suggests that age estimation with feature fusion and model selection outperforms the single feature, or the full feature model.
Cuixian Chen, Wankou Yang, Yishi Wang, Karl Ricanek, Khoa Luu
FG5
2011 Kernel spectral regression of perceived age from hybrid facial features
abstract
This paper introduces an advanced age-determination technique using hybrid facial features and Kernel Spectral Regression, a nonlinear dimensionality reduction method. In the preprocessing stage, the logarithmic nonsubsampled contourlet transform (NSCT) is conducted to denoise and amplify facial wrinkles that help to distinguish young faces from elder ones. Then the hybrid facial features that combine both local and holistic features are extracted from the preprocessed images. Our novel Uniform Local Ternary Patterns (ULTP) are used as the local features. Meanwhile the holistic features are extracted by using the Active Appearance Model (AAM) to encode each face. Kernel Spectral Regression is used to minimize inter-class distances while maximizing intra-class distances of feature sets. These reduced features are used to classify faces into two age groups (age-classification). An age-determination function is then constructed for each age group in accordance with physiological growth periods for humans - pre-adult (youth) and adult. Compared to published results, this method yields promising results in overall mean absolute error (MAE), mean absolute error per decade of life (MAE/D), and cumulative match score in various face aging corpuses.
Khoa Luu, Tien D. Bui, Ching Y. Suen
FG1
2011 Investigating age invariant face recognition based on periocular biometrics
abstract
In this paper, we will present a novel framework of utilizing periocular region for age invariant face recognition. To obtain age invariant features, we first perform preprocessing schemes, such as pose correction, illumination and periocular region normalization. And then we apply robust Walsh-Hadamard transform encoded local binary patterns (WLBP) on preprocessed periocular region only. We find the WLBP feature on periocular region maintains consistency of the same individual across ages. Finally, we use unsupervised discriminant projection (UDP) to build subspaces on WLBP featured periocular images and gain 100% rank-1 identification rate and 98% verification rate at 0.1% false accept rate on the entire FG-NET database. Compared to published results, our proposed approach yields the best recognition and identification results.
Felix Juefei-Xu, Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen
IJCB2
2011 Contourlet appearance model for facial age estimation
abstract
In this paper we propose a novel Contourlet Appearance Model (CAM) that is more accurate and faster at localizing facial landmarks than Active Appearance Models (AAMs). Our CAM also has the ability to not only extract holistic texture information, as AAMs do, but can also extract local texture information using the Nonsubsampled Contourlet Transform (NSCT). We demonstrate the efficiency of our method by applying it to the problem of facial age estimation. Compared to previously published age estimation techniques, our approach yields more accurate results when tested on various face aging databases.
Khoa Luu, Keshav Seshadri, Marios Savvides, Tien D. Bui, Ching Y. Suen
IJCB1
2010 Combined local and holistic facial features for age-determination
abstract
This paper presents an advanced age-determination technique that combines holistic and local features derived from an image of the face. A 30×1 Active Appearance Model (AAM) linear encoding of each face is produced to work as holistic features. Meanwhile, local features are extracted by using Local Ternary Patterns (LTP). These combined features are used to classify faces into one of two age groups (age-classification). An age-determination function is then constructed for each age group in accordance with physiological growth periods for humans - pre-adult (youth) and adult. Compared to published results, this method yields the highest accuracy rates in overall mean absolute error (MAE), mean absolute error per decade of life (MAE/D), and cumulative match score.
Khoa Luu, Tien D. Bui, Ching Y. Suen, Karl Ricanek
ICARCV1