Chi Nhan Duong

dblp:16/11112 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
9since 2021 · last 2024
0000-0002-5177-0393ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Multi-camera multi-object tracking on the move via single-stage global association approach
Pha A. Nguyen, Kha Gia Quach, Chi Nhan Duong, Son Lam Phung, T. Hoang Ngan Le, Khoa Luu
Pattern Recognit.3
2023 Micron-BERT: BERT-Based Facial Micro-Expression Recognition
abstract
Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT ($(\mu$-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed$\mu$-BERT significantly outperforms all previous work in various micro-expression tasks.$\mu$-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show$\mu$-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT
Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li 0005, Susan Gauch, Han-Seok Seo, Khoa Luu
CVPR2
2023 LIAAD: Lightweight attentive angular distillation for large-scale age-invariant face recognition
Thanh-Dat Truong, Chi Nhan Duong, Kha Gia Quach, T. Hoang Ngan Le, Tien D. Bui, Khoa Luu
Neurocomputing2
2022 DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition
abstract
Human action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video action recognition with competitive results. However, these methods have suffered some fundamental limitations such as lack of robustness and generalization, e.g., how does the temporal ordering of video frames affect the recognition results? This work presents a novel end-to-end Transformer-based Directed Attention (Direc-Former) framework11The implementation of DirecFormer is available at https://github.com/uark-cviu/DirecFormer for robust action recognition. The method takes a simple but novel perspective of Transformer-based approach to understand the right order of sequence actions. Therefore, the contributions of this work are three-fold. Firstly, we introduce the problem of ordered temporal learning issues to the action recognition problem. Secondly, a new Directed Attention mechanism is introduced to understand and provide attentions to human actions in the right order. Thirdly, we introduce the conditional dependency in action sequence modeling that includes orders and classes. The proposed approach consistently achieves the state-of-the-art (SOTA) results compared with the recent action recognition methods [4, 18, 72, 74]. on three standard large-scale benchmarks, i.e. Jester, Kinetics-400 and Something-Something-V2.
Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li 0005, Khoa Luu
CVPR3
2022 Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
Kha Gia Quach, T. Hoang Ngan Le, Chi Nhan Duong, Ibsa Jalata, Kaushik Roy 0003, Khoa Luu
Pattern Recognit.3
2021 Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark Recognition
abstract
The research in automatic unsupervised visual clustering has received considerable attention over the last couple years. It aims at explaining distributions of unlabeled visual images by clustering them via a parameterized model of appearance. Graph Convolutional Neural Networks (GCN) have recently been one of the most popular clustering methods. However, it has reached some limitations. Firstly, it is quite sensitive to hard or noisy samples. Secondly, it is hard to investigate with various deep network models due to its computational training time. Finally, it is hard to design an end-to-end training model between the deep feature extraction and GCN clustering modeling. This work therefore presents the Clusformer, a simple but new perspective of Transformer based approach, to automatic visual clustering via its unsupervised attention mechanism. The proposed method is able to robustly deal with noisy or hard samples. It is also flexible and effective to collaborate with different deep network models with various model sizes in an end-to-end framework. The proposed method is evaluated on two popular large-scale visual databases, i.e. Google Landmark and MS-Celeb1M face database, and outperforms prior unsupervised clustering methods. Code will be available at https://github.com/VinAIResearch/Clusformer
Xuan-Bac Nguyen, Duc Toan Bui 0002, Chi Nhan Duong, Tien D. Bui, Khoa Luu
CVPR3
2021 DyGLIP: A Dynamic Graph Model With Link Prediction for Accurate Multi-Camera Multiple Object Tracking
abstract
Multi-Camera Multiple Object Tracking (MC-MOT) is a significant computer vision problem due to its emerging applicability in several real-world applications. Despite a large number of existing works, solving the data association problem in any MC-MOT pipeline is arguably one of the most challenging tasks. Developing a robust MC-MOT system, however, is still highly challenging due to many practical issues such as inconsistent lighting conditions, varying object movement patterns, or the trajectory occlusions of the objects between the cameras. To address these problems, this work, therefore, proposes a new Dynamic Graph Model with Link Prediction (DyGLIP) approach1to solve the data association task. Compared to existing methods, our new model offers several advantages, including better feature representations and the ability to recover from lost tracks during camera transitions. Moreover, our model works gracefully regardless of the overlapping ratios between the cameras. Experimental results show that we out-perform existing MC-MOT algorithms by a large margin on several practical datasets. Notably, our model works favor-ably on online settings but can be extended to an incremental approach for large-scale datasets.
Kha Gia Quach, Pha A. Nguyen, Huu Le, Thanh-Dat Truong, Chi Nhan Duong, Minh-Triet Tran, Khoa Luu
CVPR5
2021 BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation
abstract
Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new domain well. In this work, we first introduce a new Unaligned Domain Score to measure the efficiency of a learned model on a new target domain in unsupervised manner. Then, we present the new Bijective Maximum Likelihood1(BiMaL) loss that is a generalized form of the Adversarial Entropy Minimization without any assumption about pixel independence. We have evaluated the proposed BiMaL on two domains. The proposed BiMaL approach consistently outperforms the SOTA methods on empirical experiments on "SYNTHIA to Cityscapes", "GTA5 to Cityscapes", and "SYNTHIA to Vistas".
Thanh-Dat Truong, Chi Nhan Duong, T. Hoang Ngan Le, Son Lam Phung, Chase Rainwater, Khoa Luu
ICCV2
2021 The Right to Talk: An Audio-Visual Transformer Approach
abstract
Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker’s utterances) remains a challenging task. Although some prior methods have partially addressed this task, there still remain some limitations. Firstly, a direct association of Audio and Visual features may limit the correlations to be extracted due to different modalities. Secondly, the relationship across temporal segments helping to maintain the consistency of localization, separation and conversation contexts is not effectively exploited. Finally, the interactions between speakers that usually contain the tracking and anticipatory decisions about transition to a new speaker is usually ignored. Therefore, this work introduces a new Audio-Visual Transformer approach to the problem of localization and highlighting the main speaker in both audio and visual channels of a multi-speaker conversation video in the wild. The proposed method exploits different types of correlations presented in both visual and audio signals. The temporal audio-visual relationships across spatial-temporal space are anticipated and optimized via the self-attention mechanism in a Transformer structure. Moreover, a newly collected dataset is introduced for the main speaker detection. To the best of our knowledge, it is one of the first studies that is able to automatically localize and highlight the main speaker in both visual and audio channels in multi-speaker conversation videos.
Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Bhiksha Raj, T. Hoang Ngan Le, Khoa Luu
ICCV2
2020 Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition
abstract
Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper presents a novel generative structure with Bijective Metric Learning, namely Bijective Generative Adversarial Networks in a Distillation framework (DiBiGAN), for synthesizing faces of an identity given that person's features. In order to effectively address this problem, this work firstly introduces a bijective metric so that the distance measurement and metric learning process can be directly adopted in image domain for an image reconstruction task. Secondly, a distillation process is introduced to maximize the information exploited from the blackbox face recognition engine. Then a Feature-Conditional Generator Structure with Exponential Weighting Strategy is presented for a more robust generator that can synthesize realistic faces with ID preservation. Results on several benchmarking datasets including CelebA, LFW, AgeDB, CFP-FP against matching engines have demonstrated the effectiveness of DiBiGAN on both image realism and ID preservation properties.
Chi Nhan Duong, Thanh-Dat Truong, Khoa Luu, Kha Gia Quach, Kaushik Roy 0003
CVPR1
2019 Automatic Face Aging in Videos via Deep Reinforcement Learning
abstract
This paper presents a novel approach for synthesizing automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of given subjects coherently across video frames. The approach is optimized using a long-term reward, Reinforcement Learning function with deep feature extraction from Deep Convolutional Neural Network. Unlike previous age-progression methods that are only able to synthesize an aged likeness of a face from a single input image, the proposed approach is capable of age-progressing facial likenesses in videos with consistently synthesized facial features across frames. In addition, the deep reinforcement learning method guarantees preservation of the visual identity of input faces after age-progression. Results on videos of our new collected aging face AGFW-v2 database demonstrate the advantages of the proposed solution in terms of both quality of age-progressed faces, temporal smoothness, and cross-age face verification.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Eric Patterson, Tien D. Bui, T. Hoang Ngan Le
CVPR1
2019 Deep Appearance Models: A Deep Boltzmann Machine Approach for Face Modeling
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
Int. J. Comput. Vis.1
2019 Learning from Longitudinal Face Demonstration - Where Tractable Deep Modeling Meets Inverse Reinforcement Learning
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides, Tien D. Bui
Int. J. Comput. Vis.1
2018 Deep contextual recurrent residual networks for scene labeling
T. Hoang Ngan Le, Chi Nhan Duong, Ligong Han, Khoa Luu, Kha Gia Quach, Marios Savvides
Pattern Recognit.2
2018 Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation
abstract
Variational Level Set (LS) has been a widely used method in medical segmentation. However, it is limited when dealing with multi-instance objects in the real world. In addition, its segmentation results are quite sensitive to initial settings and highly depend on the number of iterations. To address these issues and boost the classic variational LS methods to a new level of the learnable deep learning approaches, we propose a novel definition of contour evolution named Recurrent Level Set (RLS) 1 to employ Gated Recurrent Unit under the energy minimization of a variational LS functional. The curve deformation process in RLS is formed as a hidden state evolution procedure and updated by minimizing an energy functional composed of fitting forces and contour length. By sharing the convolutional features in a fully end-to-end trainable framework, we extend RLS to Contextual RLS (CRLS) to address semantic segmentation in the wild. The experimental results have shown that our proposed RLS improves both computational time and segmentation accuracy against the classic variational LS-based method whereas the fully end-to-end system CRLS achieves competitive performance compared to the state-of-the-art semantic segmentation approaches.
T. Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, Marios Savvides
IEEE Trans. Image Process.4
2017 Temporal Non-volume Preserving Approach to Facial Age-Progression and Age-Invariant Face Recognition
abstract
Modeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel generative probabilistic model, named Temporal Non-Volume Preserving (TNVP) transformation, is presented to model the facial aging process at each stage. Unlike Generative Adversarial Networks (GANs), which requires an empirical balance threshold, and Restricted Boltzmann Machines (RBM), an intractable model, our proposed TNVP approach guarantees a tractable density function, exact inference and evaluation for embedding the feature transformations between faces in consecutive stages. Our model shows its advantages not only in capturing the non-linear age related variance in each stage but also producing a smooth synthesis in age progression across faces. Our approach can model any face in the wild provided with only four basic landmark points. Moreover, the structure can be transformed into a deep convolutional network while keeping the advantages of probabilistic models with tractable log-likelihood density estimation. Our method is evaluated in both terms of synthesizing age-progressed faces and cross-age face verification and consistently shows the state-of-the-art results in various face aging databases, i.e. FG-NET, MORPH, AginG Faces in the Wild (AGFW), and Cross-Age Celebrity Dataset (CACD). A large-scale face verification on Megaface challenge 1 is also performed to further show the advantages of our proposed approach.
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides
ICCV1
2017 Non-convex online robust PCA: Enhance sparsity via ℓp-norm minimization
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
Comput. Vis. Image Underst.2
2016 Longitudinal Face Modeling via Temporal Deep Restricted Boltzmann Machines
abstract
Modeling the face aging process is a challenging task due to large and non-linear variations present in different stages of face development. This paper presents a deep model approach for face age progression that can efficiently capture the non-linear aging process and automatically synthesize a series of age-progressed faces in various age ranges. In this approach, we first decompose the long-term age progress into a sequence of short-term changes and model it as a face sequence. The Temporal Deep Restricted Boltzmann Machines based age progression model together with the prototype faces are then constructed to learn the aging transformation between faces in the sequence. In addition, to enhance the wrinkles of faces in the later age ranges, the wrinkle models are further constructed using Restricted Boltzmann Machines to capture their variations in different facial regions. The geometry constraints are also taken into account in the last step for more consistent age-progressed results. The proposed approach is evaluated using various face aging databases, i.e. FGNET, Cross-Age Celebrity Dataset (CACD) and MORPH, and our collected large-scale aging database named AginG Faces in the Wild (AGFW). In addition, when ground-truth age is not available for input image, our proposed system is able to automatically estimate the age of the input face before aging process is employed.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
CVPR1
2016 Robust Deep Appearance Models
abstract
This paper presents a novel Robust Deep Appearance Models (RDAMs) approach to learn the non-linear correlation between shape and texture of face images. In this approach, two crucial components of face images, i.e. shape and texture, are represented by Deep Boltzmann Machines and Robust Deep Boltzmann Machines (RDBM), respectively. The RDBM, an alternative form of Robust Boltzmann Machines, can separate corrupted/occluded pixels in the texture modeling to achieve better reconstruction results. The two models are connected by Restricted Boltzmann Machines at the top layer to jointly learn and capture the variations of both facial shapes and appearances. This paper also introduces new fitting algorithms with occlusion awareness through the mask obtained from the RDBM reconstruction. The proposed approach is evaluated in various applications by using challenging face datasets, i.e. Labeled Face Parts in the Wild (LFPW), Helen, EURECOM and AR databases, to demonstrate its robustness and capabilities.
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
ICPR2
2016 Depth-based 3D hand pose tracking
abstract
In this paper, we propose two new approaches using the Convolution Neural Network (CNN) and the Recurrent Neural Network (RNN) for tracking 3D hand poses. The first approach is a detection based algorithm while the second is a data driven method. Our first contribution is a new tracking-by-detection strategy extending the CNN based single frame detection method to a multiple frame tracking approach by taking into account prediction history using RNN. Our second contribution is the use of RNN to simulate the fitting of a 3D model to the input data. It helps to relax the need of a carefully designed fitting function and optimization algorithm. With such strategies, we show that our tracking frameworks can automatically correct the fail detection made in previous frames due to occlusions. Our proposed method is evaluated on two public hand datasets, i.e. NYU and ICVL, and compared against other recent hand tracking methods. Experimental results show that our approaches achieve the state-of-the-art accuracy and efficiency in the challenging problem of 3D hand pose estimation.
Kha Gia Quach, Chi Nhan Duong, Khoa Luu, Tien D. Bui
ICPR2
2015 Beyond Principal Components: Deep Boltzmann Machines for face modeling
abstract
The “interpretation through synthesis”, i.e. Active Appearance Models (AAMs) method, has received considerable attention over the past decades. It aims at “explaining” face images by synthesizing them via a parameterized model of appearance. It is quite challenging due to appearance variations of human face images, e.g. facial poses, occlusions, lighting, low resolution, etc. Since these variations are mostly non-linear, it is impossible to represent them in a linear model, such as Principal Component Analysis (PCA). This paper presents a novel Deep Appearance Models (DAMs) approach, an efficient replacement for AAMs, to accurately capture both shape and texture of face images under large variations. In this approach, three crucial components represented in hierarchical layers are modeled using the Deep Boltzmann Machines (DBM) to robustly capture the variations of facial shapes and appearances. DAMs are therefore superior to AAMs in inferring a representation for new face images under various challenging conditions. In addition, DAMs have ability to generate a compact set of parameters in higher level representation that can be used for classification, e.g. face recognition and facial age estimation. The proposed approach is evaluated in facial image reconstruction, facial super-resolution on two databases, i.e. LFPW and Helen. It is also evaluated on FG-NET database for the problem of age estimation.
Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Tien D. Bui
CVPR1
2014 Are Sparse Representation and Dictionary Learning Good for Handwritten Character Recognition?
abstract
Recently the theories of sparse representation (SR) and dictionary learning (DL) have brought much attention and become powerful tools for pattern recognition and computer vision. Due to the fact that images can be represented in a sparse and compressible way with respect to some dictionaries, these theories have shown successful applications in many different areas including face recognition, image denoising and in painting, medical imaging, image classification and registration, motion estimation, and many more. Over a relatively short time, many improvements and innovative ideas using SR and DL have been developed. However, very little published work is found in the application of these theories on handwritten character recognition. One question comes to mind is whether these theories could produce good results for handwritten character recognition as in the case of other applications. In this paper, we would like to address this question by investigating various applications of the theories to handwritten character recognition. Experiments were conducted in both handwritten digits and alphabetical characters on three benchmark databases: MNIST, USPS, and CEDAR. The results showed that while this approach can achieve good results, it cannot beat the state of the art. The main advantage of this approach is that it does not require the choice of features and hence it may reduce computational cost.
Chi Nhan Duong, Kha Gia Quach, Tien D. Bui
ICFHR1
2014 Sparse Representation and Low-Rank Approximation for Robust Face Recognition
abstract
Face recognition under various conditions such as illumination, poses, expression, and occlusion has been one of the most challenging problems in computer vision. Over the last few years there has been significant attention paid to the low-rank approximation (LRA) and sparse representation (SR) techniques. The applications of these techniques have appeared in many different areas ranging from handwritten character recognition to multi-factor face recognition. In this paper, we will review some of the most recent works using LRA and SR in the multi-factor face recognition problem, and present a novel framework to improve their performance in the recognition of faces under various affecting conditions. Our results are comparable to or better than the state-of-the-art in this area.
Kha Gia Quach, Chi Nhan Duong, Tien D. Bui
ICPR2
2012 Robust eye localization in video by combining eye detector and eye tracker
Chi Nhan Duong, Thang Cap Pham Dinh, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Bac Le, Shin'ichi Satoh 0001
ICPR1