Duc Thanh Nguyen

dblp:47/3548 · DBLP profile ↗
← Back
53ranked-venue papers
19as first author
19since 2021 · last 2026
0000-0002-2285-2066ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 13 first-author · 14 since 2021Artificial intelligence and machine learning · 29 · 10 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
abstract
Abstract Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correlation between visual and textual domains in open concepts and that diffusion-based text-to-image models can capture rich and diverse information for computer vision tasks. However, we found that those advantages do not hold for learning of features of camouflaged individuals because of the significant blending between their visual boundaries and their surroundings. In this paper, while leveraging the benefits of diffusion-based techniques and text-image models in open-vocabulary settings, we aim to address a challenging problem in computer vision: open-vocabulary camouflaged instance segmentation (OVCIS). Specifically, we propose a method built upon state-of-the-art diffusion empowered by open-vocabulary to learn multi-scale textual-visual features for camouflaged object representation learning. Such cross-domain representations are desirable in segmenting camouflaged objects where visual cues subtly distinguish the objects from the background, and in segmenting novel object classes which are not seen in training. To enable such powerful representations, we devise complementary modules to effectively fuse cross-domain features, and to engage relevant features towards respective foreground objects. We validate and compare our method with existing ones on several benchmark datasets of camouflaged and generic open-vocabulary instance segmentation. The experimental results confirm the advances of our method over existing ones. We believe that our proposed method would open a new avenue for handling camouflages such as computer vision-based surveillance systems, wildlife monitoring, and military reconnaissance.
Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo 0005, Nhat Chung, Binh-Son Hua, Ivor W. Tsang, Sai-Kit Yeung
Int. J. Comput. Vis.2
2026 MVC: a multi-task vision transformer network for COVID-19 diagnosis from chest X-ray images
abstract
Abstract Medical image analysis using computer-based algorithms has attracted considerable attention from the research community and achieved tremendous progress in the last decade. With recent advances in computing resources and availability of large-scale medical image datasets, many deep learning models have been developed for disease diagnosis from medical images. However, existing techniques focus on sub-tasks, e.g., disease classification and identification, individually, while there is a lack of a unified framework enabling multi-task diagnosis. Inspired by the capability of Vision Transformers in both patch-based and image-based representation learning, we propose in this paper a new method, namely Multi-task Vision Transformer (MVC) for simultaneously classifying chest X-ray images and identifying affected regions from the input data. Our method is built upon the Vision Transformer but extends its learning capability in a multi-task setting. We evaluated our proposed method and compared it with existing baselines on a benchmark dataset of COVID-19 chest X-ray images. Experimental results verified the superiority of the proposed method over the baselines on both the image classification and affected region identification tasks.
Huyen Tran, Duc Thanh Nguyen, John Yearwood
Neural Comput. Appl.2
2025 Color Alignment in Diffusion
abstract
Diffusion models have shown great promise in synthesizing visually appealing images. However, it remains challenging to condition the synthesis at a fine-grained level, for instance, synthesizing image pixels following some generic color pattern. Existing image synthesis methods often produce contents that fall outside the desired pixel conditions. To address this, we introduce a novel color alignment algorithm that confines the generative process in diffusion models within a given color pattern. Specifically, we project diffusion terms, either imagery samples or latent representations, into a conditional color space to align with the input color distribution. This strategy simplifies the prediction in diffusion models within a color manifold while still allowing plausible structures in generated contents, thus enabling the generation of diverse contents that comply with the target color pattern. Experimental results demonstrate our state-of-the-art performance in conditioning and controlling of color pixels, while maintaining on-par generation quality and diversity in comparison with regular diffusion models.
Ka-Chun Shum, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
CVPR3
2025 MSC: A Marine Wildlife Dataset for Video Understanding with Grounded Segmentation and Clip-Level Captions
abstract
Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captioning datasets, typically focused on generic or human-centric domains, often fail to generalize to the complexities of the marine environment and gain insights about marine life. To address these limitations, we propose a two-stage marine object-oriented video captioning pipeline. We introduce a comprehensive video understanding benchmark that leverages the triplets of video, text, and segmentation masks to facilitate visual grounding and captioning, leading to improved marine video understanding and analysis, and marine video generation. Additionally, we highlight the effectiveness of video splitting in order to detect salient object transitions in scene changes, which significantly enrich the semantics of captioning content. Our dataset and code have been released at https://msc.hkustvgd.com.
Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung
ACM Multimedia5
2025 Multi-modality guided cross-attention for visual question answering
abstract
Abstract Visual Question Answering (VQA) is a multimodality research domain that intersects the fields of computer vision and natural language processing for visual-textual data processing and understanding. Traditional VQA methods extract visual and textual features from pre-trained architectures, respectively, then combine the features from both modalities in a common feature space. The traditional methods perform well on high-level perception questions. However, attaining high accuracy on low-level perception questions still remains challenging. The difficulties include detecting relevant visual and textual information, building meaningful associations, and extracting insights from the multimodal data. To address these challenges, unlike existing approaches, we propose a novel multi-modality guided cross self-attention mechanism for building semantic relationships within individual modalities as well as between them. Specifically, we examine visual-guided cross-attention (VGCA), textual-guided cross-attention (TGCA), and multi-modality-guided cross-attention (MMGCA). We utilise convolutional neural networks (CNNs) for visual feature learning, and LSTM and FNET for textual feature learning. We evaluate our method on two benchmark datasets, including VQA 1.0 and VQA 2.0. Experimental results demonstrate the superiority of our method over existing baselines by improving the performance on various types of questions in both datasets.
Muhammad Zeeshan Khan, Duc Thanh Nguyen, Thanh Thi Nguyen 0001, Anuroop Gaddam, Muhammad Imran Razzak
Multim. Tools Appl.2
2025 An empirical study of automatic wildlife detection using drone-derived imagery and object detection
Tan Vuong, Miao Chang, Manas Palaparthi, Lachlan Howell, Alessio Bonti, Mohamed Almorsy, Duc Thanh Nguyen
Multim. Tools Appl.7
2024 Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates
abstract
Neural radiance field (NeRF) is an emerging technique for 3D scene reconstruction and modeling. However, current NeRF-based methods are limited in the capabilities of adding or removing objects. This paper fills the aforementioned gap by proposing a new language-driven method for object manipulation in NeRFs through dataset updates. Specifically, to insert an object represented by a set of multi-view images into a background NeRF, we use a text-to-image diffusion model to blend the object into the given background across views. The generated images are then used to update the NeRF so that we can render view-consistent images of the object within the background. To ensure view consistency, we propose a dataset update strategy that prioritizes the radiance field training based on camera poses in a pose-ordered manner. We validate our method in two case studies: object insertion and object removal. Experimental results show that our method can generate photo-realistic results and achieves state-of-the-art performance in NeRF editing.
Ka-Chun Shum, Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
CVPR4
2024 Deep cross-domain transfer for emotion recognition via joint learning
abstract
Abstract Deep learning has been applied to achieve significant progress in emotion recognition from multimedia data. Despite such substantial progress, existing approaches are hindered by insufficient training data, leading to weak generalisation under mismatched conditions. To address these challenges, we propose a learning strategy which jointly transfers emotional knowledge learnt from rich datasets to source-poor datasets. Our method is also able to learn cross-domain features, leading to improved recognition performance. To demonstrate the robustness of the proposed learning strategy, we conducted extensive experiments on several benchmark datasets including eNTERFACE, SAVEE, EMODB, and RAVDESS. Experimental results show that the proposed method surpassed existing transfer learning schemes by a significant margin.
Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Mohamed Almorsy, Simon Denman, Son N. Tran, Clinton Fookes
Multim. Tools Appl.2
2024 SAWIT: A small-sized animal wild image dataset with annotations
abstract
Abstract Computer vision has found many applications in automatic wildlife data analytics and biodiversity monitoring. Automating tasks like animal recognition or animal detection usually require machine learning models (e.g., deep neural networks) trained on annotated datasets. However, image datasets built for general purposes fail to capture realistic conditions of ecological studies, and existing datasets collected with camera-traps mainly focus on medium to large-sized animals. There is a lack of annotated small-sized animal datasets in the field. Small-sized animals (e.g., small mammals, frogs, lizards, arthropods) play an important role in ecosystems but are difficult to capture on camera-traps. They also present additional challenges: small animals can be more difficult to identify and blend more easily with their surroundings. To fill this gap, we introduce in this paper a new dataset dedicated to ecological studies of small-sized animals, and provide benchmark results of computer vision-based wildlife monitoring. The novelty of our work lies on SAWIT ( s mall-sized a nimal w ild i mage da t aset), the first real-world dataset of small-sized animals, collected from camera traps and in realistic conditions. Our dataset consists of 34,434 images and is annotated by experts in the field with object-level annotations (bounding boxes) providing 34,820 annotated animals for seven animal categories. The dataset encompasses a wide range of challenging scenarios, such as occlusions, blurriness, and instances where animals blend into the dense vegetation. Based on the dataset, we benchmark two prevailing object detection algorithms: Faster RCNN and YOLO, and their variants. Experimental results show that all the variants of YOLO (version 5) perform similarly, ranging from 59.3% to 62.6% for the overall mean Average Precision (mAP) across all the animal categories. Faster RCNN with ResNet50 and HRNet backbone achieve 61.7% mAP and 58.5% mAP respectively. Through experiments, we indicate challenges and suggest research directions for computer vision-based wildlife monitoring. We provide both the dataset and the animal detection code at https://github.com/dtnguyen0304/sawit .
Thi Thu Thuy Nguyen, Anne C. Eichholtzer, Don A. Driscoll, Nathan I. Semianiw, Dean M. Corva, Abbas Z. Kouzani, Thanh Thi Nguyen 0001, Duc Thanh Nguyen
Multim. Tools Appl.8
2023 Conditional 360-degree Image Synthesis for Immersive Indoor Scene Decoration
abstract
In this paper, we address the problem of conditional scene decoration for 360° images. Our method takes a 360° background photograph of an indoor scene and generates decorated images of the same scene in the panorama view. To do this, we develop a 360-aware object layout generator that learns latent object vectors in the 360° view to enable a variety of furniture arrangements for an input 360° background image. We use this object layout to condition a generative adversarial network to synthesize images of an input scene. To further reinforce the generation capability of our model, we develop a simple yet effective scene emptier that removes the generated furniture and produces an emptied scene for our model to learn a cyclic constraint. We train the model on the Structure3D dataset and show that our model can generate diverse decorations with controllable object layout. Our method achieves state-of-the-art performance on the Structure3D dataset and generalizes well to the Zillow indoor scene dataset. Our user study confirms the immersive experiences provided by the realistic image quality and furniture layout in our generation results. Our implementation is available at https://github.com/kcshum/neural_360_decoration.git.
Ka-Chun Shum, Hong-Wing Pang, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
ICCV4
2023 Causal Inference via Style Transfer for Out-of-distribution Generalisation
abstract
Out-of-distribution (OOD) generalisation aims to build a model that can generalise well on an unseen target domain using knowledge from multiple source domains. To this end, the model should seek the causal dependence between inputs and labels, which may be determined by the semantics of inputs and remain invariant across domains. However, statistical or non-causal methods often cannot capture this dependence and perform poorly due to not considering spurious correlations learnt from model training via unobserved confounders. A well-known existing causal inference method like back-door adjustment cannot be applied to remove spurious correlations as it requires the observation of confounders. In this paper, we propose a novel method that effectively deals with hidden confounders by successfully implementing front-door adjustment (FA). FA requires the choice of a mediator, which we regard as the semantic information of images that helps access the causal mechanism without the need for observing confounders. Further, we propose to estimate the combination of the mediator with other observed images in the front-door formula via style transfer algorithms. Our use of style transfer to estimate FA is novel and sensible for OOD generalisation, which we justify by extensive experimental results on widely used benchmark datasets.
Toan Nguyen 0004, Kien Do, Duc Thanh Nguyen, Bao Duong, Thin Nguyen
KDD3
2023 PointInverter: Point Cloud Reconstruction and Editing via a Generative Model with Shape Priors
abstract
In this paper, we propose a new method for mapping a 3D point cloud to the latent space of a 3D generative adversarial network. Our generative model for 3D point clouds is based on SP-GAN, a state-of-the-art sphere-guided 3D point cloud generator. We derive an efficient way to encode an input 3D point cloud to the latent space of the SP-GAN. Our point cloud encoder can resolve the point ordering issue during inversion, and thus can determine the correspondences between points in the generated 3D point cloud and those in the canonical sphere used by the generator. We show that our method outperforms previous GAN inversion methods for 3D point clouds, achieving state-of-the-art results both quantitatively and qualitatively. Our code is available at https://github.com/hkust-vgd/point_inverter.
Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
WACV3
2023 Meta-transfer learning for emotion recognition
abstract
Abstract Deep learning has been widely adopted in automatic emotion recognition and has lead to significant progress in the field. However, due to insufficient training data, pre-trained models are limited in their generalisation ability, leading to poor performance on novel test sets. To mitigate this challenge, transfer learning performed by fine-tuning pr-etrained models on novel domains has been applied. However, the fine-tuned knowledge may overwrite and/or discard important knowledge learnt in pre-trained models. In this paper, we address this issue by proposing a PathNet-based meta-transfer learning method that is able to (i) transfer emotional knowledge learnt from one visual/audio emotion domain to another domain and (ii) transfer emotional knowledge learnt from multiple audio emotion domains to one another to improve overall emotion recognition accuracy. To show the robustness of our proposed method, extensive experiments on facial expression-based emotion recognition and speech emotion recognition are carried out on three bench-marking data sets: SAVEE, EMODB, and eNTERFACE. Experimental results show that our proposed method achieves superior performance compared with existing transfer learning methods.
Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Simon Denman, Thanh Thi Nguyen 0001, David Dean, Clinton Fookes
Neural Comput. Appl.2
2022 Neural Scene Decoration from a Single Photograph
Hong-Wing Pang, Yingshu Chen, Phuoc-Hieu Le, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
ECCV (23)5
2022 RFNet-4D: Joint Object Reconstruction and Flow Estimation from 4D Point Clouds
Tuan-Anh Vu, Duc Thanh Nguyen, Binh-Son Hua, Quang-Hieu Pham, Sai-Kit Yeung
ECCV (23)2
2022 Deep learning for deepfakes creation and detection: A survey
Thanh Thi Nguyen 0001, Nguyen Quoc Viet Hung, Duc Thanh Nguyen, Thien Huynh-The, Saeid Nahavandi, Thanh Tam Nguyen, Quoc-Viet Pham, Cuong M. Nguyen
Comput. Vis. Image Underst.4
2022 Deep Auto-Encoders With Sequential Learning for Multimodal Dimensional Emotion Recognition
abstract
Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions still remain unanswered for most of existing approaches including: (i) how to simultaneously learn compact yet representative features from multimodal data, (ii) how to effectively capture complementary features from multimodal streams, and (iii) how to perform all the tasks in an end-to-end manner. To address these challenges, in this paper, we propose a novel deep neural network architecture consisting of a two-stream auto-encoder and a long short term memory for effectively integrating visual and audio signal streams for emotion recognition. To validate the robustness of our proposed architecture, we carry out extensive experiments on the multimodal emotion in the wild dataset: RECOLA. Experimental results show that the proposed method achieves state-of-the-art recognition performance.
Dung Nguyen 0001, Duc Thanh Nguyen, Thanh Thi Nguyen 0001, Son N. Tran, Thin Nguyen, Sridha Sridharan, Clinton Fookes
IEEE Trans. Multim.2
2021 Minimal Adversarial Examples for Deep Learning on 3D Point Clouds
abstract
With recent developments of convolutional neural net-works, deep learning for 3D point clouds has shown significant progress in various 3D scene understanding tasks, e.g., object recognition, semantic segmentation. In a safety-critical environment, it is however not well understood how such deep learning models are vulnerable to adversarial examples. In this work, we explore adversarial attacks for point cloud-based neural networks. We propose a unified formulation for adversarial point cloud generation that can generalise two different attack strategies. Our method generates adversarial examples by attacking the classification ability of point cloud-based networks while considering the perceptibility of the examples and ensuring the minimal level of point manipulations. Experimental results show that our method achieves the state-of-the-art performance with higher than 89% and 90% of attack success rate on synthetic and real-world data respectively, while manipulating only about 4% of the total points.
Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
ICCV3
2021 A graph-based approach for population health analysis using Geo-tagged tweets
Thin Nguyen, Duc Thanh Nguyen
Multim. Tools Appl.3
2020 LCD: Learned Cross-Domain Descriptors for 2D-3D Matching
abstract
In this work, we present a novel method to learn a local cross-domain descriptor for 2D image and 3D point cloud matching. Our proposed method is a dual auto-encoder neural network that maps 2D and 3D input into a shared latent space representation. We show that such local cross-domain descriptors in the shared embedding are more discriminative than those obtained from individual training in 2D and 3D domains. To facilitate the training process, we built a new dataset by collecting ≈ 1.4 millions of 2D-3D correspondences with various lighting conditions and settings from publicly available RGB-D scenes. Our descriptor is evaluated in three main experiments: 2D-3D matching, cross-domain retrieval, and sparse-to-dense depth estimation. Experimental results confirm the robustness of our approach as well as its competitive performance not only in solving cross-domain tasks but also in being able to generalize to solve sole 2D and 3D tasks. Our dataset and code are released publicly at https://hkust-vgd.github.io/lcd.
Quang-Hieu Pham, Mikaela Angelina Uy, Binh-Son Hua, Duc Thanh Nguyen, Gemma Roig, Sai-Kit Yeung
AAAI4
2020 SideInfNet: A Deep Neural Network for Semi-Automatic Semantic Segmentation with Side Information
Jing Yu Koh, Duc Thanh Nguyen, Quang-Trung Truong, Sai-Kit Yeung, Alexander Binder
ECCV (24)2
2020 MrPC: Causal Structure Learning in Distributed Systems
Thin Nguyen, Duc Thanh Nguyen, Thuc Duy Le, Svetha Venkatesh
ICONIP (4)2
2020 Using spatiotemporal distribution of geocoded Twitter data to predict US county-level health indices
Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, John Yearwood, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen
Future Gener. Comput. Syst.5
2019 JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds With Multi-Task Pointwise Networks and Multi-Value Conditional Random Fields
abstract
Deep learning techniques have become the to-go models for most vision-related tasks on 2D images. However, their power has not been fully realised on several tasks in 3D space, e.g., 3D scene understanding. In this work, we jointly address the problems of semantic and instance segmentation of 3D point clouds. Specifically, we develop a multi-task pointwise network that simultaneously performs two tasks: predicting the semantic classes of 3D points and embedding the points into high-dimensional vectors so that points of the same object instance are represented by similar embeddings. We then propose a multi-value conditional random field model to incorporate the semantic and instance labels and formulate the problem of semantic and instance segmentation as jointly optimising labels in the field model. The proposed method is thoroughly evaluated and compared with existing methods on different indoor scene datasets including S3DIS and SceneNN. Experimental results showed the robustness of the proposed joint semantic-instance segmentation scheme over its single components. Our method also achieved state-of-the-art performance on semantic segmentation.
Quang-Hieu Pham, Duc Thanh Nguyen, Binh-Son Hua, Gemma Roig, Sai-Kit Yeung
CVPR2
2019 Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World Data
abstract
Deep learning techniques for point cloud data have demonstrated great potentials in solving classical problems in 3D computer vision such as 3D object classification and segmentation. Several recent 3D object classification methods have reported state-of-the-art performance on CAD model datasets such as ModelNet40 with high accuracy (~92\%). Despite such impressive results, in this paper, we argue that object classification is still a challenging task when objects are framed with real-world settings. To prove this, we introduce ScanObjectNN, a new real-world point cloud object dataset based on scanned indoor scene data. From our comprehensive benchmark, we show that our dataset poses great challenges to existing point cloud classification techniques as objects from real-world scans are often cluttered with background and/or are partial due to occlusions. We identify three key open problems for point cloud object classification, and propose new point cloud classification neural networks that achieve state-of-the-art performance on classifying objects with cluttered background. Our dataset and code are publicly available in our project page https://hkust-vgd.github.io/scanobjectnn/.
Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
ICCV4
2019 Creating and understanding 3D annotated scene meshes
abstract
Deep learning requires availability of massive 3D data.
Duc Thanh Nguyen, Quang-Hieu Pham, Binh-Son Hua
SIGGRAPH Asia1
2019 Real-Time Progressive 3D Semantic Segmentation for Indoor Scenes
abstract
The widespread adoption of autonomous systems such as drones and assistant robots has created a need for real-time high-quality semantic scene segmentation. In this paper, we propose an efficient yet robust technique for on-the-fly dense reconstruction and semantic segmentation of 3D indoor scenes. To guarantee (near) real-time performance, our method is built atop an efficient super-voxel clustering method and a conditional random field with higher-order constraints from structural and object cues, enabling progressive dense semantic segmentation without any precomputation. We extensively evaluate our method on different indoor scenes including kitchens, offices, and bedrooms in the SceneNN and ScanNet datasets and show that our technique consistently produces state-of-the-art segmentation results in both qualitative and quantitative experiments.
Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
WACV3
2019 An empirical study on prediction of population health through social media
Thin Nguyen, Duc Thanh Nguyen
J. Biomed. Informatics3
2018 Urban Zoning Using Higher-Order Markov Random Fields on Multi-View Imagery Data
Tian Feng 0001, Quang-Trung Truong, Duc Thanh Nguyen, Jing Yu Koh, Lap-Fai Yu, Alexander Binder, Sai-Kit Yeung
ECCV (8)3
2018 Jointly Predicting Affective and Mental Health Scores Using Deep Neural Networks of Visual Cues on the Web
Van Nguyen 0002, Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen
WISE (2)6
2018 Improving Chamfer Template Matching Using Image Segmentation
abstract
This letter proposes an effective method to improve object location in Chamfer template matching (CTM) based object detection using image segmentation. In our method, object bounding boxes are iteratively adjusted to fit with the object images obtained from image segmentation in a probabilistic model. The proposed method was tested with state-of-the-art CTM-based object detectors. Experimental results have shown the proposed method improved the location accuracy of the object detectors and reduce the false alarms rate.
Duc Thanh Nguyen, Ngoc-Son Vu, Thanh-Toan Do, Thin Nguyen, John Yearwood
IEEE Signal Process. Lett.1
2018 A Robust 3D-2D Interactive Tool for Scene Segmentation and Annotation
abstract
Recent advances of 3D acquisition devices have enabled large-scale acquisition of 3D scene data. Such data, if completely and well annotated, can serve as useful ingredients for a wide spectrum of computer vision and graphics works such as data-driven modeling and scene understanding, object detection and recognition. However, annotating a vast amount of 3D scene data remains challenging due to the lack of an effective tool and/or the complexity of 3D scenes (e.g. clutter, varying illumination conditions). This paper aims to build a robust annotation tool that effectively and conveniently enables the segmentation and annotation of massive 3D data. Our tool works by coupling 2D and 3D information via an interactive framework, through which users can provide high-level semantic annotation for objects. We have experimented our tool and found that a typical indoor scene could be well segmented and annotated in less than 30 minutes by using the tool, as opposed to a few hours if done manually. Along with the tool, we created a dataset of over a hundred 3D scenes associated with complete annotations using our tool. Both the tool and dataset will be available at http://scenenn.net.
Duc Thanh Nguyen, Binh-Son Hua, Lap-Fai Yu, Sai-Kit Yeung
IEEE Trans. Vis. Comput. Graph.1
2017 Kernel-based features for predicting population health indices from geocoded social media data
Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, John Yearwood, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen
Decis. Support Syst.4
2016 SceneNN: A Scene Meshes Dataset with aNNotations
abstract
Several RGB-D datasets have been publicized over the past few years for facilitating research in computer vision and robotics. However, the lack of comprehensive and fine-grained annotation in these RGB-D datasets has posed challenges to their widespread usage. In this paper, we introduce SceneNN, an RGB-D scene dataset consisting of 100 scenes. All scenes are reconstructed into triangle meshes and have per-vertex and per-pixel annotation. We further enriched the dataset with fine-grained information such as axis-aligned bounding boxes, oriented bounding boxes, and object poses. We used the dataset as a benchmark to evaluate the state-of-the-art methods on relevant research problems such as intrinsic decomposition and shape completion. Our dataset and annotation tools are available at http://www.scenenn.net.
Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, Sai-Kit Yeung
3DV3
2016 A Field Model for Repairing 3D Shapes
abstract
This paper proposes a field model for repairing 3D shapes constructed from multi-view RGB data. Specifically, we represent a 3D shape in a Markov random field (MRF) in which the geometric information is encoded by random binary variables and the appearance information is retrieved from a set of RGB images captured at multiple viewpoints. The local priors in the MRF model capture the local structures of object shapes and are learnt from 3D shape templates using a convolutional deep belief network. Repairing a 3D shape is formulated as the maximum a posteriori (MAP) estimation in the corresponding MRF. Variational mean field approximation technique is adopted for the MAP estimation. The proposed method was evaluated on both artificial data and real data obtained from reconstruction of practical scenes. Experimental results have shown the robustness and efficiency of the proposed method in repairing noisy and incomplete 3D shapes.
Duc Thanh Nguyen, Binh-Son Hua, Minh-Khoi Tran, Quang-Hieu Pham, Sai-Kit Yeung
CVPR1
2016 Binary Hashing with Semidefinite Relaxation and Augmented Lagrangian
Thanh-Toan Do, Anh-Dzung Doan, Duc Thanh Nguyen, Ngai-Man Cheung
ECCV (2)3
2016 Human detection from images and videos: A survey
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
Pattern Recognit.1
2015 An MRF-Poselets Model for Detecting Highly Articulated Humans
abstract
Detecting highly articulated objects such as humans is a challenging problem. This paper proposes a novel part-based model built upon poselets, a notion of parts, and Markov Random Field (MRF) for modelling the human body structure under the variation of human poses and viewpoints. The problem of human detection is then formulated as maximum a posteriori (MAP) estimation in the MRF model. Variational mean field method, a robust statistical inference, is adopted to approximate the MAP estimation. The proposed method was evaluated and compared with existing methods on different test sets including H3D and PASCAL VOC 2007-2009. Experimental results have favourbly shown the robustness of the proposed method in comparison to the state-of-the-art.
Duc Thanh Nguyen, Minh-Khoi Tran, Sai-Kit Yeung
ICCV1
2014 A Novel Chamfer Template Matching Method Using Variational Mean Field
abstract
This paper proposes a novel mean field-based Chamfer template matching method. In our method, each template is represented as a field model and matching a template with an input image is formulated as estimation of a maximum of posteriori in the field model. Variational approach is then adopted to approximate the estimation. The proposed method was applied for two different variants of Chamfer template matching and evaluated through the task of object detection. Experimental results on benchmark datasets including ETHZShapeClass and INRIAHorse have shown that the proposed method could significantly improve the accuracy of template matching while not sacrificing much of the efficiency. Comparisons with other recent template matching algorithms have also shown the robustness of the proposed method.
Duc Thanh Nguyen
CVPR1
2014 Food image classification using local appearance and global structural information
Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Yasmine C. Probst, Wanqing Li 0001
Neurocomputing1
2013 Inter-occlusion reasoning for human detection based on variational mean field
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
Neurocomputing1
2013 A novel shape-based non-redundant local binary pattern descriptor for object detection
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
Pattern Recognit.1
2011 Detecting humans under occlusion using variational mean field method
abstract
This paper proposes a human detection method using variational mean field approximation for occlusion reasoning. In the method, parts of human objects are detected individually using template matching. Initial detection hypotheses with spatial layout information are represented in a graphical model and refined through a Bayesian estimation. In this paper, mean field method is employed for such an estimation. The proposed method was evaluated on the popular CAVIAR-INRIA dataset. Experimental results show that the proposed algorithm is able to detect humans in severe occlusion within reasonable processing time.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICIP1
2011 Human detection with contour-based local motion binary patterns
abstract
This paper presents a human detection method using contour- based local motion features. The local motion is encoded using a variant of the popular Local Binary Pattern (LBP) called Non-Redundant Local Binary Pattern (NRLBP) descriptor computed on the difference image of two consecutive frames. In addition, the local motion features are extracted along the human's boundary contour. Localising features on the contours has the advantage of utilizing a precise human shape description. A motivation of the proposed method is that most of informative movements are performed on boundary contours of the body parts, e.g. legs of pedestrians. Evaluation of the proposed method was conducted on the INRIA and ETH datasets. Apart from showing the importance of motion information, experimental results also showed that localising features along the object boundary contours improves the detection performance.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICIP1
2011 Smoke detection in videos using Non-Redundant Local Binary Pattern-based features
abstract
This paper presents a novel and low complexity method for real-time video-based smoke detection. As a local texture operator, Non-Redundant Local Binary Pattern (NRLBP) is more discriminative and robust to illumination changes in comparison with original Local Binary Pattern (LBP), thus is employed to encode the appearance information of smoke. Non-Redundant Local Motion Binary Pattern (NRLMBP), which is computed on the difference image of consecutive frames, is introduced to capture the motion information of smoke. Experimental results show that NRLBP outperforms the original LBP in the smoke detection task. Furthermore, the combination of NRLBP and NRLMBP, which can be considered as a spatial-temporal descriptor of smoke, can lead to remarkable improvement on detection performance.
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Duc Thanh Nguyen, Ce Zhan
MMSP4
2010 Human detection using local shape and Non-Redundant binary patterns
abstract
Motivated by the advantages of using shape matching technique in detecting objects in various postures and viewpoints and the discriminative power of local patterns in object recognition, this paper proposes a human detection method combining both shape and appearance cues. In particular, local shapes of the body parts are detected using template matching. Based on body parts' shapes, local appearance features are extracted. We introduce a novel local binary pattern (LBP) descriptor, called Non-Redundant LBP (NRLBP), to encode local appearance of human. The proposed method was evaluated and compared with other state-of-the-art human detection methods on two commonly used datasets: MIT and INRIA pedestrian test sets. We also performed extensive experiments on selecting appropriate parameters as well as verifying the improvement of the proposed method through all stages of the framework.
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ICARCV1
2010 Object detection using Non-Redundant Local Binary Patterns
abstract
Local Binary Pattern (LBP) as a descriptor, has been successfully used in various object recognition tasks because of its discriminative property and computational simplicity. In this paper a variant of the LBP referred to as Non-Redundant Local Binary Pattern (NRLBP) is introduced and its application for object detection is demonstrated. Compared with the original LBP descriptor, the NRLBP has advantage of providing a more compact description of object's appearance. Furthermore, the NRLBP is more discriminative since it reflects the relative contrast between the background and foreground. The proposed descriptor is employed to encode human's appearance in a human detection task. Experimental results show that the NRLBP is robust and adaptive with changes of the background and foreground and also outperforms the original LBP in detection task.
Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Wanqing Li 0001
ICIP1
2010 On the Combination of Local Texture and Global Structure for Food Classification
abstract
This paper proposes a food image classification method using local textural patterns and their global structure to describe the food image. In this paper, a visual codebook of local textural patterns is created by employing Scale Invariant Feature Transformation (SIFT) interest point detector with the Local Binary Pattern (LBP) feature. In addition to describing the food image using local texture, the global structure of the food object is represented as the spatial distribution of the local textural structures and encoded using shape context. We evaluated the proposed method on the Pittsburgh Fast-Food Image (PFI) dataset. Experimental results showed that the proposed method could obtain better performance than the baseline experiment on the PFI dataset.
Zhimin Zong, Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ISM2
2009 An Improved Template Matching Method for Object Detection
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ACCV (3)1
2009 A novel template matching method for human detection
abstract
This paper proposes a novel weighted template matching method. It employs a generalized distance transform (GDT) and an orientation map (OM). The GDT allows us to weight the distance transform more on the strong edge points and the OM provides supplementary local orientation information for matching. Based on the matching method, a two-stage human detection method consisting of template matching and Bayesian verification is developed. Experimental results have shown that the proposed method can effectively reduce the false positive and false negative detection rates and perform superiorly in comparison to the conventional Chamfer matching method.
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ICIP1
2009 Human detection based on weighted template matching
abstract
This paper proposes a new two-stage human detection method involving matching and verification. A Bayesian framework is developed to verify the matching score obtained from a weighted distance measure. Performance evaluation indicates that the proposed method is able to utilize the flexible matching scheme and produce superior true positive, true negative and low misclassification rates.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICME1
2008 A rotation method for binary document images using DDA algorithm
abstract
DDA (Digital Differential Analyzer) is a famous algorithm used commonly in computer graphics to interpolate integer coordinate pixels of a straight line. In this paper, we introduce a method of image rotation for binary document images using DDA algorithm with assumption that the true skew angles of the documents have already been computed. The proposed method applies the main idea of DDA algorithm with some modifications for the skew scanning lines along to the inverse direction of the skew angle. In this method the ratios between the length of black runs and the whole scan line are guaranteed. Thus the algorithm can overcome disadvantages of mathematical rotation such as white holes and over segmentation. Moreover, using DDA algorithm to approximate integer points helps this method reduce the number of rotation operations.
Duc Thanh Nguyen
ACM Symposium on Document Engineering1
2007 A Robust Document Skew Estimation Algorithm Using Mathematical Morphology
abstract
The most common problem of all document skew estimation methods using morphological operators is the limitation of skew angles and choosing the appropriate size for structuring elements. This paper introduces an improvement for methods based on mathematical morphology to estimate the skew of document images with arbitrary angles of orientation. Our proposed method can take advantages in being free with user parameters and adaptive with arbitrary skew angles of document images. Moreover, the experimental results and comparison with other methods reveal the potential of our proposed method in skew estimation not only for Latin documents but also for documents of other languages such as Japanese, Chinese, Thai, etc.
Duc Thanh Nguyen, Dai Binh Vo, Tu Mi Nguyen, Thuy Giang Nguyen
ICTAI (1)1