Thanh-Toan Do

dblp:28/8692 · DBLP profile ↗
← Back
83ranked-venue papers
18as first author
42since 2021 · last 2026
0000-0002-6249-0848ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 16 first-author · 17 since 2021Artificial intelligence and machine learning · 50 · 8 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Coverage-Constrained Human-AI Cooperation with Multiple Experts
abstract
Human-AI cooperative classification (HAI-CC) aims to develop hybrid intelligent systems that enhance decision-making in various high-stakes real-world scenarios by leveraging both human expertise and AI capabilities. Current HAI-CC methods primarily focus on learning-to-defer (L2D), where decisions are deferred to human experts when AI is not confident, and learning-to-complement (L2C), where AI and human experts make predictions cooperatively. However, existing research in both L2D and L2C has not effectively been explored under diverse expert knowledge to improve decision-making, particularly when constrained by the operation cost of human involvement. In this paper, we address this research gap by proposing the Coverage-constrained Learning to Defer and Complement with Specific Experts (CL2DC) method. In particular, CL2DC assesses input data before making final decisions through either AI prediction alone or by deferring to or complementing a specific human expert. Furthermore, we propose a coverage-constrained optimisation to control the cooperation cost, ensuring it approximates a target probability for AI-only selection. This approach enables an effective assessment of system performance within a specified budget. Comprehensive evaluations on both synthetic and real-world datasets demonstrate that CL2DC achieves superior performance compared to state-of-the-art HAI-CC methods.
Zheng Zhang 0046, Cuong Nguyen 0006, Kevin Wells, Thanh-Toan Do, David Rosewarne, Gustavo Carneiro 0001
AAAI4
2026 Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
abstract
Cuong Pham, Anh Dung Hoang, Cuong C. Nguyen, Trung Le, Gustavo Carneiro, Thanh-Toan Do. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Cuong Pham 0007, Dung Anh Hoang, Cuong Nguyen 0006, Trung Le 0001, Gustavo Carneiro 0001, Thanh-Toan Do
ACL (1)6
2026 ShowFlow: From robust single concept to condition-free multi-concept generation
Trong-Vu Hoang, Quang-Binh Nguyen, Thanh-Toan Do, Tam V. Nguyen 0002, Minh-Triet Tran, Trung-Nghia Le
Neurocomputing3
2026 PASS: Peer-agreement based sample selection for training with instance dependent noisy labels
abstract
Deep learning encounters significant challenges in the form of noisy-label samples, which can cause the overfitting of trained models. A primary challenge in learning with noisy-label (LNL) techniques is their ability to differentiate between hard samples (clean-label samples near the decision boundary) and instance-dependent noisy (IDN) label samples to allow these samples to be treated differently during training. Existing methodologies to identify IDN samples, including the small-loss hypothesis and feature-based selection, have demonstrated limited efficacy, thus impeding their effectiveness in dealing with real-world label noise. We present Peer-Agreement-based Sample Selection (PASS), a novel approach that utilises three classifiers, where a consensus-driven agreement between two models accurately differentiates between clean and noisy-label IDN samples to train the third model. In contrast to current techniques, PASS is specifically designed to address the complexities of IDN, where noise patterns are correlated with instance features. Our approach seamlessly integrates with existing LNL algorithms to enhance the accuracy of detecting both noisy and clean samples. Comprehensive experiments conducted on simulated benchmarks (CIFAR-100 and Red mini-ImageNet) and real-world datasets (Animal-10N, CIFAR-N, Clothing1M, and mini-WebVision) demonstrated that PASS substantially improved the performance of multiple state-of-the-art methods. This technique achieves superior classification accuracy, particularly in scenarios with high noise levels. 1
Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001
Image Vis. Comput.4
2026 Learning to complement with multiple humans
Zheng Zhang 0046, Cuong Nguyen 0006, Kevin Wells, Thanh-Toan Do, Gustavo Carneiro 0001
Pattern Recognit.4
2025 Smart Glasses for CVI: Co-Designing Extended Reality Solutions to Support Environmental Perception by People with Cerebral Visual Impairment
abstract
Cerebral Visual Impairment (CVI) is the set to be the leading cause of vision impairment, yet remains underrepresented in assistive technology research. Unlike ocular conditions, CVI affects higher-order visual processing-impacting object recognition, facial perception, and attention in complex environments. This paper presents a co-design study with two adults with CVI investigating how smart glasses, i.e. head-mounted extended reality displays, can support understanding and interaction with the immediate environment. Guided by the Double Diamond design framework, we conducted a two-week diary study, two ideation workshops, and ten iterative development sessions using the Apple Vision Pro. Our findings demonstrate that smart glasses can meaningfully address key challenges in locating objects, reading text, recognising people, engaging in conversations, and managing sensory stress. With the rapid advancement of smart glasses and increasing recognition of CVI as a distinct form of vision impairment, this research addresses a timely and under-explored intersection of technology and need.
Bhanuka Gamage, Nicola McDowell, Dijana Kovacic, Leona Holloway, Thanh-Toan Do, Arthur James Lowery, Nicholas Price, Kim Marriott
ASSETS5
2025 Enhancing Dataset Distillation via Non-Critical Region Refinement
abstract
Dataset distillation has become a popular method for compressing large datasets into smaller, more efficient representations while preserving critical information for model training. Data features are broadly categorized into two types: instance-specific features, which capture unique, fine-grained details of individual examples, and class-general features, which represent shared, broad patterns across a class. However, previous approaches often struggle to balance these features—some focus solely on class-general patterns, neglecting finer instance details, while others prioritize instance-specific features, overlooking the shared characteristics essential for class-level understanding. In this paper, we introduce the Non-Critical Region Refinement Dataset Distillation (NRR-DD) method, which preserves instance-specific details and fine-grained regions in synthetic data while enriching non-critical regions with class-general information. This approach enables models to leverage all pixel information, capturing both feature types and enhancing overall performance. Additionally, we present Distance-Based Representative (DBR) knowledge transfer, which eliminates the need for soft labels in training by relying on the distance between synthetic data predictions and one-hot encoded labels. Experimental results show that NRR-DD achieves state-of-the-art performance on both small- and large-scale datasets. Furthermore, by storing only two distances per instance, our method delivers comparable results across various settings. The code is available at https://github.com/tmtuan1307/NRR-DD.
Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Thanh-Toan Do, Dinh Q. Phung
CVPR4
2025 Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation
abstract
Recent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-training on diverse image-text datasets. While these methods achieve state-of-the-art performance across benchmarks, much of the improvement stems from base pseudolabels (CLIP zero-shot predictions) and self-training mechanisms. Thus, the training mechanism exhibits a key limitation wherein the visual embedding distribution in target domains can deviate from the visual embedding distribution in the pre-trained model, leading to misguided signals from class descriptions. This work introduces a fresh solution to reinforce these pseudo-labels and facilitate target-prompt learning, by exploiting the geometry of visual and text embeddings - an aspect that is overlooked by existing methods. We first propose to directly leverage the reference predictions (from source prompts) based on the relationship between source and target visual embeddings. We later show that there is a strong clustering behavior observed between visual and text embeddings in pre-trained multi-modal models. Building on optimal transport theory, we transform this insight into a novel strategy to enforce the clustering property in text embeddings, further enhancing the alignment in the target domain. Our experiments and ablation studies validate the effectiveness of the proposed approach, demonstrating superior performance and improved quality of target prompts in terms of representation.
Tung Long Vuong, Hoang Phan, Vy Vo, Anh Bui, Thanh-Toan Do, Trung Le 0001, Dinh Q. Phung
CVPR5
2025 MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
abstract
Continually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints.We propose MixLoRA-DSI, a novel framework that combines an expandable mixture of Low-Rank Adaptation experts with a layer-wise out-of-distribution (OOD)driven expansion strategy.Instead of allocating new experts for each new corpus, our proposed expansion strategy enables sublinear parameter growth by selectively introducing new experts only when significant number of OOD documents are detected.Experiments on NQ320k and MS MARCO Passage demonstrate that MixLoRA-DSI outperforms full-model update baselines, with minimal parameter overhead and substantially lower training costs. 1
Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do
EMNLP7
2025 Self-supervised Learning for Acoustic Few-Shot Classification
abstract
Labelled data are limited and self-supervised learning is one of the most important approaches for reducing labelling requirements. While it has been extensively explored in the image domain, it has so far not received the same amount of attention in the acoustic domain. Yet, reducing labelling is a key requirement for many acoustic applications. Specifically in bioacoustic, there are rarely sufficient labels for fully supervised learning available. This has led to the widespread use of acoustic recognisers that have been pre-trained on unrelated data for bioacoustic tasks. We posit that training on the actual task data and combining self-supervised pre-training with few-shot classification is a superior approach that has the ability to deliver high accuracy even when only a few labels are available. To this end, we introduce and evaluate a new architecture that combines CNN-based preprocessing with feature extraction based on state space models (SSMs). This combination is motivated by the fact that CNN-based networks alone struggle to capture temporal information effectively, which is crucial for classifying acoustic signals. SSMs, specifically S4 and Mamba, on the other hand, have been shown to have an excellent ability to capture long-range dependencies in sequence data. We pre-train this architecture using contrastive learning on the actual task data and subsequent fine-tuning with an extremely small amount of labelled data. We evaluate the performance of this proposed architecture for (n-shot, n-way) classification on standard benchmarks as well as real-world data. Our evaluation shows that it outperforms state-of-the-art architectures on the few-shot classification problem.
Jingyong Liang, Bernd Meyer 0001, Issac Ning Lee, Thanh-Toan Do
ICASSP4
2025 Probabilistic Learning to Defer: Handling Missing Expert Annotations and Controlling Workload Distribution
abstract
Recent progress in machine learning research is gradually shifting its focus towards *human-AI cooperation* due to the advantages of exploiting the reliability of human experts and the efficiency of AI models. One of the promising approaches in human-AI cooperation is *learning to defer* (L2D), where the system analyses the input data and decides to make its own decision or defer to human experts. Although L2D has demonstrated state-of-the-art performance, in its standard setting, L2D entails a severe limitation: all human experts must annotate the whole training dataset of interest, resulting in a time-consuming and expensive annotation process that can subsequently influence the size and diversity of the training set. Moreover, the current L2D does not have a principled way to control workload distribution among human experts and the AI classifier, which is critical to optimise resource allocation. We, therefore, propose a new probabilistic modelling approach inspired by the mixture-of-experts, where the Expectation - Maximisation algorithm is leverage to address the issue of missing expert's annotations. Furthermore, we introduce a constraint, which can be solved efficiently during the E-step, to control the workload distribution among human experts and the AI classifier. Empirical evaluation on synthetic and real-world datasets shows that our proposed probabilistic approach performs competitively, or surpasses previously proposed methods assessed on the same benchmarks.
Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001
ICLR2
2025 Multi-Perspective Data Augmentation for Few-shot Object Detection
abstract
Recent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of $17.5\%$ in nAP50 over the baseline on PASCAL VOC.
Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002
ICLR5
2025 Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
Shreya Ghosh 0001, Tom Gedeon, Thanh-Toan Do, Abhinav Dhall
ACM Multimedia4
2025 Geometry-Aware Collaborative Multi-Solutions Optimizer for Model Fine-Tuning with Parameter Efficiency
abstract
We propose a framework grounded in gradient flow theory and informed by geometric structure that provides multiple diverse solutions for a given task, ensuring collaborative results that enhance performance and adaptability across different tasks. This framework enables flexibility, allowing for efficient task-specific fine-tuning while preserving the knowledge of the pre-trained foundation models. Extensive experiments across transfer learning, few-shot learning, and domain generalization show that our proposed approach consistently outperforms existing Bayesian methods, delivering strong performance with affordable computational overhead and offering a practical solution by updating only a small subset of parameters.
Van-Anh Nguyen, Trung Le 0001, Mehrtash Harandi, Ehsan Abbasnejad, Thanh-Toan Do, Dinh Q. Phung
NeurIPS5
2025 PromptDSI: Prompt-Based Rehearsal-Free Continual Learning for Document Retrieval
Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Yinwei Wei, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do
ECML/PKDD (7)8
2025 CamoFA: A Learnable Fourier-Based Augmentation for Camouflage Segmentation
abstract
Camouflaged object detection (COD) and camouflaged instance segmentation (CIS) aim to recognize and segment objects that are blended into their surroundings, respectively. While several deep neural network models have been proposed to tackle those tasks, augmentation methods for COD and CIS have not been thoroughly explored. strategies can help improve models' performance by increasing the size and diversity of the training data and exposing the model to a wider range of variations in the data. Besides, we aim to automatically learn transformations that help to reveal the underlying structure of camou-flaged objects and allow the model to learn to identify better and segment camouflaged objects. To achieve this, we pro-pose a learnable augmentation method in the frequency domain for COD and CIS via the Fourier transform approach, dubbed CamoFA. Our method leverages a conditional generative adversarial network and cross-attention mechanism to generate a reference image and an adaptive hybrid swapping with parameters to mix the low-frequency component of the reference image and the high-frequency component of the input image. This approach aims to make camouflaged objects more visible for detection and segmentation models. Without bells and whistles, our proposed augmentation method boosts the performance of camouflaged object detectors and instance segmenters by large margins.
Minh-Quan Le, Minh-Triet Tran, Trung-Nghia Le, Tam V. Nguyen 0002, Thanh-Toan Do
WACV5
2025 Few-shot object detection via synthetic features with optimal transport
Anh-Khoa Nguyen Vu, Thanh-Toan Do, Vinh-Tiep Nguyen, Tam Le, Minh-Triet Tran, Tam V. Nguyen 0002
Comput. Vis. Image Underst.2
2025 RoTipBot: Robotic Handling of Thin and Flexible Objects Using Rotatable Tactile Sensors
abstract
This paper introduces RoTipBot, a novel robotic system for handling thin, flexible objects. Different from previous works that are limited to singulating them using suction cups or soft grippers, RoTipBot can count multiple layers and then grasp them simultaneously in a single grasp closure. Specifically, we first develop a vision-based tactile sensor named RoTip that can rotate and sense contact information around its tip. Equipped with two RoTip sensors, RoTipBot rolls and feeds multiple layers of thin, flexible objects into the centre between its fingers, enabling effective grasping. Moreover, we design a tactile-based grasping strategy that uses RoTip's sensing ability to ensure both fingers maintain secure contact with the object while accurately counting the number of fed objects. Extensive experiments demonstrate the efficacy of the RoTip sensor and the RoTipBot approach. The results show that RoTipBot not only achieves a higher success rate but also grasps and counts multiple layers simultaneously – capabilities not possible with previous methods. Furthermore, RoTipBot operates up to three times faster than state-of-the-art methods. The success of RoTipBot paves the way for future research in object manipulation using mobilised tactile sensors. All the materials used in this paper are available athttps://sites.google.com/view/rotipbot.
Daniel Fernandes Gomes, Thanh-Toan Do, Shan Luo 0001
IEEE Trans. Robotics4
2024 MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
abstract
Few-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, known as point estimation. However, this mechanism depends on prototypes (e.g. mean of K-shot) for prediction, leading to performance instability. To overcome the disadvantage of the point estimation mechanism, we propose a novel approach, dubbed MaskDiff, which models the underlying conditional distribution of a binary mask, which is conditioned on an object region and K-shot information. Inspired by augmentation approaches that perturb data with Gaussian noise for populating low data density regions, we model the mask distribution with a diffusion probabilistic model. We also propose to utilize classifier-free guided mask sampling to integrate category information into the binary mask generation process. Without bells and whistles, our proposed method consistently outperforms state-of-the-art methods on both base and novel classes of the COCO dataset while simultaneously being more stable than existing methods. The source code is available at: https://github.com/minhquanlecs/MaskDiff.
Minh-Quan Le, Tam V. Nguyen 0002, Trung-Nghia Le, Thanh-Toan Do, Minh N. Do, Minh-Triet Tran
AAAI4
2024 Conditional Distribution Modelling for Few-Shot Image Synthesis with Diffusion Models
Munawar Hayat, Abhinav Dhall, Thanh-Toan Do
ACCV (5)4
2024 Vision-Based Assistive Technologies for People with Cerebral Visual Impairment: A Review and Focus Study
abstract
Over the past decade, considerable research has investigated Vision-Based Assistive Technologies (VBAT) to support people with vision impairments to understand and interact with their immediate environment using machine learning, computer vision, image enhancement, and/or augmented/virtual reality. However, this has almost totally overlooked a growing demographic: people with Cerebral Visual Impairment (CVI). Unlike ocular vision impairments, CVI arises from damage to the brain’s visual processing centres. Through a scoping review, this paper reveals a significant research gap in addressing the needs of this demographic. Three focus studies involving 7 participants with CVI explored the challenges, current strategies, and opportunities for VBAT. We also discussed the assistive technology needs of people with CVI compared with ocular low vision. Our findings highlight the opportunity for the Human-Computer Interaction and Assistive Technologies research community to explore and address this underrepresented domain, thereby enhancing the quality of life for people with CVI.
Bhanuka Gamage, Leona Holloway, Nicola McDowell, Thanh-Toan Do, Nicholas Price, Arthur James Lowery, Kim Marriott
ASSETS4
2024 Instance-Dependent Noisy-Label Learning with Graphical Model Based Noise-Rate Estimation
Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001
ECCV (4)4
2024 MetaAug: Meta-data Augmentation for Post-training Quantization
Cuong Pham 0007, Hoang Anh Dung, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do
ECCV (27)7
2024 Learning to Complement and to Defer to Multiple Users
Zheng Zhang 0046, Wenjie Ai, Kevin Wells, David Rosewarne, Thanh-Toan Do, Gustavo Carneiro 0001
ECCV (56)5
2024 Sharpness-Aware Data Generation for Zero-shot Quantization
abstract
Zero-shot quantization aims to learn a quantized model from a pre-trained full-precision model with no access to original real training data. The common idea in zero-shot quantization approaches is to generate synthetic data for quantizing the full-precision model. While it is well-known that deep neural networks with low sharpness have better generalization ability, none of the previous zero-shot quantization works considers the sharpness of the quantized model as a criterion for generating training data. This paper introduces a novel methodology that takes into account quantized model sharpness in synthetic data generation to enhance generalization. Specifically, we first demonstrate that sharpness minimization can be attained by maximizing gradient matching between the reconstruction loss gradients computed on synthetic and real validation data, under certain assumptions. We then circumvent the problem of the gradient matching without real validation set by approximating it with the gradient matching between each generated sample and its neighbors. Experimental evaluations on CIFAR-100 and ImageNet datasets demonstrate the superiority of the proposed method over the state-of-the-art techniques in low-bit quantization settings.
Hoang Anh Dung, Cuong Pham 0007, Trung Le 0001, Jianfei Cai 0001, Thanh-Toan Do
ICML5
2024 Frequency Attention for Knowledge Distillation
abstract
Knowledge distillation is an attractive approach for learning compact deep neural networks, which learns a lightweight student model by distilling knowledge from a complex teacher model. Attention-based knowledge distillation is a specific form of intermediate feature-based knowledge distillation that uses attention mechanisms to encourage the student to better mimic the teacher. However, most of the previous attention-based distillation approaches perform attention in the spatial domain, which primarily affects local regions in the input image. This may not be sufficient when we need to capture the broader context or global information necessary for effective knowledge transfer. In frequency domain, since each frequency is determined from all pixels of the image in spatial domain, it can contain global information about the image. Inspired by the benefits of the frequency domain, we propose a novel module that functions as an attention mechanism in the frequency domain. The module consists of a learnable global filter that can adjust the frequencies of student’s features under the guidance of the teacher’s features, which encourages the student’s features to have patterns similar to the teacher’s features. We then propose an enhanced knowledge review-based distillation model by leveraging the proposed frequency attention module. The extensive experiments with various teacher and student architectures on image classification and object detection benchmark datasets show that the proposed approach outperforms other knowledge distillation methods.
Cuong Pham 0007, Van-Anh Nguyen, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do
WACV6
2023 What do Blind and Low-Vision People Really Want from Assistive Smart Devices? Comparison of the Literature with a Focus Study
abstract
Over the last decade there has been considerable research into how artificial intelligence (AI), specifically computer vision, can assist people who are blind or have low-vision (BLV) to understand their environment. However, there has been almost no research into whether the tasks (object detection, image captioning, text recognition etc.) and devices (smartphones, smart-glasses etc.) investigated by researchers align with the needs and preferences of BLV people. We identified 646 studies published in the last two and a half years that have investigated such assistive AI techniques. We analysed these papers to determine the task, device and participation by BLV individuals. We then interviewed 24 BLV people and asked for their top five AI-based applications and to rank the applications found in the literature. We found only a weak positive correlation between BLV participants’ perceived importance of tasks and researchers’ focus and that participants prefer conversational agent interface and head-mounted devices.
Bhanuka Gamage, Thanh-Toan Do, Nicholas Price, Arthur James Lowery, Kim Marriott
ASSETS2
2023 Optimal Transport Model Distributional Robustness
abstract
Distributional robustness is a promising framework for training deep learning models that are less vulnerable to adversarial examples and data distribution shifts. Previous works have mainly focused on exploiting distributional robustness in the data space. In this work, we explore an optimal transport-based distributional robustness framework in model spaces. Specifically, we examine a model distribution within a Wasserstein ball centered on a given model distribution that maximizes the loss. We have developed theories that enable us to learn the optimal robust center model distribution. Interestingly, our developed theories allow us to flexibly incorporate the concept of sharpness awareness into training, whether it's a single model, ensemble models, or Bayesian Neural Networks, by considering specific forms of the center model distribution. These forms include a Dirac delta distribution over a single model, a uniform distribution over several models, and a general Bayesian Neural Network. Furthermore, we demonstrate that Sharpness-Aware Minimization (SAM) is a specific case of our framework when using a Dirac delta distribution over a single model, while our framework can be seen as a probabilistic extension of SAM. To validate the effectiveness of our framework in the aforementioned settings, we conducted extensive experiments, and the results reveal remarkable improvements compared to the baselines.
Van-Anh Nguyen, Trung Le 0001, Anh Tuan Bui, Thanh-Toan Do, Dinh Q. Phung
NeurIPS4
2023 Flat Seeking Bayesian Neural Networks
abstract
Bayesian Neural Networks (BNNs) provide a probabilistic interpretation for deep learning models by imposing a prior distribution over model parameters and inferring a posterior distribution based on observed data. The model sampled from the posterior distribution can be used for providing ensemble predictions and quantifying prediction uncertainty. It is well-known that deep learning models with lower sharpness have better generalization ability. However, existing posterior inferences are not aware of sharpness/flatness in terms of formulation, possibly leading to high sharpness for the models sampled from them. In this paper, we develop theories, the Bayesian setting, and the variational inference approach for the sharpness-aware posterior. Specifically, the models sampled from our sharpness-aware posterior, and the optimal approximate posterior estimating this sharpness-aware posterior, have better flatness, hence possibly possessing higher generalization ability. We conduct experiments by leveraging the sharpness-aware posterior with state-of-the-art Bayesian Neural Networks, showing that the flat-seeking counterparts outperform their baselines in all metrics of interest.
Van-Anh Nguyen, Tung Long Vuong, Hoang Phan, Thanh-Toan Do, Dinh Q. Phung, Trung Le 0001
NeurIPS4
2023 Model and Feature Diversity for Bayesian Neural Networks in Mutual Learning
abstract
Bayesian Neural Networks (BNNs) offer probability distributions for model parameters, enabling uncertainty quantification in predictions. However, they often underperform compared to deterministic neural networks. Utilizing mutual learning can effectively enhance the performance of peer BNNs. In this paper, we propose a novel approach to improve BNNs performance through deep mutual learning. The proposed approaches aim to increase diversity in both network parameter distributions and feature distributions, promoting peer networks to acquire distinct features that capture different characteristics of the input, which enhances the effectiveness of mutual learning. Experimental results demonstrate significant improvements in the classification accuracy, negative log-likelihood, and expected calibration error when compared to traditional mutual learning for BNNs.
Van Cuong Pham, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do
NeurIPS6
2023 Instance-Dependent Noisy Label Learning via Graphical Modelling
abstract
Noisy labels are unavoidable yet troublesome in the ecosystem of deep learning because models can easily overfit them. There are many types of label noise, such as symmetric, asymmetric and instance-dependent noise (IDN), with IDN being the only type that depends on image information. Such dependence on image information makes IDN a critical type of label noise to study, given that labelling mistakes are caused in large part by insufficient or ambiguous information about the visual classes present in images. Aiming to provide an effective technique to address IDN, we present a new graphical modelling approach called InstanceGM, that combines discriminative and generative models. The main contributions of InstanceGM are: i) the use of the continuous Bernoulli distribution to train the generative model, offering significant training advantages, and ii) the exploration of a state-of-the-art noisy-label discriminative classifier to generate clean labels from instance-dependent noisy-label samples. InstanceGM is competitive with current noisy-label learning approaches, particularly in IDN benchmarks using synthetic and real-world datasets, where our method shows better accuracy than the competitors in most experiments1.
Arpit Garg, Cuong Nguyen 0006, Rafael Felix, Thanh-Toan Do, Gustavo Carneiro 0001
WACV4
2023 Collaborative Multi-Teacher Knowledge Distillation for Learning Low Bit-width Deep Neural Networks
abstract
Knowledge distillation which learns a lightweight student model by distilling knowledge from a cumbersome teacher model is an attractive approach for learning compact deep neural networks (DNNs). Recent works further improve student network performance by leveraging multiple teacher networks. However, most of the existing knowledge distillation-based multi-teacher methods use separately pretrained teachers. This limits the collaborative learning between teachers and the mutual learning between teachers and student. Network quantization is another attractive approach for learning compact DNNs. However, most existing network quantization methods are developed and evaluated without considering multi-teacher support to enhance the performance of quantized student model. In this paper, we propose a novel framework that leverages both multi-teacher knowledge distillation and network quantization for learning low bit-width DNNs. The proposed method encourages both collaborative learning between quantized teachers and mutual learning between quantized teachers and quantized student. During learning process, at corresponding layers, knowledge from teachers will form an importance-aware shared knowledge which will be used as input for teachers at subsequent layers and also be used to guide student. Our experimental results on CIFAR-100 and ImageNet datasets show that the compact quantized student models trained with our method achieve competitive results compared to other state-of-the-art methods, and in some cases, surpass the full precision models.
Cuong Pham 0007, Tuan Hoang, Thanh-Toan Do
WACV3
2023 PAC-Bayes Meta-Learning With Implicit Task-Specific Posteriors
abstract
We introduce a new and rigorously-formulated PAC-Bayes meta-learning algorithm that solves few-shot learning. Our proposed method extends the PAC-Bayes framework from a single-task setting to the meta-learning multiple-task setting to upper-bound the error evaluated on any, even unseen, tasks and samples. We also propose a generative-based approach to estimate the posterior of task-specific model parameters more expressively compared to the usual assumption based on a multivariate normal distribution with a diagonal covariance matrix. We show that the models trained with our proposed meta-learning algorithm are well-calibrated and accurate, with state-of-the-art calibration errors while still being competitive on classification results on few-shot classification (mini-ImageNet and tiered-ImageNet) and regression (multi-modal task-distribution regression) benchmarks.
Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Instance-Level Few-Shot Learning With Class Hierarchy Mining
abstract
Few-shot learning is proposed to tackle the problem of scarce training data in novel classes. However, prior works in instance-level few-shot learning have paid less attention to effectively utilizing the relationship between categories. In this paper, we exploit the hierarchical information to leverage discriminative and relevant features of base classes to effectively classify novel objects. These features are extracted from abundant data of base classes, which could be utilized to reasonably describe classes with scarce data. Specifically, we propose a novel superclass approach that automatically creates a hierarchy considering base and novel classes as fine-grained classes for few-shot instance segmentation (FSIS). Based on the hierarchical information, we design a novel framework called Soft Multiple Superclass (SMS) to extract relevant features or characteristics of classes in the same superclass. A new class assigned to the superclass is easier to classify by leveraging these relevant features. Besides, in order to effectively train the hierarchy-based-detector in FSIS, we apply the label refinement to further describe the associations between fine-grained classes. The extensive experiments demonstrate the effectiveness of our method on FSIS benchmarks. The source code is available here: https://github.com/nvakhoa/superclass-FSIS.
Anh-Khoa Nguyen Vu, Thanh-Toan Do, Nhat-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Tam V. Nguyen 0002
IEEE Trans. Image Process.2
2023 Multimodal Mutual Information Maximization: A Novel Approach for Unsupervised Deep Cross-Modal Hashing
abstract
In this article, we adopt the maximizing mutual information (MI) approach to tackle the problem of unsupervised learning of binary hash codes for efficient cross-modal retrieval. We proposed a novel method, dubbed cross-modal info-max hashing (CMIMH). First, to learn informative representations that can preserve both intramodal and intermodal similarities, we leverage the recent advances in estimating variational lower bound of MI to maximizing the MI between the binary representations and input features and between binary representations of different modalities. By jointly maximizing these MIs under the assumption that the binary representations are modeled by multivariate Bernoulli distributions, we can learn binary representations, which can preserve both intramodal and intermodal similarities, effectively in a mini-batch manner with gradient descent. Furthermore, we find out that trying to minimize the modality gap by learning similar binary representations for the same instance from different modalities could result in less informative representations. Hence, balancing between reducing the modality gap and losing modality-private information is important for the cross-modal retrieval tasks. Quantitative evaluations on standard benchmark datasets demonstrate that the proposed method consistently outperforms other state-of-the-art cross-modal retrieval methods.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Neural Networks Learn. Syst.2
2022 Logic Rules Meet Deep Learning: A Novel Approach for Ship Type Classification (Extended Abstract)
abstract
The shipping industry is an important component of the global trade and economy. In order to ensure law compliance and safety, it needs to be monitored. In this paper, we present a novel ship type classification model that combines vessel transmitted data from the Automatic Identification System, with vessel imagery. The main components of our approach are the Faster R-CNN Deep Neural Network and a Neuro-Fuzzy system with IF-THEN rules. We evaluate our model using real world data and showcase the advantages of this combination while also compare it with other methods. Results show that our model can increase prediction scores by up to 15.4% when compared with the next best model we considered, while also maintaining a level of explainability as opposed to common black box approaches.
Manolis Pitsikalis, Thanh-Toan Do, Alexei Lisitsa 0001, Shan Luo 0001
IJCAI2
2022 Few-shot object detection via baby learning
Anh-Khoa Nguyen Vu, Nhat-Duy Nguyen, Khanh-Duy Nguyen, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen 0002
Image Vis. Comput.6
2022 Big data directed acyclic graph model for real-time COVID-19 twitter stream detection
Bakhtiar Amen, Syahirul Faiz, Thanh-Toan Do
Pattern Recognit.3
2022 Camouflaged Instance Segmentation In-the-Wild: Dataset, Method, and Benchmark Suite
abstract
This paper pushes the envelope on decomposing camouflaged regions in an image into meaningful components, namely, camouflaged instances. To promote the new task of camouflaged instance segmentation of in-the-wild images, we introduce a dataset, dubbed CAMO++, that extends our preliminary CAMO dataset (camouflaged object segmentation) in terms of quantity and diversity. The new dataset substantially increases the number of images with hierarchical pixel-wise ground truths. We also provide a benchmark suite for the task of camouflaged instance segmentation. In particular, we present an extensive evaluation of state-of-the-art instance segmentation methods on our newly constructed CAMO++ dataset in various scenarios. We also present a camouflage fusion learning (CFL) framework for camouflaged instance segmentation to further improve the performance of state-of-the-art methods. The dataset, model, evaluation suite, and benchmark will be made publicly available on our project page.
Trung-Nghia Le, Yubo Cao, Tan-Cong Nguyen, Minh-Quan Le, Khanh-Duy Nguyen, Thanh-Toan Do, Minh-Triet Tran, Tam V. Nguyen 0002
IEEE Trans. Image Process.6
2021 Probabilistic task modelling for meta-learning
abstract
We propose probabilistic task modelling – a generative probabilistic model for collections of tasks used in meta-learning. The proposed model combines variational auto-encoding and latent Dirichlet allocation to model each task as a mixture of Gaussian distribution in an embedding space. Such modelling provides an explicit representation of a task through its task-theme mixture. We present an efficient approximation inference technique based on variational inference method for empirical Bayes parameter estimation. We perform empirical evaluations to validate the task uncertainty and task distance produced by the proposed method through correlation diagrams of the prediction accuracy on testing tasks. We also carry out experiments of task selection in meta-learning to demonstrate how the task relatedness inferred from the proposed model help to facilitate meta-learning algorithms.
Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001
UAI2
2021 Visual localization under appearance change: filtering approaches
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Shin-Fang Ch'ng, Thanh-Toan Do, Ian D. Reid 0001
Neural Comput. Appl.6
2021 Deterministic Approximate Methods for Maximum Consensus Robust Fitting
abstract
Maximum consensus estimation plays a critically important role in several robust fitting problems in computer vision. Currently, the most prevalent algorithms for consensus maximization draw from the class of randomized hypothesize-and-verify algorithms, which are cheap but can usually deliver only rough approximate solutions. On the other extreme, there are exact algorithms which are exhaustive search in nature and can be costly for practical-sized inputs. This paper fills the gap between the two extremes by proposing deterministic algorithms to approximately optimize the maximum consensus criterion. Our work begins by reformulating consensus maximization with linear complementarity constraints. Then, we develop two novel algorithms: one based on non-smooth penalty method with a Frank-Wolfe style optimization scheme, the other based on the Alternating Direction Method of Multipliers (ADMM). Both algorithms solve convex subproblems to efficiently perform the optimization. We demonstrate the capability of our algorithms to greatly improve a rough initial estimate, such as those obtained using least squares or a randomized algorithm. Compared to the exact algorithms, our approach is much more practical on realistic input sizes. Further, our approach is naturally applicable to estimation problems with geometric residuals. Matlab code and demo program for our methods can be downloaded from https://goo.gl/FQcxpi.
Huu Le, Tat-Jun Chin, Anders P. Eriksson, Thanh-Toan Do, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Deep Metric Learning Meets Deep Clustering: An Novel Unsupervised Approach for Feature Embedding
Binh X. Nguyen, Binh D. Nguyen, Gustavo Carneiro 0001, Erman Tjiputra, Quang D. Tran, Thanh-Toan Do
BMVC6
2020 Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
abstract
This paper proposes two novel techniques to train deep convolutional neural networks with low bit-width weights and activations. First, to obtain low bit-width weights, most existing methods obtain the quantized weights by performing quantization on the full-precision network weights. However, this approach would result in some mismatch: the gradient descent updates full-precision weights, but it does not update the quantized weights. To address this issue, we propose a novel method that enables direct updating of quantized weights with learnable quantization levels to minimize the cost function using gradient descent. Second, to obtain low bit-width activations, existing works consider all channels equally. However, the activation quantizers could be biased toward a few channels with high-variance. To address this issue, we propose a method to take into account the quantization errors of individual channels. With this approach, we can learn activation quantizers that minimize the quantization errors in the majority of channels. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the image classification task, using AlexNet, ResNet and MobileNetV2 architectures on CIFAR-100 and ImageNet datasets.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IJCAI2
2020 Sensor Data for Human Activity Recognition: Feature Representation and Benchmarking
abstract
The field of Human Activity Recognition (HAR) focuses on obtaining and analysing data captured from monitoring devices (e.g. sensors). There is a wide range of applications within the field; for instance, assisted living, security surveillance, and intelligent transportation. In HAR, the development of Activity Recognition models is dependent upon the data captured by these devices and the methods used to analyse them, which directly affect performance metrics. In this work, we address the issue of accurately recognising human activities using different Machine Learning (ML) techniques. We propose a new feature representation based on consecutive occurring observations and compare it against previously used feature representations using a wide range of classification methods. Experimental results demonstrate that techniques based on the proposed representation outperform the baselines and a better accuracy was achieved for both highly and less frequent actions. We also investigate how the addition of further features and their pre-processing techniques affect performance results leading to state-of-the-art accuracy on a Human Activity Recognition dataset.
Flávia Alves, Martin Gairing, Frans A. Oliehoek, Thanh-Toan Do
IJCNN4
2020 Reproducibility Companion Paper: Selective Deep Convolutional Features for Image Retrieval
abstract
In this companion paper, firstly, we briefly summarize the contributions of our main manuscript: Selective Deep Convolutional Features for Image Retrieval, published in ACM MultiMedia 2017. In addition, we provide detail instructions together with pre-configured MATLAB scripts which allow experiments to be executed and to reproduce the results reported in our main manuscript effortlessly. The source code is available at https://github.com/hnanhtuan/selectiveConvFeatures_ACMMM_reproducibility.
Tuan Hoang, Thanh-Toan Do, Ngai-Man Cheung, Michael Riegler 0001, Jan Zahálka
ACM Multimedia2
2020 Uncertainty in Model-Agnostic Meta-Learning using Variational Inference
abstract
We introduce a new, rigorously-formulated Bayesian meta-learning algorithm that learns a probability distribution of model parameter prior for few-shot learning. The proposed algorithm employs a gradient-based variational inference to infer the posterior of model parameters for a new task. Our algorithm can be applied to any model architecture and can be implemented in various machine learning paradigms, including regression and classification. We show that the models trained with our proposed meta-learning algorithm are well calibrated and accurate, with state-of-the-art calibration and classification results on three few-shot classification benchmarks (Om- niglot, mini-ImageNet and tiered-ImageNet), and competitive results in a multi-modal task-distribution regression.
Cuong Nguyen 0006, Thanh-Toan Do, Gustavo Carneiro 0001
WACV2
2020 Simultaneous compression and quantization: A joint approach for efficient unsupervised hashing
Tuan Hoang, Thanh-Toan Do, Huu Le, Dang-Khoa Le Tan, Ngai-Man Cheung
Comput. Vis. Image Underst.2
2020 Unsupervised Deep Cross-modality Spectral Hashing
abstract
This paper presents a novel framework, namely Deep Cross-modality Spectral Hashing (DCSH), to tackle the unsupervised learning problem of binary hash codes for efficient cross-modal retrieval. The framework is a two-step hashing approach which decouples the optimization into (1) binary optimization and (2) hashing function learning. In the first step, we propose a novel spectral embedding-based algorithm to simultaneously learn single-modality and binary cross-modality representations. While the former is capable of well preserving the local structure of each modality, the latter reveals the hidden patterns from all modalities. In the second step, to learn mapping functions from informative data inputs (images and word embeddings) to binary codes obtained from the first step, we leverage the powerful CNN for images and propose a CNN-based deep architecture to learn text modality. Quantitative evaluations on three standard benchmark datasets demonstrate that the proposed DCSH method consistently outperforms other state-of-the-art methods.
Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Image Process.2
2020 Compact Hash Code Learning With Binary Deep Neural Network
abstract
Learning compact binary codes for image retrieval problem using deep neural networks has recently attracted increasing attention. However, training deep hashing networks is challenging due to the binary constraints on the hash codes. In this paper, we propose deep network models and learning algorithms for learning binary hash codes given image representations under both unsupervised and supervised manners. The novelty of our network design is that we constrain one hidden layer to directly output the binary codes. This design has overcome a challenging problem in some previous works: optimizing non-smooth objective functions because of binarization. In addition, we propose to incorporate independence and balance properties in the direct and strict forms into the learning schemes. We also include a similarity preserving property in our objective functions. The resulting optimizations involving these binary, independence, and balance constraints are difficult to solve. To tackle this difficulty, we propose to learn the networks with alternating optimization and careful relaxation. Furthermore, by leveraging the powerful capacity of convolutional neural networks, we propose an end-to-end architecture that jointly learns to extract visual features and produce binary hash codes. Experimental results for the benchmark datasets show that the proposed methods compare favorably or outperform the state of the art.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Anh-Dzung Doan, Ngai-Man Cheung
IEEE Trans. Multim.1
2019 A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning
abstract
We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training process based on a na¨ıve optimization of the triplet loss function has a run-time complexity O(N^3), where N is the number of training samples. Such optimization scales poorly, and the most common approach proposed to address this high complexity issue is based on sub-sampling the set of triplets needed for the training process. Another approach explored in the field relies on an ad-hoc linearization (in terms of N) of the triplet loss that introduces class centroids, which must be optimized using the whole training set for each mini-batch – this means that a na¨ıve implementation of this approach has run-time complexity O(N^2). This complexity issue is usually mitigated with poor, but computationally cheap, approximate centroid optimization methods. In this paper, we first propose a solid theory on the linearization of the triplet loss with the use of class centroids, where the main conclusion is that our new linear loss represents a tight upper-bound to the triplet loss. Furthermore, based on the theory above, we propose a training algorithm that no longer requires the centroid optimization step, which means that our approach is the first in the field with a guaranteed linear run-time complexity. We show that the training of deep distance metric learning methods using the proposed upper-bound is substantially faster than triplet-based methods, while producing competitive retrieval accuracy results on benchmark datasets (CUB-200-2011 and CAR196).
Thanh-Toan Do, Toan Tran 0002, Ian D. Reid 0001, Tuan Hoang, Gustavo Carneiro 0001
CVPR1
2019 SDRSAC: Semidefinite-Based Randomized Approach for Robust Point Cloud Registration Without Correspondences
abstract
This paper presents a novel randomized algorithm for robust point cloud registration without correspondences. Most existing registration approaches require a set of putative correspondences obtained by extracting invariant descriptors. However, such descriptors could become unreliable in noisy and contaminated settings. In these settings, methods that directly handle input point sets are preferable. Without correspondences, however, conventional randomized techniques require a very large number of samples in order to reach satisfactory solutions. In this paper, we propose a novel approach to address this problem. In particular, our work enables the use of randomized methods for point cloud registration without the need of putative correspondences. By considering point cloud alignment as a special instance of graph matching and employing an efficient semi-definite relaxation, we propose a novel sampling mechanism, in which the size of the sampled subsets can be larger-than-minimal. Our tight relaxation scheme enables fast rejection of the outliers in the sampled sets, resulting in high quality hypotheses. We conduct extensive experiments to demonstrate that our approach outperforms other state-of-the-art methods. Importantly, our proposed method serves as a generic framework which can be extended to problems with known correspondences.
Huu Le, Thanh-Toan Do, Tuan Hoang, Ngai-Man Cheung
CVPR2
2019 Compact Trilinear Interaction for Visual Question Answering
abstract
In Visual Question Answering (VQA), answers have a great correlation with question meaning and visual contents. Thus, to selectively utilize image, question and answer information, we propose a novel trilinear interaction model which simultaneously learns high level associations between these three inputs. In addition, to overcome the interaction complexity, we introduce a multimodal tensor-based PARALIND decomposition which efficiently parameterizes trilinear teraction between the three inputs. Moreover, knowledge distillation is first time applied in Free-form Opened-ended VQA. It is not only for reducing the computational cost and required memory but also for transferring knowledge from trilinear interaction model to bilinear interaction model. The extensive experiments on benchmarking datasets TDIUC, VQA-2.0, and Visual7W show that the proposed compact trilinear interaction model achieves state-of-the-art results when using a single model on all three datasets.
Tuong KL. Do, Thanh-Toan Do, Erman Tjiputra, Quang D. Tran
ICCV3
2019 Scalable Place Recognition Under Appearance Change for Autonomous Driving
abstract
A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetation growth, etc.) environmental variations. A promising solution is to continuously accumulate images to maintain an adequate sample of the conditions and incorporate new changes into the place recognition decision. However, this demands a place recognition technique that is scalable on an ever growing dataset. To this end, we propose a novel place recognition technique that can be efficiently retrained and compressed, such that the recognition of new queries can exploit all available data (including recent changes) without suffering from visible growth in computational cost. Underpinning our method is a novel temporal image matching technique based on Hidden Markov Models. Our experiments show that, compared to state-of-the-art techniques, our method has much greater potential for large-scale place recognition for autonomous driving.
Anh-Dzung Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 0029, Thanh-Toan Do, Ian D. Reid 0001
ICCV5
2019 Bayesian Generative Active Deep Learning
abstract
Deep learning models have demonstrated outstanding performance in several problems, but their training process tends to require immense amounts of computational and human resources for training and labeling, constraining the types of problems that can be tackled. Therefore, the design of effective training methods that require small labeled training sets is an important research direction that will allow a more effective use of resources. Among current approaches designed to address this issue, two are particularly interesting: data augmentation and active learning. Data augmentation achieves this goal by artificially generating new training points, while active learning relies on the selection of the “most informative” subset of unlabeled training samples to be labelled by an oracle. Although successful in practice, data augmentation can waste computational resources because it indiscriminately generates samples that are not guaranteed to be informative, and active learning selects a small subset of informative samples (from a large un-annotated set) that may be insufficient for the training process. In this paper, we propose a Bayesian generative active deep learning approach that combines active learning with data augmentation – we provide theoretical and empirical evidence (MNIST, CIFAR-$\{10,100\}$, and SVHN) that our approach has more efficient training and better classification results than data augmentation and active learning.
Toan Tran 0002, Thanh-Toan Do, Ian D. Reid 0001, Gustavo Carneiro 0001
ICML2
2019 Overcoming Data Limitation in Medical Visual Question Answering
Binh D. Nguyen, Thanh-Toan Do, Binh X. Nguyen, Tuong KL. Do, Erman Tjiputra, Quang D. Tran
MICCAI (4)2
2019 Binary Constrained Deep Hashing Network for Image Retrieval Without Manual Annotation
abstract
Learning compact binary codes for image retrieval task using deep neural networks has attracted increasing attention recently. However, training deep hashing networks for the task is challenging due to the binary constraints on the hash codes, the similarity preserving property, and the requirement for a vast amount of labelled images. To the best of our knowledge, none of the existing methods has tackled all of these challenges completely in a unified framework. In this work, we propose a novel end-to-end deep learning approach for the task, in which the network is trained to produce binary codes directly from image pixels without the need o f manual annotation. In particular, to deal with the non-smoothness of binary constraints, we propose a novel pairwise constrained loss function, which simultaneously encodes the distances between pairs of hash codes, and the binary quantization error. In order to train the network with the proposed loss function, we propose an efficient parameter learning algorithm. In addition, to provide similar / dissimilar training images to train the network, we exploit 3D models reconstructed from unlabelled images for automatic generation of enormous training image pairs. The extensive experiments on image retrieval benchmark datasets demonstrate the improvements of the proposed method over the state-of-the-art compact representation methods on the image retrieval problem.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Trung Pham, Huu Le, Ngai-Man Cheung, Ian D. Reid 0001
WACV1
2019 Simultaneous Feature Aggregating and Hashing for Compact Binary Code Learning
abstract
Representing images by compact hash codes is an attractive approach for large-scale content-based image retrieval. In most state-of-the-art hashing-based image retrieval systems, for each image, local descriptors are first aggregated as a global representation vector. This global vector is then subjected to a hashing function to generate a binary hash code. In previous works, the aggregating and the hashing processes are designed independently. Hence, these frameworks may generate suboptimal hash codes. In this paper, we first propose a novel unsupervised hashing framework in which feature aggregating and hashing are designed simultaneously and optimized jointly. Specifically, our joint optimization generates aggregated representations that can be better reconstructed by some binary codes. This leads to more discriminative binary hash codes and improved retrieval accuracy. In addition, the proposed method is flexible. It can be extended for supervised hashing. When the data label is available, the framework can be adapted to learn binary codes which minimize the reconstruction loss with respect to label vectors. Furthermore, we also propose a fast version of the state-of-the-art hashing method Binary Autoencoder to be used in our proposed frameworks. Extensive experiments on benchmark datasets under various settings show that the proposed methods outperform the state-of-the-art unsupervised and supervised hashing methods.
Thanh-Toan Do, Khoa Le, Tuan Hoang, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung
IEEE Trans. Image Process.1
2019 Semantic Prior Analysis for Salient Object Detection
abstract
Salient object detection aims to detect the main objects in the given image. In this paper, we proposed an approach that integrates semantic priors into the salient object detection process. The method first obtains an explicit saliency map that is refined by the explicit semantic priors learned from data. Then an implicit saliency map is constructed using a trained model that maps the implicit semantic priors embedded into superpixel features with the saliency values. Next, the fusion saliency map is computed by adaptively fusing both the explicit and implicit semantic maps. The final saliency map is eventually computed via the post-processing refinement step. Experimental results have demonstrated the effectiveness of the proposed method, particularly, it achieves competitive performance with the state-of-the-art baselines on three challenging datasets, namely, ECSSD, HKUIS, and iCoSeg.
Tam V. Nguyen 0002, Khanh-Duy Nguyen, Thanh-Toan Do
IEEE Trans. Image Process.3
2019 On-Device Scalable Image-Based Localization via Prioritized Cascade Search and Fast One-Many RANSAC
abstract
We present the design of an entire on-device system for large-scale urban localization using images. The proposed design integrates compact image retrieval and 2D-3D correspondence search to estimate the location in extensive city regions. Our design is GPS agnostic and does not require network connection. In order to overcome the resource constraints of mobile devices, we propose a system design that leverages the scalability advantage of image retrieval and accuracy of 3D model-based localization. Furthermore, we propose a new hashing-based cascade search for fast computation of 2D-3D correspondences. In addition, we propose a new one-many RANSAC for accurate pose estimation. The new one-many RANSAC addresses the challenge of repetitive building structures (e.g. windows and balconies) in urban localization. Extensive experiments demonstrate that our 2D-3D correspondence search achieves the state-of-the-art localization accuracy on multiple benchmark datasets. Furthermore, our experiments on a large Google street view image dataset show the potential of large-scale localization entirely on a typical mobile device.
Ngoc-Trung Tran, Dang-Khoa Le Tan, Anh-Dzung Doan, Thanh-Toan Do, Tuan-Anh Bui, Mengxuan Tan, Ngai-Man Cheung
IEEE Trans. Image Process.4
2019 From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
abstract
In the large-scale image retrieval task, the two most important requirements are the discriminability of image representations and the efficiency in computation and storage of representations. Regarding the former requirement, Convolutional Neural Network is proven to be a very powerful tool to extract highly discriminative local descriptors for effective image search. Additionally, to further improve the discriminative power of the descriptors, recent works adopt fine-tuned strategies. In this article, taking a different approach, we propose a novel, computationally efficient, and competitive framework. Specifically, we first propose various strategies to compute masks, namely, SIFT-masks , SUM-mask , and MAX-mask , to select a representative subset of local convolutional features and eliminate redundant features. Our in-depth analyses demonstrate that proposed masking schemes are effective to address the burstiness drawback and improve retrieval accuracy. Second, we propose to employ recent embedding and aggregating methods that can significantly boost the feature discriminability. Regarding the computation and storage efficiency, we include a hashing module to produce very compact binary image representations. Extensive experiments on six image retrieval benchmarks demonstrate that our proposed framework achieves the state-of-the-art retrieval performances.
Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Huu Le, Tam V. Nguyen 0002, Ngai-Man Cheung
ACM Trans. Multim. Comput. Commun. Appl.1
2018 A Binary Optimization Approach for Constrained K-Means Clustering
Huu Le, Anders P. Eriksson, Thanh-Toan Do, Michael Milford
ACCV (4)3
2018 LieNet: Real-time Monocular Object Instance 6D Pose Estimation
Thanh-Toan Do, Trung Pham, Ian D. Reid 0001
BMVC1
2018 Non-smooth M-estimator for Maximum Consensus Estimation
Huu Le, Anders P. Eriksson, Thanh-Toan Do, Tat-Jun Chin, David Suter
BMVC3
2018 Bayesian Semantic Instance Segmentation in Open Set World
Trung Pham, Thanh-Toan Do, Gustavo Carneiro 0001, Ian D. Reid 0001
ECCV (10)3
2018 Supervised Hashing with End-to-End Binary Deep Neural Network
abstract
Image hashing is a popular technique applied to large scale content-based visual retrieval due to its compact and efficient binary codes. Our work proposes a new end-to-end deep network architecture for supervised hashing which directly learns binary codes from input images and maintains hashing properties, namely similarity preservation, independence, and balancing. Furthermore, we also propose a new learning scheme that copes with the binary constrained loss function. The proposed algorithm not only is scalable for learning over large-scale datasets but also outperforms state-of-the-art supervised hashing methods, which are illustrated throughout extensive experiments from various image retrieval benchmarks.
Dang-Khoa Le Tan, Thanh-Toan Do, Ngai-Man Cheung
ICIP2
2018 AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
abstract
We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the object to its most probable affordance label. The proposed framework employs three key components for effectively handling the multiclass problem in the affordance mask: a sequence of deconvolutional layers, a robust resizing strategy, and a multi-task loss function. The experimental results on the public datasets show that our AffordanceNet outperforms recent state-of-the-art methods by a fair margin, while its end-to-end architecture allows the inference at the speed of 150ms per image. This makes our AffordanceNet well suitable for real-time robotic applications. Furthermore, we demonstrate the effectiveness of AffordanceNet in different testing environments and in real robotic applications. The source code is available at https://github.com/nqanh/affordance-net.
Thanh-Toan Do, Anh Nguyen 0003, Ian D. Reid 0001
ICRA1
2018 SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes
abstract
This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over scene semantics and geometry allows a robot to detect and segment object instances in complex scenes where modern deep learning-based methods either fail to separate object instances, or fail to detect objects that were not seen during training. SceneCut automatically decomposes a scene into meaningful regions which either represent objects or scene surfaces. The decomposition is qualified by an unified energy function over objectness and geometric fitting. We show how this energy function can be optimized efficiently by utilizing hierarchical segmentation trees. Moreover, we leverage a pre-trained convolutional oriented boundary network to predict accurate boundaries from images, which are used to construct high-quality region hierarchies. We evaluate SceneCut on several different indoor environments, and the results show that SceneCut significantly outperforms all the existing methods.
Trung T. Pham, Thanh-Toan Do, Niko Sünderhauf, Ian D. Reid 0001
ICRA2
2018 Embedding Based on Function Approximation for Large Scale Image Search
abstract
The objective of this paper is to design an embedding method that maps local features describing an image (e.g., SIFT) to a higher dimensional representation useful for the image retrieval problem. First, motivated by the relationship between the linear approximation of a nonlinear function in high dimensional space and the state-of-the-art feature representation used in image retrieval, i.e., VLAD, we propose a new approach for the approximation. The embedded vectors resulted by the function approximation process are then aggregated to form a single representation for image retrieval. Second, in order to make the proposed embedding method applicable to large scale problem, we further derive its fast version in which the embedded vectors can be efficiently computed, i.e., in the closed-form. We compare the proposed embedding methods with the state of the art in the context of image search under various settings: when the images are represented by medium length vectors, short vectors, or binary vectors. The experimental results show that the proposed embedding methods outperform existing the state of the art on the standard public image retrieval benchmarks.
Thanh-Toan Do, Ngai-Man Cheung
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Improving Chamfer Template Matching Using Image Segmentation
abstract
This letter proposes an effective method to improve object location in Chamfer template matching (CTM) based object detection using image segmentation. In our method, object bounding boxes are iteratively adjusted to fit with the object images obtained from image segmentation in a probabilistic model. The proposed method was tested with state-of-the-art CTM-based object detectors. Experimental results have shown the proposed method improved the location accuracy of the object detectors and reduce the false alarms rate.
Duc Thanh Nguyen, Ngoc-Son Vu, Thanh-Toan Do, Thin Nguyen, John Yearwood
IEEE Signal Process. Lett.3
2018 Computation and Memory Efficient Image Segmentation
abstract
In this paper, we address the segmentation problem under limited computation and memory resources. Given a segmentation algorithm, we propose a framework that can reduce its computation time and memory requirementsimultaneously, while preserving its accuracy. The proposed framework uses standard pixel-domain downsampling and includes two main steps.Coarse segmentationis first performed on the downsampled image.Refinementis then applied to the coarse segmentation results. We make two novel contributions to enable competitive accuracy using this simple framework. First, we rigorously examine the effect of downsampling on segmentation using a signal processing analysis. The analysis helps to determine theuncertain regions, which are small image regions where pixel labels are uncertain after the coarse segmentation. Second, we propose an efficient minimum spanning tree-based algorithm to propagate the labels into the uncertain regions. We perform extensive experiments using several standard data sets. The experimental results show that our segmentation accuracy is comparable to state-of-the-art methods, while requiring much less computation time and memory than those methods.
Yiren Zhou, Thanh-Toan Do, Haitian Zheng, Ngai-Man Cheung, Lu Fang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Accessible Melanoma Detection Using Smartphones and Mobile Image Analysis
abstract
We investigate the design of an entire mobile imaging system for early detection of melanoma. Different from previous work, we focus on smartphone-captured visible light images. Our design addresses two major challenges. First, images acquired using a smartphone under loosely-controlled environmental conditions may be subject to various distortions, and this makes melanoma detection more difficult. Second, processing performed on a smartphone is subject to stringent computation and memory constraints. In our work, we propose a detection system that is optimized to run entirely on the resource-constrained smartphone. Our system intends to localize the skin lesion by combining a lightweight method for skin detection with a hierarchical segmentation approach using two fast segmentation methods. Moreover, we study an extensive set of image features and propose new numerical features to characterize a skin lesion. Furthermore, we propose an improved feature selection algorithm to determine a small set of discriminative features used by the final lightweight system. In addition, we study the human-computer interface (HCI) design to understand the usability and acceptance issues of the proposed system. Our extensive evaluation on an image dataset provided by National Skin Center - Singapore (117 benign nevi and 67 malignant melanoma) confirms the effectiveness of the proposed system for melanoma detection: 89.09% sensitivity at specificity ≥90%.
Thanh-Toan Do, Tuan Hoang, Victor Pomponiu, Yiren Zhou, Zhao Chen 0005, Ngai-Man Cheung, Dawn Chin-Ing Koh, Aaron Tan, Suat-Hoon Tan
IEEE Trans. Multim.1
2017 Simultaneous Feature Aggregating and Hashing for Large-Scale Image Search
Thanh-Toan Do, Dang-Khoa Le Tan, Trung T. Pham, Ngai-Man Cheung
CVPR1
2017 Enhancing feature discrimination for unsupervised hashing
abstract
We introduce a novel approach to improve unsupervised hashing. Specifically, we propose a very efficient embedding method: Gaussian Mixture Model embedding (Gemb). The proposed method, using Gaussian Mixture Model, embeds feature vector into a low-dimensional vector and, simultaneously, enhances the discriminative property of features before passing them into hashing. Our experiment shows that the proposed method boosts the hashing performance of many state-of-the-art, e.g. Binary Autoencoder (BA) [1], Iterative Quantization (ITQ) [2], in standard evaluation metrics for the three main benchmark datasets.
Tuan Hoang, Thanh-Toan Do, Dang-Khoa Le Tan, Ngai-Man Cheung
ICIP2
2017 Selective Deep Convolutional Features for Image Retrieval
abstract
Convolutional Neural Network (CNN) is a very powerful approach to extract discriminative local descriptors for effective image search. Recent work adopts fine-tuned strategies to further improve the discriminative power of the descriptors. Taking a different approach, in this paper, we propose a novel framework to achieve competitive retrieval performance. Firstly, we propose various masking schemes, namely SIFT-mask, SUM-mask, and MAX-mask, to select a representative subset of local convolutional features and remove a large number of redundant features. We demonstrate that this can effectively address the burstiness issue and improve retrieval accuracy. Secondly, we propose to employ recent embedding and aggregating methods to further enhance feature discriminability. Extensive experiments demonstrate that our proposed framework achieves state-of-the-art retrieval accuracy.
Tuan Hoang, Thanh-Toan Do, Dang-Khoa Le Tan, Ngai-Man Cheung
ACM Multimedia2
2016 Learning to Hash with Binary Deep Neural Network
Thanh-Toan Do, Anh-Dzung Doan, Ngai-Man Cheung
ECCV (5)1
2016 Binary Hashing with Semidefinite Relaxation and Augmented Lagrangian
Thanh-Toan Do, Anh-Dzung Doan, Duc Thanh Nguyen, Ngai-Man Cheung
ECCV (2)1
2015 FAemb: A function approximation-based embedding method for image retrieval
abstract
The objective of this paper is to design an embedding method mapping local features describing image (e.g. SIFT) to a higher dimensional representation used for image retrieval problem. By investigating the relationship between the linear approximation of a nonlinear function in high dimensional space and state-of-the-art feature representation used in image retrieval, i.e., VLAD, we first introduce a new approach for the approximation. The embedded vectors resulted by the function approximation process are then aggregated to form a single representation used in the image retrieval framework. The evaluation shows that our embedding method gives a performance boost over the state of the art in image retrieval, as demonstrated by our experiments on the standard public image retrieval benchmarks.
Thanh-Toan Do, Quang D. Tran, Ngai-Man Cheung
CVPR1
2012 Face recognition using Co-occurrence Histograms of Oriented Gradients
abstract
Recently, Histogram of Oriented Gradient (HOG) is applied in face recognition. In this paper, we apply Co-occurrence of Oriented Gradient (CoHOG), which is an extension of HOG, on the face recognition problem. Some weighted functions for magnitude gradient are tested. We also proposed a weighted approach for CoHOG, where a weight value is set for each subregion of face image. Numerical experiments performed on Yale and ORL datasets show that 1) CoHOG has recognition accuracy higher than HOG; 2) using gradient magnitude in CoHOG improves recognition results; and 3) weighted CoHOG approach improves accuracy recognition rate. The recognition results using CoHOG are competitive with some of the state of the art methods. This proves the effectiveness of CoHOG descriptor for face recognition.
Thanh-Toan Do, Ewa Kijak
ICASSP1
2012 Enlarging hacker's toolbox: Deluding image recognition by attacking keypoint orientations
abstract
Content-Based Image Retrieval Systems (CBIRS) used in forensics related contexts require very good image recognition capabilities. Whereas the robustness criterion has been extensively covered by Computer Vision or Multimedia literature, none of these communities explored the security of CBIRS. Recently, preliminary studies have shown real systems can be deluded by applying transformations to images that are very specific to the SIFT local description scheme commonly used for recognition. The work presented in this paper adds one strategy for attacking images, and somehow enlarges the box of tools hackers can use for deluding systems. This paper shows how the orientation of keypoints can be tweaked, which in turn lowers matches since this deeply changes the final SIFT feature vectors. The method learns what visual patch should be applied to change the orientation of keypoints thanks to an SVM-based process. Experiments with a database made of 100,000 real world images confirms the effectiveness of this keypoint-orientation attacking scheme.
Thanh-Toan Do, Ewa Kijak, Laurent Amsaleg, Teddy Furon
ICASSP1
2012 Security-oriented picture-in-picture visual modifications
abstract
The performance of Content-Based Image Retrieval Systems (CBIRS) is typically evaluated via benchmarking their capacity to match images despite various generic distortions such as crops, rescalings or Picture in Picture (PiP) attacks, which are the most challenging. Distortions are made in a very generic manner, by applying a set of transformations that are completely independent from the systems later performing recognition tasks. Recently, studies have shown that exploiting the finest details of the various techniques used in a CBIRS offers the opportunity to create distortions that dramatically reduce the recognition performance. Such a security perspective is taken in this paper. Instead of creating generic PiP distortions, it proposes a creation scheme able to delude the recognition capabilities of a CBIRS that is representative of state of the art techniques as it relies on SIFT, high-dimensional k-nearest neighbors searches and geometrical robustification steps. Experiments using 100,000 real-world images confirm the effectiveness of these security-oriented PiP visual modifications.
Thanh-Toan Do, Ewa Kijak, Laurent Amsaleg, Teddy Furon
ICMR1
2010 Understanding the security and robustness of SIFT
abstract
Many content-based retrieval systems (CBIRS) describe images using the SIFT local features because of their very robust recognition capabilities. While SIFT features proved to cope with a wide spectrum of general purpose image distortions, its security has not fully been assessed yet. In one of their scenario, Hsu et al. in [2] show that very specific anti-SIFT attacks can jeopardize the keypoint detection. These attacks can delude systems using SIFT targeting application such as image authentication and (pirated) copy detection.
Thanh-Toan Do, Ewa Kijak, Teddy Furon, Laurent Amsaleg
ACM Multimedia1
2010 Challenging the security of Content-Based Image Retrieval systems
abstract
Content-Based Image Retrieval (CBIR) has been recently used as a filtering mechanism against the piracy of multimedia contents. Many publications in the last few years have proposed very robust schemes where pirated contents are detected despite severe modifications. As none of these systems have addressed the piracy problem from a security perspective, it is time to check whether they are secure: Can pirates mount violent attacks against CBIR systems by carefully studying the technology they use? This paper is an initial analysis of the security flaws of the typical technology blocks used in state-of-the-art CBIR systems. It is so far too early to draw any definitive conclusion about their inherent security, but it motivates and encourages further studies on this topic.
Thanh-Toan Do, Ewa Kijak, Teddy Furon, Laurent Amsaleg
MMSP1