Soan Thi Minh Duong

dblp:234/4271 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-2092-0088ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021
YearPublicationVenuePosition
2026 NaviFormer: Multimodal scene segmentation for assistive navigation
abstract
Safe and independent navigation is an essential part of daily life for individuals with visual impairments. Recent advances in assistive navigation have leveraged deep learning-based semantic scene segmentation with promising results by using RGB images. In comparison, RGB-D images provide richer geometric and appearance information, but their integration into deep learning frameworks for assistive navigation could be further explored to enhance segmentation accuracy. In modeling cross-modal interactions, the existing RGB-D segmentation methods tend to lose fine-grained spatial details or discard useful information, which hinders performance on dense prediction tasks. To address these gaps, we propose NaviFormer, a novel RGB-D semantic segmentation architecture tailored for assistive navigation. NaviFormer features a dual-stream transformer encoder with shared weights to efficiently extract latent features from RGB and depth modalities. It also incorporates a Local-Global Cross-Modal Fusion module, which facilitates effective information exchange between the two modalities across both local and global feature levels. In training NaviFormer, we further employ a pixel-wise contrastive loss to enhance the separability of pixel-level embeddings in the RGB-D feature space. Extensive experiments on TrueSight and Cityscapes datasets indicate that NaviFormer achieves superior performance compared to existing RGB-D segmentation methods. Our findings highlight the importance of leveraging RGB-D data for enhancing semantic understanding in assistive navigation systems, and establish NaviFormer as a solid baseline for future research in this domain.
Ly Bui, Son Lam Phung, Yang Di, Soan Thi Minh Duong, Abdesselam Bouzerdoum
Comput. Vis. Image Underst.4
2025 A Confidence-Based Sampling Strategy for Dense Temporal Token Learning in Thermal Infrared Object Tracking
abstract
Single object tracking (SOT) aims to locate and follow a target object in a video sequence. While most existing SOT methods rely on color images due to the widespread use of color cameras, challenging weather (e.g., nighttime, rain, fog) and rapid object pose variations often cause lost or missed tracks, degrading performance. These issues can be miti-gated by utilizing thermal infrared (TIR) data, which is less susceptible to adverse weather conditions, and leveraging temporal information to handle the significant object pose changes across frames. This paper introduces a confidence-based sampling strategy for dense temporal token SOT, using the TIR data to address tracking failures. Our approach generates both the object’s state, as in conventional methods, and an additional confidence score that quantifies the tracking reliability. Obtained from a simple yet effective multi-layer perceptron-based scoring head, which processes dense temporal tokens capturing the target’s appearance, spatio-temporal location, and trajectory, the confidence score guides the sampling strategy to suggest optimal reference frames for subsequent tracking. Intensive experiments on LSOTB-TIR, the largest TIR dataset, demonstrate that our proposed method achieves state-of-the-art performance, improving upon the baseline by 1.17% in success rate, 1.50% in precision, and 1.07% in normalized precision.
Quynh T. X. Hoang, Soan Thi Minh Duong, Ly Bui, Duy Q. Tran
ICIP2
2024 DAP: A dataset-agnostic predictor of neural network performance
abstract
Training a deep neural network on a large dataset to convergence is a time-demanding task. This task often must be repeated many times, especially when developing a new deep learning algorithm or performing a neural architecture search. This problem can be mitigated if a deep neural network’s performance can be estimated without actually training it. In this work, we investigate the feasibility of two tasks: (i) predicting a deep neural network’s performance accurately given only its architectural descriptor, and (ii) generalizing the predictor across different datasets without re-training. To this end, we propose a dataset-agnostic regression framework that uses a novel dual-LSTM model and a new dataset difficulty feature. The experimental results show that both tasks above are indeed feasible, and the proposed method outperforms the existing techniques in all experimental cases. Additionally, we also demonstrate several practical use-cases of the proposed predictor.
Sui Paul Ang, Soan Thi Minh Duong, Son Lam Phung, Abdesselam Bouzerdoum
Neurocomputing2
2023 Revisiting Reverse Distillation for Anomaly Detection
abstract
Anomaly detection is an important application in large-scale industrial manufacturing. Recent methods for this task have demonstrated excellent accuracy but come with a latency trade-off. Memory based approaches with dominant performances like PatchCore or Coupled-hypersphere-based Feature Adaptation (CFA) require an external memory bank, which significantly lengthens the execution time. Another approach that employs Reversed Distillation (RD) can perform well while maintaining low latency. In this paper, we revisit this idea to improve its performance, establishing a new state-of-the-art benchmark on the challenging MVTec dataset for both anomaly detection and localization. The proposed method, called RD++, runs six times faster than PatchCore, and two times faster than CFA but introduces a negligible latency compared to RD. We also experiment on the BTAD and Retinal OCT datasets to demonstrate our method's generalizability and conduct important ablation experiments to provide insights into its configurations. Source code will be available at https://github.com/tientrandinh/Revisiting-Reverse-Distillation.
Tran Dinh Tien, Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Steven Quoc Hung Truong
CVPR5
2023 Logovit: Local-Global Vision Transformer for Object Re-Identification
abstract
Object re-identification (ReID) is prone to errors under variations in scale, illumination, complex background, and object occlusion scenarios. To overcome these challenges, attention mechanisms are employed to focus on the object's characteristics, thereby extracting better discriminative features. This paper introduces a local-global vision transformer (LoGoViT) for object re-identification by learning a hierarchical-level representation from fine-grained (local) to general (global) context features. It comprises two components: (i) shift and shuffle operations to generate robust local features and (ii) local-global module to aggregate the multi-level hierarchy features of an object. Extensive experiments show that our method achieves state-of-the-art on the ReID benchmarks. We further investigate effective augmentation operations and discuss how the patch modifications improve the proposed model's generalization under occlusion scenarios. The source code is available at https://github.com/nguyenphan99/LoGoViT.
Nguyen Phan, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Hoang Tran, Sam Tran, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong
ICASSP3
2023 MSD-NAS: multi-scale dense neural architecture search for real-time pedestrian lane detection
abstract
Abstract Accurate detection of pedestrian lanes is a crucial criterion for vision-impaired people to navigate freely and safely. The current deep learning methods have achieved reasonable accuracy at this task. However, they lack practicality for real-time pedestrian lane detection due to non-optimal accuracy, speed, and model size trade-off. Hence, an optimized deep neural network (DNN) for pedestrian lane detection is required. Designing a DNN from scratch is a laborious task that requires significant experience and time. This paper proposes a novel neural architecture search (NAS) algorithm, named MSD-NAS, to automate this laborious task. The proposed method designs an optimized deep network with multi-scale input branches, allowing the derived network to utilize local and global contexts for predictions. The search is also performed in a large and generic space that includes many existing hand-designed network architectures as candidates. To further boost performance, we propose a Short-term Visual Memory mechanism to improve information facilitation within the derived networks. Evaluated on the PLVP3 dataset of 10,000 images, the DNN designed by MSD-NAS achieves state-of-the-art accuracy (0.9781) and mIoU (0.9542), while being 20.16 times faster and 2.56 times smaller than the current best deep learning model.
Sui Paul Ang, Son Lam Phung, Soan Thi Minh Duong, Abdesselam Bouzerdoum
Appl. Intell.3
2022 Dual consistency assisted multi-confident learning for the hepatic vessel segmentation using noisy labels
Nam Nguyen Phuong, Tuan Van Vo, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong
BMVC3
2022 Improving Local Features with Relevant Spatial Information by Vision Transformer for Crowd Counting
Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Phan, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong
BMVC3
2022 Adaptive Proxy Anchor Loss for Deep Metric Learning
abstract
Deep metric learning (or simply called metric learning) uses the deep neural network to learn the representation of images, leading to widely used in many applications, e.g. image retrieval and face recognition. In the metric learning approaches, proxy anchor takes advantage of proxy-based and pair-based approaches to enable fast convergence time and robustness to noisy labels. However, in training the proxy anchor, selecting the hyperparameter margin is important to achieve a good performance. This selection requires expertise and is time-consuming. This paper proposes a novel method to learn the margin while training the proxy anchor approach adaptively. The proposed adaptive proxy anchor simplifies the hyperparameter tuning process while advancing the proxy anchor. We achieve state of the art on three public datasets with a noticeably faster convergence time. Our code is available at https: //github.com/tks1998/Adaptive-Proxy-Anchor
Nguyen Phan, Sen Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong
ICIP4
2021 ReSORT: an ID-recovery multi-face tracking method for surveillance cameras
abstract
As an improvement over the standard simple online real-time tracking (SORT) method, DeepSORT introduces a cascade matching mechanism to track objects during a certain period of occlusion, effectively reducing the number of identity (ID) switches. However, DeepSORT lacks the capability of lost-identities recovery, which enables robustness and performance in face recognition systems. To address the issue, we propose a novel multi-face tracking method, named ReSORT, that can recover lost identities. Our method removes the cascade matching block in DeepSORT and extends a similarity matching (SM) block after the Kalman filter to assign uncertain tracks to their probable tracking IDs. Such arrangement significantly reduces the processing time while maintaining the longevity of tracking IDs. The SM block functions by storing existing facial features and comparing the similarity between the new and the existing facial features, enabling ReSORT to recover the lost IDs or IDs from other cameras. To benchmark the ID-recovery ability, we introduce three new metrics, calling IDnew, TIDRate, and TReRate. We also produce face tracking annotations for three public surveillance camera datasets, i.e., LAB, MSU-AVIS, and ChokePoint. Extensive experiments conducted on the three datasets with various resolutions and frame-rates settings demonstrate the superiority of ReSORT over DeepSORT, i.e. reducing the identity switches by average 36.38%, and the processing time by 5.19 times. Source code and annotations of all three datasets are available at https://github.com/tantm97/ReSORT.
Tan M. Tran, Nguyen Hoang Tran, Soan Thi Minh Duong, Ta Duc Huy, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong
FG3
2020 Real-time Pedestrian Lane Detection for Assistive Navigation using Neural Architecture Search
abstract
Pedestrian lane detection is a core component in many assistive and autonomous navigation systems. These systems are usually deployed in environments that require realtime processing. Many state-of-the-art deep neural networks only focus on detection accuracy but not inference speed. Without further modifications, they are not suitable for real-time applications. Furthermore, the task of designing a high-performing deep neural network is time-consuming and requires experience. To tackle these issues, we propose a neural architecture search algorithm that can find the best deep network for pedestrian lane detection automatically. The proposed method searches in a network-level space using the gradient descent algorithm. Evaluated on a dataset of 5,000 images, the deep network found by the proposed algorithm achieves comparable segmentation accuracy, while being significantly faster than other state-of-the-art methods. The proposed method has been successfully implemented as a real-time pedestrian lane detection tool.
Sui Paul Ang, Son Lam Phung, Abdesselam Bouzerdoum, Thi Nhat Anh Nguyen, Soan Thi Minh Duong, Mark M. Schira
ICPR5
2018 Anatomy-Guided Inverse-Gradient Susceptibility Artifact Correction Method for High-Resolution FMRI
abstract
Functional Magnetic Resonance Imaging (fMRI) is a widely used and non-invasive technique for recording changes in brain activity. However, susceptibility artifacts are ubiquitous distortions in fMRI, especially strong in high-resolution images, causing the misrepresentation of brain function and structure in the affected regions. Here, we present a novel method for correcting these distortions in high-resolution fMRI images based on the hyper-elastic susceptibility artifact correction (HySCO) method. The novelty of the proposed method is the utilization of the easily-acquired T1-weighted (T1w) anatomy image as a ground-truth measurement to regularize deformations, thereby obtaining meaningful corrections. The performance of the new method is compared to that of HySCO. Results from high-resolution (1mm) EPI data are presented demonstrating the robustness of the new method for image correction and its suitability for subsequent fMRI analysis.
Soan Thi Minh Duong, Mark M. Schira, Son Lam Phung, Abdesselam Bouzerdoum, H. G. B. Taylor
ICASSP1