Van-Tuong-Lan Le

dblp:381/5940 · also Tuong-Lan Le Van · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0008-2687-5346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 A Lightweight CNN-Transformer Hybrid Network for Efficient Cancer Detection Using Ultrasound Images
Thanh-An Pham, Van-Dung Hoang, Van-Tuong-Lan Le, Dung Nguyen 0006
ACIIDS (2)3
2026 Early-Stage Integration of Attention Mechanisms in OmniPose for Human Pose Estimation
Khac-Anh Phu, Van-Dung Hoang, Van-Tuong-Lan Le, Quang-Khai Tran
ACIIDS (1)3
2026 Towards enhancing learning on imbalanced data: A novel adaptive weighting strategy
Dung Nguyen 0006, Van-Dung Hoang, Van-Tuong-Lan Le
Neurocomputing3
2025 CerMixer: An Efficient Model for Cervical Cancer Classification Based on Patching and Multi-scale Depthwise Convolutional Fusion
Thanh-An Pham, Van-Dung Hoang, Doan-Hieu Tran, Van-Tuong-Lan Le
ACIIDS (1)4
2025 A Real-Time Object Detection and Tracking Framework Based on RT-DETR and DeepSORT
abstract
Object detection and tracking are two critical tasks in computer vision, widely applied in security surveillance, autonomous vehicles, and behavioral analysis. Strong performance in object detection has been demonstrated by recent Transformer-based models, such as RT-DETR (Real-Time Detection Transformer), due to their global context modeling and high accuracy. However, an inherent tracking mechanism is lacking in RT-DETR, which requires additional components to maintain identity consistency across frames. To address this limitation, an integration of RT-DETR with DeepSORT is proposed, leveraging the strengths of both models to enhance real-time object detection and tracking. A comparative evaluation with YOLOv8, a widely used real-time detector, is conducted to highlight the advantages of the proposed approach in tracking accuracy and robustness. Experiments show that effective performance is achieved in challenging scenarios such as object occlusion and intersection. Specifically, an IDF1 score of 60.0%, a MOTA of 42.4%, and a MOTP of 43.3% are obtained by RTl+DeepSORT on the MOT17-02-DPM dataset, outperforming YOLOv8x+DeepSORT. These results indicate that significant improvements in tracking accuracy are attained while real-time efficiency is maintained, making the proposed approach well-suited for applications such as intelligent surveillance and autonomous navigation.
Dung Nguyen 0006, Van-Dung Hoang, Van-Tuong-Lan Le, Tri-Cong Pham, Quang-Khai Tran
HSI3
2025 WDCViT: Enhancing Monkeypox Prediction via Lightweight Vision Transformer with Window Attention and Dilated Convolutions
abstract
Monkeypox was declared a public health emergency of international concern by WHO in July 2022 due to the unprecedented global spread of the disease outside of previously endemic countries in Africa. Previously proposed classification models based mainly on pre-trained CNNs are ineffective in detecting Monkeypox. In this paper, a novel model is presented for the Monkeypox virus detection hybrid window attention and depthwise asymmetric dilation convolution. This study's novelty lies in the proposed lightweight and robust multi-stage model, utilizing depthwise dilation convolution and pointwise convolution in the two first stages with a large spatial dimension. Window attention combination with depthwise asymmetric dilation convolution is applied at the two last stages to reduce model complexity. We have evaluated the proposed approach with ViT, Swin Transformer, and MaxViT, DenseNet201, and ResNet50 on the public dataset of “Monkeypox Skin Images Dataset” (MSID) on Mendeley. The model performance was evaluated using metrics such as accuracy, recall, precision, specificity, and F1-score. The proposed method achieves the best results with an average of 99.30%, 98.65%, 98.60%, 99.56%, and 98.60% in accuracy, precision, recall, and specificity, F1-score., respectively. Additionally, the proposed model has only 19M parameters, which is smaller than MaxViT and Swin Transformer.
Thanh-An Pham, Van-Tuong-Lan Le, Van-Dung Hoang
HSI2
2025 A Survey on Vehicle Damage Detection using Deep Learning Towards Intelligent Insurance
abstract
Vehicle damage detection plays a vital role in assessing damage severity and estimating repair costs, which are essential for efficient insurance claim processing. However, this task is still predominantly manual, making it time-consuming and prone to errors or fraud due to human involvement. This paper presents an experimental comparative analysis of state-of-the-art object detection models for identifying and classifying vehicle damage. The evaluated models include RT-DETR, YOLOv12n, YOLOv11n, and YOLOv8n. Experimental results indicate that the models perform differently in detecting small and complex damage areas and in maintaining stable performance under varying lighting and viewing conditions. RT-DETR achieves the highest mean Average Precision (mAP) of 54.1%, outperforming YOLOv11n (47.0%), YOLOv8n (46.2%), and YOLOv12n (48.1%). This analysis highlights the potential of object detection models in the domain of vehicle damage assessment, contributing to cost reduction, improved transparency, and more efficient insurance claim workflows.
Doan-Hieu Tran, Van-Dung Hoang, Dung Nguyen 0006, Thanh-An Pham, Van-Tuong-Lan Le
HSI5
2025 Predicting occluded skeletal joints via tracking-based feature extraction
Khac-Anh Phu, Van-Dung Hoang, Van-Tuong-Lan Le
Neurocomputing3
2024 V-DETR: Pure Transformer for End-to-End Object Detection
Dung Nguyen 0006, Van-Dung Hoang, Van-Tuong-Lan Le
ACIIDS (2)3
2024 Enhancing Human Pose Estimation with SE-Block in the OmniPose Model
abstract
The interaction and communication between humans and computers have brought diversity and richness to the field of computer vision research, offering numerous potentials and challenges in developing human action recognition applications. In this domain, recognizing human actions from image or video data plays a crucial role in various practical applications, from security surveillance to interactive control. Although the OmniPose model has demonstrated its effectiveness, there is still potential to improve its performance. In the scope of this study, we focus on enhancing the OmniPose model, which is used for extracting skeleton data from input image data. We propose two improvement methods: utilizing the Self-Attention mechanism and employing Squeeze-and-Excitation to enhance the skeleton data extraction capability of the OmniPose model. Through this approach, we aim to contribute to enhancing the performance of the OmniPose model in skeleton data extraction and human pose recognition, while opening doors to advancements in human action recognition in computer vision.
Khac-Anh Phu, Van-Dung Hoang, Van-Tuong-Lan Le, Thinh Vinh Le
HSI3