Khang Nguyen 0001

dblp:82/10435-1 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-6571-7075ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
abstract
Lifelong learning on Whole Slide Images (WSIs) aims to train or fine-tune a unified model sequentially on cancer-related tasks, reducing the resources and effort required for data transfer and processing, especially given the gigabyte-scale size of WSIs. In this paper, we introduce MergeSlide, a simple yet effective framework that treats lifelong learning as a model-merging problem by leveraging a vision–language pathology foundation model. When a new task arrives, it is ❶ defined with class-aware prompts, ❷ fine-tuned for a few epochs in a classifier-free manner, and ❸ merged into a unified model using an orthogonal continual-merging strategy that preserves performance and mitigates catastrophic forgetting. For inference under the class-incremental learning (CLASS-IL) setting, where task identity is unknown, we introduce Task-to-Class Prompt-aligned (TCP) inference. Specifically, TCP first identifies the most relevant task using task-level prompts and then applies the corresponding class-aware prompts to generate predictions. To evaluate MergeSlide, we conduct experiments on a stream of six TCGA datasets. The results show that MergeSlide outperforms both rehearsal-based continual learning and vision-language zero-shot baselines. Code and data are available at https://github.com/caodoanh2001/MergeSlide.
Doanh C. Bui, Ba Hung Ngo, Hoai Luan Pham, Khang Nguyen 0001, Maï K. Nguyen, Yasuhiko Nakashima
WACV4
2026 UIT-OpenViIC: An open-domain benchmark for evaluating image captioning in Vietnamese
Doanh C. Bui, Nghia Hieu Nguyen, Khang Nguyen 0001
Signal Process. Image Commun.3
2025 Small object detection in aerial traffic imagery: A benchmark for motorbike-dominated road scenes
Dung Truong, Khanh-Duy Nguyen, Tam V. Nguyen 0002, Khang Nguyen 0001
J. Vis. Commun. Image Represent.5
2024 Transformer with multi-level grid features and depth pooling for image captioning
Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001
Mach. Vis. Appl.3
2024 Transformer-Based Spatio-Temporal Unsupervised Traffic Anomaly Detection in Aerial Videos
abstract
Anomaly detection is an area of video analysis and plays an increasing role in ensuring safety, preventing risks, and guaranteeing quick response in intelligent surveillance systems. It has become a popular research topic and has piqued the interest of researchers in different communities, such as computer vision, machine learning, remote sensing, and data mining, in recent years. This promotes novel mobile systems where drones are equipped with cameras to help people find better and more efficient solutions to automatically detect anomalies (e.g., car accidents, traffic congestion, street fighting) in traffic surveillance videos. However, anomaly detection methods are still rarely studied and developed in the remote sensing community due to anomalous events rarely occurring in real life, along with the high similarities between the objects of interest with small sizes, multi-scale objects, complex backgrounds of great variations, and high overlap between objects. Therefore, in order to fully exploit the spatio-temporal information for anomaly detection in traffic surveillance circumstances, we propose a future frame prediction network based on transformer architectures to detect abnormal events from drone videography in an unsupervised way. Our model treats consecutive video frames from an input clip and feeds features to a transformer encoder to capture spatial and temporal representations from the sequence. Then, it leverages a decoder to predict the next frame. Furthermore, an event with high reconstruction error is identified as an anomaly in the test phase. Thoroughly empirical studies demonstrate that our method achieves superior performance on the UIT-ADrone dataset and largely outperforms the state-of-the-art anomaly methods on the Drone-Anomaly dataset in aerial surveillance. The source code is available online at https://github.com/Tungufm/ASTT.
Tung Minh Tran, Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 A brief review of state-of-the-art object detectors on benchmark document images datasets
Hai Le, Nguyen D. Vo, Khang Nguyen 0001
Int. J. Document Anal. Recognit.5
2023 Improving human-object interaction with auxiliary semantic information and enhanced instance representation
Khang Nguyen 0001, Thinh V. Le, Huyen Ngoc N. Van, Doanh C. Bui
Pattern Recognit. Lett.1
2022 CDeRSNet: Towards High Performance Object Detection in Vietnamese Document Images
Thuan Q. Nguyen, Long Duong, Nguyen D. Vo, Khang Nguyen 0001
MMM (2)5
2021 Parsing Digitized Vietnamese Paper Documents
Linh Truong Dieu, Nguyen D. Vo, Tam V. Nguyen 0002, Khang Nguyen 0001
CAIP (1)5
2019 You always look again: Learning to detect the unseen objects
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002
J. Vis. Commun. Image Represent.2
2019 YADA: you always dream again for better object detection
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002
Multim. Tools Appl.2
2016 Exploiting generic multi-level convolutional neural networks for scene understanding
abstract
In this paper, we introduce the application of generic multi-level Convolutional Neural Networks (CNN) approach into the scene understanding or image parsing task. Given an input image, first, a set of similar images from the training set are retrieved based on global-level CNN feature matching similarities. Then, the input test image and the similar images are oversegmented into superpixels. Next, the class of each test image's superpixel is initialized by the majority vote of the k-nearest-neighbor superpixels based on regional-level CNN features and hand-crafted features matching. The initial superpixel parsing is later combined with per-exemplar sliding windows to roughly form the pixel labels. Eventually, the final labels are further refined by the contextual smoothing. Extensive experiments on different challenging datasets demonstrate the potentials of the proposed method.
Tam V. Nguyen 0002, Luoqi Liu, Khang Nguyen 0001
ICARCV3