VLDB 2026 Research / reviewers in the wild / expert
Khang Nguyen 0001
dblp:82/10435-1
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-6571-7075ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide ImagesabstractLifelong learning on Whole Slide Images (WSIs) aims to train or fine-tune a unified model sequentially on cancer-related tasks, reducing the resources and effort required for data transfer and processing, especially given the gigabyte-scale size of WSIs. In this paper, we introduce MergeSlide, a simple yet effective framework that treats lifelong learning as a model-merging problem by leveraging a vision–language pathology foundation model. When a new task arrives, it is ❶ defined with class-aware prompts, ❷ fine-tuned for a few epochs in a classifier-free manner, and ❸ merged into a unified model using an orthogonal continual-merging strategy that preserves performance and mitigates catastrophic forgetting. For inference under the class-incremental learning (CLASS-IL) setting, where task identity is unknown, we introduce Task-to-Class Prompt-aligned (TCP) inference. Specifically, TCP first identifies the most relevant task using task-level prompts and then applies the corresponding class-aware prompts to generate predictions. To evaluate MergeSlide, we conduct experiments on a stream of six TCGA datasets. The results show that MergeSlide outperforms both rehearsal-based continual learning and vision-language zero-shot baselines. Code and data are available at https://github.com/caodoanh2001/MergeSlide. Doanh C. Bui, Ba Hung Ngo, Hoai Luan Pham, Khang Nguyen 0001, Maï K. Nguyen, Yasuhiko Nakashima |
WACV | 4 |
| 2026 | UIT-OpenViIC: An open-domain benchmark for evaluating image captioning in Vietnamese
Doanh C. Bui, Nghia Hieu Nguyen, Khang Nguyen 0001 |
Signal Process. Image Commun. | 3 |
| 2025 | Small object detection in aerial traffic imagery: A benchmark for motorbike-dominated road scenes
Dung Truong, Khanh-Duy Nguyen, Tam V. Nguyen 0002, Khang Nguyen 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Transformer with multi-level grid features and depth pooling for image captioning
Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001 |
Mach. Vis. Appl. | 3 |
| 2024 | Transformer-Based Spatio-Temporal Unsupervised Traffic Anomaly Detection in Aerial VideosabstractAnomaly detection is an area of video analysis and plays an increasing role in ensuring safety, preventing risks, and guaranteeing quick response in intelligent surveillance systems. It has become a popular research topic and has piqued the interest of researchers in different communities, such as computer vision, machine learning, remote sensing, and data mining, in recent years. This promotes novel mobile systems where drones are equipped with cameras to help people find better and more efficient solutions to automatically detect anomalies (e.g., car accidents, traffic congestion, street fighting) in traffic surveillance videos. However, anomaly detection methods are still rarely studied and developed in the remote sensing community due to anomalous events rarely occurring in real life, along with the high similarities between the objects of interest with small sizes, multi-scale objects, complex backgrounds of great variations, and high overlap between objects. Therefore, in order to fully exploit the spatio-temporal information for anomaly detection in traffic surveillance circumstances, we propose a future frame prediction network based on transformer architectures to detect abnormal events from drone videography in an unsupervised way. Our model treats consecutive video frames from an input clip and feeds features to a transformer encoder to capture spatial and temporal representations from the sequence. Then, it leverages a decoder to predict the next frame. Furthermore, an event with high reconstruction error is identified as an anomaly in the test phase. Thoroughly empirical studies demonstrate that our method achieves superior performance on the UIT-ADrone dataset and largely outperforms the state-of-the-art anomaly methods on the Drone-Anomaly dataset in aerial surveillance. The source code is available online at https://github.com/Tungufm/ASTT. Tung Minh Tran, Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A brief review of state-of-the-art object detectors on benchmark document images datasets
Hai Le, Nguyen D. Vo, Khang Nguyen 0001 |
Int. J. Document Anal. Recognit. | 5 |
| 2023 | Improving human-object interaction with auxiliary semantic information and enhanced instance representation
Khang Nguyen 0001, Thinh V. Le, Huyen Ngoc N. Van, Doanh C. Bui |
Pattern Recognit. Lett. | 1 |
| 2022 | CDeRSNet: Towards High Performance Object Detection in Vietnamese Document Images
Thuan Q. Nguyen, Long Duong, Nguyen D. Vo, Khang Nguyen 0001 |
MMM (2) | 5 |
| 2021 | Parsing Digitized Vietnamese Paper Documents
Linh Truong Dieu, Nguyen D. Vo, Tam V. Nguyen 0002, Khang Nguyen 0001 |
CAIP (1) | 5 |
| 2019 | You always look again: Learning to detect the unseen objects
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | YADA: you always dream again for better object detection
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 2 |
| 2016 | Exploiting generic multi-level convolutional neural networks for scene understandingabstractIn this paper, we introduce the application of generic multi-level Convolutional Neural Networks (CNN) approach into the scene understanding or image parsing task. Given an input image, first, a set of similar images from the training set are retrieved based on global-level CNN feature matching similarities. Then, the input test image and the similar images are oversegmented into superpixels. Next, the class of each test image's superpixel is initialized by the majority vote of the k-nearest-neighbor superpixels based on regional-level CNN features and hand-crafted features matching. The initial superpixel parsing is later combined with per-exemplar sliding windows to roughly form the pixel labels. Eventually, the final labels are further refined by the contextual smoothing. Extensive experiments on different challenging datasets demonstrate the potentials of the proposed method. Tam V. Nguyen 0002, Luoqi Liu, Khang Nguyen 0001 |
ICARCV | 3 |