Doanh C. Bui

dblp:309/9322 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0003-1310-5808ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
abstract
Lifelong learning on Whole Slide Images (WSIs) aims to train or fine-tune a unified model sequentially on cancer-related tasks, reducing the resources and effort required for data transfer and processing, especially given the gigabyte-scale size of WSIs. In this paper, we introduce MergeSlide, a simple yet effective framework that treats lifelong learning as a model-merging problem by leveraging a vision–language pathology foundation model. When a new task arrives, it is ❶ defined with class-aware prompts, ❷ fine-tuned for a few epochs in a classifier-free manner, and ❸ merged into a unified model using an orthogonal continual-merging strategy that preserves performance and mitigates catastrophic forgetting. For inference under the class-incremental learning (CLASS-IL) setting, where task identity is unknown, we introduce Task-to-Class Prompt-aligned (TCP) inference. Specifically, TCP first identifies the most relevant task using task-level prompts and then applies the corresponding class-aware prompts to generate predictions. To evaluate MergeSlide, we conduct experiments on a stream of six TCGA datasets. The results show that MergeSlide outperforms both rehearsal-based continual learning and vision-language zero-shot baselines. Code and data are available at https://github.com/caodoanh2001/MergeSlide.
Doanh C. Bui, Ba Hung Ngo, Hoai Luan Pham, Khang Nguyen 0001, Maï K. Nguyen, Yasuhiko Nakashima
WACV1
2026 Welcome new doctor: Continual learning with expert consultation and autoregressive inference for whole slide image analysis
abstract
Whole Slide Image (WSI) analysis, with its ability to reveal detailed tissue structures in magnified views, plays a crucial role in cancer diagnosis and prognosis. Due to their giga-sized nature, WSIs require substantial storage and computational resources for processing and training predictive models. With the rapid increase in WSIs used in clinics and hospitals, there is a growing need for a continual learning system that can efficiently process and adapt existing models to new tasks without retraining or fine-tuning on previous tasks. Such a system must balance resource efficiency with high performance. In this study, we introduce COSFormer, a Transformer-based continual learning framework tailored for multi-task WSI analysis. COSFormer is designed to learn sequentially from new tasks wile avoiding the need to revisit full historical datasets. We evaluate COSFormer on a sequence of seven WSI datasets covering seven organs and six WSI-related tasks under both class-incremental and task-incremental settings. The results demonstrate COSFormer's superior generalizability and effectiveness compared to existing continual learning frameworks, establishing it as a robust solution for continual WSI analysis in clinical applications. The code is released at https://github.com/QuIIL/COSFormer.
Doanh C. Bui, Jin Tae Kwak
Medical Image Anal.1
2026 UIT-OpenViIC: An open-domain benchmark for evaluating image captioning in Vietnamese
Doanh C. Bui, Nghia Hieu Nguyen, Khang Nguyen 0001
Signal Process. Image Commun.1
2025 HiGDA: Hierarchical Graph of Nodes to Learn Local-to-Global Topology for Semi-Supervised Domain Adaptation
abstract
The enhanced representational power and broad applicability of deep learning models have attracted significant interest from the research community in recent years. However, these models often struggle to perform effectively under domain shift conditions, where the training data (the source domain) is related to but exhibits different distributions from the testing data (the target domain). To address this challenge, previous studies have attempted to reduce the domain gap between source and target data by incorporating a few labeled target samples during training—a technique known as semi-supervised domain adaptation (SSDA). While this strategy has demonstrated notable improvements in classification performance, the network architectures used in these approaches primarily focus on exploiting the features of individual images, leaving room for improvement in capturing rich representations. In this study, we introduce a Hierarchical Graph of Nodes designed to simultaneously present representations at both feature and category levels. At the feature level, we introduce a local graph to identify the most relevant patches within an image, facilitating adaptability to defined main object representations. At the category level, we employ a global graph to aggregate the features from samples within the same category, thereby enriching overall representations. Extensive experiments on widely used SSDA benchmark datasets, including Office-Home, DomainNet, and VisDA2017, demonstrate that both quantitative and qualitative results substantiate the effectiveness of HiGDA, establishing it as a new state-of-the-art method.
Ba Hung Ngo, Doanh C. Bui, Nhat-Tuong Do-Tran, Tae Jong Choi
AAAI2
2025 How to enrich cross-domain representations? Data augmentation, cycle-pseudo labeling, and category-aware graph learning
Ba Hung Ngo, Doanh C. Bui, Tae Jong Choi
Expert Syst. Appl.2
2025 CLEAR: Cross-Transformers With Pre-Trained Language Model for Person Attribute Recognition and Retrieval
Doanh C. Bui, Thinh V. Le, Ba Hung Ngo, Tae Jong Choi
Pattern Recognit.1
2025 Spatially-Constrained and -Unconstrained Bi-Graph Interaction Network for Multi-Organ Pathology Image Classification
abstract
In computational pathology, graphs have shown to be promising for pathology image analysis. There exist various graph structures that can discover differing features of pathology images. However, the combination and interaction between differing graph structures have not been fully studied and utilized for pathology image analysis. In this study, we propose a parallel, bi-graph neural network, designated as SCUBa-Net, equipped with both graph convolutional networks and Transformers, that processes a pathology image as two distinct graphs, including a spatially-constrained graph and a spatially-unconstrained graph. For efficient and effective graph learning, we introduce two inter-graph interaction blocks and an intra-graph interaction block. The inter-graph interaction blocks learn the node-to-node interactions within each graph. The intra-graph interaction block learns the graph-to-graph interactions at both global- and local-levels with the help of the virtual nodes that collect and summarize the information from the entire graphs. SCUBa-Net is systematically evaluated on four multi-organ datasets, including colorectal, prostate, gastric, and bladder cancers. The experimental results demonstrate the effectiveness of SCUBa-Net in comparison to the state-of-the-art convolutional neural networks, Transformer, and graph neural networks.
Doanh C. Bui, Boram Song, Kyungeun Kim, Jin Tae Kwak
IEEE Trans. Medical Imaging1
2024 MECFormer: Multi-task Whole Slide Image Classification with Expert Consultation Network
Doanh C. Bui, Jin Tae Kwak
ACCV (2)1
2024 FALFormer: Feature-Aware Landmarks Self-attention for Whole-Slide Image Classification
Doanh C. Bui, Trinh Thi Le Vuong, Jin Tae Kwak
MICCAI (4)1
2024 Transformer with multi-level grid features and depth pooling for image captioning
Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001
Mach. Vis. Appl.1
2024 Transformer-Based Spatio-Temporal Unsupervised Traffic Anomaly Detection in Aerial Videos
abstract
Anomaly detection is an area of video analysis and plays an increasing role in ensuring safety, preventing risks, and guaranteeing quick response in intelligent surveillance systems. It has become a popular research topic and has piqued the interest of researchers in different communities, such as computer vision, machine learning, remote sensing, and data mining, in recent years. This promotes novel mobile systems where drones are equipped with cameras to help people find better and more efficient solutions to automatically detect anomalies (e.g., car accidents, traffic congestion, street fighting) in traffic surveillance videos. However, anomaly detection methods are still rarely studied and developed in the remote sensing community due to anomalous events rarely occurring in real life, along with the high similarities between the objects of interest with small sizes, multi-scale objects, complex backgrounds of great variations, and high overlap between objects. Therefore, in order to fully exploit the spatio-temporal information for anomaly detection in traffic surveillance circumstances, we propose a future frame prediction network based on transformer architectures to detect abnormal events from drone videography in an unsupervised way. Our model treats consecutive video frames from an input clip and feeds features to a transformer encoder to capture spatial and temporal representations from the sequence. Then, it leverages a decoder to predict the next frame. Furthermore, an event with high reconstruction error is identified as an anomaly in the test phase. Thoroughly empirical studies demonstrate that our method achieves superior performance on the UIT-ADrone dataset and largely outperforms the state-of-the-art anomaly methods on the Drone-Anomaly dataset in aerial surveillance. The source code is available online at https://github.com/Tungufm/ASTT.
Tung Minh Tran, Doanh C. Bui, Tam V. Nguyen 0002, Khang Nguyen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Improving human-object interaction with auxiliary semantic information and enhanced instance representation
Khang Nguyen 0001, Thinh V. Le, Huyen Ngoc N. Van, Doanh C. Bui
Pattern Recognit. Lett.4