EDBT 2026 Demo / reviewers in the wild / expert
Ta Duc Huy
dblp:293/7541 · also Huy D. Ta
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0001-6181-0478ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Vision-Language Foundations for No-Reference Image Quality AssessmentabstractLarge-scale vision–language pre-training has recently shown promise for no-reference image-quality assessment (NR-IQA), yet the relative merits of modern Vision Transformer foundations remain poorly understood. In this work, we present the first systematic evaluation of six prominent pretrained backbones, CLIP, SigLIP2, DINOv2, DINOv3, Perception, and ResNet, for the task of No-Reference Image Quality Assessment (NR-IQA), each fine-tuned using an identical lightweight MLP head. Our study uncovers two previously overlooked factors: (1) SigLIP2 consistently achieves strong performance; and (2) the choice of activation function plays a surprisingly crucial role, particularly for enhancing the generalization ability of image quality assessment models. Notably, we find that simple sigmoid activations outperform commonly used ReLU and GELU on several benchmarks. Motivated by this finding, we introduce a learnable activation selection mechanism that adaptively determines the nonlinearity for each channel, eliminating the need for manual activation design, and achieving new state-of-the-art SRCC on CLIVE, KADID10K, and AGIQA3K. Extensive ablations confirm the benefits across architectures and regimes, establishing strong, resource-efficient NR-IQA baselines. Code is available at https://github.com/drkkgy/NR_IQA_AGM Ta Duc Huy, Lingqiao Liu |
WACV | 2 |
| 2025 | ProjectedEx: Enhancing Generation in Explainable AI for Prostate CancerabstractProstate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy. Recent advancements in explainable AI and representation learning have significantly improved prostate cancer diagnosis by enabling automated and precise lesion classification. However, existing explainable AI methods, particularly those based on frameworks like generative adversarial networks (GANs), are predominantly developed for natural image generation, and their application to medical imaging often leads to suboptimal performance due to the unique characteristics and complexity of medical image. To address these challenges, our paper introduces three key contributions. First, we propose ProjectedEx, a generative framework that provides interpretable, multi-attribute explanations, effectively linking medical image features to classifier decisions. Second, we enhance the encoder module by incorporating feature pyramids, which enables multiscale feedback to refine the latent space and improves the quality of generated explanations. Additionally, we conduct comprehensive experiments on both the generator and classifier, demonstrating the clinical relevance and effectiveness of ProjectedEx in enhancing interpretability and supporting the adoption of AI in medical settings. Code will be released at https://github.com/Richardqiyi/ProjectedEx. Xuyin Qi, Zeyu Zhang 0006, Aaron Berliano Handoko, Huazhan Zheng, Mingxi Chen, Ta Duc Huy, Vu Minh Hieu Phan, Linqi Cheng, Zhibin Liao, Yang Zhao 0019, Minh-Son To |
CBMS | 6 |
| 2025 | Interactive Medical Image Analysis with Concept-based Similarity ReasoningabstractThe ability to interpret and intervene model decisions is important for the adoption of computer-aided diagnosis methods in clinical workflows. Recent concept-based methods link the model predictions with interpretable concepts and modify their activation scores to interact with the model. However, these concepts are at the image level, which hinders the model from pinpointing the exact patches the concepts are activated. Alternatively, prototype-based methods learn representations from training image patches and compare these with test image patches, using the similarity scores for final class prediction. However, interpreting the underlying concepts of these patches can be challenging and often necessitates post-hoc guesswork. To address this issue, this paper introduces the novel Concept-based Similarity Reasoning network (CSR), which offers (i) patch-level prototype with intrinsic concept interpretation, and (ii) spatial interactivity. First, the proposed CSR provides localized explanation by grounding prototypes of each concept on image regions. Second, our model introduces novel spatial-level interaction, allowing doctors to engage directly with specific image areas, making it an intuitive and transparent tool for medical imaging. CSR improves upon prior state-of-the-art interpretable methods by up to 4.5% across three biomedical datasets. Our code is released at https://github.com/tadeephuy/InteractCSR. Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan Verjans, Minh-Son To, Vu Minh Hieu Phan |
CVPR | 1 |
| 2025 | Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingabstractVisual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets. Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan |
ICCV | 1 |
| 2025 | Localizing Before Answering: A Benchmark for Grounded Medical Visual Question AnsweringabstractMedical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA. Minh Khoi Ho, Ta Duc Huy, Thanh Tam Nguyen, Qi Chen 0014, Kumar Rav, Quy Duong Dang, Satwik Ramchandre, Son Lam Phung, Zhibin Liao, Minh-Son To, Johan Verjans, Phi-Le Nguyen, Vu Minh Hieu Phan |
IJCAI | 3 |
| 2025 | PedCLIP: A Vision-Language Model for Pediatric X-Rays with Mixture of Body Part Experts
Ta Duc Huy, Abin Shoby, Sen Kim Tran, Yutong Xie 0001, Qi Chen 0014, Phi-Le Nguyen, Akshay Gole, Lingqiao Liu, Antonios Perperidis, Mark Friswell, Rebecca Linke, Andrea Glynn, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao, Minh Hieu Phan |
MICCAI (5) | 1 |
| 2024 | A Multi-phase Multi-graph Approach for Focal Liver Lesion Classification on CT Scans
Tran Bao Sam, Ta Duc Huy, Cong Tuyen Dao, Thanh Tin Lam, Van Ha Tang, Steven Quoc Hung Truong |
ACCV (10) | 2 |
| 2023 | Revisiting Reverse Distillation for Anomaly DetectionabstractAnomaly detection is an important application in large-scale industrial manufacturing. Recent methods for this task have demonstrated excellent accuracy but come with a latency trade-off. Memory based approaches with dominant performances like PatchCore or Coupled-hypersphere-based Feature Adaptation (CFA) require an external memory bank, which significantly lengthens the execution time. Another approach that employs Reversed Distillation (RD) can perform well while maintaining low latency. In this paper, we revisit this idea to improve its performance, establishing a new state-of-the-art benchmark on the challenging MVTec dataset for both anomaly detection and localization. The proposed method, called RD++, runs six times faster than PatchCore, and two times faster than CFA but introduces a negligible latency compared to RD. We also experiment on the BTAD and Retinal OCT datasets to demonstrate our method's generalizability and conduct important ablation experiments to provide insights into its configurations. Source code will be available at https://github.com/tientrandinh/Revisiting-Reverse-Distillation. Tran Dinh Tien, Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Steven Quoc Hung Truong |
CVPR | 4 |
| 2023 | Logovit: Local-Global Vision Transformer for Object Re-IdentificationabstractObject re-identification (ReID) is prone to errors under variations in scale, illumination, complex background, and object occlusion scenarios. To overcome these challenges, attention mechanisms are employed to focus on the object's characteristics, thereby extracting better discriminative features. This paper introduces a local-global vision transformer (LoGoViT) for object re-identification by learning a hierarchical-level representation from fine-grained (local) to general (global) context features. It comprises two components: (i) shift and shuffle operations to generate robust local features and (ii) local-global module to aggregate the multi-level hierarchy features of an object. Extensive experiments show that our method achieves state-of-the-art on the ReID benchmarks. We further investigate effective augmentation operations and discuss how the patch modifications improve the proposed model's generalization under occlusion scenarios. The source code is available at https://github.com/nguyenphan99/LoGoViT. Nguyen Phan, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Hoang Tran, Sam Tran, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
ICASSP | 2 |
| 2022 | Improving Local Features with Relevant Spatial Information by Vision Transformer for Crowd Counting
Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Phan, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
BMVC | 2 |
| 2022 | Adaptive Proxy Anchor Loss for Deep Metric LearningabstractDeep metric learning (or simply called metric learning) uses the deep neural network to learn the representation of images, leading to widely used in many applications, e.g. image retrieval and face recognition. In the metric learning approaches, proxy anchor takes advantage of proxy-based and pair-based approaches to enable fast convergence time and robustness to noisy labels. However, in training the proxy anchor, selecting the hyperparameter margin is important to achieve a good performance. This selection requires expertise and is time-consuming. This paper proposes a novel method to learn the margin while training the proxy anchor approach adaptively. The proposed adaptive proxy anchor simplifies the hyperparameter tuning process while advancing the proxy anchor. We achieve state of the art on three public datasets with a noticeably faster convergence time. Our code is available at https: //github.com/tks1998/Adaptive-Proxy-Anchor Nguyen Phan, Sen Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
ICIP | 3 |
| 2022 | Vietnamese Capitalization and Punctuation Recovery Models
Hoang Thi Thu Uyen, Nguyen Anh Tu, Ta Duc Huy |
INTERSPEECH | 3 |
| 2022 | ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text MiningabstractPre-trained language models have become crucial to achieving competitive results across many Natural Language Processing (NLP) problems. For monolingual pre-trained models in low-resource languages, the quantity has been significantly increased. However, most of them relate to the general domain, and there are limited strong baseline language models for domain-specific. We introduce ViHealthBERT, the first domain-specific pre-trained language model for Vietnamese healthcare. The performance of our model shows strong results while outperforming the general domain language models in all health-related datasets. Moreover, we also present Vietnamese datasets for the healthcare domain for two tasks are Acronym Disambiguation (AD) and Frequently Asked Questions (FAQ) Summarization. We release our ViHealthBERT to facilitate future research and downstream application for Vietnamese NLP in domain-specific. Our dataset and code are available in https://github.com/demdecuong/vihealthbert. Nguyen Phuc Minh, Tran Hoang Vu, Vu Hoang, Ta Duc Huy, Trung Huu Bui, Steven Quoc Hung Truong |
LREC | 4 |
| 2021 | Diffeomorphism Matching for Fast Unsupervised Pretraining on Radiographs
Huynh Minh Thanh, Chanh D. Tr. Nguyen, Ta Duc Huy, Hoang Cao Huyen, Trung H. Bui, Steven Quoc Hung Truong |
BMVC | 3 |
| 2021 | ReSORT: an ID-recovery multi-face tracking method for surveillance camerasabstractAs an improvement over the standard simple online real-time tracking (SORT) method, DeepSORT introduces a cascade matching mechanism to track objects during a certain period of occlusion, effectively reducing the number of identity (ID) switches. However, DeepSORT lacks the capability of lost-identities recovery, which enables robustness and performance in face recognition systems. To address the issue, we propose a novel multi-face tracking method, named ReSORT, that can recover lost identities. Our method removes the cascade matching block in DeepSORT and extends a similarity matching (SM) block after the Kalman filter to assign uncertain tracks to their probable tracking IDs. Such arrangement significantly reduces the processing time while maintaining the longevity of tracking IDs. The SM block functions by storing existing facial features and comparing the similarity between the new and the existing facial features, enabling ReSORT to recover the lost IDs or IDs from other cameras. To benchmark the ID-recovery ability, we introduce three new metrics, calling IDnew, TIDRate, and TReRate. We also produce face tracking annotations for three public surveillance camera datasets, i.e., LAB, MSU-AVIS, and ChokePoint. Extensive experiments conducted on the three datasets with various resolutions and frame-rates settings demonstrate the superiority of ReSORT over DeepSORT, i.e. reducing the identity switches by average 36.38%, and the processing time by 5.19 times. Source code and annotations of all three datasets are available at https://github.com/tantm97/ReSORT. Tan M. Tran, Nguyen Hoang Tran, Soan Thi Minh Duong, Ta Duc Huy, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
FG | 4 |
| 2021 | ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
Ta Duc Huy, Nguyen Anh Tu, Tran Hoang Vu, Nguyen Phuc Minh, Nguyen Phan, Trung H. Bui, Steven Quoc Hung Truong |
ICONIP (6) | 1 |