Minh-Son To

dblp:215/8051 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-8060-6218ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unpaired multi-modal multi-label learning for detecting endometriosis signs
abstract
Endometriosis is a widespread gynecological disorder causing severe pain and infertility, with diagnosis currently relying on slow, costly, and risky laparoscopy. This highlights the critical need for non-invasive imaging diagnostics using transvaginal ultrasound (TVUS) and magnetic resonance imaging (MRI). A key challenge is that patients typically receive only one scan modality in practice, despite TVUS and MRI offering differing diagnostic strengths for endometriosis signs like Pouch of Douglas (POD) obliteration and bowel nodules (BN). Previous work partially addressed this challenge by leveraging unpaired multi-modal data for detecting a single marker: Pouch of Douglas (POD) obliteration. However, this is restrictive because endometriosis signs, such as POD obliteration and bowel nodules (BN), often provide correlated diagnostic cues. Capturing these correlations is essential for accurate detection of endometriosis imaging signs, particularly when combined with multi-modal learning, as each modality offers complementary strengths for different signs. To overcome these limitations, we propose EndoFusion, a novel unpaired multi-modal, multi-label learning framework that enables the detection of POD and BN from TVUS and MRI. Our approach introduces three key innovations: (1) label-based pairing, mixup, and cross-modal feature exchange for robust single-modality inference; (2) Dynamic Mutual Knowledge Distillation (DMKD), which adaptively selects teachers using a worst-student-oriented strategy for effective cross-modal transfer; and (3) label correlations modeling with multi-head attention and a specialized loss to handle imbalance and boost accuracy. This design ensures that knowledge from the superior modality and from co-occurring signs is effectively transferred, mitigating modality-specific weaknesses and improving robustness in imaging sign detection. Experiments on our endometriosis dataset show that our method significantly outperforms all comparison methods, achieving an average AUC of 0.827 (95% CI: 0.790-0.861) when evaluated using single-modality inference. These results represent an initial proof-of-concept toward multi-modal, non-invasive assessment of selected endometriosis imaging signs from MRI and TVUS.
Hu Wang 0005, Yutong Xie 0001, Minh-Son To, Steven Knox, Mathew Leonardi, George Condous, Jodie Avery, Louise Hull, Gustavo Carneiro 0001
Artif. Intell. Medicine4
2025 ProjectedEx: Enhancing Generation in Explainable AI for Prostate Cancer
abstract
Prostate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy. Recent advancements in explainable AI and representation learning have significantly improved prostate cancer diagnosis by enabling automated and precise lesion classification. However, existing explainable AI methods, particularly those based on frameworks like generative adversarial networks (GANs), are predominantly developed for natural image generation, and their application to medical imaging often leads to suboptimal performance due to the unique characteristics and complexity of medical image. To address these challenges, our paper introduces three key contributions. First, we propose ProjectedEx, a generative framework that provides interpretable, multi-attribute explanations, effectively linking medical image features to classifier decisions. Second, we enhance the encoder module by incorporating feature pyramids, which enables multiscale feedback to refine the latent space and improves the quality of generated explanations. Additionally, we conduct comprehensive experiments on both the generator and classifier, demonstrating the clinical relevance and effectiveness of ProjectedEx in enhancing interpretability and supporting the adoption of AI in medical settings. Code will be released at https://github.com/Richardqiyi/ProjectedEx.
Xuyin Qi, Zeyu Zhang 0006, Aaron Berliano Handoko, Huazhan Zheng, Mingxi Chen, Ta Duc Huy, Vu Minh Hieu Phan, Linqi Cheng, Zhibin Liao, Yang Zhao 0019, Minh-Son To
CBMS14
2025 Interactive Medical Image Analysis with Concept-based Similarity Reasoning
abstract
The ability to interpret and intervene model decisions is important for the adoption of computer-aided diagnosis methods in clinical workflows. Recent concept-based methods link the model predictions with interpretable concepts and modify their activation scores to interact with the model. However, these concepts are at the image level, which hinders the model from pinpointing the exact patches the concepts are activated. Alternatively, prototype-based methods learn representations from training image patches and compare these with test image patches, using the similarity scores for final class prediction. However, interpreting the underlying concepts of these patches can be challenging and often necessitates post-hoc guesswork. To address this issue, this paper introduces the novel Concept-based Similarity Reasoning network (CSR), which offers (i) patch-level prototype with intrinsic concept interpretation, and (ii) spatial interactivity. First, the proposed CSR provides localized explanation by grounding prototypes of each concept on image regions. Second, our model introduces novel spatial-level interaction, allowing doctors to engage directly with specific image areas, making it an intuitive and transparent tool for medical imaging. CSR improves upon prior state-of-the-art interpretable methods by up to 4.5% across three biomedical datasets. Our code is released at https://github.com/tadeephuy/InteractCSR.
Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan Verjans, Minh-Son To, Vu Minh Hieu Phan
CVPR9
2025 Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification Models
abstract
Counterfactual explanations (CFE) for deep image classifiers aim to reveal how minimal input changes lead to different model decisions, providing critical insights for model interpretation and improvement. However, existing CFE methods often rely on additional image encoders and generative models to create plausible images, neglecting the classifier's own feature space and decision boundaries. As such, they do not explain the intrinsic feature space and decision boundaries learned by the classifier. To address this limitation, we propose Mirror-CFE, a novel method that generates faithful counterfactual explanations by operating directly in the classifier's feature space, treating decision boundaries as mirrors that ``reflect'' feature representations in the mirror. Mirror-CFE learns a mapping function from feature space to image space while preserving distance relationships, enabling smooth transitions between source images and their counterfactuals. Through extensive experiments on four image datasets, we demonstrate that Mirror-CFE achieves superior performance in validity while maintaining input resemblance compared to state-of-the-art explanation methods. Finally, mirror-CFE provides interpretable visualization of the classifier's decision process by generating step-wise transitions that reveal how features evolve as classification confidence changes.
Townim F. Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao
ICCV5
2025 Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
abstract
Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets.
Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan
ICCV11
2025 Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering
abstract
Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA.
Minh Khoi Ho, Ta Duc Huy, Thanh Tam Nguyen, Qi Chen 0014, Kumar Rav, Quy Duong Dang, Satwik Ramchandre, Son Lam Phung, Zhibin Liao, Minh-Son To, Johan Verjans, Phi-Le Nguyen, Vu Minh Hieu Phan
IJCAI11
2025 MedConv: Convolutions Beat Transformers on Long-Tailed Bone Density Prediction
abstract
Bone density prediction via CT scans to estimate T-scores is crucial, providing a more precise assessment of bone health compared to traditional methods like X-ray bone density tests, which lack spatial resolution and the ability to detect localized changes. However, CT-based prediction faces two major challenges: the high computational complexity of transformer-based architectures, which limits their deployment in portable and clinical settings, and the imbalanced, long-tailed distribution of real-world hospital data that skews predictions. To address these issues, we introduce MedConv, a convolutional model for bone density prediction that outperforms transformer models with lower computational demands. We also adapt Bal-CE loss and post-hoc logit adjustment to improve class balance. Extensive experiments on our AustinSpine dataset shows that our approach achieves up to 21% improvement in accuracy and 20% in ROC AUC over previous state-of-the-art methods. Code will be available at https://github.com/Richardqiyi/MedConv.
Xuyin Qi, C. Zeyu Zhang, Huazhan Zheng, Mingxi Chen, Numan Kutaiba, Ruth Lim, Cherie Chiang, Zi En Tham, Xuan Ren, Wenxin Zhang 0005, Wenbing Lv, Guangzhen Yao, Renda Han, Kangsheng Wang, Hongtao Mao, Yu Li 0047, Zhibin Liao, Yang Zhao 0019, Minh-Son To
IJCNN22
2025 PedCLIP: A Vision-Language Model for Pediatric X-Rays with Mixture of Body Part Experts
Ta Duc Huy, Abin Shoby, Sen Kim Tran, Yutong Xie 0001, Qi Chen 0014, Phi-Le Nguyen, Akshay Gole, Lingqiao Liu, Antonios Perperidis, Mark Friswell, Rebecca Linke, Andrea Glynn, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao, Minh Hieu Phan
MICCAI (5)13
2024 Act Like a Radiologist: Radiology Report Generation Across Anatomical Regions
Qi Chen 0014, Yutong Xie 0001, Biao Wu 0006, Minh-Son To, Xiaojun Chang, Qi Wu 0001
ACCV (6)6
2024 CAPE: CAM as a Probabilistic Ensemble for Enhanced DNN Interpretation
abstract
Deep Neural Networks (DNNs) are widely used for visual classification tasks, but their complex computation process and black-box nature hinder decision transparency and interpretability. Class activation maps (CAMs) and recent variants provide ways to visually explain the DNN decision-making process by displaying ‘attention’ heatmaps of the DNNs. Nevertheless, the CAM explanation only offers relative attention information, that is, on an attention heatmap, we can interpret which image region is more or less important than the others. However, these regions cannot be meaningfully compared across classes, and the contribution of each region to the model's class prediction is not revealed. To address these challenges that ultimately lead to better DNN Interpretation, in this paper, we propose CAPE, a novel reformulation of CAM that provides a unified and probabilistically meaningful assessment of the contributions of image regions. We quantitatively and qualitatively compare CAPE with state-of-the-art CAM methods on CUB and ImageNet benchmark datasets to demonstrate enhanced interpretability. We also test on a cytology imaging dataset depicting a challenging Chronic Myelomonocytic Leukemia (CMML) diagnosis problem. Code is available at: https://github.com/AIML-MED/CAPE.
Townim F. Chowdhury, Kewen Liao, Vu Minh Hieu Phan, Minh-Son To, Yutong Xie 0001, Kevin Hung, Anton van den Hengel, Johan Verjans, Zhibin Liao
CVPR4
2024 Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training Framework
abstract
Medical vision language pre-training (VLP) has emerged as a frontier of research, enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts, current methods struggle to align medical images with key pathological findings in un-structured reports. This leads to the misalignment with the target disease's textual representation. In this paper, we introduce a novel VLP framework designed to dissect disease descriptions into their fundamental aspects, leveraging prior knowledge about the visual manifestations of pathologies. This is achieved by consulting a large language model and medical experts. Integrating a Transformer module, our approach aligns an input image with the diverse elements of a disease, generating aspect-centric image representations. By consolidating the matches from each aspect, we improve the compatibility between an image and its associated disease. Additionally, capitalizing on the aspect-oriented representations, we present a dual-head Transformer tailored to process known and unknown diseases, optimizing the comprehensive detection efficacy. Conducting experiments on seven downstream datasets, ours improves the accuracy of recent methods by up to 8.56% and 17.26% for seen and unseen categories, respectively. Our code is released at https://github.com/HieuPhan33/MAVL.
Vu Minh Hieu Phan, Yutong Xie 0001, Yuankai Qi, Lingqiao Liu, Liyang Liu, Bowen Zhang 0009, Zhibin Liao, Qi Wu 0001, Minh-Son To, Johan Verjans
CVPR9
2024 PairAug: What Can Augmented Image-Text Pairs Do for Radiology?
abstract
Current vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets, a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data scarcity, however, most augmentation methods exhibit a limited focus, prioritising either image or text augmentation exclusively. Acknowledging this limitation, our objective is to devise a framework capable of concurrently augmenting medical image and text data. We design a Pairwise Augmentation (PairAug) approach that contains an Inter-patient Augmentation (InterAug) branch and an Intra-patient Augmentation (IntraAug) branch. Specifically, the InterAug branch of our approach generates radiology images using synthesised yet plausible reports derived from a Large Language Model (LLM). The generated pairs can be considered a collection of new patient cases since they are artificially created and may not exist in the original dataset. In contrast, the IntraAug branch uses newly generated reports to manipulate images. This process allows us to create new paired data for each individual with diverse medical conditions. Our extensive experiments on various downstream tasks covering medical image classification zero-shot and fine-tuning analysis demonstrate that our PairAug, concurrently expanding both image and text data, substantially outperforms image-/text-only expansion baselines and advanced medical VLP baselines. Our code is released at https://github.com/YtongXie/PairAug.
Yutong Xie 0001, Qi Chen 0014, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia 0001, Qi Wu 0001
CVPR4
2024 AdaCBM: An Adaptive Concept Bottleneck Model for Explainable and Accurate Diagnosis
Townim F. Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Minh-Son To, Yutong Xie 0001, Anton van den Hengel, Johan Verjans, Zhibin Liao
MICCAI (10)4
2024 Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis
Vu Minh Hieu Phan, Yutong Xie 0001, Bowen Zhang 0009, Yuankai Qi, Zhibin Liao, Antonios Perperidis, Son Lam Phung, Johan Verjans, Minh-Son To
MICCAI (7)9
2023 Structure-Preserving Synthesis: MaskGAN for Unpaired MR-CT Translation
Minh-Hieu Phan, Zhibin Liao, Johan Verjans, Minh-Son To
MICCAI (10)4
2023 Improved Flexibility and Interpretability of Large Vessel Stroke Prognostication Using Image Synthesis and Multi-task Learning
Minyan Zeng, Yutong Xie 0001, Minh-Son To, Lauren Oakden-Rayner, Luke Whitbread, Stephen Bacchi, Alix Bird, Luke Smith, Rebecca Scroop, Timothy Kleinig, Jim Jannes, Lyle John Palmer, Mark Jenkinson
MICCAI (5)3
2021 Self-Supervised Lesion Change Detection and Localisation in Longitudinal Multiple Sclerosis Brain Imaging
Minh-Son To, Ian G. Sarno, Chee Chong, Mark Jenkinson, Gustavo Carneiro 0001
MICCAI (7)1
2017 Mechanisms underlying a thalamocortical transformation during active tactile sensation
abstract
During active somatosensation, neural signals expected from movement of the sensors are suppressed in the cortex, whereas information related to touch is enhanced. This tactile suppression underlies low-noise encoding of relevant tactile features and the brain's ability to make fine tactile discriminations. Layer (L) 4 excitatory neurons in the barrel cortex, the major target of the somatosensory thalamus (VPM), respond to touch, but have low spike rates and low sensitivity to the movement of whiskers. Most neurons in VPM respond to touch and also show an increase in spike rate with whisker movement. Therefore, signals related to self-movement are suppressed in L4. Fast-spiking (FS) interneurons in L4 show similar dynamics to VPM neurons. Stimulation of halorhodopsin in FS interneurons causes a reduction in FS neuron activity and an increase in L4 excitatory neuron activity. This decrease of activity of L4 FS neurons contradicts the "paradoxical effect" predicted in networks stabilized by inhibition and in strongly-coupled networks. To explain these observations, we constructed a model of the L4 circuit, with connectivity constrained by in vitro measurements. The model explores the various synaptic conductance strengths for which L4 FS neurons actively suppress baseline and movement-related activity in layer 4 excitatory neurons. Feedforward inhibition, in concert with recurrent intracortical circuitry, produces tactile suppression. Synaptic delays in feedforward inhibition allow transmission of temporally brief volleys of activity associated with touch. Our model provides a mechanistic explanation of a behavior-related computation implemented by the thalamocortical circuit.
Diego A. Gutnisky, Samuel Andrew Hires, Minh-Son To, Michael Ross Bale, Karel Svoboda, David Golomb
PLoS Comput. Biol.4