EDBT 2026 Demo / reviewers in the wild / expert
Johan Verjans
dblp:228/8303 · also Johan W. Verjans
· DBLP profile ↗
23ranked-venue papers
0as first author
21since 2021 · last 2025
0000-0002-8336-6774ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 11 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interactive Medical Image Analysis with Concept-based Similarity ReasoningabstractThe ability to interpret and intervene model decisions is important for the adoption of computer-aided diagnosis methods in clinical workflows. Recent concept-based methods link the model predictions with interpretable concepts and modify their activation scores to interact with the model. However, these concepts are at the image level, which hinders the model from pinpointing the exact patches the concepts are activated. Alternatively, prototype-based methods learn representations from training image patches and compare these with test image patches, using the similarity scores for final class prediction. However, interpreting the underlying concepts of these patches can be challenging and often necessitates post-hoc guesswork. To address this issue, this paper introduces the novel Concept-based Similarity Reasoning network (CSR), which offers (i) patch-level prototype with intrinsic concept interpretation, and (ii) spatial interactivity. First, the proposed CSR provides localized explanation by grounding prototypes of each concept on image regions. Second, our model introduces novel spatial-level interaction, allowing doctors to engage directly with specific image areas, making it an intuitive and transparent tool for medical imaging. CSR improves upon prior state-of-the-art interpretable methods by up to 4.5% across three biomedical datasets. Our code is released at https://github.com/tadeephuy/InteractCSR. Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan Verjans, Minh-Son To, Vu Minh Hieu Phan |
CVPR | 8 |
| 2025 | Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsabstractCounterfactual explanations (CFE) for deep image classifiers aim to reveal how minimal input changes lead to different model decisions, providing critical insights for model interpretation and improvement. However, existing CFE methods often rely on additional image encoders and generative models to create plausible images, neglecting the classifier's own feature space and decision boundaries. As such, they do not explain the intrinsic feature space and decision boundaries learned by the classifier. To address this limitation, we propose Mirror-CFE, a novel method that generates faithful counterfactual explanations by operating directly in the classifier's feature space, treating decision boundaries as mirrors that ``reflect'' feature representations in the mirror. Mirror-CFE learns a mapping function from feature space to image space while preserving distance relationships, enabling smooth transitions between source images and their counterfactuals. Through extensive experiments on four image datasets, we demonstrate that Mirror-CFE achieves superior performance in validity while maintaining input resemblance compared to state-of-the-art explanation methods. Finally, mirror-CFE provides interpretable visualization of the classifier's decision process by generating step-wise transitions that reveal how features evolve as classification confidence changes. Townim F. Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao |
ICCV | 7 |
| 2025 | Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingabstractVisual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets. Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan |
ICCV | 12 |
| 2025 | Localizing Before Answering: A Benchmark for Grounded Medical Visual Question AnsweringabstractMedical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA. Minh Khoi Ho, Ta Duc Huy, Thanh Tam Nguyen, Qi Chen 0014, Kumar Rav, Quy Duong Dang, Satwik Ramchandre, Son Lam Phung, Zhibin Liao, Minh-Son To, Johan Verjans, Phi-Le Nguyen, Vu Minh Hieu Phan |
IJCAI | 12 |
| 2025 | PedCLIP: A Vision-Language Model for Pediatric X-Rays with Mixture of Body Part Experts
Ta Duc Huy, Abin Shoby, Sen Kim Tran, Yutong Xie 0001, Qi Chen 0014, Phi-Le Nguyen, Akshay Gole, Lingqiao Liu, Antonios Perperidis, Mark Friswell, Rebecca Linke, Andrea Glynn, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao, Minh Hieu Phan |
MICCAI (5) | 15 |
| 2025 | Adaptive Quantization in Generative Flow Networks for Probabilistic Sequential PredictionabstractProbabilistic time series forecasting, essential in domains like healthcare and neuroscience, requires models capable of capturing uncertainty and intricate temporal dependencies. While deep learning has advanced forecasting, generating calibrated probability distributions over continuous future values remains challenging. We introduce Temporal Generative Flow Networks (Temporal GFNs), adapting Generative Flow Networks (GFNs) – a powerful framework for generating compositional objects – to this sequential prediction task. GFNs learn policies to construct objects (eg. forecast trajectories) step-by-step, sampling final objects proportionally to a reward signal. However, applying GFNs directly to continuous time series necessitates addressing their inherently discrete action spaces and ensuring differentiability. Our framework tackles this by representing time series segments as states and sequentially generating future values via quantized actions chosen by a forward policy. We introduce two key innovations: (1) An adaptive, curriculum-based quantization strategy that dynamically adjusts the number of discretization bins based on reward improvement and policy entropy, balancing precision and exploration throughout training. (2) A straight-through estimator mechanism enabling the forward policy to output both discrete (hard) samples for trajectory construction and continuous (soft) samples for stable gradient propagation. Training utilizes a trajectory balance loss objective, ensuring flow consistency, augmented by an entropy regularizer. We provide rigorous theoretical bounds on the quantization error's impact and the adaptive factor's range. We demonstrate how Temporal GFNs offer a principled way to leverage the structured generation capabilities of GFNs for probabilistic forecasting in continuous domains. Nadhir Vincent Hassen, Johan Verjans |
NeurIPS | 3 |
| 2024 | CAPE: CAM as a Probabilistic Ensemble for Enhanced DNN InterpretationabstractDeep Neural Networks (DNNs) are widely used for visual classification tasks, but their complex computation process and black-box nature hinder decision transparency and interpretability. Class activation maps (CAMs) and recent variants provide ways to visually explain the DNN decision-making process by displaying ‘attention’ heatmaps of the DNNs. Nevertheless, the CAM explanation only offers relative attention information, that is, on an attention heatmap, we can interpret which image region is more or less important than the others. However, these regions cannot be meaningfully compared across classes, and the contribution of each region to the model's class prediction is not revealed. To address these challenges that ultimately lead to better DNN Interpretation, in this paper, we propose CAPE, a novel reformulation of CAM that provides a unified and probabilistically meaningful assessment of the contributions of image regions. We quantitatively and qualitatively compare CAPE with state-of-the-art CAM methods on CUB and ImageNet benchmark datasets to demonstrate enhanced interpretability. We also test on a cytology imaging dataset depicting a challenging Chronic Myelomonocytic Leukemia (CMML) diagnosis problem. Code is available at: https://github.com/AIML-MED/CAPE. Townim F. Chowdhury, Kewen Liao, Vu Minh Hieu Phan, Minh-Son To, Yutong Xie 0001, Kevin Hung, Anton van den Hengel, Johan Verjans, Zhibin Liao |
CVPR | 9 |
| 2024 | Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training FrameworkabstractMedical vision language pre-training (VLP) has emerged as a frontier of research, enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts, current methods struggle to align medical images with key pathological findings in un-structured reports. This leads to the misalignment with the target disease's textual representation. In this paper, we introduce a novel VLP framework designed to dissect disease descriptions into their fundamental aspects, leveraging prior knowledge about the visual manifestations of pathologies. This is achieved by consulting a large language model and medical experts. Integrating a Transformer module, our approach aligns an input image with the diverse elements of a disease, generating aspect-centric image representations. By consolidating the matches from each aspect, we improve the compatibility between an image and its associated disease. Additionally, capitalizing on the aspect-oriented representations, we present a dual-head Transformer tailored to process known and unknown diseases, optimizing the comprehensive detection efficacy. Conducting experiments on seven downstream datasets, ours improves the accuracy of recent methods by up to 8.56% and 17.26% for seen and unseen categories, respectively. Our code is released at https://github.com/HieuPhan33/MAVL. Vu Minh Hieu Phan, Yutong Xie 0001, Yuankai Qi, Lingqiao Liu, Liyang Liu, Bowen Zhang 0009, Zhibin Liao, Qi Wu 0001, Minh-Son To, Johan Verjans |
CVPR | 10 |
| 2024 | Occluded Person Retrieval with Hierarchical Feature OptimizationabstractOccluded person retrieval aims to match images from occluded pedestrians. It pushes forward progress of person retrieval towards applications in real-world scenarios, thus attracting increasing attention in recent years. A key challenge is to learn discriminative representation within limited informative regions due to obstacle or pedestrian occlusion. To that end, we propose a hierarchical feature optimization model (HFO) that jointly optimizes image-level, object-level and part-level features for improved occluded person retrieval. A hierarchical discriminative feature grouping (HDFG) module is developed to generate hierarchical object/part masks for comprehensive feature extraction. Via learning a set of part prototypes, HDFG localizes hierarchical informative object/parts by grouping intermediate feature vectors based on their similarity to these prototypes. The proposed HFO is trained in an end-to-end manner using only identity labels, making it a practical solution for occluded person retrieval. We verify the effectiveness of the proposed method on three challenging occluded datasets and two holistic datasets, i.e., Occluded-DukeMTMC, Occluded-REID, P-DukeMTMC-reID, Market1501, and DukeMTMC-reID. Extensive experiments and ablation studies demonstrate superior or comparable performance of the proposed method over the state-of-the-art methods. The code is available at https://github.com/Patrickzad/HFO. Yang Zhao 0019, Pengcheng Zhang 0003, Xiaohan Yu 0001, Zhibin Liao, Johan Verjans, Xiao Bai 0001 |
FG | 5 |
| 2024 | AdaCBM: An Adaptive Concept Bottleneck Model for Explainable and Accurate Diagnosis
Townim F. Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Minh-Son To, Yutong Xie 0001, Anton van den Hengel, Johan Verjans, Zhibin Liao |
MICCAI (10) | 7 |
| 2024 | Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis
Vu Minh Hieu Phan, Yutong Xie 0001, Bowen Zhang 0009, Yuankai Qi, Zhibin Liao, Antonios Perperidis, Son Lam Phung, Johan Verjans, Minh-Son To |
MICCAI (7) | 8 |
| 2024 | ReFs: A hybrid pre-training paradigm for 3D medical image segmentation
Yutong Xie 0001, Lingqiao Liu, Hu Wang 0005, Yiwen Ye, Johan Verjans, Yong Xia 0001 |
Medical Image Anal. | 6 |
| 2023 | Multi-Head Multi-Loss Model Calibration
Adrian Galdran, Johan Verjans, Gustavo Carneiro 0001, Miguel Ángel González Ballester |
MICCAI (3) | 2 |
| 2023 | Structure-Preserving Synthesis: MaskGAN for Unpaired MR-CT Translation
Minh-Hieu Phan, Zhibin Liao, Johan Verjans, Minh-Son To |
MICCAI (10) | 3 |
| 2023 | Self-supervised pseudo multi-class pre-training for unsupervised anomaly detection and segmentation in medical images
Yu Tian 0001, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan Verjans, Rajvinder Singh, Gustavo Carneiro 0001 |
Medical Image Anal. | 6 |
| 2022 | Contrastive Transformer-Based Multiple Instance Learning for Weakly Supervised Polyp Frame Detection
Yu Tian 0001, Guansong Pang, Fengbei Liu, Yuyuan Liu, Chong Wang 0012, Yuanhong Chen, Johan Verjans, Gustavo Carneiro 0001 |
MICCAI (3) | 7 |
| 2022 | Intra- and Inter-Pair Consistency for Semi-Supervised Gland SegmentationabstractAccurate gland segmentation in histology tissue images is a critical but challenging task. Although deep models have demonstrated superior performance in medical image segmentation, they commonly require a large amount of annotated data, which are hard to obtain due to the extensive labor costs and expertise required. In this paper, we propose an intra- and inter-pair consistency-based semi-supervised (I2CS) model that can be trained on both labeled and unlabeled histology images for gland segmentation. Considering that each image contains glands and hence different images could potentially share consistent semantics in the feature space, we introduce a novel intra- and inter-pair consistency module to explore such consistency for learning with unlabeled data. It first characterizes the pixel-level relation between a pair of images in the feature space to create an attention map that highlights the regions with the same semantics but on different images. Then, it imposes a consistency constraint on the attention maps obtained from multiple image pairs, and thus filters low-confidence attention regions to generate refined attention maps that are then merged with original features to improve their representation ability. In addition, we also design an object-level loss to address the issues caused by touching glands. We evaluated our model against several recent gland segmentation methods and three typical semi-supervised methods on the GlaS and CRAG datasets. Our results not only demonstrate the effectiveness of the proposed due consistency module and Obj-Dice loss, but also indicate that the proposed I2CS model achieves state-of-the-art gland segmentation performance on both benchmarks. Yutong Xie 0001, Zhibin Liao, Johan Verjans, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | CNN Attention Guidance for Improved Orthopedics Radiographic Fracture ClassificationabstractConvolutional neural networks (CNNs) have gained significant popularity in orthopedic imaging in recent years due to their ability to solve fracture classification problems. A common criticism of CNNs is their opaque learning and reasoning process, making it difficult to trust machine diagnosis and the subsequent adoption of such algorithms in clinical setting. This is especially true when the CNN is trained with limited amount of medical data, which is a common issue as curating sufficiently large amount of annotated medical imaging data is a long and costly process. While interest has been devoted to explaining CNN learnt knowledge by visualizing network attention, the utilization of the visualized attention to improve network learning has been rarely investigated. This paper explores the effectiveness of regularizing CNN network with human-provided attention guidance on where in the image the network should look for answering clues. On two orthopedics radiographic fracture classification datasets, through extensive experiments we demonstrate that explicit human-guided attention indeed can direct correct network attention and consequently significantly improve classification performance. The development code for the proposed attention guidance is publicly available on https://github.com/zhibinliao89/fracture_attention_guidance. Zhibin Liao, Kewen Liao, Haifeng Shen, Marouska F. van Boxel, Jasper Prijs, Ruurd L. Jaarsma, Job N. Doornberg, Anton van den Hengel, Johan Verjans |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningabstractAnomaly detection with weakly supervised video-level labels is typically formulated as a multiple instance learning (MIL) problem, in which we aim to identify snippets containing abnormal events, with each video represented as a bag of video snippets. Although current methods show effective detection performance, their recognition of the positive instances, i.e., rare abnormal snippets in the abnormal videos, is largely biased by the dominant negative instances, especially when the abnormal events are subtle anomalies that exhibit only small differences compared with normal events. This issue is exacerbated in many methods that ignore important video temporal dependencies. To address this issue, we introduce a novel and theoretically sound method, named Robust Temporal Feature Magnitude learning (RTFM), which trains a feature magnitude learning function to effectively recognise the positive instances, substantially improving the robustness of the MIL approach to the negative instances from abnormal videos. RTFM also adapts dilated convolutions and self-attention mechanisms to capture long- and short-range temporal dependencies to learn the feature magnitude more faithfully. Extensive experiments show that the RTFM-enabled MIL model (i) outperforms several state-of-the-art methods by a large margin on four benchmark data sets (ShanghaiTech, UCF-Crime, XD-Violence and UCSD-Peds) and (ii) achieves significantly improved subtle anomaly discriminability and sample efficiency. Yu Tian 0001, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan Verjans, Gustavo Carneiro 0001 |
ICCV | 5 |
| 2021 | Constrained Contrastive Distribution Learning for Unsupervised Anomaly Detection and Localisation in Medical Images
Yu Tian 0001, Guansong Pang, Fengbei Liu, Yuanhong Chen, Seon-Ho Shin, Johan Verjans, Rajvinder Singh, Gustavo Carneiro 0001 |
MICCAI (5) | 6 |
| 2021 | Viral Pneumonia Screening on Chest X-Rays Using Confidence-Aware Anomaly DetectionabstractClusters of viral pneumonia occurrences over a short period may be a harbinger of an outbreak or pandemic. Rapid and accurate detection of viral pneumonia using chest X-rays can be of significant value for large-scale screening and epidemic prevention, particularly when other more sophisticated imaging modalities are not readily accessible. However, the emergence of novel mutated viruses causes a substantial dataset shift, which can greatly limit the performance of classification-based approaches. In this paper, we formulate the task of differentiating viral pneumonia from non-viral pneumonia and healthy controls into a one-class classification-based anomaly detection problem. We therefore propose the confidence-aware anomaly detection (CAAD) model, which consists of a shared feature extractor, an anomaly detection module, and a confidence prediction module. If the anomaly score produced by the anomaly detection module is large enough, or the confidence score estimated by the confidence prediction module is small enough, the input will be accepted as an anomaly case (i.e., viral pneumonia). The major advantage of our approach over binary classification is that we avoid modeling individual viral pneumonia classes explicitly and treat all known viral pneumonia cases as anomalies to improve the one-class model. The proposed model outperforms binary classification models on the clinical X-VIRAL dataset that contains 5,977 viral pneumonia (no COVID-19) cases, 37,393 non-viral pneumonia or healthy cases. Moreover, when directly testing on the X-COVID dataset that contains 106 COVID-19 cases and 107 normal controls without any fine-tuning, our model achieves an AUC of 83.61% and sensitivity of 71.70%, which is comparable to the performance of radiologists reported in the literature. Yutong Xie 0001, Guansong Pang, Zhibin Liao, Johan Verjans, Wenxing Li, Zongji Sun, Chunhua Shen, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Few-Shot Anomaly Detection for Polyp Frames from Colonoscopy
Yu Tian 0001, Gabriel Maicas, Leonardo Z. C. T. Pu, Rajvinder Singh, Johan Verjans, Gustavo Carneiro 0001 |
MICCAI (6) | 5 |
| 2020 | Pairwise Relation Learning for Semi-supervised Gland Segmentation
Yutong Xie 0001, Zhibin Liao, Johan Verjans, Chunhua Shen, Yong Xia 0001 |
MICCAI (5) | 4 |