Junlin Hou

dblp:275/7810 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-7989-5913ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Benchmarking real-world medical image classification with noisy labels: Challenges, practice, and outlook
abstract
Learning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research on learning with noisy labels (LNL), the robustness of existing methods in medical imaging has not been systematically assessed. To address this gap, we introduce LNMBench, a comprehensive benchmark for Label Noise in Medical imaging. LNMBench encompasses \textbf{10} representative methods evaluated across 7 datasets, 6 imaging modalities, and 3 noise patterns, establishing a unified and reproducible framework for robustness evaluation under realistic conditions. Comprehensive experiments reveal that the performance of existing LNL methods degrades substantially under high and real-world noise, highlighting the persistent challenges of class imbalance and domain variability in medical data. Motivated by these findings, we further propose a simple yet effective improvement to enhance model robustness under such conditions. The LNMBench codebase is publicly released to facilitate standardized evaluation, promote reproducible research, and provide practical insights for developing noise-resilient algorithms in both research and real-world medical applications.The codebase is publicly available on https://github.com/myyy777/LNMBench.
Junlin Hou, Chao Zhang 0030, ZongYuan Ge, Haoran Xie 0002, Lie Ju
Pattern Recognit.2
2026 Expert-Guided Cross-View Fusion With Self-Derived Lesion Proposals for Multi-View Diabetic Retinopathy Grading
abstract
Recent advances in multi-view fundus imaging show great promise for automated diabetic retinopathy (DR) grading. However, mainstream end-to-end CNN/Transformer pipelines rely on striding or tokenization that compresses spatial detail, causing small, low-contrast lesions (e.g., microaneurysms) to be under-represented and creating performance ceilings. Prior efforts have mitigated this by incorporating external lesion- or vessel-level annotations into models. However, such labels are costly to acquire, break the end-to-end training, and make performance over-reliant on the annotation quality. To reduce dependence on expensive annotations, we propose an end-to-end framework that generates lesion proposals on the fly during training and inference, providing self-derived cues for grading. First, we introduce a Grade-Activated Lesion Proposal (GALP) module that derives grade-conditioned evidence maps (GEMs) from stage-wise auxiliary classifiers and selects the top-K high-evidence regions per view as lesion proposals. Second, we propose a Cross-View Lesion Expert Guided Regional Fusion (LGRF) module, which selectively activates experts for a view's lesion proposals based on contextual guidance from other views, ensuring that only the most relevant feature extractors contribute to fusion. Experimental results on two multi-view DR datasets show that our method matches or surpasses strong baselines without external annotations, confirming that self-generated proposals can substantially reduce annotation needs.
Wai Keung Wong, Xueling Zhou, Junlin Hou, Jie Wen 0001
IEEE Trans. Image Process.4
2026 Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging Frontiers
abstract
Over the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows.
Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis
IEEE J. Biomed. Health Informatics5
2025 EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
abstract
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate future frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop a fully automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the public Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
Jilan Xu, Yifei Huang 0002, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
ICLR4
2025 Explain Any Pathological Concept: Discovering Hierarchical Explanations for Pathology Foundation Models
Shuting Xu, Junlin Hou, Hao Chen 0011
MICCAI (6)2
2025 QMix: Quality-Aware Learning With Mixed Noise for Robust Retinal Disease Diagnosis
abstract
Due to the complex nature of medical image acquisition and annotation, medical datasets inevitably contain noise. This adversely affects the robustness and generalization of deep neural networks. Previous noise learning methods mainly considered noise arising from images being mislabeled, i.e., label noise, assuming all mislabeled images were of high quality. However, medical images can also suffer from severe data quality issues, i.e., data noise, where discriminative visual features for disease diagnosis are missing. In this paper, we propose QMix, a noise learning framework that learns a robust disease diagnosis model under mixed noise scenarios. QMix alternates between sample separation and quality-aware semi-supervised training in each epoch. The sample separation phase uses a joint uncertainty-loss criterion to effectively separate (1) correctly labeled images, (2) mislabeled high-quality images, and (3) mislabeled low-quality images. The semi-supervised training phase then learns a robust disease diagnosis model from the separated samples. Specifically, we propose a sample-reweighing loss to mitigate the effect of mislabeled low-quality images during training, and a contrastive enhancement loss to further distinguish them from correctly labeled images. QMix achieved state-of-the-art performance on six public retinal image datasets and exhibited significant improvements in robustness against mixed noise. Code will be available upon acceptance.
Junlin Hou, Jilan Xu, Rui Feng 0001, Hao Chen 0011
IEEE Trans. Medical Imaging1
2024 Retrieval-Augmented Egocentric Video Captioning
abstract
Understanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos, (2) for training the cross-view retrieval module, we devise an au-tomatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets, (3) we train the cross-view retrieval module with a novel EgoEx-oNCE loss that pulls egocentric and exocentric video features closer, by aligning them to shared text features that describe similar actions, (4) through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocen-tric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references.
Jilan Xu, Yifei Huang 0002, Junlin Hou, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
CVPR3
2024 Cross-Image Distillation for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation approaches have drawn much more attention in recent years, which aim to exploit a large amount of unlabeled data together with a small number of labeled data. However, existing models usually regarded segmentation as pixel-wise classification, neglecting global semantic relations among pixels across various images. Moreover, scarce annotated data usually exhibits a biased distribution against the desired one, hindering performance improvement. To address these challenging problems, we propose a novel cross-image distillation framework for semi-supervised semantic segmentation. Specifically, we introduce a relation distillation module to model inter-channel correlations between features of labeled samples and unlabeled samples. In addition, we propose a style distillation strategy to explicitly calibrate the learned feature distributions of labeled and unlabeled data to be aligned. Experimental results on two popular benchmarks demonstrate that our proposed approach achieves superior performance over other state-of-the-art methods. We will release the code soon.
Junlin Hou, Rui Feng 0001
ICASSP3
2024 Concept-Attention Whitening for Interpretable Skin Lesion Diagnosis
Junlin Hou, Jilan Xu
MICCAI (10)1
2023 Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision
abstract
This paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled imagetext pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slotattention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Yi Wang 0074, Yu Qiao 0001, Weidi Xie
CVPR2
2023 Diabetic Retinopathy Grading with Weakly-Supervised Lesion Priors
abstract
Explicit information of lesions can provide visual instructions for diabetic retinopathy (DR) grading on fundus images. However, pixel-level lesion annotations are extremely difficult and time-consuming to acquire. In this work, we propose a novel weakly-supervised lesion-aware network for DR grading, which enhances the discriminative features with lesion priors by only image-level supervision. Specifically, we design a lesion attention module that generates lesion activation maps by introducing an auxiliary task of binary DR identification. Lesion activation maps are utilized to assist the network to focus on the most relevant regions for boosting DR grading performance. Besides, we particularly devise an adaptive joint loss to balance the DR identification and DR grading tasks dynamically. Extensive results on the public DR dataset demonstrate the superiority and generality of our proposed lesion-aware network. The interpretability of generated lesion activation maps is also verified by the comparison with ground truth segmentation masks.
Junlin Hou, Jilan Xu, Rui Feng 0001, Yue Zhang 0004, Haidong Zou, Lina Lu, Wenwen Xue
ICASSP1
2022 Cross-Field Transformer for Diabetic Retinopathy Grading on Two-field Fundus Images
abstract
Automatic diabetic retinopathy (DR) grading based on fundus photography has been widely explored to benefit the routine screening and early treatment. Existing researches generally focus on single-field fundus images, which have limited field of view for precise eye examinations. In clinical applications, ophthalmologists adopt two-field fundus photography as the dominating tool, where the information from each field (i.e., macula-centric and optic disc-centric) is highly correlated and complementary, and benefits comprehensive decisions. However, automatic DR grading based on two-field fundus photography remains a challenging task due to the lack of publicly available datasets and effective fusion strategies. In this work, we first construct a new benchmark dataset (DRTiD) for DR grading, consisting of 3,100 two-field fundus images. To the best of our knowledge, it is the largest public DR dataset with diverse and high-quality two-field images. Then, we propose a novel DR grading approach, namely Cross-Field Transformer (CrossFiT), to capture the correspondence between two fields as well as the long-range spatial correlations within each field. Considering the inherent two-field geometric constraints, we particularly define aligned position embeddings to preserve relative consistent position in fundus. Besides, we perform masked cross-field attention during interaction to filter the noisy relations between fields. Extensive experiments on our DRTiD dataset and a public DeepDRiD dataset demonstrate the effectiveness of our CrossFiT network. The new dataset and the source code of CrossFiT will be publicly available at https://github.com/DU-VTS/DRTiD.
Junlin Hou, Jilan Xu, Yuejie Zhang, Haidong Zou, Lina Lu, Wenwen Xue, Rui Feng 0001
BIBM1
2022 CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping
abstract
Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
CVPR2
2021 Semi-supervised Medical Image Segmentation with Distribution Calibration and Non-local Semantic Constraint
abstract
The available medical images with accurate segmentation masks are usually limited due to the expensive and time-consuming annotation cost. Many semi-supervised approaches tried to exploit a large amount of unlabeled data together with the small number of labeled data. However, their learned segmentation models can easily become overfitted on the biased labeled data, mainly because of the misalignment between the labeled and unlabeled data distribution. To address this challenging problem, we propose a novel semi-supervised model with distribution calibration and non-local semantic constraint for medical image segmentation. In specific, we explicitly calibrate the learned feature distributions of the labeled and unlabeled data to make them aligned. Meanwhile, we add a special nonlocal semantic loss to encourage the learned features to be more discriminative for the segmentation task at the same time. Consequently, our final segmentation networks have the advantage to better generalize on the unlabeled data in both training and test set. Experimental results on three popular medical image segmentation benchmarks demonstrate that our proposed model achieves superior performance over other state-of-the-art methods. We will release our source code in this URL.
Junlin Hou, Rui Feng 0001, Yuejie Zhang
BIBM2
2021 Exploiting Deep Cross-Slice Features From CT Images For Multi-Class Pneumonia Classification
abstract
Computed Tomography (CT) scanning is widely used for chest diseases detection including pneumonia due to its diagnostic efficacy and efficiency. Recent studies have shown that the distribution of infection regions caused by COVID19 in CT images is different from other pneumonia, and the COVID-19 cases are more likely to suffer severe and large-area infections. However, most deep learning methods only focus on intra-slice features and ignore cross-slice features. In this paper, we propose a novel two-stage method to fully exploit deep cross-slice features from volumetric CT data, including a dual-task supervised CNN and a context-aware Bi-LSTM. To further demonstrate the effectiveness of our model, we conduct extensive experiments on a chest CT imaging dataset with a total of 801 patients (250 healthy people, 238 COVID-19 patients, 191 H1N1 patients, and 122 CAP patients). The experimental results indicate the superiority of our proposed model on the multi-class pneumonia classification task.
Jiawang Cao, Junlin Hou, Longquan Jiang 0003, Weiya Shi, Rui Feng 0001
ICIP3
2021 Periphery-aware COVID-19 diagnosis with contrastive representation enhancement
Junlin Hou, Jilan Xu, Longquan Jiang 0003, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Xiangyang Xue 0001
Pattern Recognit.1
2020 Data-Efficient Histopathology Image Analysis with Deformation Representation Learning
abstract
Histopathological examination of tissue biopsies plays a fundamental role in disease assessment. Automatic histopathology image analysis requires substantial task-specific annotations, which are often expensive and laborious in realworld scenarios. This insufficient annotation of data limits the generalization ability of supervised learning models. To address this challenge, we propose a self-supervised Deformation Representation Learning (DRL) framework to learn semantic features from unlabeled data. As a novel paradigm, our approach utilizes deformation as supervisory signals based on two critical features, i.e., local structure heterogeneity and global context homogeneity. Given an original histopathology image and its deformed counterpart, there exists a moderate difference in local structures. In contrast, due to the transformation-invariance, both images share a similar global context compared with other images. Specifically, an encoder network is trained to distinguish the local inconsistency by measuring the mutual information and maintain the global consistency with noise contrastive estimation. Extensive experiments on public histopathology image datasets show that the learned representations are generalizable for various downstream tasks, such as transfer learning on segmentation and semi-supervised classification. Our approach achieves superior results over other self-supervised methods and the ImageNet pre-trained model, and it reveals the ability as a novel pre-training scheme in histopathology image analysis.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Chunyang Ruan, Tao Zhang 0022, Weiguo Fan
BIBM2
2020 Fuzzy Graph Neural Network for Few-Shot Learning
abstract
Recent works have shown that graph neural net-works (GNNs) can substantially improve the performance of few-shot learning benefitting from their natural ability to learn inter-class uniqueness and intra-class commonality. However, previous GNN methods have not achieved satisfactory performance due to the absence of a strong relational inductive bias which determines how entities interact and are isolated. In this paper, inspired by the fuzzy theory, we propose a novel meta-learning method called Fuzzy GNN (FGNN), which obtains superior relational inductive biases in each episode, for few-shot learning. Specifically, we employ an edge-focused GNN to perform the edge prediction by iteratively updating the edge-labels. According to the output of edge prediction, we design a fuzzy membership function to achieve more exact relationship representations for node classification. The parameters of the FGNN are learned by episodic training with mixed loss including node-label and edge-label. Extensive experimental evaluation clearly demonstrates the effectiveness of FGNN. The results show that our method achieves state-of-the-art performance and a significant improvement over other GNN methods on two few-shot learning benchmarks.
Junlin Hou, Rui Feng 0001
IJCNN2