VLDB 2026 Research / reviewers in the wild / expert
Sheng Wang 0014
dblp:85/1868-14
· DBLP profile ↗
32ranked-venue papers
3as first author
32since 2021 · last 2026
0000-0002-5472-6929ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdLER: Adversarial training with label error rectification for one-shot medical image segmentation
Xiangyu Zhao 0003, Sheng Wang 0014, Zhiyun Song, Zhenrong Shen 0001, Linlin Yao, Haolei Yuan, Qian Wang 0001, Lichi Zhang |
Expert Syst. Appl. | 2 |
| 2026 | Enhancing Knee Disease Diagnosis via Multi-View Graph Representation With Multi-Task Pre-TrainingabstractMagnetic resonance imaging (MRI) is an indispensable tool for clinical knee examination, which often scans 2D stacked slices from multiple views. Radiologists typically locate lesion regions in one view, and then refer to other views to formulate a comprehensive diagnosis. However, existing computer-aided diagnosis methods fall short of identifying and fusing local regions in multi-view scans, leading to a decline in diagnostic performance and a heavy reliance on extensively annotated data. This paper introduces a novel framework that represents multi-view MRI scans as a knee graph, and conducts diagnosis using the proposed Knee Graph Network (KGNet). Moreover, KGNet is greatly enhanced by multi-task pre-training, which requires KGNet to reconstruct masked knee local patches and segment unmasked ones working alongside corresponding decoders. Experimental evaluations on public and in-house clinical datasets confirm that our framework outperforms existing approaches in diagnosing cartilage defects, anterior cruciate ligament tears, and knee abnormalities. In conclusion, our framework demonstrates the potential of enhancing knee disease diagnosis by representing multi-view MRI scans as a graph and employing multi-task pre-training in the graph network. The code is publicly available at https://github.com/zixuzhuang/KGNet. Zixu Zhuang, Dongdong Chen 0003, Sheng Wang 0014, Kai Xuan, Xiangyu Zhao 0003, Zhong Xue, Dinggang Shen, Lichi Zhang, Weiwu Yao, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body ReconstructionabstractMultiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenges posed by self-occlusion under multiple views and the continuity of human body shape estimation. In this study, we introduce a method to reconstruct the 3D human body from multiple uncalibrated camera views. Initially, we utilize a pre-trained human body encoder to process each camera view individually, enabling the reconstruction of human body models and parameters for each view along with predicted camera positions. Rather than merely averaging the models across views, we develop a neural network trained to assign weights to individual views for all human body joints, based on the estimated distribution of joint distances from each camera. Additionally, we focus on the mesh surface of the human body for dynamic fusion, allowing for the seamless integration of facial expressions and body shape into a unified human body model. Our method has shown excellent performance in reconstructing the human body on two public datasets, advancing beyond previous work from the SMPL model to the SMPL-X model. This extension incorporates more complex hand poses and facial expressions, enhancing the detail and accuracy of the reconstructions. Crucially, it supports the flexible ad-hoc deployment of any number of cameras, offering significant potential for various applications. Yitao Zhu, Sheng Wang 0014, Mengjie Xu, Zixu Zhuang, Zhixin Wang, Kaidong Wang, Han Zhang 0002, Qian Wang 0001 |
AAAI | 2 |
| 2025 | MITracker: Multi-View Integration for Visual Object TrackingabstractMulti-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progress has been limited by the lack of comprehensive multi-view datasets and effective cross-view integration methods. To overcome these limitations, we compiled a Multi-View object Tracking (MVTrack) dataset of 234K high-quality annotated frames featuring 27 distinct objects across various scenes. In conjunction with this dataset, we introduce a novel MVOT method, Multi-View Integration Tracker (MITracker), to efficiently integrate multi-view object features and provide stable tracking outcomes. MI-Tracker can track any object in video frames of arbitrary length from arbitrary viewpoints. The key advancements of our method over traditional single-view approaches come from two aspects: (1) MITracker transforms 2D image features into a 3D feature volume and compresses it into a bird’s eye view (BEV) plane, facilitating inter-view information fusion; (2) we propose an attention mechanism that leverages geometric information from fused 3D feature volume to refine the tracking results at each view. MI-Tracker outperforms existing methods on the MVTrack and GMTD datasets, achieving state-of-the-art performance. The code and the new dataset will be available at mii-laboratory.github.io/MITracker. 1 Mengjie Xu, Yitao Zhu, Jiaming Li 0012, Zhenrong Shen 0001, Sheng Wang 0014, Haolin Huang, Han Zhang 0002, Qian Wang 0001 |
CVPR | 6 |
| 2025 | MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language ModelsabstractArtificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities for interactive diagnostic tools. However, these models often suffer from factual hallucination, which can lead to incorrect diagnoses. Fine-tuning and retrieval-augmented generation (RAG) have emerged as methods to address these issues. However, the amount of high-quality data and distribution shifts between training data and deployment data limit the application of fine-tuning methods. Although RAG is lightweight and effective, existing RAG-based approaches are not sufficiently general to different medical domains and can potentially cause misalignment issues, both between modalities and between the model and the ground truth. In this paper, we propose a versatile multimodal RAG system, MMed-RAG, designed to enhance the factuality of Med-LVLMs. Our approach introduces a domain-aware retrieval mechanism, an adaptive retrieved contexts selection, and a provable RAG-based preference fine-tuning strategy. These innovations make the RAG process sufficiently general and reliable, significantly improving alignment when introducing retrieved contexts. Experimental results across five medical datasets (involving radiology, ophthalmology, pathology) on medical VQA and report generation demonstrate that MMed-RAG can achieve an average improvement of 43.8% in factual accuracy in the factual accuracy of Med-LVLMs. Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Tianze Wang, Sheng Wang 0014, Linjun Zhang, James Zou 0001, Huaxiu Yao |
ICLR | 6 |
| 2025 | MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference OptimizationabstractThe advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the models prioritize textual knowledge over visual input, leading to hallucinations that contradict information in medical images. Previous attempts to enhance modality alignment in Med-LVLMs through preference optimization have inadequately addressed clinical relevance in preference data, making these samples easily distinguishable and reducing alignment effectiveness. In response, we propose MMedPO, a novel multimodal medical preference optimization approach that considers the clinical relevance of preference samples to enhance Med-LVLM alignment. MMedPO curates multimodal preference data by introducing two types of dispreference: (1) plausible hallucinations injected through target Med-LVLMs or GPT-4o to produce medically inaccurate responses, and (2) lesion region neglect achieved through local lesion-noising, disrupting visual understanding of critical areas. We then calculate clinical relevance for each sample based on scores from multiple Med-LLMs and visual tools, enabling effective alignment. Our experiments demonstrate that MMedPO significantly enhances factual accuracy in Med-LVLMs, achieving substantial improvements over existing preference optimization methods by 14.2% and 51.7% on the Med-VQA and report generation tasks, respectively. Our code are available in https://github.com/aiming-lab/MMedPO}{https://github.com/aiming-lab/MMedPO. Kangyu Zhu, Peng Xia 0005, Yun Li 0010, Hongtu Zhu, Sheng Wang 0014, Huaxiu Yao |
ICML | 5 |
| 2025 | Query-Level Alignment for End-to-End Lesion Detection with Human Gaze
Yan Kong, Zhixiang Peng, Yonghao Li, Jiangdong Cai, Sheng Wang 0014, Qian Wang 0001, Yuqi Fang, Caifeng Shan |
MICCAI (13) | 6 |
| 2025 | Med-LEGO: Editing and Adapting Toward Generalist Medical Image Diagnosis
Yitao Zhu, Jiaming Li 0012, Mengjie Xu, Zihao Zhao 0002, Honglin Xiong, Sheng Wang 0014, Qian Wang 0001 |
MICCAI (6) | 7 |
| 2025 | Uni-COAL: A unified framework for cross-modality synthesis and super-resolution of MR images
Zhiyun Song, Zengxin Qi, Xin Wang 0125, Xiangyu Zhao 0003, Zhenrong Shen 0001, Sheng Wang 0014, Manman Fei, Di Zang, Dongdong Chen 0003, Linlin Yao, Mengjun Liu, Qian Wang 0001, Xuehai Wu, Lichi Zhang |
Expert Syst. Appl. | 6 |
| 2025 | Guiding fusion of dynamic functional and effective connectivity in spatio-temporal graph neural network for brain disorder classification
Dongdong Chen 0003, Mengjun Liu, Sheng Wang 0014, Zheren Li, Lu Bai 0001, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Knowl. Based Syst. | 3 |
| 2025 | ReactDiff: Latent Diffusion for Facial Reaction Generation
Jiaming Li 0012, Sheng Wang 0014, Yitao Zhu, Honglin Xiong, Zixu Zhuang, Qian Wang 0001 |
Neural Networks | 2 |
| 2025 | Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 1 |
| 2025 | Improving Self-Supervised Medical Image Pre-Training by Early Alignment With Human Eye Gaze InformationabstractAlignment between human knowledge and machine learning models is crucial for achieving efficient and interpretable AI systems. However, conventional self-supervised pre-training methods often suffer from low efficiency, as they do not incorporate human knowledge during the pre-training process and instead rely mainly on post-hoc alignment techniques. We propose Gaze Pre-Training (GzPT), a novel approach that introduces early alignment with human eye gaze information during the pre-training process to enhance both the learning efficiency and performance of self-supervised models. By leveraging contrastive learning to pull together images with similar gaze patterns, GzPT can effectively align the model with human attention during the pre-training. We demonstrate the effectiveness of our approach on three diverse medical image datasets, showing that GzPT can consistently outperform baseline methods and learn more meaningful and interpretable representations. Our findings also highlight the potential of incorporating human eye gaze as a form of passive knowledge to bridge the gap between human and machine learning in the self-supervised pre-training. Our code is available at Github. Sheng Wang 0014, Zihao Zhao 0002, Zhenrong Shen 0001, Bin Wang 0068, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Mining Gaze for Contrastive Learning toward Computer-Assisted DiagnosisabstractObtaining large-scale radiology reports can be difficult for medical images due to ethical concerns, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In this paper, we propose eye-tracking as an alternative to text reports, as it allows for the passive collection of gaze signals without ethical issues. By tracking the gaze of radiologists as they read and diagnose medical images, we can understand their visual attention and clinical reasoning. When a radiologist has similar gazes for two medical images, it may indicate semantic similarity for diagnosis, and these images should be treated as positive pairs when pre-training a computer-assisted diagnosis (CAD) network through contrastive learning. Accordingly, we introduce the Medical contrastive Gaze Image Pre-training (McGIP) as a plug-and-play module for contrastive learning frameworks. McGIP uses radiologist gaze to guide contrastive pre-training. We evaluate our method using two representative types of medical images and two common types of gaze data. The experimental results demonstrate the practicality of McGIP, indicating its high potential for various clinical scenarios and applications. Zihao Zhao 0002, Sheng Wang 0014, Qian Wang 0001, Dinggang Shen |
AAAI | 2 |
| 2024 | Gaze-DETR: Using Expert Gaze to Reduce False Positives in Vulvovaginal Candidiasis Screening
Yan Kong, Sheng Wang 0014, Jiangdong Cai, Zihao Zhao 0002, Zhenrong Shen 0001, Yonghao Li, Manman Fei, Qian Wang 0001 |
MICCAI (4) | 2 |
| 2024 | Spatial attention-based implicit neural representation for arbitrary reduction of MRI slice spacing
Xin Wang 0125, Sheng Wang 0014, Honglin Xiong, Kai Xuan, Zixu Zhuang, Mengjun Liu, Zhenrong Shen 0001, Xiangyu Zhao 0003, Lichi Zhang, Qian Wang 0001 |
Medical Image Anal. | 2 |
| 2024 | RCPS: Rectified Contrastive Pseudo Supervision for Semi-Supervised Medical Image SegmentationabstractMedical image segmentation methods are generally designed as fully-supervised to guarantee model performance, which requires a significant amount of expert annotated samples that are high-cost and laborious. Semi-supervised image segmentation can alleviate the problem by utilizing a large number of unlabeled images along with limited labeled images. However, learning a robust representation from numerous unlabeled images remains challenging due to potential noise in pseudo labels and insufficient class separability in feature space, which undermines the performance of current semi-supervised segmentation approaches. To address the issues above, we propose a novel semi-supervised segmentation method named as Rectified Contrastive Pseudo Supervision (RCPS), which combines a rectified pseudo supervision and voxel-level contrastive learning to improve the effectiveness of semi-supervised segmentation. Particularly, we design a novel rectification strategy for the pseudo supervision method based on uncertainty estimation and consistency regularization to reduce the noise influence in pseudo labels. Furthermore, we introduce a bidirectional voxel contrastive loss in the network to ensure intra-class consistency and inter-class contrast in feature space, which increases class separability in the segmentation. The proposed RCPS segmentation method has been validated on two public datasets and an in-house clinical dataset. Experimental results reveal that the proposed method yields better segmentation performance compared with the state-of-the-art methods in semi-supervised medical image segmentation. The source code is available at https://github.com/hsiangyuzhao/RCPS. Xiangyu Zhao 0003, Zengxin Qi, Sheng Wang 0014, Qian Wang 0001, Xuehai Wu, Ying Mao 0002, Lichi Zhang |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | ChatCAD+: Toward a Universal and Reliable Interactive CAD Using LLMsabstractThe integration of Computer-Aided Diagnosis (CAD) with Large Language Models (LLMs) presents a promising frontier in clinical applications, notably in automating diagnostic processes akin to those performed by radiologists and providing consultations similar to a virtual family doctor. Despite the promising potential of this integration, current works face at least two limitations: (1) From the perspective of a radiologist, existing studies typically have a restricted scope of applicable imaging domains, failing to meet the diagnostic needs of different patients. Also, the insufficient diagnostic capability of LLMs further undermine the quality and reliability of the generated medical reports. (2) Current LLMs lack the requisite depth in medical expertise, rendering them less effective as virtual family doctors due to the potential unreliability of the advice provided during patient consultations. To address these limitations, we introduce ChatCAD+, to be universal and reliable. Specifically, it is featured by two main modules: (1) Reliable Report Generation and (2) Reliable Interaction. The Reliable Report Generation module is capable of interpreting medical images from diverse domains and generate high-quality medical reports via our proposed hierarchical in-context learning. Concurrently, the interaction module leverages up-to-date information from reputable medical websites to provide reliable medical advice. Together, these designed modules synergize to closely align with the expertise of human medical professionals, offering enhanced consistency and reliability for interpretation and advice. The source code is available at GitHub. Zihao Zhao 0002, Sheng Wang 0014, Jinchen Gu, Yitao Zhu, Lanzhuju Mei, Zixu Zhuang, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen |
MICCAI (7) | 4 |
| 2023 | CellGAN: Conditional Cervical Cell Synthesis for Augmenting Cytopathological Image Classification
Zhenrong Shen 0001, Maosong Cao, Sheng Wang 0014, Lichi Zhang, Qian Wang 0001 |
MICCAI (6) | 3 |
| 2023 | Alias-Free Co-modulated Network for Cross-Modality Synthesis and Super-Resolution of MR Images
Zhiyun Song, Xin Wang 0125, Xiangyu Zhao 0003, Sheng Wang 0014, Zhenrong Shen 0001, Zixu Zhuang, Mengjun Liu, Qian Wang 0001, Lichi Zhang |
MICCAI (10) | 4 |
| 2023 | One-Shot Traumatic Brain Segmentation with Adversarial Training and Uncertainty Rectification
Xiangyu Zhao 0003, Zhenrong Shen 0001, Dongdong Chen 0003, Sheng Wang 0014, Zixu Zhuang, Qian Wang 0001, Lichi Zhang |
MICCAI (4) | 4 |
| 2023 | HENet: Hierarchical Enhancement Network for Pulmonary Vessel Segmentation in Non-contrast CT Images
Xiao Zhang 0028, Dongdong Gu, Sheng Wang 0014, Jiayu Huo, Zhihao Jiang 0001, Feng Shi 0001, Zhong Xue, Yiqiang Zhan, Xi Ouyang, Dinggang Shen |
MICCAI (3) | 4 |
| 2023 | CAS-Net: Cross-View Aligned Segmentation by Graph Representation of Knees
Zixu Zhuang, Xin Wang 0125, Sheng Wang 0014, Zhenrong Shen 0001, Xiangyu Zhao 0003, Mengjun Liu, Zhong Xue, Dinggang Shen, Lichi Zhang, Qian Wang 0001 |
MICCAI (4) | 3 |
| 2023 | Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut LearningabstractLearning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | TaG-Net: Topology-Aware Graph Network for Centerline-Based Vessel LabelingabstractAnatomical labeling of head and neck vessels is a vital step for cerebrovascular disease diagnosis. However, it remains challenging to automatically and accurately label vessels in computed tomography angiography (CTA) since head and neck vessels are tortuous, branched, and often spatially close to nearby vasculature. To address these challenges, we propose a novel topology-aware graph network (TaG-Net) for vessel labeling. It combines the advantages of volumetric image segmentation in the voxel space and centerline labeling in the line space, wherein the voxel space provides detailed local appearance information, and line space offers high-level anatomical and topological information of vessels through the vascular graph constructed from centerlines. First, we extract centerlines from the initial vessel segmentation and construct a vascular graph from them. Then, we conduct vascular graph labeling using TaG-Net, in which techniques of topology-preserving sampling, topology-aware feature grouping, and multi-scale vascular graph are designed. After that, the labeled vascular graph is utilized to improve volumetric segmentation via vessel completion. Finally, the head and neck vessels of 18 segments are labeled by assigning centerline labels to the refined segmentation. We have conducted experiments on CTA images of 401 subjects, and experimental results show superior vessel segmentation and labeling of our method compared to other state-of-the-art methods. Linlin Yao, Feng Shi 0001, Sheng Wang 0014, Xiao Zhang 0028, Zhong Xue, Xiaohuan Cao, Yiqiang Zhan, Lizhou Chen, Yuntian Chen, Bin Song 0002, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Knee Cartilage Defect Assessment by Graph Representation and Surface ConvolutionabstractKnee osteoarthritis (OA) is the most common osteoarthritis and a leading cause of disability. Cartilage defects are regarded as major manifestations of knee OA, which are visible by magnetic resonance imaging (MRI). Thus early detection and assessment for knee cartilage defects are important for protecting patients from knee OA. In this way, many attempts have been made on knee cartilage defect assessment by applying convolutional neural networks (CNNs) to knee MRI. However, the physiologic characteristics of the cartilage may hinder such efforts: the cartilage is a thin curved layer, implying that only a small portion of voxels in knee MRI can contribute to the cartilage defect assessment; heterogeneous scanning protocols further challenge the feasibility of the CNNs in clinical practice; the CNN-based knee cartilage evaluation results lack interpretability. To address these challenges, we model the cartilages structure and appearance from knee MRI into a graph representation, which is capable of handling highly diverse clinical data. Then, guided by the cartilage graph representation, we design a non-Euclidean deep learning network with the self-attention mechanism, to extract cartilage features in the local and global, and to derive the final assessment with a visualized result. Our comprehensive experiments show that the proposed method yields superior performance in knee cartilage defect assessment, plus its convenient 3D visualization for interpretability. Zixu Zhuang, Liping Si, Sheng Wang 0014, Kai Xuan, Xi Ouyang, Yiqiang Zhan, Zhong Xue, Lichi Zhang, Dinggang Shen, Weiwu Yao, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Whole Slide Cervical Cancer Screening Using Graph Attention Network and Supervised Contrastive Learning
Xin Zhang 0013, Maosong Cao, Sheng Wang 0014, Jiayin Sun, Xiangshan Fan, Qian Wang 0001, Lichi Zhang |
MICCAI (2) | 3 |
| 2022 | Local Graph Fusion of Multi-view MR Images for Knee Osteoarthritis Diagnosis
Zixu Zhuang, Sheng Wang 0014, Liping Si, Kai Xuan, Zhong Xue, Dinggang Shen, Lichi Zhang, Weiwu Yao, Qian Wang 0001 |
MICCAI (3) | 2 |
| 2022 | Automatic Grading Assessments for Knee MRI Cartilage Defects via Self-ensembling Semi-supervised Learning with Dual-Consistency
Jiayu Huo, Xi Ouyang, Liping Si, Kai Xuan, Sheng Wang 0014, Weiwu Yao, Dahong Qian, Zhong Xue, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Medical Image Anal. | 5 |
| 2022 | Follow My Eye: Using Gaze to Supervise Computer-Aided DiagnosisabstractWhen deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a properly functioning DNN. This demand for supervision data and labels is a major bottleneck in current medical image analysis, since collecting a large number of annotations from experienced experts can be time-consuming and expensive. In this paper, we demonstrate that the eye movement of radiologists reading medical images can be a new form of supervision to train the DNN-based computer-aided diagnosis (CAD) system. Particularly, we record the tracks of the radiologists' gaze when they are reading images. The gaze information is processed and then used to supervise the DNN's attention via an Attention Consistency module. To the best of our knowledge, the above pipeline is among the earliest efforts to leverage expert eye movement for deep-learning-based CAD. We have conducted extensive experiments on knee X-ray images for osteoarthritis assessment. The results show that our method can achieve considerable improvement in diagnosis performance, with the help of gaze supervision. Sheng Wang 0014, Xi Ouyang, Tianming Liu 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Domain Generalization for Mammography Detection via Multi-style and Multi-view Contrastive Learning
Zheren Li, Zhiming Cui 0001, Sheng Wang 0014, Yuji Qi, Xi Ouyang, Qitian Chen, Yuezhi Yang, Zhong Xue, Dinggang Shen, Jie-Zhi Cheng |
MICCAI (7) | 3 |