EDBT 2026 Demo / reviewers in the wild / expert
Xiaosong Wang 0001
dblp:34/5737-1
· DBLP profile ↗
35ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-3840-5658ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 7 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACDA: Anatomically constrained distribution alignment for robust medical image segmentation
Yifan Gao 0004, Xiaosong Wang 0001, Xin Gao 0003 |
Knowl. Based Syst. | 2 |
| 2026 | SCALAR: Spatial-concept alignment for robust vision in harsh open world
Xiaoyu Yang 0007, Lijian Xu, Xingyu Zeng, Xiaosong Wang 0001, Hongsheng Li 0001, Shaoting Zhang 0001 |
Pattern Recognit. | 4 |
| 2025 | Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion ModelsabstractWe explore Generalizable Tumor Segmentation, aiming to train a single model for zero-shot tumor segmentation across diverse anatomical regions. Existing methods face limitations related to segmentation quality, scalability, and the range of applicable imaging modalities. In this paper, we uncover the potential of the internal representations within frozen medical foundation diffusion models as highly efficient zero-shot learners for tumor segmentation by introducing a novel framework named DiffuGTS. DiffuGTS creates anomaly-aware open-vocabulary attention maps based on text prompts to enable generalizable anomaly segmentation without being restricted by a predefined training category list. To further improve and refine anomaly segmentation masks, DiffuGTS leverages the diffusion model, transforming pathological regions into high-quality pseudo-healthy counterparts through latent space inpainting, and applies a novel pixel-level and feature-level residual learning approach, resulting in segmentation masks with significantly enhanced quality and generalization. Comprehensive experiments on four datasets and seven tumor categories demonstrate the superior performance of our method, surpassing current state-of-the-art models across multiple zero-shot settings. Codes are available at https://github.com/Yankai96/DiffuGTS. Yankai Jiang 0003, Donglin Yang, Yuan Tian 0017, Hai Lin 0003, Xiaosong Wang 0001 |
CVPR | 6 |
| 2025 | Multi-modal Vision Pre-training for Medical Image AnalysisabstractSelf-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby neglecting the inter-modal correlations essential for effective learning of cross-modal image representations. This limitation is particularly significant for naturally grouped multi-modal data, e.g., multi-parametric MRI scans for a patient undergoing various functional imaging protocols in the same study. To bridge this gap, we conduct a novel multi-modal image pre-training with three proxy tasks to facilitate the learning of cross-modality representations and correlations using multi-modal brain MRI scans (over 2.4 million images in 16,022 scans of 3,755 patients), i.e., cross-modal image reconstruction, modality-aware contrastive learning, and modality template distillation. To demonstrate the generalizability of our pre-trained model, we conduct extensive experiments on various benchmarks with ten downstream tasks. The superior performance of our method is reported in comparison to state-of-the-art pre-training methods, with Dice Score improvement of 0.28%-14.47% across six segmentation benchmarks and a consistent accuracy boost of 0.65%-18.07% in four individual image classification tasks. Shaohao Rui, Lingzhi Chen, Zhenyu Tang 0005, Lilong Wang, Mianxin Liu, Shaoting Zhang 0001, Xiaosong Wang 0001 |
CVPR | 7 |
| 2025 | Towards All-in-One Medical Image Re-IdentificationabstractMedical image re-identification (MedReID) is underexplored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem. First, to handle various medical modalities, we propose a novel Continuous Modality-based Parameter Adapter (ComPA). ComPA condenses medical content into a continuous modality representation and dynamically adjusts the modality-agnostic model with modalityspecific parameters at runtime. This allows a single model to adaptively learn and process diverse modality data. Furthermore, we integrate medical priors into our model by aligning it with a bag of pre-trained medical foundation models, in terms of the differential features. Compared to single-image feature, modeling the inter-image difference better fits the re-identification problem, which involves discriminating multiple images. We evaluate the proposed model against 25 foundation models and 8 large multimodal language models across 11 image datasets, demonstrating consistently superior performance. Additionally, we deploy the proposed MedReID technique to two realworld applications, i.e., history-augmented personalized diagnosis and medical privacy protection. Codes and model is available at https://github.com/tianyuan168326/All-inOne-MedReID-Pytorch. Yuan Tian 0017, Kaiyuan Ji, Rongzhao Zhang, Yankai Jiang 0003, Chunyi Li 0001, Xiaosong Wang 0001, Guangtao Zhai |
CVPR | 6 |
| 2025 | Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-Identification
Yuan Tian 0017, Rongzhao Zhang, Zijian Chen 0001, Yankai Jiang 0003, Chunyi Li 0001, Fang Yan 0002, Qiang Hu 0003, Xiaosong Wang 0001, Guangtao Zhai |
ICCV | 10 |
| 2025 | A Composite Alignment-Aware Framework for Myocardial Lesion Segmentation in Multi-sequence CMR Images
Yifan Gao 0004, Shaohao Rui, Haoyang Su 0001, Jinyi Xiang, Lian-Ming Wu, Xiaosong Wang 0001 |
MICCAI (1) | 6 |
| 2025 | CTSL: Codebook-Based Temporal-Spatial Learning for Accurate Non-contrast Cardiac Risk Prediction Using Cine MRIs
Haoyang Su 0001, Shaohao Rui, Jinyi Xiang, Lian-Ming Wu, Xiaosong Wang 0001 |
MICCAI (15) | 5 |
| 2025 | Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative SearchabstractMultimodal large language models (MLLMs) have begun to demonstrate robust reasoning capabilities on general tasks, yet their application in the medical domain remains in its early stages. Constructing chain-of-thought (CoT) training data is essential for bolstering the reasoning abilities of medical MLLMs. However, existing approaches exhibit a deficiency in offering a comprehensive framework for searching and evaluating effective reasoning paths towards critical diagnosis. To address this challenge, we propose Mentor-Intern Collaborative Search (MICS), a novel reasoning-path searching scheme to generate rigorous and effective medical CoT data. MICS first leverages mentor models to initialize the reasoning, one step at a time, then prompts each intern model to continue the thinking along those initiated paths, and finally selects the optimal reasoning path according to the overall reasoning performance of multiple intern models. The reasoning performance is determined by an MICS-Score, which assesses the quality of generated reasoning paths. Eventually, we construct MMRP, a multi-task medical reasoning dataset with ranked difficulty, and Chiron-o1, a new medical MLLM devised via a curriculum learning strategy, with robust visual question-answering and generalizable reasoning capabilities. Extensive experiments demonstrate that Chiron-o1, trained on our CoT dataset constructed using MICS, achieves state-of-the-art performance across a list of medical visual question answering and reasoning benchmarks. Codes are available at https://github.com/Yankai96/Chiron-o1 Yankai Jiang 0003, Wenjie Lou, Lilong Wang, Mianxin Liu, Lei Liu 0029, Xiaosong Wang 0001 |
NeurIPS | 9 |
| 2025 | Editorial for Special Issue on Foundation Models for Medical Image Analysis
Xiaosong Wang 0001, Dequan Wang, Jens Rittscher, Dimitris N. Metaxas, Shaoting Zhang 0001 |
Medical Image Anal. | 1 |
| 2024 | Pathology-Knowledge Enhanced Multi-instance Prompt Learning for Few-Shot Whole Slide Image Classification
Linhao Qu, Dingkang Yang, Qinhao Guo, Rongkui Luo, Shaoting Zhang 0001, Xiaosong Wang 0001 |
ECCV (11) | 7 |
| 2024 | Multi-modal Data Binding for Survival Analysis Modeling with Incomplete Data and Annotations
Linhao Qu, Shaoting Zhang 0001, Xiaosong Wang 0001 |
MICCAI (5) | 4 |
| 2024 | BrainSCK: Brain Structure and Cognition Alignment via Knowledge Injection and Reactivation for Diagnosing Brain Disorders
Lilong Wang, Mianxin Liu, Shaoting Zhang 0001, Xiaosong Wang 0001 |
MICCAI (2) | 4 |
| 2024 | Learning Quality Labels for Robust Image ClassificationabstractSupervised learning paradigms largely benefit from the tremendous amount of annotated data. However, the quality of the annotations often varies among labelers. Multi-observer studies have been conducted to examine the annotation variances (by labeling the same data multiple times) to see how it affects critical applications like medical image analysis. In this paper, we demonstrate how multiple sets of annotations (either hand-labeled or algorithm-generated) can be utilized together and mutually benefit the learning of classification tasks. A scheme of learning-to-vote is introduced to sample quality label sets for each data entry on-the-fly during the training. Specifically, a label-sampling module is designed to achieve refined labels (weighted sum of attended ones) that benefit the model learning the most through additional back-propagations. We apply the learning-to-vote scheme on the classification task of a synthetic noisy CIFAR-10 to prove the concept and then demonstrate superior results (3-5% increase on average in multiple disease classification AUCs) on the chest x-ray images from a hospital-scale dataset (MIMIC-CXR) and hand-labeled dataset (OpenI) in comparison to regular training paradigms. Xiaosong Wang 0001, Ziyue Xu 0001, Dong Yang 0005, Leo K. Tam, Holger Roth, Daguang Xu |
WACV | 1 |
| 2023 | Text-Guided Foundation Model Adaptation for Pathological Image Classification
Yunkun Zhang, Mu Zhou, Xiaosong Wang 0001, Yu Qiao 0001, Shaoting Zhang 0001, Dequan Wang |
MICCAI (5) | 4 |
| 2022 | RemixFormer: A Transformer Model for Precision Skin Tumor Differential Diagnosis via Multi-modal Imaging and Non-imaging Data
Yuan Gao 0017, Wei Liu 0127, Kai Huang 0008, Le Lu 0001, Xiaosong Wang 0001, Xian-Sheng Hua 0001, Yu Wang 0108 |
MICCAI (3) | 7 |
| 2022 | Clinical-Realistic Annotation for Histopathology Images with Probabilistic Semi-supervision: A Worst-Case Study
Ziyue Xu 0001, Andriy Myronenko, Dong Yang 0005, Holger Roth, Can Zhao 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (2) | 6 |
| 2021 | T-AutoML: Automated Machine Learning for Lesion Segmentation using Transformers in 3D Medical ImagingabstractLesion segmentation in medical imaging has been an important topic in clinical research. Researchers have proposed various detection and segmentation algorithms to address this task. Recently, deep learning-based approaches have significantly improved the performance over conventional methods. However, most state-of-the-art deep learning methods require the manual design of multiple network components and training strategies. In this paper, we propose a new automated machine learning algorithm, T-AutoML, which not only searches for the best neural architecture, but also finds the best combination of hyper-parameters and data augmentation strategies simultaneously. The proposed method utilizes the modern transformer model, which is introduced to adapt to the dynamic length of the search space embedding and can significantly improve the ability of the search. We validate T-AutoML on several large-scale public lesion segmentation data-sets and achieve state-of-the-art performance. Dong Yang 0005, Andriy Myronenko, Xiaosong Wang 0001, Ziyue Xu 0001, Holger Roth, Daguang Xu |
ICCV | 3 |
| 2021 | Improving Pneumonia Localization via Cross-Attention on Medical Images and Reports
Riddhish Bhalodia, Ali Hatamizadeh, Leo K. Tam, Ziyue Xu 0001, Xiaosong Wang 0001, Evrim Turkbey, Daguang Xu |
MICCAI (2) | 5 |
| 2021 | Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures
Holger Roth, Dong Yang 0005, Wenqi Li 0001, Andriy Myronenko, Wentao Zhu 0001, Ziyue Xu 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (3) | 7 |
| 2021 | Federated semi-supervised learning for COVID region segmentation in chest CT using multi-national data from China, Italy, Japan
Dong Yang 0005, Ziyue Xu 0001, Wenqi Li 0001, Andriy Myronenko, Holger Roth, Stephanie A. Harmon, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Xiaosong Wang 0001, Wentao Zhu 0001, Gianpaolo Carrafiello, Francesca Patella, Maurizio Cariati, Hirofumi Obinata, Hitoshi Mori, Kaku Tamura, Peng An 0002, Bradford J. Wood, Daguang Xu |
Medical Image Anal. | 10 |
| 2021 | Multi-Domain Image Completion for Random Missing Input DataabstractMulti-domain data are widely leveraged in vision applications taking advantage of complementary information from different modalities, e.g., brain tumor segmentation from multi-parametric magnetic resonance imaging (MRI). However, due to possible data corruption and different imaging protocols, the availability of images for each domain could vary amongst multiple data sources in practice, which makes it challenging to build a universal model with a varied set of input data. To tackle this problem, we propose a general approach to complete the random missing domain(s) data in real applications. Specifically, we develop a novel multi-domain image completion method that utilizes a generative adversarial network (GAN) with a representational disentanglement scheme to extract shared content encoding and separate style encoding across multiple domains. We further illustrate that the learned representation in multi-domain image completion could be leveraged for high-level tasks, e.g., segmentation, by introducing a unified framework consisting of image completion and segmentation with a shared content encoder. The experiments demonstrate consistent performance improvement on three datasets for brain tumor segmentation, prostate segmentation, and facial expression image completion respectively. Liyue Shen, Wentao Zhu 0001, Xiaosong Wang 0001, Lei Xing 0001, John M. Pauly, Baris Turkbey, Stephanie A. Harmon, Thomas Sanford, Sherif Mehralivand, Peter L. Choyke, Bradford J. Wood, Daguang Xu |
IEEE Trans. Medical Imaging | 3 |
| 2020 | When Radiology Report Generation Meets Knowledge GraphabstractAutomatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully adapted to generating radiology reports. However, radiology image reporting is different from the natural image captioning task in two aspects: 1) the accuracy of positive disease keyword mentions is critical in radiology image reporting in comparison to the equivalent importance of every single word in a natural image caption; 2) the evaluation of reporting quality should focus more on matching the disease keywords and their associated attributes instead of counting the occurrence of N-gram. Based on these concerns, we propose to utilize a pre-constructed graph embedding module (modeled with a graph convolutional neural network) on multiple disease findings to assist the generation of reports in this work. The incorporation of knowledge graph allows for dedicated feature learning for each disease finding and the relationship modeling between them. In addition, we proposed a new evaluation metric for radiology image reporting with the assistance of the same composed graph. Experimental results demonstrate the superior performance of the methods integrated with the proposed graph embedding module on a publicly accessible dataset (IU-RR) of chest radiographs compared with previous approaches using both the conventional evaluation metrics commonly adopted for image captioning and our proposed ones. Yixiao Zhang 0001, Xiaosong Wang 0001, Ziyue Xu 0001, Qihang Yu, Alan L. Yuille, Daguang Xu |
AAAI | 2 |
| 2020 | Weakly Supervised One-Stage Vision and Language Disease Detection Using Large Scale Pneumonia and Pneumothorax Studies
Leo K. Tam, Xiaosong Wang 0001, Evrim Turkbey, Yuhong Wen, Daguang Xu |
MICCAI (4) | 2 |
| 2020 | Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient DataabstractPrognostic tumor growth modeling via volumetric medical imaging observations can potentially lead to better outcomes of tumor treatment management and surgical planning. Recent advances of convolutional networks (ConvNets) have demonstrated higher accuracy than traditional mathematical models can be achieved in predicting future tumor volumes. This indicates that deep learning based data-driven techniques may have great potentials on addressing such problem. However, current 2D image patch based modeling approaches can not make full use of the spatio-temporal imaging context of the tumor's longitudinal 4D (3D + time) patient data. Moreover, they are incapable to predict clinically-relevant tumor properties, other than the tumor volumes. In this paper, we exploit to formulate the tumor growth process through convolutional Long Short-Term Memory (ConvLSTM) that extract tumor's static imaging appearances and simultaneously capture its temporal dynamic changes within a single network. We extend ConvLSTM into the spatio-temporal domain (ST-ConvLSTM) by jointly learning the inter-slice 3D contexts and the longitudinal or temporal dynamics from multiple patient studies. Our approach can incorporate other non-imaging patient information in an end-to-end trainable manner. Experiments are conducted on the largest 4D longitudinal tumor dataset of 33 patients to date. Results validate that the proposed ST-ConvLSTM model produces a Dice score of 83.2%±5.1% and a RVD of 11.2%±10.8%, both statistically significantly outperforming (p < 0.05) other compared methods of traditional linear model, ConvLSTM, and generative adversarial network (GAN) under the metric of predicting future tumor volumes. Additionally, our new method enables the prediction of both cell density and CT intensity numbers. Last, we demonstrate the generalizability of ST-ConvLSTM by employing it in 4D medical image segmentation task, which achieves an averaged Dice score of 86.3%±1.2% for left-ventricle segmentation in 4D ultrasound with 3 seconds per patient case. Ling Zhang 0002, Le Lu 0001, Xiaosong Wang 0001, Robert Zhu, Mohammadhadi Bagheri, Ronald M. Summers, Jianhua Yao 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked TransformationabstractRecent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. However, application of these models in clinically realistic environments can result in poor generalization and decreased accuracy, mainly due to the domain shift across different hospitals, scanner vendors, imaging protocols, and patient populations etc. Common transfer learning and domain adaptation techniques are proposed to address this bottleneck. However, these solutions require data (and annotations) from the target domain to retrain the model, and is therefore restrictive in practice for widespread model deployment. Ideally, we wish to have a trained (locked) model that can work uniformly well across unseen domains without further training. In this paper, we propose a deep stacked transformation approach for domain generalization. Specifically, a series of n stacked transformations are applied to each image during network training. The underlying assumption is that the "expected" domain shift for a specific medical imaging modality could be simulated by applying extensive data augmentation on a single source domain, and consequently, a deep model trained on the augmented "big" data (BigAug) could generalize well on unseen domains. We exploit four surprisingly effective, but previously understudied, image-based characteristics for data augmentation to overcome the domain generalization problem. We train and evaluate the BigAug model (with n=9 transformations) on three different 3D segmentation tasks (prostate gland, left atrial, left ventricle) covering two medical imaging modalities (MRI and ultrasound) involving eight publicly available challenge datasets. The results show that when training on relatively small dataset (n = 10~32 volumes, depending on the size of the available datasets) from a single source domain: (i) BigAug models degrade an average of 11%(Dice score change) from source to unseen domain, substantially better than conventional augmentation (degrading 39%) and CycleGAN-based domain adaptation method (degrading 25%), (ii) BigAug is better than "shallower" stacked transforms (i.e. those with fewer transforms) on unseen domains and demonstrates modest improvement to conventional augmentation on the source domain, (iii) after training with BigAug on one source domain, performance on an unseen domain is similar to training a model from scratch on that domain when using the same number of training samples. When training on large datasets (n = 465 volumes) with BigAug, (iv) application to unseen domains reaches the performance of state-of-the-art fully supervised models that are trained and tested on their source domains. These findings establish a strong benchmark for the study of domain generalization in medical imaging, and can be generalized to the design of highly robust deep segmentation models for clinical deployment. Ling Zhang 0002, Xiaosong Wang 0001, Dong Yang 0005, Thomas Sanford, Stephanie A. Harmon, Baris Turkbey, Bradford J. Wood, Holger Roth, Andriy Myronenko, Daguang Xu, Ziyue Xu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2018 | TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-RaysabstractChest X-rays are one of the most common radiological examinations in daily clinical routines. Reporting thorax diseases using chest X-rays is often an entry-level task for radiologist trainees. Yet, reading a chest X-ray image remains a challenging job for learning-oriented machine intelligence, due to (1) shortage of large-scale machine-learnable medical image datasets, and (2) lack of techniques that can mimic the high-level reasoning of human radiologists that requires years of knowledge accumulation and professional training. In this paper, we show the clinical free-text radiological reportscan be utilized as a priori knowledge for tackling these two key problems. We propose a novel Text-Image Embedding network (TieNet) for extracting the distinctive image and text representations. Multi-level attention models are integrated into an end-to-end trainable CNN-RNN architecture for highlighting the meaningful text words and image regions. We first apply TieNet to classify the chest X-rays by using both image features and text embeddings extracted from associated reports. The proposed auto-annotation framework achieves high accuracy (over 0.9 on average in AUCs) in assigning disease labels for our hand-label evaluation dataset. Furthermore, we transform the TieNet into a chest X-ray reporting system. It simulates the reporting process and can output disease classification and a preliminary report together. The classification results are significantly improved (6% increase on average in AUCs) compared to the state-of-the-art baseline on an unseen and hand-labeled dataset (OpenI). Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Ronald M. Summers |
CVPR | 1 |
| 2018 | Deep Lesion Graphs in the Wild: Relationship Learning and Organization of Significant Radiology Image Findings in a Diverse Large-Scale Lesion DatabaseabstractRadiologists in their daily work routinely find and annotate significant abnormalities on a large number of radiology images. Such abnormalities, or lesions, have collected over years and stored in hospitals' picture archiving and communication systems. However, they are basically unsorted and lack semantic annotations like type and location. In this paper, we aim to organize and explore them by learning a deep feature representation for each lesion. A large-scale and comprehensive dataset, DeepLesion, is introduced for this task. DeepLesion contains bounding boxes and size measurements of over 32K lesions. To model their similarity relationship, we leverage multiple supervision information including types, self-supervised location coordinates, and sizes. They require little manual annotation effort but describe useful attributes of the lesions. Then, a triplet network is utilized to learn lesion embeddings with a sequential sampling strategy to depict their hierarchical similarity structure. Experiments show promising qualitative and quantitative results on lesion retrieval, clustering, and classification. The learned embeddings can be further employed to build a lesion graph for various clinically useful applications. An algorithm for intra-patient lesion matching is proposed and validated with experiments. Ke Yan 0006, Xiaosong Wang 0001, Le Lu 0001, Ling Zhang 0002, Adam P. Harrison, Mohammadhadi Bagheri, Ronald M. Summers |
CVPR | 2 |
| 2017 | Text Mining Radiology Reports for Deep Learning Radiology Images
Yifan Peng 0002, Xiaosong Wang 0001, Le Lu 0001, Mohammadhadi Bagheri, Ronald M. Summers, Zhiyong Lu |
AMIA | 2 |
| 2017 | ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax DiseasesabstractThe chest X-ray is one of the most commonly accessible radiological examinations for screening and diagnosis of many lung diseases. A tremendous number of X-ray imaging studies accompanied by radiological reports are accumulated and stored in many modern hospitals Picture Archiving and Communication Systems (PACS). On the other side, it is still an open question how this type of hospital-size knowledge database containing invaluable imaging informatics (i.e., loosely labeled) can be used to facilitate the data-hungry deep learning paradigms in building truly large-scale high precision computer-aided diagnosis (CAD) systems. In this paper, we present a new chest X-ray database, namely ChestX-ray8, which comprises 108,948 frontal-view X-ray images of 32,717 unique patients with the text-mined eight disease image labels (where each image can have multi-labels), from the associated radiological reports using natural language processing. Importantly, we demonstrate that these commonly occurring thoracic diseases can be detected and even spatially-located via a unified weakly-supervised multi-label image classification and disease localization framework, which is validated using our proposed dataset. Although the initial quantitative results are promising as reported, deep convolutional neural network based reading chest X-rays (i.e., recognizing and locating the common disease patterns trained with only image-level labels) remains a strenuous task for fully-automated high precision CAD systems. Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Mohammadhadi Bagheri, Ronald M. Summers |
CVPR | 1 |
| 2017 | Unsupervised Joint Mining of Deep Features and Image Labels for Large-Scale Radiology Image Categorization and Scene RecognitionabstractThe recent rapid and tremendous success of deep convolutional neural networks (CNN) on many challenging computer vision tasks largely derives from the accessibility of the well-annotated ImageNet and PASCAL VOC datasets. Nevertheless, unsupervised image categorization (i.e., without the ground-truth labeling) is much less investigated, yet critically important and difficult when annotations are extremely hard to obtain in the conventional way of "Google Search" and crowd sourcing. We address this problem by presenting a looped deep pseudo-task optimization (LDPO) framework for joint mining of deep CNN features and image labels. Our method is conceptually simple and rests upon the hypothesized "convergence" of better labels leading to better trained CNN models which in turn feed more discriminative image representations to facilitate more meaningful clusters/labels. Our proposed method is validated in tackling two important applications: 1) Large-scale medical image annotation has always been a prohibitively expensive and easily-biased task even for well-trained radiologists. Significantly better image categorization results are achieved via our proposed approach compared to the previous state-of-the-art method. 2) Unsupervised scene recognition on representative and publicly available datasets with our proposed technique is examined. The LDPO achieves excellent quantitative scene classification results. On the MIT indoor scene dataset, it attains a clustering accuracy of 75:3%, compared to the state-of-the-art supervised classification accuracy of 81:0% (when both are based on the VGG-VD model). Xiaosong Wang 0001, Le Lu 0001, Hoo-Chang Shin, Lauren Kim, Mohammadhadi Bagheri, Isabella Nogues, Jianhua Yao 0001, Ronald M. Summers |
WACV | 1 |
| 2016 | Automatic Lymph Node Cluster Segmentation Using Holistically-Nested Neural Networks and Structured Optimization in CT Images
Isabella Nogues, Le Lu 0001, Xiaosong Wang 0001, Holger Roth, Gedas Bertasius, Nathan Lay, Jianbo Shi, Yohannes Tsehay, Ronald M. Summers |
MICCAI (2) | 3 |
| 2012 | Archive Film Defect Detection and Removal: An Automatic Restoration FrameworkabstractIn this paper, we present an automatic restoration system targeting on dirt and blotches in digitized archive films. The system is composed of mainly two modules: defect detection and defect removal. In defect detection, we locate the defects by combing temporal and spatial information across a number of frames. An HMM is trained for normal observation sequences and then applied within a framework to detect defective pixels. The resulting defect maps are refined in a two-stage false alarm elimination process and then passed over to the defect removal procedure. A labelled (degraded) pixels is restored in a multiscale framework by first searching the optimal replacement in its dynamically generated, random walk based region of candidate pixel-exemplars and then updating all its features (intensity, motion and texture). Finally, the proposed system is compared against state-of-the-art methods to demonstrate improved accuracy in both detection and restoration using synthetic and real degraded image sequences. Xiaosong Wang 0001, Majid Mirmehdi |
IEEE Trans. Image Process. | 1 |
| 2010 | Archive Film Restoration Based on Spatiotemporal Random Walks
Xiaosong Wang 0001, Majid Mirmehdi |
ECCV (5) | 1 |
| 2009 | HMM based Archive Film Defect Detection with Spatial and Temporal ConstraintsabstractWe propose a novel probabilistic approach to detect defects in digitized archive film, by combining temporal and spatial information across a number of frames. An HMM is trained for normal observation sequences and then applied within a framework to detect defective pixels by examining each new observation sequence and its subformations via a leave-one-out process. A two-stage false alarm elimination process is then applied on the resulting defect maps, comprising MRF modelling and localised feature tracking, which impose spatial and temporal constraints respectively. The proposed method is compared against state-of-the-art and industry-standard methods to demonstrate its superior detection rate. Xiaosong Wang 0001, Majid Mirmehdi |
BMVC | 1 |