Yan Xu 0001

dblp:03/4702-1 · DBLP profile ↗
← Back
53ranked-venue papers
15as first author
27since 2021 · last 2026
0000-0002-2636-7594ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 10 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 UniPET: A universal network for high-quality PET image denoising across varied dose reduction factors
Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
Medical Image Anal.7
2026 Restore-RWKV: Efficient and Effective Medical Image Restoration With RWKV
abstract
Transformers have revolutionized medical image restoration, but the quadratic complexity still poses limitations for their application to high-resolution medical images. The recent advent of the Receptance Weighted Key Value (RWKV) model in the natural language processing field has attracted much attention due to its ability to process long sequences efficiently. To leverage its advanced design, we propose Restore-RWKV, the first RWKV-based model for medical image restoration. Since the original RWKV model is designed for 1D sequences, we make two necessary modifications for modeling spatial relations in 2D medical images. First, we present a recurrent WKV (Re-WKV) attention mechanism that captures global dependencies with linear computational complexity. Re-WKV incorporates bidirectional attention as basic for a global receptive field and recurrent attention to effectively model 2D dependencies from various scan directions. Second, we develop an omnidirectional token shift (Omni-Shift) layer that enhances local dependencies by shifting tokens from all directions and across a wide context range. These adaptations make the proposed Restore-RWKV an efficient and effective model for medical image restoration. Even a lightweight variant of Restore-RWKV, with only 1.16 million parameters, achieves comparable or even superior results compared to existing state-of-the-art (SOTA) methods. Extensive experiments demonstrate that the resulting Restore-RWKV achieves SOTA performance across a range of medical image restoration tasks, including PET image synthesis, CT image denoising, MRI image super-resolution, and all-in-one medical image restoration.
Zhiwen Yang 0001, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
IEEE J. Biomed. Health Informatics6
2026 VQPET: Leveraging Vector-Quantized Codebook Prior for PET Image Synthesis
abstract
Positron emission tomography (PET) image synthesis is a highly ill-posed problem that requires auxiliary priors to 1) alleviate the loss of high-quality (HQ) information in low-quality (LQ) inputs, and 2) impose additional constraints to reduce mapping uncertainty. However, existing auxiliary priors in PET image synthesis often provide inadequate guidance due to inaccurate prior information or limited prior expressiveness. To overcome the aforementioned limitations, the vector-quantized (VQ) codebook prior is employed as a promising solution. By learning discrete latent feature representations of HQ images through deep models, the VQ codebook prior encompasses accurate HQ information and possesses great expressiveness. Building upon this, we propose a novel two-stage framework, VQPET, that introduces the VQ codebook prior for PET image synthesis. In the first stage, it pretrains a VQGAN on an additional large-scale HQ PET dataset, encoding intrinsic HQ features as code items in the VQ codebook. The VQ codebook prior is thus derived from the high-level features obtained from the pretrained VQGAN and serves as an additional constraint for downstream synthesis. In the second stage, it develops a codebook-prior-guided network (CPGNet) that effectively exploits the VQ codebook prior to produce realistic outputs. Specifically, CPGNet progressively incorporates the VQ codebook prior at multiple decoding levels, providing reliable guidance for HQ synthesis. Compared to previous works, VQPET innovatively leverages additional large-scale HQ datasets to transfer pretrained prior knowledge for enhanced synthesis and functions as a general framework applicable to any encoder-decoder network. Extensive experiments demonstrate the substantial effect and robust generalizability of VQPET.
Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
IEEE Trans. Medical Imaging9
2025 CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology
abstract
Multimodal large language models (MLLMs) have demonstrated strong performance in patch-level pathological image analysis; however, they often lack the holistic perceptual capability necessary for comprehensive Whole Slide Image (WSI) interpretation. Recent approaches have explored constructing slide-level MLLMs using VQA datasets that are entirely generated from pathology reports by large language models (LLMs). However, these datasets suffer from critical limitations: hallucinated content, information leakage in question stems, clinically irrelevant or visual independent questions, and the omission of essential diagnostic features-issues that undermine both data quality and clinical validity. In this paper, we introduce a clinical diagnosis template-based pipeline to collect pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA-BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS-Align, a dataset of 80k slide-description pairs from 804 WSIs for vision-language alignment training, and CTISBench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question-answer pairs. CTIS-Bench emphasizes clinically grounded, closed-ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non-visual reasoning, and require genuine slide understanding. We further propose CTIS-QA, a Slide-level Question Answering model, featuring a dual-stream architecture that mimics pathologists' diagnostic approach. One stream captures global slidelevel context via clustering-based feature aggregation, while the other focuses on salient local regions through attention-guided patch perception module. Extensive experiments on WSI-VQA, CTIS-Bench, and slide-level diagnostic tasks show that CTIS-QA consistently outperforms existing state-of-the-art models across multiple metrics. We will fully release both CTIS-Bench and CTIS-QA as open-source resources.
Ziniu Qian, Yang Zhou 0036, Bingzheng Wei, Yan Xu 0001
BIBM6
2025 Visual Textualization for Image Prompted Object Detection
abstract
We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are difficult to describe textually and nearly absent from their pre-training data, while preserving their pre-trained object-text alignment. Specifically, VisTex-OVLM leverages multi-scale textualizing blocks and a multi-stage fusion strategy to integrate visual information from visual exemplars, generating textualized visual tokens that effectively guide OVLMs alongside text prompts. Unlike previous methods, our method maintains the original architecture of OVLM, maintaining its generalization capabilities while enhancing performance in few-shot settings. VisTex-OVLM demonstrates superior performance across open-set datasets which have minimal overlap with OVLM's pre-training data and achieves state-of-the-art results on few-shot benchmarks PASCAL VOC and MSCOCO. The code will be released at https://github.com/WitGotFlg/VisTex-OVLM.
Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Yan Xu 0001
ICCV5
2025 All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
Zhiwen Yang 0001, Haotian Hou, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
MICCAI (16)7
2025 FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
Huihan Wang, Zhiwen Yang 0001, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
MICCAI (9)6
2025 TAT: Task-Adaptive Transformer for All-in-One Medical Image Restoration
Zhiwen Yang 0001, Jiaju Zhang, Bingzheng Wei, Yan Xu 0001
MICCAI (16)6
2025 TDFormer: Top-Down Token Generation for 3D Medical Image Segmentation
abstract
Accurate medical image segmentation is critical to effective treatment strategies. Existing transformer-based methods for image segmentation mostly split the input image into a fixed and regular grid and regard cells in the grid as the vision tokens. However, not all tokens are of equal importance in the medical segmentation tasks, e.g., the tokens in tumor areas must be processed in a higher resolution than the background tokens which can be easily predicted with fewer transformer layers. In this paper, we propose a simple yet efficient segmentation framework called Top-Down Transformer (TDFormer), which incorporates a spatially adaptive token generation scheme into the transformer. The proposed top-down token generation comprises the following three components: attentiveness calculation, token splitting, and token fusion, where the collaboration of these components gradually fuses redundant background tokens and focuses only on the most critical areas. This allows for allocating more computation to process tokens containing delicate details in a finer resolution. Extensive experiments are conducted to demonstrate the robustness and effectiveness of the proposed TDFormer, that our method are superior to other state-of-the-art methods on the following publicly accessible datasets: BTCV Challenge, LiTS and BraTS 2020. We also dissect our method and evaluate the performance of each component.
Hao Du 0006, Qihua Dong, Yan Xu 0001, Jing Liao 0001
IEEE J. Biomed. Health Informatics3
2025 Unsupervised Non-Rigid Histological Image Registration Guided by Keypoint Correspondences Based on Learnable Deep Features With Iterative Training
abstract
Histological image registration is a fundamental task in histological image analysis. It is challenging because of substantial appearance differences due to multiple staining. Keypoint correspondences, i.e., matched keypoint pairs, have been introduced to guide unsupervised deep learning (DL) based registration methods to handle such a registration task. This paper proposes an iterative keypoint correspondence-guided (IKCG) unsupervised network for non-rigid histological image registration. Fixed deep features and learnable deep features are introduced as keypoint descriptors to automatically establish keypoint correspondences, the distance between which is used as a loss function to train the registration network. Fixed deep features extracted from DL networks that are pre-trained on natural image datasets are more discriminative than handcrafted ones, benefiting from the deep and hierarchical nature of DL networks. The intermediate layer outputs of the registration networks trained on histological image datasets are extracted as learnable deep features, which reveal unique information for histological images. An iterative training strategy is adopted to train the registration network and optimize learnable deep features jointly. Benefiting from the excellent matching ability of learnable deep features optimized with the iterative training strategy, the proposed method can solve the local non-rigid large displacement problem, an inevitable problem usually caused by misoperation, such as tears in producing tissue slices. The proposed method is evaluated on the Automatic Non-rigid Histology Image Registration (ANHIR) website and AutomatiC Registration Of Breast cAncer Tissue (ACROBAT) website. It ranked 1st on both websites as of August 6th, 2024.
Xingyue Wei, Jianwen Luo 0001, Yan Xu 0001
IEEE Trans. Medical Imaging5
2025 AttriPrompter: Auto-Prompting With Attribute Semantics for Zero-Shot Nuclei Detection via Visual-Language Pre-Trained Models
abstract
Large-scale visual-language pre-trained models (VLPMs) have demonstrated exceptional performance in downstream object detection through text prompts for natural scenes. However, their application to zero-shot nuclei detection on histopathology images remains relatively unexplored, mainly due to the significant gap between the characteristics of medical images and the web-originated text-image pairs used for pre-training. This paper aims to investigate the potential of the object-level VLPM, Grounded Language-Image Pre-training (GLIP), for zero-shot nuclei detection. Specifically, we propose an innovative auto-prompting pipeline, named AttriPrompter, comprising attribute generation, attribute augmentation, and relevance sorting, to avoid subjective manual prompt design. AttriPrompter utilizes VLPMs' text-to-image alignment to create semantically rich text prompts, which are then fed into GLIP for initial zero-shot nuclei detection. Additionally, we propose a self-trained knowledge distillation framework, where GLIP serves as the teacher with its initial predictions used as pseudo labels, to address the challenges posed by high nuclei density, including missed detections, false positives, and overlapping instances. Our method exhibits remarkable performance in label-free nuclei detection, outperforming all existing unsupervised methods and demonstrating excellent generality. Notably, this work highlights the astonishing potential of VLPMs pre-trained on natural image-text pairs for downstream tasks in the medical field as well. Code will be released at github.com/AttriPrompter.
Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yan Xu 0001
IEEE Trans. Medical Imaging7
2024 Tuning Stable Rank Shrinkage: Aiming at the Overlooked Structural Risk in Fine-tuning
abstract
Existing finetuning methods for computer vision tasks primarily focus on re-weighting the knowledge learned from the source domain during pre-training. They aim to retain beneficial knowledge for the target domain while suppressing unfavorable knowledge. During the pre-training and fine-tuning stages, there is a notable disparity in the data scale. Consequently, it is theoretically necessary to employ a model with reduced complexity to mitigate the potential structural risk. However, our empirical investigation in this paper reveals that models finetuned using existing methods still manifest a high level of model complexity inherited from the pre-training stage, leading to a suboptimal stability and generalization ability. This phenomenon indicates an issue that has been overlooked in fine-tuning: Structural Risk Minimization. To address this issue caused by data scale disparity during the fine-tuning stage, we propose a simple yet effective approach called Tuning Stable Rank Shrinkage (TSRS). TSRS mitigates the structural risk during the fine-tuning stage by constraining the noise sensitivity of the target model based on stable rank theories. Through extensive experiments, we demonstrate that incorporating TSRS into fine-tuning methods leads to improved generalization ability on various tasks, regardless of whether the neural networks are based on convolution or transformer architectures. Additionally, empirical analysis reveals that TSRS enhances the robustness, convexity, and smoothness of the loss landscapes in fine-tuned models. Code is available at https://github.com/WitGotFlg/TSRS.
Sicong Shen, Yang Zhou 0036, Bingzheng Wei, Eric I-Chao Chang, Yan Xu 0001
CVPR5
2024 SDPT: Synchronous Dual Prompt Tuning for Fusion-Based Visual-Language Pre-trained Models
Yang Zhou 0036, Yongjian Wu 0002, Jiya Saiyin, Bingzheng Wei, Maode Lai, Eric Chang, Yan Xu 0001
ECCV (49)7
2024 All-In-One Medical Image Restoration via Task-Adaptive Routing
Zhiwen Yang 0001, Ziniu Qian, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
MICCAI (7)8
2024 Region Attention Transformer for Medical Image Restoration
Zhiwen Yang 0001, Ziniu Qian, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yan Xu 0001
MICCAI (7)8
2024 Nucleus-Aware Self-Supervised Pretraining Using Unpaired Image-to-Image Translation for Histopathology Images
abstract
Self-supervised pretraining attempts to enhance model performance by obtaining effective features from unlabeled data, and has demonstrated its effectiveness in the field of histopathology images. Despite its success, few works concentrate on the extraction of nucleus-level information, which is essential for pathologic analysis. In this work, we propose a novel nucleus-aware self-supervised pretraining framework for histopathology images. The framework aims to capture the nuclear morphology and distribution information through unpaired image-to-image translation between histopathology images and pseudo mask images. The generation process is modulated by both conditional and stochastic style representations, ensuring the reality and diversity of the generated histopathology images for pretraining. Further, an instance segmentation guided strategy is employed to capture instance-level information. The experiments on 7 datasets show that the proposed pretraining method outperforms supervised ones on Kather classification, multiple instance learning, and 5 dense-prediction tasks with the transfer learning protocol, and yields superior results than other self-supervised approaches on 8 semi-supervised tasks. Our project is publicly available at https://github.com/zhiyuns/UNITPathSSL.
Zhiyun Song, Penghui Du, Junpeng Yan, Kailu Li, Jianzhong Shou, Maode Lai, Yubo Fan, Yan Xu 0001
IEEE Trans. Medical Imaging8
2024 Learning a Single Network for Robust Medical Image Segmentation With Noisy Labels
abstract
Robust segmenting with noisy labels is an important problem in medical imaging due to the difficulty of acquiring high-quality annotations. Despite the enormous success of recent developments, these developments still require multiple networks to construct their frameworks and focus on limited application scenarios, which leads to inflexibility in practical applications. They also do not explicitly consider the coarse boundary label problem, which results in sub-optimal results. To overcome these challenges, we propose a novel Simultaneous Edge Alignment and Memory-Assisted Learning (SEAMAL) framework for noisy-label robust segmentation. It achieves single-network robust learning, which is applicable for both 2D and 3D segmentation, in both Set-HQ-knowable and Set-HQ-agnostic scenarios. Specifically, to achieve single-model noise robustness, we design a Memory-assisted Selection and Correction module (MSC) that utilizes predictive history consistency from the Prediction Memory Bank to distinguish between reliable and non-reliable labels pixel-wisely, and that updates the reliable ones at the superpixel level. To overcome the coarse boundary label problem, which is common in practice, and to better utilize shape-relevant information at the boundary, we propose an Edge Detection Branch (EDB) that explicitly learns the boundary via an edge detection layer with only slight additional computational cost, and we improve the sharpness and precision of the boundary with a thinning loss. Extensive experiments verify that SEAMAL outperforms previous works significantly.
Shuquan Ye, Yan Xu 0001, Dongdong Chen 0001, Songfang Han, Jing Liao 0001
IEEE Trans. Medical Imaging2
2023 Preserving Tumor Volumes for Unsupervised Medical Image Registration
abstract
Medical image registration is a critical task that estimates the spatial correspondence between pairs of images. However, current traditional and deep-learning-based methods rely on similarity measures to generate a deforming field, which often results in disproportionate volume changes in dissimilar regions, especially in tumor regions. These changes can significantly alter the tumor size and underlying anatomy, which limits the practical use of image registration in clinical diagnosis. To address this issue, we have formulated image registration with tumors as a constraint problem that preserves tumor volumes while maximizing image similarity in other normal regions. Our proposed strategy involves a two-stage process. In the first stage, we use similarity-based registration to identify potential tumor regions by their volume change, generating a soft tumor mask accordingly. In the second stage, we propose a volume-preserving registration with a novel adaptive volume-preserving loss that penalizes the change in size adaptively based on the masks calculated from the previous stage. Our approach balances image similarity and volume preservation in different regions, i.e., normal and tumor regions, by using soft tumor masks to adjust the imposition of volume-preserving loss on each one. This ensures that the tumor volume is preserved during the registration process. We have evaluated our strategy on various datasets and network architectures, demonstrating that our method successfully preserves the tumor volume while achieving comparable registration results with state-of-the-art methods. Our codes is available at: https://dddraxxx.github.io/Volume-Preserving-Registration/.
Qihua Dong, Hao Du 0006, Yan Xu 0001, Jing Liao 0001
ICCV4
2023 Zero-Shot Nuclei Detection via Visual-Language Pre-trained Models
Yongjian Wu 0002, Yang Zhou 0036, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yubo Fan, Yan Xu 0001
MICCAI (6)8
2023 DRMC: A Generalist Model with Dynamic Routing for Multi-center PET Image Synthesis
Zhiwen Yang 0001, Yang Zhou 0036, Hui Zhang 0099, Bingzheng Wei, Yubo Fan, Yan Xu 0001
MICCAI (3)6
2023 Weakly supervised histopathology image segmentation with self-attention
Kailu Li, Ziniu Qian, Yingnan Han, Eric I-Chao Chang, Bingzheng Wei, Maode Lai, Jing Liao 0001, Yubo Fan, Yan Xu 0001
Medical Image Anal.9
2023 Weakly-Supervised 3D Medical Image Segmentation Using Geometric Prior and Contrastive Similarity
abstract
Medical image segmentation is almost the most important pre-processing procedure in computer-aided diagnosis but is also a very challenging task due to the complex shapes of segments and various artifacts caused by medical imaging, (i.e., low-contrast tissues, and non-homogenous textures). In this paper, we propose a simple yet effective segmentation framework that incorporates the geometric prior and contrastive similarity into the weakly-supervised segmentation framework in a loss-based fashion. The proposed geometric prior built on point cloud provides meticulous geometry to the weakly-supervised segmentation proposal, which serves as better supervision than the inherent property of the bounding-box annotation (i.e., height and width). Furthermore, we propose the contrastive similarity to encourage organ pixels to gather around in the contrastive embedding space, which helps better distinguish low-contrast tissues. The proposed contrastive embedding space can make up for the poor representation of the conventionally-used gray space. Extensive experiments are conducted to verify the effectiveness and the robustness of the proposed weakly-supervised segmentation framework. The proposed framework are superior to state-of-the-art weakly-supervised methods on the following publicly accessible datasets: LiTS 2017 Challenge, KiTS 2021 Challenge and LPBA40. We also dissect our method and evaluate the performance of each component.
Hao Du 0006, Qihua Dong, Yan Xu 0001, Jing Liao 0001
IEEE Trans. Medical Imaging3
2023 Cyclic Learning: Bridging Image-Level Labels and Nuclei Instance Segmentation
abstract
Nuclei instance segmentation on histopathology images is of great clinical value for disease analysis. Generally, fully-supervised algorithms for this task require pixel-wise manual annotations, which is especially time-consuming and laborious for the high nuclei density. To alleviate the annotation burden, we seek to solve the problem through image-level weakly supervised learning, which is underexplored for nuclei instance segmentation. Compared with most existing methods using other weak annotations (scribble, point, etc.) for nuclei instance segmentation, our method is more labor-saving. The obstacle to using image-level annotations in nuclei instance segmentation is the lack of adequate location information, leading to severe nuclei omission or overlaps. In this paper, we propose a novel image-level weakly supervised method, called cyclic learning, to solve this problem. Cyclic learning comprises a front-end classification task and a back-end semi-supervised instance segmentation task to benefit from multi-task learning (MTL). We utilize a deep learning classifier with interpretability as the front-end to convert image-level labels to sets of high-confidence pseudo masks and establish a semi-supervised architecture as the back-end to conduct nuclei instance segmentation under the supervision of these pseudo masks. Most importantly, cyclic learning is designed to circularly share knowledge between the front-end classifier and the back-end semi-supervised part, which allows the whole system to fully extract the underlying information from image-level labels and converge to a better optimum. Experiments on three datasets demonstrate the good generality of our method, which outperforms other image-level weakly supervised methods for nuclei instance segmentation, and achieves comparable performance to fully-supervised methods.
Yang Zhou 0036, Yongjian Wu 0002, Zihua Wang, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yubo Fan, Yan Xu 0001
IEEE Trans. Medical Imaging8
2022 Transformer Based Multiple Instance Learning for Weakly Supervised Histopathology Image Segmentation
Ziniu Qian, Kailu Li, Maode Lai, Eric I-Chao Chang, Bingzheng Wei, Yubo Fan, Yan Xu 0001
MICCAI (2)7
2022 Unsupervised Histological Image Registration Using Structural Feature Guided Convolutional Neural Network
abstract
Registration of multiple stained images is a fundamental task in histological image analysis. In supervised methods, obtaining ground-truth data with known correspondences is laborious and time-consuming. Thus, unsupervised methods are expected. Unsupervised methods ease the burden of manual annotation but often at the cost of inferior results. In addition, registration of histological images suffers from appearance variance due to multiple staining, repetitive texture, and section missing during making tissue sections. To deal with these challenges, we propose an unsupervised structural feature guided convolutional neural network (SFG). Structural features are robust to multiple staining. The combination of low-resolution rough structural features and high-resolution fine structural features can overcome repetitive texture and section missing, respectively. SFG consists of two components of structural consistency constraints according to the formations of structural features, i.e., dense structural component and sparse structural component. The dense structural component uses structural feature maps of the whole image as structural consistency constraints, which represent local contextual information. The sparse structural component utilizes the distance of automatically obtained matched key points as structural consistency constraints because the matched key points in an image pair emphasize the matching of significant structures, which imply global information. In addition, a multi-scale strategy is used in both dense and sparse structural components to make full use of the structural information at low resolution and high resolution to overcome repetitive texture and section missing. The proposed method was evaluated on a public histological dataset (ANHIR) and ranked first as of Jan 18th, 2022.
Xingyue Wei, Yayu Hao, Jianwen Luo 0001, Yan Xu 0001
IEEE Trans. Medical Imaging5
2022 3D Segmentation Guided Style-Based Generative Adversarial Networks for PET Synthesis
abstract
Potential radioactive hazards in full-dose positron emission tomography (PET) imaging remain a concern, whereas the quality of low-dose images is never desirable for clinical use. So it is of great interest to translate low-dose PET images into full-dose. Previous studies based on deep learning methods usually directly extract hierarchical features for reconstruction. We notice that the importance of each feature is different and they should be weighted dissimilarly so that tiny information can be captured by the neural network. Furthermore, the synthesis on some regions of interest is important in some applications. Here we propose a novel segmentation guided style-based generative adversarial network (SGSGAN) for PET synthesis. (1) We put forward a style-based generator employing style modulation, which specifically controls the hierarchical features in the translation process, to generate images with more realistic textures. (2) We adopt a task-driven strategy that couples a segmentation task with a generative adversarial network (GAN) framework to improve the translation performance. Extensive experiments show the superiority of our overall framework in PET synthesis, especially on those regions of interest.
Yang Zhou 0036, Zhiwen Yang 0001, Hui Zhang 0099, Eric I-Chao Chang, Yubo Fan, Yan Xu 0001
IEEE Trans. Medical Imaging6
2021 Large Scale Image Completion via Co-Modulated Generative Adversarial Networks
Shengyu Zhao, Jonathan Cui, Yilun Sheng, Eric I-Chao Chang, Yan Xu 0001
ICLR7
2020 MaskFlownet: Asymmetric Feature Matching With Learnable Occlusion Mask
abstract
Feature warping is a core technique in optical flow estimation; however, the ambiguity caused by occluded areas during warping is a major problem that remains unsolved. In this paper, we propose an asymmetric occlusion-aware feature matching module, which can learn a rough occlusion mask that filters useless (occluded) areas immediately after feature warping without any explicit supervision. The proposed module can be easily integrated into end-to-end network architectures and enjoys performance gains while introducing negligible computational cost. The learned occlusion mask can be further fed into a subsequent network cascade with dual feature pyramids with which we achieve state-of-the-art performance. At the time of submission, our method, called MaskFlownet, surpasses all published optical flow methods on the MPI Sintel, KITTI 2012 and 2015 benchmarks. Code is available at https://github.com/microsoft/MaskFlownet.
Shengyu Zhao, Yilun Sheng, Eric I-Chao Chang, Yan Xu 0001
CVPR5
2020 Microscopic Fine-Grained Instance Classification Through Deep Attention
Mengran Fan, Tapabrata Chakraborti, Eric I-Chao Chang, Yan Xu 0001, Jens Rittscher
MICCAI (5)4
2020 Unsupervised 3D End-to-End Medical Image Registration With Volume Tweening Network
abstract
3D medical image registration is of great clinical importance. However, supervised learning methods require a large amount of accurately annotated corresponding control points (or morphing), which are very difficult to obtain. Unsupervised learning methods ease the burden of manual annotation by exploiting unlabeled data without supervision. In this article, we propose a new unsupervised learning method using convolutional neural networks under an end-to-end framework, Volume Tweening Network (VTN), for 3D medical image registration. We propose three innovative technical components: (1) An end-to-end cascading scheme that resolves large displacement; (2) An efficient integration of affine registration network; and (3) An additional invertibility loss that encourages backward consistency. Experiments demonstrate that our algorithm is 880x faster (or 3.3x faster without GPU acceleration) than traditional optimization-based methods and achieves state-of-the-art performance in medical image registration.
Shengyu Zhao, Ting Fung Lau, Ji Luo 0002, Eric I-Chao Chang, Yan Xu 0001
IEEE J. Biomed. Health Informatics5
2020 ANHIR: Automatic Non-Rigid Histological Image Registration Challenge
abstract
Automatic Non-rigid Histological Image Registration (ANHIR) challenge was organized to compare the performance of image registration algorithms on several kinds of microscopy histology images in a fair and independent manner. We have assembled 8 datasets, containing 355 images with 18 different stains, resulting in 481 image pairs to be registered. Registration accuracy was evaluated using manually placed landmarks. In total, 256 teams registered for the challenge, 10 submitted the results, and 6 participated in the workshop. Here, we present the results of 7 well-performing methods from the challenge together with 6 well-known existing methods. The best methods used coarse but robust initial alignment, followed by non-rigid registration, used multiresolution, and were carefully tuned for the data at hand. They outperformed off-the-shelf methods, mostly by being more robust. The best methods could successfully register over 98% of all landmarks and their mean landmark registration accuracy (TRE) was 0.44% of the image diagonal. The challenge remains open to submissions and all images are available for download.
Jirí Borovec, Jan Kybic, Ignacio Arganda-Carreras, Dmitry V. Sorokin, Gloria Bueno García, Alexander V. Khvostikov, Spyridon Bakas, Eric I-Chao Chang, Stefan Heldmann, Kimmo Kartasalo, Leena Latonen, Johannes Lotz 0002, Michelle Noga, Sarthak Pati, Kumaradevan Punithakumar, Pekka Ruusuvuori, Andrzej Skalski, Nazanin Tahmasebi, Masi Valkonen, Ludovic Venet, Nick Weiss, Marek Wodzinski, Yan Xu 0001, Paul A. Yushkevich, Shengyu Zhao, Arrate Muñoz-Barrutia
IEEE Trans. Medical Imaging25
2019 Recursive Cascaded Networks for Unsupervised Medical Image Registration
abstract
We present recursive cascaded networks, a general architecture that enables learning deep cascades, for deformable image registration. The proposed architecture is simple in design and can be built on any base network. The moving image is warped successively by each cascade and finally aligned to the fixed image; this procedure is recursive in a way that every cascade learns to perform a progressive deformation for the current warped image. The entire system is end-to-end and jointly trained in an unsupervised manner. In addition, enabled by the recursive architecture, one cascade can be iteratively applied for multiple times during testing, which approaches a better fit between each of the image pairs. We evaluate our method on 3D medical images, where deformable registration is most commonly applied. We demonstrate that recursive cascaded networks achieve consistent, significant gains and outperform state-of-the-art methods. The performance reveals an increasing trend as long as more cascades are trained, while the limit is not observed.
Shengyu Zhao, Eric I-Chao Chang, Yan Xu 0001
ICCV4
2019 Wound area measurement with 3D transformation and smartphone images
abstract
BACKGROUND: Quantitative areas is of great measurement of wound significance in clinical trials, wound pathological analysis, and daily patient care. 2D methods cannot solve the problems caused by human body curvatures and different camera shooting angles. Our objective is to simply collect wound areas, accurately measure wound areas and overcome the shortcomings of 2D methods. RESULTS: We propose a method with 3D transformation to measure wound area on a human body surface, which combines structure from motion (SFM), least squares conformal mapping (LSCM), and image segmentation. The method captures 2D images of wound, which is surrounded by adhesive tape scale next to it, by smartphone and implements 3D reconstruction from the images based on SFM. Then it uses LSCM to unwrap the UV map of the 3D model. In the end, it utilizes image segmentation by interactive method for wound extraction and measurement. Our system yields state-of-the-art results on a dataset of 118 wounds on 54 patients, and performs with an accuracy of 0.97. The Pearson correlation, standardized regression coefficient and adjusted R square of our method are 0.999, 0.895 and 0.998 respectively. CONCLUSIONS: A smartphone is used to capture wound images, which lowers costs, lessens dependence on hardware, and avoids the risk of infection. The quantitative calculation of the 3D wound area is realized, solving the challenges that 2D methods cannot and achieving a good accuracy.
Xingyu Fan, Zhizhi Guo, Zhongjun Mo, Eric I-Chao Chang, Yan Xu 0001
BMC Bioinform.6
2019 Mapping anatomical related entities to human body parts based on wikipedia in discharge summaries
abstract
*: Background Consisting of dictated free-text documents such as discharge summaries, medical narratives are widely used in medical natural language processing. Relationships between anatomical entities and human body parts are crucial for building medical text mining applications. To achieve this, we establish a mapping system consisting of a Wikipedia-based scoring algorithm and a named entity normalization method (NEN). The mapping system makes full use of information available on Wikipedia, which is a comprehensive Internet medical knowledge base. We also built a new ontology, Tree of Human Body Parts (THBP), from core anatomical parts by referring to anatomical experts and Unified Medical Language Systems (UMLS) to make the mapping system efficacious for clinical treatments. *: Result The gold standard is derived from 50 discharge summaries from our previous work, in which 2,224 anatomical entities are included. The F1-measure of the baseline system is 70.20%, while our algorithm based on Wikipedia achieves 86.67% with the assistance of NEN. *: Conclusions We construct a framework to map anatomical entities to THBP ontology using normalization and a scoring algorithm based on Wikipedia. The proposed framework is proven to be much more effective and efficient than the main baseline system.
Xingyu Fan, Luoxin Chen, Eric I-Chao Chang, Sophia Ananiadou, Jun'ichi Tsujii, Yan Xu 0001
BMC Bioinform.7
2019 Unsupervised Learning for Cell-Level Visual Representation in Histopathology Images With Generative Adversarial Networks
abstract
The visual attributes of cells, such as the nuclear morphology and chromatin openness, are critical for histopathology image analysis. By learning cell-level visual representation, we can obtain a rich mix of features that are highly reusable for various tasks, such as cell-level classification, nuclei segmentation, and cell counting. In this paper, we propose a unified generative adversarial networks architecture with a new formulation of loss to perform robust cell-level visual representation learning in an unsupervised setting. Our model is not only label-free and easily trained but also capable of cell-level unsupervised classification with interpretable visualization, which achieves promising results in the unsupervised classification of bone marrow cellular components. Based on the proposed cell-level visual representation learning, we further develop a pipeline that exploits the varieties of cellular elements to perform histopathology image classification, the advantages of which are demonstrated on bone marrow datasets.
Eric I-Chao Chang, Yubo Fan, Maode Lai, Yan Xu 0001
IEEE J. Biomed. Health Informatics6
2018 End-to-end subtitle detection and recognition for videos in East Asian languages via CNN ensemble
Yan Xu 0001, Siyuan Shan, Ziming Qiu, Zhipeng Jia, Zhengyang Shen, Mengfei Shi, Eric I-Chao Chang
Signal Process. Image Commun.1
2017 Large scale tissue histopathology image classification, segmentation, and visualization via deep convolutional activation features
abstract
BACKGROUND: Histopathology image analysis is a gold standard for cancer recognition and diagnosis. Automatic analysis of histopathology images can help pathologists diagnose tumor and cancer subtypes, alleviating the workload of pathologists. There are two basic types of tasks in digital histopathology image analysis: image classification and image segmentation. Typical problems with histopathology images that hamper automatic analysis include complex clinical representations, limited quantities of training images in a dataset, and the extremely large size of singular images (usually up to gigapixels). The property of extremely large size for a single image also makes a histopathology image dataset be considered large-scale, even if the number of images in the dataset is limited. RESULTS: In this paper, we propose leveraging deep convolutional neural network (CNN) activation features to perform classification, segmentation and visualization in large-scale tissue histopathology images. Our framework transfers features extracted from CNNs trained by a large natural image database, ImageNet, to histopathology images. We also explore the characteristics of CNN features by visualizing the response of individual neuron components in the last hidden layer. Some of these characteristics reveal biological insights that have been verified by pathologists. According to our experiments, the framework proposed has shown state-of-the-art performance on a brain tumor dataset from the MICCAI 2014 Brain Tumor Digital Pathology Challenge and a colon cancer histopathology image dataset. CONCLUSIONS: The framework proposed is a simple, efficient and effective system for histopathology image automatic analysis. We successfully transfer ImageNet knowledge as deep convolutional activation features to the classification and segmentation of histopathology images with little training data. CNN features are significantly more powerful than expert-designed features.
Yan Xu 0001, Zhipeng Jia, Liang-Bo Wang, Yuqing Ai, Maode Lai, Eric I-Chao Chang
BMC Bioinform.1
2017 Parallel multiple instance learning for extremely large histopathology image analysis
abstract
BACKGROUND: Histopathology images are critical for medical diagnosis, e.g., cancer and its treatment. A standard histopathology slice can be easily scanned at a high resolution of, say, 200,000×200,000 pixels. These high resolution images can make most existing imaging processing tools infeasible or less effective when operated on a single machine with limited memory, disk space and computing power. RESULTS: In this paper, we propose an algorithm tackling this new emerging "big data" problem utilizing parallel computing on High-Performance-Computing (HPC) clusters. Experimental results on a large-scale data set (1318 images at a scale of 10 billion pixels each) demonstrate the efficiency and effectiveness of the proposed algorithm for low-latency real-time applications. CONCLUSIONS: The framework proposed an effective and efficient system for extremely large histopathology image analysis. It is based on the multiple instance learning formulation for weakly-supervised learning for image classification, segmentation and clustering. When a max-margin concept is adopted for different clusters, we obtain further improvement in clustering performance.
Yan Xu 0001, Yeshu Li, Zhengyang Shen, Teng Gao, Yubo Fan, Maode Lai, Eric I-Chao Chang
BMC Bioinform.1
2017 Learning multi-level features for sensor-based human action recognition
Yan Xu 0001, Zhengyang Shen, Yifan Gao 0001, Shujian Deng, Yubo Fan, Eric I-Chao Chang
Pervasive Mob. Comput.1
2017 Constrained Deep Weak Supervision for Histopathology Image Segmentation
abstract
In this paper, we develop a new weakly supervised learning algorithm to learn to segment cancerous regions in histopathology images. This paper is under a multiple instance learning (MIL) framework with a new formulation, deep weak supervision (DWS); we also propose an effective way to introduce constraints to our neural networks to assist the learning process. The contributions of our algorithm are threefold: 1) we build an end-to-end learning system that segments cancerous regions with fully convolutional networks (FCNs) in which image-to-image weakly-supervised learning is performed; 2) we develop a DWS formulation to exploit multi-scale learning under weak supervision within FCNs; and 3) constraints about positive instances are introduced in our approach to effectively explore additional weakly supervised information that is easy to obtain and enjoy a significant boost to the learning process. The proposed algorithm, abbreviated as DWS-MIL, is easy to implement and can be trained efficiently. Our system demonstrates the state-of-the-art results on large-scale histopathology image data sets and can be applied to various applications in medical imaging beyond histopathology images, such as MRI, CT, and ultrasound images.
Zhipeng Jia, Xingyi Huang, Eric I-Chao Chang, Yan Xu 0001
IEEE Trans. Medical Imaging4
2016 Gland Instance Segmentation by Deep Multichannel Side Supervision
Yan Xu 0001, Yang Li 0075, Mingyuan Liu 0002, Maode Lai, Eric I-Chao Chang
MICCAI (2)1
2015 Deep convolutional activation features for large scale Brain Tumor histopathology image classification and segmentation
abstract
We propose a simple, efficient and effective method using deep convolutional activation features (CNNs) to achieve stat- of-the-art classification and segmentation for the MICCAI 2014 Brain Tumor Digital Pathology Challenge. Common traits of such medical image challenges are characterized by large image dimensions (up to the gigabyte size of an image), a limited amount of training data, and significant clinical feature representations. To tackle these challenges, we transfer the features extracted from CNNs trained with a very large general image database to the medical image challenge. In this paper, we used CNN activations trained by ImageNet to extract features (4096 neurons, 13.3% active). In addition, feature selection, feature pooling, and data augmentation are used in our work. Our system obtained 97.5% accuracy on classification and 84% accuracy on segmentation, demonstrating a significant performance gain over other participating teams.
Yan Xu 0001, Zhipeng Jia, Yuqing Ai, Maode Lai, Eric I-Chao Chang
ICASSP1
2015 Bilingual term alignment from comparable corpora in English discharge summary and Chinese discharge summary
abstract
BACKGROUND: Electronic medical record (EMR) systems have become widely used throughout the world to improve the quality of healthcare and the efficiency of hospital services. A bilingual medical lexicon of Chinese and English is needed to meet the demand for the multi-lingual and multi-national treatment. We make efforts to extract a bilingual lexicon from English and Chinese discharge summaries with a small seed lexicon. The lexical terms can be classified into two categories: single-word terms (SWTs) and multi-word terms (MWTs). For SWTs, we use a label propagation (LP; context-based) method to extract candidates of translation pairs. For MWTs, which are pervasive in the medical domain, we propose a term alignment method, which firstly obtains translation candidates for each component word of a Chinese MWT, and then generates their combinations, from which the system selects a set of plausible translation candidates. RESULTS: We compare our LP method with a baseline method based on simple context-similarity. The LP based method outperforms the baseline with the accuracies: 4.44% Acc1, 24.44% Acc10, and 62.22% Acc100, where AccN means the top N accuracy. The accuracy of the LP method drops to 5.41% Acc10 and 8.11% Acc20 for MWTs. Our experiments show that the method based on term alignment improves the performance for MWTs to 16.22% Acc10 and 27.03% Acc20. CONCLUSIONS: We constructed a framework for building an English-Chinese term dictionary from discharge summaries in the two languages. Our experiments have shown that the LP-based method augmented with the term alignment method will contribute to reduction of manual work required to compile a bilingual sydictionary of clinical terms.
Yan Xu 0001, Luoxin Chen, Junsheng Wei, Sophia Ananiadou, Yubo Fan, Eric I-Chao Chang, Jun'ichi Tsujii
BMC Bioinform.1
2015 Unsupervised Object Class Discovery via Saliency-Guided Multiple Class Learning
abstract
In this paper, we tackle the problem of common object (multiple classes) discovery from a set of input images, where we assume the presence of one object class in each image. This problem is, loosely speaking, unsupervised since we do not know a priori about the object type, location, and scale in each image. We observe that the general task of object class discovery in a fully unsupervised manner is intrinsically ambiguous; here we adopt saliency detection to propose candidate image windows/patches to turn an unsupervised learning problem into a weakly-supervised learning problem. In the paper, we propose an algorithm for simultaneously localizing objects and discovering object classes via bottom-up (saliency-guided) multiple class learning (bMCL). Our contributions are three-fold: (1) we adopt saliency detection to convert unsupervised learning into multiple instance learning, formulated as bottom-up multiple class learning (bMCL); (2) we propose an integrated framework that simultaneously performs object localization, object class discovery, and object detector training; (3) we demonstrate that our framework yields significant improvements over existing methods for multi-class object discovery and possess evident advantages over competing methods in computer vision. In addition, although saliency detection has recently attracted much attention, its practical usage for high-level vision tasks has yet to be justified. Our method validates the usefulness of saliency detection to output "noisy input" for a top-down method to extract common patterns.
Jun-Yan Zhu, Jiajun Wu 0001, Yan Xu 0001, Eric I-Chao Chang, Zhuowen Tu
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Deep learning of feature representation with multiple instance learning for medical image analysis
abstract
This paper studies the effectiveness of accomplishing high-level tasks with a minimum of manual annotation and good feature representations for medical images. In medical image analysis, objects like cells are characterized by significant clinical features. Previously developed features like SIFT and HARR are unable to comprehensively represent such objects. Therefore, feature representation is especially important. In this paper, we study automatic extraction of feature representation through deep learning (DNN). Furthermore, detailed annotation of objects is often an ambiguous and challenging task. We use multiple instance learning (MIL) framework in classification training with deep learning features. Several interesting conclusions can be drawn from our work: (1) automatic feature learning outperforms manual feature; (2) the unsupervised approach can achieve performance that's close to fully supervised approach (93.56%) vs. (94.52%); and (3) the MIL performance of coarse label (96.30%) outweighs the supervised performance of fine label (95.40%) in supervised deep learning features.
Yan Xu 0001, Tao Mo, Qiwei Feng, Peilin Zhong, Maode Lai, Eric I-Chao Chang
ICASSP1
2014 Weakly supervised histopathology cancer image segmentation and classification
Yan Xu 0001, Jun-Yan Zhu, Eric I-Chao Chang, Maode Lai, Zhuowen Tu
Medical Image Anal.1
2013 An end-to-end system to identify temporal relation in discharge summaries: 2012 i2b2 challenge
abstract
OBJECTIVE: To create an end-to-end system to identify temporal relation in discharge summaries for the 2012 i2b2 challenge. The challenge includes event extraction, timex extraction, and temporal relation identification. DESIGN: An end-to-end temporal relation system was developed. It includes three subsystems: an event extraction system (conditional random fields (CRF) name entity extraction and their corresponding attribute classifiers), a temporal extraction system (CRF name entity extraction, their corresponding attribute classifiers, and context-free grammar based normalization system), and a temporal relation system (10 multi-support vector machine (SVM) classifiers and a Markov logic networks inference system) using labeled sequential pattern mining, syntactic structures based on parse trees, and results from a coordination classifier. Micro-averaged precision (P), recall (R), averaged P&R (P&R), and F measure (F) were used to evaluate results. RESULTS: For event extraction, the system achieved 0.9415 (P), 0.8930 (R), 0.9166 (P&R), and 0.9166 (F). The accuracies of their type, polarity, and modality were 0.8574, 0.8585, and 0.8560, respectively. For timex extraction, the system achieved 0.8818, 0.9489, 0.9141, and 0.9141, respectively. The accuracies of their type, value, and modifier were 0.8929, 0.7170, and 0.8907, respectively. For temporal relation, the system achieved 0.6589, 0.7129, 0.6767, and 0.6849, respectively. For end-to-end temporal relation, it achieved 0.5904, 0.5944, 0.5921, and 0.5924, respectively. With the F measure used for evaluation, we were ranked first out of 14 competing teams (event extraction), first out of 14 teams (timex extraction), third out of 12 teams (temporal relation), and second out of seven teams (end-to-end temporal relation). CONCLUSIONS: The system achieved encouraging results, demonstrating the feasibility of the tasks defined by the i2b2 organizers. The experiment result demonstrates that both global and local information is useful in the 2012 challenge.
Yan Xu 0001, Tianren Liu, Jun'ichi Tsujii, Eric I-Chao Chang
J. Am. Medical Informatics Assoc.1
2012 Multiple clustered instance learning for histopathology cancer image classification, segmentation and clustering
abstract
Cancer tissues in histopathology images exhibit abnormal patterns; it is of great clinical importance to label a histopathology image as having cancerous regions or not and perform the corresponding image segmentation. However, the detailed annotation of cancer cells is often an ambiguous and challenging task. In this paper, we propose a new learning method, multiple clustered instance learning (MCIL), to classify, segment and cluster cancer cells in colon histopathology images. The proposed MCIL method simultaneously performs image-level classification (cancer vs. non-cancer image), pixel-level segmentation (cancer vs. non-cancer tissue), and patch-level clustering (cancer subclasses). We embed the clustering concept into the multiple instance learning (MIL) setting and derive a principled solution to perform the above three tasks in an integrated framework. Experimental results demonstrate the efficiency and effectiveness of MCIL in analyzing colon cancers.
Yan Xu 0001, Jun-Yan Zhu, Eric I-Chao Chang, Zhuowen Tu
CVPR1
2012 Context-Constrained Multiple Instance Learning for Histopathology Image Segmentation
Yan Xu 0001, Eric I-Chao Chang, Maode Lai, Zhuowen Tu
MICCAI (3)1
2012 Feature engineering combined with machine learning and rule-based methods for structured information extraction from narrative clinical discharge summaries
abstract
OBJECTIVE: A system that translates narrative text in the medical domain into structured representation is in great demand. The system performs three sub-tasks: concept extraction, assertion classification, and relation identification. DESIGN: The overall system consists of five steps: (1) pre-processing sentences, (2) marking noun phrases (NPs) and adjective phrases (APs), (3) extracting concepts that use a dosage-unit dictionary to dynamically switch two models based on Conditional Random Fields (CRF), (4) classifying assertions based on voting of five classifiers, and (5) identifying relations using normalized sentences with a set of effective discriminating features. MEASUREMENTS: Macro-averaged and micro-averaged precision, recall and F-measure were used to evaluate results. RESULTS: The performance is competitive with the state-of-the-art systems with micro-averaged F-measure of 0.8489 for concept extraction, 0.9392 for assertion classification and 0.7326 for relation identification. CONCLUSIONS: The system exploits an array of common features and achieves state-of-the-art performance. Prudent feature engineering sets the foundation of our systems. In concept extraction, we demonstrated that switching models, one of which is especially designed for telegraphic sentences, improved extraction of the treatment concept significantly. In assertion classification, a set of features derived from a rule-based classifier were proven to be effective for the classes such as conditional and possible. These classes would suffer from data scarcity in conventional machine-learning methods. In relation identification, we use two-staged architecture, the second of which applies pairwise classifiers to possible candidate classes. This architecture significantly improves performance.
Yan Xu 0001, Kai Hong, Jun'ichi Tsujii, Eric I-Chao Chang
J. Am. Medical Informatics Assoc.1
2012 A classification approach to coreference in discharge summaries: 2011 i2b2 challenge
abstract
OBJECTIVE: To create a highly accurate coreference system in discharge summaries for the 2011 i2b2 challenge. The coreference categories include Person, Problem, Treatment, and Test. DESIGN: An integrated coreference resolution system was developed by exploiting Person attributes, contextual semantic clues, and world knowledge. It includes three subsystems: Person coreference system based on three Person attributes, Problem/Treatment/Test system based on numerous contextual semantic extractors and world knowledge, and Pronoun system based on a multi-class support vector machine classifier. The three Person attributes are patient, relative and hospital personnel. Contextual semantic extractors include anatomy, position, medication, indicator, temporal, spatial, section, modifier, equipment, operation, and assertion. The world knowledge is extracted from external resources such as Wikipedia. MEASUREMENTS: Micro-averaged precision, recall and F-measure in MUC, BCubed and CEAF were used to evaluate results. RESULTS: The system achieved an overall micro-averaged precision, recall and F-measure of 0.906, 0.925, and 0.915, respectively, on test data (from four hospitals) released by the challenge organizers. It achieved a precision, recall and F-measure of 0.905, 0.920 and 0.913, respectively, on test data without Pittsburgh data. We ranked the first out of 20 competing teams. Among the four sub-tasks on Person, Problem, Treatment, and Test, the highest F-measure was seen for Person coreference. CONCLUSIONS: This system achieved encouraging results. The Person system can determine whether personal pronouns and proper names are coreferent or not. The Problem/Treatment/Test system benefits from both world knowledge in evaluating the similarity of two mentions and contextual semantic extractors in identifying semantic clues. The Pronoun system can automatically detect whether a Pronoun mention is coreferent to that of the other four types. This study demonstrates that it is feasible to accomplish the coreference task in discharge summaries.
Yan Xu 0001, Jiahua Liu, Jiajun Wu 0001, Yue Wang 0035, Zhuowen Tu, Jian-Tao Sun, Jun'ichi Tsujii, Eric I-Chao Chang
J. Am. Medical Informatics Assoc.1
2012 Named entity recognition of follow-up and time information in 20 000 radiology reports
abstract
OBJECTIVE: To develop a system to extract follow-up information from radiology reports. The method may be used as a component in a system which automatically generates follow-up information in a timely fashion. METHODS: A novel method of combining an LSP (labeled sequential pattern) classifier with a CRF (conditional random field) recognizer was devised. The LSP classifier filters out irrelevant sentences, while the CRF recognizer extracts follow-up and time phrases from candidate sentences presented by the LSP classifier. MEASUREMENTS: The standard performance metrics of precision (P), recall (R), and F measure (F) in the exact and inexact matching settings were used for evaluation. RESULTS: Four experiments conducted using 20,000 radiology reports showed that the CRF recognizer achieved high performance without time-consuming feature engineering and that the LSP classifier further improved the performance of the CRF recognizer. The performance of the current system is P=0.90, R=0.86, F=0.88 in the exact matching setting and P=0.98, R=0.93, F=0.95 in the inexact matching setting. CONCLUSION: The experiments demonstrate that the system performs far better than a baseline rule-based system and is worth considering for deployment trials in an alert generation system. The LSP classifier successfully compensated for the inherent weakness of CRF, that is, its inability to use global information.
Yan Xu 0001, Jun'ichi Tsujii, Eric I-Chao Chang
J. Am. Medical Informatics Assoc.1
2010 Volumetric Topological Analysis: A Novel Approach for Trabecular Bone Classification on the Continuum Between Plates and Rods
abstract
Trabecular bone (TB) is a complex quasi-random network of interconnected plates and rods. TB constantly remodels to adapt to the stresses to which it is subjected (Wolff's Law). In osteoporosis, this dynamic equilibrium between bone formation and resorption is perturbed, leading to bone loss and structural deterioration. Both bone loss and structural deterioration increase fracture risk. Bone's mechanical behavior can only be partially explained by variations in bone mineral density, which led to the notion of bone structural quality. Previously, we developed digital topological analysis (DTA) which classifies plates, rods, profiles, edges, and junctions in a TB skeletal representation. Although the method has become quite popular, a major limitation of DTA is that it provides only hard classifications of different topological entities, failing to distinguish between narrow and wide plates. Here, we present a new method called volumetric topological analysis (VTA) for regional quantification of TB topology. At each TB location, the method uniquely classifies its topology on the continuum between perfect plates and perfect rods, facilitating early detections of TB alterations from plates to rods according to the known etiology of osteoporotic bone loss. Several new ideas, including manifold distance transform, manifold scale, and feature propagation have been introduced here and combined with existing DTA and distance transform methods, leading to the new VTA technology. This method has been applied to multidetector computed tomography (CT) and micro-computed tomography ( μCT) images of four cadaveric distal tibia and five distal radius specimens. Both intra- and inter-modality reproducibility of the method has been examined using repeat CT and μCT scans of distal tibia specimens. Also, the method's ability to predict experimental biomechanical properties of TB via CT imaging under in vivo conditions has been quantitatively examined and the results found are very encouraging.
Punam K. Saha, Yan Xu 0001, Hong Duan, Anneliese Heiner, Guoyuan Liang
IEEE Trans. Medical Imaging2