VLDB 2026 Research / reviewers in the wild / expert
Lin Yang 0002
dblp:20/2970-2
· DBLP profile ↗
123ranked-venue papers
11as first author
46since 2021 · last 2026
0009-0005-0742-0280ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 71 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 60 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 44 · 2 first-author · 19 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional RadiologyabstractWe present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun 0002, Yixuan Si, Lin Yang 0002, Liang Xiao 0001 |
AAAI | 9 |
| 2026 | Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide ImagesabstractNucleus detection in histopathology whole slide images (WSIs) is crucial for a broad spectrum of clinical applications. The gigapixel size of WSIs necessitates the use of sliding window methodology for nucleus detection. However, mainstream methods process each sliding window independently, which overlooks broader contextual information and easily leads to inaccurate predictions. To address this limitation, recent studies additionally crop a large Filed-of-View (LFoV) patch centered on each sliding window to extract contextual features. However, such methods substantially increase whole-slide inference latency. In this work, we propose an effective and efficient context-aware nucleus detection approach. Specifically, instead of using lFoV patches, we aggregate contextual clues from off-the-shelf features of historically visited sliding windows, which greatly enhances the inference efficiency. Moreover, compared to lFoV patches used in previous works, the sliding window patches have higher magnification and provide finer-grained tissue details, thereby enhancing the classification accuracy. To develop the proposed context-aware model, we utilize annotated patches along with their surrounding unlabeled patches for training. Beyond exploiting high-level tissue context from these surrounding regions, we design a post-training strategy that leverages abundant unlabeled nucleus samples within them to enhance the model's context adaptability. Extensive experimental results on three challenging benchmarks demonstrate the superiority of our method. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Yiwen Ye, Pingyi Chen, Ruizhe Guo, Lei Cui 0004, Chenglu Zhu, Lin Yang 0002 |
AAAI | 10 |
| 2026 | DFFormer: Dual Frequency-Driven Transformer for real-world image deblurring
Ruizhe Guo, Shichuan Zhang, Jingxiong Li, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
Comput. Vis. Image Underst. | 6 |
| 2025 | CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyabstractThe emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs and resulting in redundant models. In this work, we introduce CPath-Omni, the first 15B parameter LMM that unifies patch and WSI analysis, consolidating a variety of tasks at both levels, including classification, visual question answering, captioning, and visual referring prompting. Extensive experiments demonstrate that CPath-Omni achieves state-of-the-art (SOTA) performance across seven diverse tasks on 39 out of 42 datasets, outperforming or matching task-specific models trained for individual tasks. Additionally, we develop a specialized pathology CLIP-based visual processor for CPath-Omni, CPath-CLIP, which, for the first time, integrates different vision models and incorporates a large language model as a text encoder to build a more powerful CLIP model, which achieves SOTA performance on nine zero-shot and four few-shot datasets. Our findings highlight CPath-Omni’s ability to unify diverse pathology tasks, demonstrating its potential to streamline and advance the field of foundation model in pathology. The code and model are available at CPath-Omni. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Pingyi Chen, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
CVPR | 10 |
| 2025 | Stable Test-Time Training for Semantic Segmentation with Output Contrastive LossabstractDeep learning-based models have achieved impressive performance on public segmentation benchmarks, yet generalizing to unseen environments remains challenging. Test-time training (TTT) addresses this by adapting source-pretrained models during evaluation. While existing TTT methods have shown promise in image classification, they often exhibit instability with small test batches and class imbalance—challenges that intensify in semantic segmentation tasks. To tackle this issue, we present Output Contrastive Loss (OCL) to improve the stability of contrastive loss when applied to TTT for segmentation. OCL applies contrastive loss directly to the output space, avoiding the need for extra regularization, and employs a high temperature to prevent model collapse. To further stabilize the TTT process, we integrate BN statistics Modulation and Stochastic Restoration techniques. Extensive experiments across diverse datasets, settings, architectures, and pretrained methods demonstrate consistent performance improvements, achieving a 7.5 mIoU gain on the GTA→CS benchmark and showing effectiveness even with domain adaptation pretraining. Code is available at https://github.com/dazhangyul23/OCL. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Chenglu Zhu, Lin Yang 0002 |
ICASSP | 6 |
| 2025 | Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingabstractArtificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities, leading to ambiguous patient-level pairings. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease (including several most deadly cancers) diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively. Additionally, on the publicly available CT-RATE and Rad-ChestCT benchmarks, our fVLM outperformed the current state-of-the-art methods with absolute AUC gains of 7.4% and 4.8%, respectively. Zhongyi Shui, Sinuo Wang, Ruizhe Guo, Le Lu 0001, Lin Yang 0002, Xianghua Ye, Tingbo Liang, Ling Zhang 0002 |
ICLR | 7 |
| 2025 | PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationabstractVision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as vision encoders when combined with large language models (LLMs) to support broader capabilities. Current efforts to train pathology VLMs rely on pathology image-text pairs from platforms like PubMed, YouTube, and Twitter, which provide limited, unscalable data with generally suboptimal image quality. In this work, we leverage large-scale WSI datasets like TCGA to extract numerous high-quality image patches. We then train a large multimodal model (LMM) to generate captions for extracted images, creating PathGen-1.6M, a dataset containing 1.6 million high-quality image-caption pairs. Our approach involves multiple agent models collaborating to extract representative WSI patches, generating and refining captions to obtain high-quality image-text pairs. Extensive experiments show that integrating these generated pairs with existing datasets to train a pathology-specific CLIP model, PathGen-CLIP, significantly enhances its ability to analyze pathological images, with substantial improvements across nine pathology-related zero-shot image classification tasks and three whole-slide image tasks. Furthermore, we construct 200K instruction-tuning data based on PathGen-1.6M and integrate PathGen-CLIP with the Vicuna LLM to create more powerful multimodal models through instruction tuning. Overall, we provide a scalable pathway for high-quality data generation in pathology, paving the way for next-generation general pathology models. Our dataset, code, and model are open-access at https://github.com/PathFoundation/PathGen-1.6M. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ICLR | 11 |
| 2025 | AEM: Attention Entropy Maximization for Multiple Instance Learning Based Whole Slide Image Classification
Honglin Li 0001, Yuxuan Sun 0002, Zhongyi Shui, Jingxiong Li, Chenglu Zhu, Lin Yang 0002 |
MICCAI (7) | 7 |
| 2025 | PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationabstractPathology whole slide image (WSI) analysis is vital for disease diagnosis and understanding. While foundation models (FMs) have driven recent advances, their scalability in pathology remains a key challenge. In particular, vision-language (VL) pathology FMs align visual features with language annotation for downstream tasks, but they rely heavily on large-scale image-text paired data, which is scarce thus limiting generalization. On the other hand, vision-only pathology FMs can leverage abundant unlabeled data via self-supervised learning (SSL). However, current approaches often use the [CLS] token from tile-level ViTs as slide-level input for efficiency (a tile with 224×224 pixels composed of 196 patches with 16×16 pixels). This SSL pretrained [CLS] token lacks alignment with downstream objectives, limiting effectiveness. We find that spatial patch tokens retain a wealth of informative features beneficial for downstream tasks, but utilizing all of them incurs up to 200× higher computation and storage costs compared [CLS] token only (e.g., 196 tokens per ViT$_{224}$). This highlights a fundamental trade-off between efficiency and representational richness to build scalable pathology FMs. To address this, we propose a feature distillation framework via vector-quantization (VQ) that compresses patch tokens into discrete indices and reconstructs them via a decoder, achieving 64× compression (1024 → 16 dimensions) while preserving fidelity. We further introduce a multi-scale VQ (MSVQ) strategy, enhancing both reconstruction and providing SSL supervision for slide-level pretraining. Built upon MSVQ features and supervision signals, we design a progressive convolutional module and a slide-level SSL objective to learn spatially rich representations for downstream WSI tasks. Extensive experiments across multiple datasets demonstrate that our approach achieves state-of-the-art performance, offering a scalable and effective solution for high-performing pathology FMs in WSI analysis. Honglin Li 0001, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 5 |
| 2025 | CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicabstractRecent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply multimodal approaches to generate reports directly from images. However, these models cannot emulate the diagnostic approach of pathologists, who systematically examine slides at low magnification to obtain an overview before progressively zooming in on suspicious regions to formulate comprehensive diagnoses. Instead, existing models directly output final diagnoses without revealing the underlying reasoning process.
To address this gap, we introduce CPathAgent, an innovative agent-based approach that mimics pathologists' diagnostic workflow by autonomously navigating across WSI through zoom-in/out and move operations based on observed visual features, thereby generating substantially more transparent and interpretable diagnostic summaries. To achieve this, we develop a multi-stage training strategy that unifies patch-level, region-level, and WSI-level capabilities within a single model, which is essential for replicating how pathologists understand and reason across diverse image scales.
Additionally, we construct PathMMU-HR², the first expert-validated benchmark for large region analysis. This represents a critical intermediate scale between patches and whole slides, reflecting a key clinical reality where pathologists typically examine several key large regions rather than entire slides at once. Extensive experiments demonstrate that CPathAgent consistently outperforms existing approaches across benchmarks at three different image scales, validating the effectiveness of our agent-based diagnostic approach and highlighting a promising direction for computational pathology. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
NeurIPS | 8 |
| 2025 | Exploring Unbiased Activation Maps for Weakly Supervised Tissue Segmentation of Histopathological ImagesabstractTissue segmentation in histopathological images plays a crucial role in computational pathology, owing to its significant potential to indicate the prognosis of cancer patients. Presently, numerous Weakly Supervised Semantic Segmentation (WSSS) methods strive to utilize image-level labels to achieve pixel-level segmentation, aiming to minimize the need for detailed annotations. Most of these methods rely on Class Activation Maps (CAM) extracted from classification models, frequently leading to poor coverage of objects. The major cause is attributed to the strong inductive bias of the classification model, focusing primarily on discriminative feature of objects, rather than non-discriminative features. Inspired by this, we propose a simple yet effective method that introduces a self-supervised task by exploiting both the discriminative and non-discriminative features, and generate Unbiased Activation Maps (UAM) to encompass the whole object. Specifically, our method entails clustering all spatial features of an object class to derive semantic centers. Each center then works as a spatial filter that amplifies similar feature and suppresses dissimilar feature, and extract high-quality pseudo-labels (some noise at object boundaries). Moreover, we further propose a Noise-Reduced (NR) Learning method to train the segmentation network towards credible signals and lessen the impact of false predictions. Comprehensive experimental results on two public histopathology image datasets demonstrate the superior performance of our method over the state-of-the-art weakly supervised segmentation methods. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Xiao Zhang 0028, Yaqiong Xing, Yuting Wen, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 10 |
| 2025 | ToPoFM: Topology-Guided Pathology Foundation Model for High-Resolution Pathology Image Synthesis With Cellular-Level ControlabstractSynthetic data generation emerges as a strategy to mitigate data scarcity in digital pathology, where complicated tissue and cellular features are correlated with cancer diagnosis. The synthesis of such visuals, however, suffers from limited inter class diversity and scarcity of cellular annotations. Current methodologies struggle with capturing the broad spectrum of pathology features, causing unpredictable objects and defected fidelity. Moreover, discrepancies in image resolution across developmental and operational phases can amplify the distribution shifts, undermining the precision of diagnosis. To address these challenges, we introduce TOpology guided PathOlogy Foundation Model (ToPoFM), a visual foundation model designed for the synthesis of high-resolution pathology images with cellular-level control. Our approach integrates a topology-informed cell arrangement generator to steer large language models for crafting synthetic cell arrangements. We correlate cell arrangement guidance with diffusion model for pathology content generation, then further implement a random sliding inference strategy, merging discrete low-resolution samplings into single high-resolution representation. Our model requires only small patches for training. The efficacy of ToPoFM is demonstrated through extensive experiments, complemented by expert validations, showing high fidelity on data synthesis. Additionally, we underscore the utility of our generated imagery as an augmentation tool, enhancing the performance of downstream tasks, including cancer subtype classification and segmentation. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Pingyi Chen, Yuxuan Sun 0002, Honglin Li 0001, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | PathBench: Advancing the Benchmark of Large Multimodal Models for Pathology Image Understanding at Patch and Whole Slide LevelabstractRapid advancements in large multimodal models (LMMs) have significantly enhanced their applications in pathology, particularly in image classification, pathology image description, and whole slide image (WSI) classification. In pathology, WSIs represent gigapixel-scale images composed of thousands of image patches. Therefore, both patch-level and WSI-level evaluations are essential and inherently interconnected for assessing LMM capabilities. In this work, we propose PathBench, which comprises three subsets at both patch and WSI levels, to refine and enhance the validation of LMMs. At the patch-level, evaluations using existing multi-choice Q&A datasets reveal that some LMMs can predict answers without genuine image analysis. To address this, we introduce PatchVQA, a large-scale visual question answering (VQA) dataset containing 5,382 images and 6,335 multiple-choice questions designed with distractor options to prevent shortcut learning. These new questions are rigorously validated by professional pathologists to ensure reliable model assessments. At the WSI-level, current efforts primarily focus on image classification tasks and lack diverse validation datasets for multimodal models. To address this, we generate a detailed WSI report dataset through an innovative approach that integrates detailed patch descriptions generated by foundational models into comprehensive WSI reports. These are then combined with physician-written reports corresponding to TCGA WSIs, resulting in WSICap, a detailed report dataset containing 7,000 samples. Based on WSICap, we further develop a WSI-level VQA dataset, WSIVQA, to serve as a validation set for WSI LMMs. Using these PathBench subsets, we conduct extensive experiments to benchmark the performance of state-of-the-art LMMs at both the patch and WSI levels. The proposed dataset is available at https://github.com/superjamessyx/PathBench. Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Yixuan Si, Qizi Chen, Kai Zhang 0033, Jingxiong Li, Jiatong Cai, Lin Sun 0006, Tao Lin 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 13 |
| 2024 | DPA-P2PNet: Deformable Proposal-Aware P2PNet for Accurate Point-Based Cell DetectionabstractPoint-based cell detection (PCD), which pursues high-performance cell sensing under low-cost data annotation, has garnered increased attention in computational pathology community. Unlike mainstream PCD methods that rely on intermediate density map representations, the Point-to-Point network (P2PNet) has recently emerged as an end-to-end solution for PCD, demonstrating impressive cell detection accuracy and efficiency. Nevertheless, P2PNet is limited to decoding from a single-level feature map due to the scale-agnostic property of point proposals, which is insufficient to leverage multi-scale information. Moreover, the spatial distribution of pre-set point proposals is biased from that of cells, leading to inaccurate cell localization. To lift these limitations, we present DPA-P2PNet in this work. The proposed method directly extracts multi-scale features for decoding according to the coordinates of point proposals on hierarchical feature maps. On this basis, we further devise deformable point proposals to mitigate the positional bias between proposals and potential cells to promote cell localization. Inspired by practical pathological diagnosis that usually combines high-level tissue structure and low-level cell morphology for accurate cell classification, we propose a multi-field-of-view (mFoV) variant of DPA-P2PNet to accommodate additional large FoV images with tissue information as model input. Finally, we execute the first self-supervised pre-training on immunohistochemistry histopathology image data and evaluate the suitability of four representative self-supervised methods on the PCD task. Experimental results on three benchmarks and a large-scale and real-world interval dataset demonstrate the superiority of our proposed models over the state-of-the-art counterparts. Codes and pre-trained weights are available at https://github.com/windygoo/DPA-P2PNet. Zhongyi Shui, Sunyi Zheng, Chenglu Zhu, Shichuan Zhang, Xiaoxuan Yu, Honglin Li 0001, Jingxiong Li, Pingyi Chen, Lin Yang 0002 |
AAAI | 9 |
| 2024 | PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyabstractAs advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and designing comprehensive model frameworks. To bridge the gap in pathology MLLMs, we present PathAsst, a multimodal generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. The development of PathAsst involves three pivotal steps: data acquisition, CLIP model adaptation, and the training of PathAsst's multimodal generative capabilities. Firstly, we collect over 207K high-quality pathology image-text pairs from authoritative sources. Leveraging the advanced power of ChatGPT, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data specifically tailored for invoking eight pathology-specific sub-models we prepared, allowing the PathAsst to effectively collaborate with these models, enhancing its diagnostic ability. Secondly, by leveraging the collected data, we construct PathCLIP, a pathology-dedicated CLIP, to enhance PathAsst's capabilities in interpreting pathology images. Finally, we integrate PathCLIP with the Vicuna-13b and utilize pathology-specific instruction-tuning data to enhance the multimodal generation capacity of PathAsst and bolster its synergistic interactions with sub-models. The experimental results of PathAsst show the potential of harnessing AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We open-source our dataset, as well as a comprehensive toolkit for extensive pathology data collection and preprocessing at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Kai Zhang 0033, Lin Sun 0006, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
AAAI | 9 |
| 2024 | WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ECCV (36) | 5 |
| 2024 | Unleashing the Power of Prompt-Driven Nucleus Instance Segmentation
Zhongyi Shui, Chenglu Zhu, Sunyi Zheng, Jingxiong Li, Honglin Li 0001, Yuxuan Sun 0002, Ruizhe Guo, Lin Yang 0002 |
ECCV (27) | 10 |
| 2024 | PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology
Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Sunyi Zheng, Qizi Chen, Kai Zhang 0033, Dan Wan, Xiaoxiao Lan, Mengyue Zheng, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ECCV (62) | 14 |
| 2024 | Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
Honglin Li 0001, Yunxuan Sun, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
ECCV (53) | 6 |
| 2024 | HLS-FGVC: Hierarchical Label Semantics Enhanced Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) intends to confirm the sub-classes of a specific object category, e.g., identifying the species of dogs or birds. It is a challenging problem with the inter-class similarity among these sub-categories and intra-class variance in every fine-grained class. Most of the recent works intend to learn discriminative representations and class-consistency features. However, they only take the finest labels into account. We argue that the hierarchical label structure (HLS) implied in the category names can enhance the FGVC task. In this paper, we proposed two modules to leverage the hierarchical label structure. (i) We build a weighted graph in each batch based on the hierarchical label structure, the nodes of which are image features. The messages are passed among graph nodes for feature interaction. (ii) A hierarchy-aware ranking loss is proposed to regularize the distribution in feature space. The ablation study and experimental results show that our proposed modules achieve significant improvements over previous works. Shichuan Zhang, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002 |
ICASSP | 4 |
| 2024 | Context-Aware Text-Assisted Multimodal Framework for Cervical Cytology Cell Diagnosis and ChattingabstractRecent advancements underscore the potential of deep learning-based Computer-Assisted Diagnosis (CAD) systems for cervical cytology image analysis. However, traditional methods focusing solely on single-view of cells fall short in performance due to the lack of contextual information. Moreover, the unclear reasoning behind model’s classification hinders their interpretability. To overcome these issues, we present Cervi-CAT, a context-aware, text-assisted multimodal framework for cervical cytology cell classification. CerviCAT captures visual cell representations from both global and local perspectives and subsequently generates textual descriptions based on the visual representation. A multimodal transformer then integrates these descriptions with visual features for interpretable and accurate cell classification. Additionally, we introduce Cyto-Vicuna, a cytology-specific large language model fine-tuned based on Vicuna-7b using collected cytology-specific data. When integrated into CerviCAT, it produces more detailed diagnostic reports while simultaneously fostering interaction between the model and cytologists, promoting collaborative diagnosis. Our results demonstrate that CerviCAT not only surpasses traditional CAD methods in performance but also provides interpretable diagnosis. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ICME | 6 |
| 2024 | WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
Pingyi Chen, Honglin Li 0001, Chenglu Zhu, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002 |
MICCAI (4) | 6 |
| 2024 | PathUp: Patch-wise Timestep Tracking for Multi-class Large Pathology Image Synthesising Diffusion ModelabstractIn digital pathology, cancer lesions are identified by analyzing the spatial context within pathology images. Synthesizing such complex spatial context is challenging as pathology whole slide images typically exhibit high resolution, low inter-class variety, and are sparsely labeled. To address these challenges, we propose PathUp, a novel diffusion model tailored for the synthesis of multi-class high-resolution pathology images. Our approach includes a latent space patch-wise timestep tracking, which helps to generate high-quality images without tiling artifacts. Pathology knowledge is integrated through our patho-align. The robust generation of lesion subtypes and scale information is ensured by introducing a feature entropy loss. The effectiveness of our method is evaluated through extensive experiments, supplemented by assessments from human experts, demonstrating the authenticity of the synthetic data produced. Furthermore, we highlight the potential utility of our generated images as an augmentation method, thereby enhancing the performance of downstream tasks such as cancer subtype classification. Jingxiong Li, Sunyi Zheng, Chenglu Zhu, Yuxuan Sun 0002, Pingyi Chen, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
ACM Multimedia | 9 |
| 2024 | Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisabstractHistopathology Whole Slide Image (WSI) analysis serves as the gold standard for clinical cancer diagnosis in the daily routines of doctors. To develop computer-aided diagnosis model for histopathology WSIs, previous methods typically employ Multi-Instance Learning to enable slide-level prediction given only slide-level labels.
Among these models, vanilla attention mechanisms without pairwise interactions have traditionally been employed but are unable to model contextual information. More recently, self-attention models have been utilized to address this issue. To alleviate the computational complexity of long sequences in large WSIs, methods like HIPT use region-slicing, and TransMIL employs Nystr\"{o}mformer as an approximation of full self-attention. Both approaches suffer from suboptimal performance due to the loss of key information. Moreover, their use of absolute positional embedding struggles to effectively handle long contextual dependencies in shape-varying WSIs.
In this paper, we first analyze how the low-rank nature of the long-sequence attention matrix constrains the representation ability of WSI modelling. Then, we demonstrate that the rank of attention matrix can be improved by focusing on local interactions via a local attention mask. Our analysis shows that the local mask aligns with the attention patterns in the lower layers of the Transformer. Furthermore, the local attention mask can be implemented during chunked attention calculation, reducing the quadratic computational complexity to linear with a small local bandwidth. Additionally, this locality helps the model generalize to unseen or under-fitted positions more easily.
Building on this, we propose a local-global hybrid Transformer for both computational acceleration and local-global information interactions modelling. Our method, Long-contextual MIL (LongMIL), is evaluated through extensive experiments on various WSI tasks to validate its superiority in: 1) overall performance, 2) memory usage and speed, and 3) extrapolation ability compared to previous methods. Honglin Li 0001, Pingyi Chen, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 6 |
| 2024 | Gradient-aware learning for joint biases: Label noise and class imbalance
Shichuan Zhang, Chenglu Zhu, Honglin Li 0001, Jiatong Cai, Lin Yang 0002 |
Neural Networks | 5 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 12 |
| 2023 | Addressing Sparse Annotation: a Novel Semantic Energy Loss for Tumor Cell Detection from Histopathologic ImagesabstractTumor cell detection plays a vital role in immunohistochemistry (IHC) quantitative analysis. While recent remarkable developments in fully-supervised deep learning have greatly contributed to the efficiency of this task, the necessity for manually annotating all cells of specific detection types remains impractical. Obviously, if we directly use full supervision to train these datasets, it can cause error in loss calculation due to the misclassification of unannotated cells as background. To address this issue, we observe that although some cells are omitted during the annotation process, these unannotated cells have a significant feature similarity with the annotated ones. Leveraging this characteristic, we propose a novel calibrated loss named Semantic Energy Loss (SEL). Specifically, our SEL automatically adjusts the loss to be lower for unannotated regions with similar semantic to the labeled ones, while penalizing regions with lager semantic difference. Besides, to prevent all regions from having similar semantics during training, we propose Stretched Feature Loss (SFL) that widen the semantic distance. We evaluate our method on two different IHC datasets and achieve significant performance improvements in both sparse and exhaustive annotation scenarios. Furthermore, we also validate that our method holds significant potential for detecting multiple types of cells. Our code is available at here. Xianglong Du, Yuxin Kang, Hong Lv, Lei Cui 0004, Hansheng Li, Yaqiong Xing, Jun Feng 0003, Lin Yang 0002 |
BIBM | 11 |
| 2023 | Multi-modal Learning with Missing Modality in Predicting Axillary Lymph Node MetastasisabstractMulti-modal Learning has attracted widespread attention in medical image analysis. Using multi-modal data, whole slide images (WSIs) and clinical information, can improve the performance of deep learning models in the diagnosis of axillary lymph node metastasis. However, clinical information is not easy to collect in clinical practice due to privacy concerns, limited resources, lack of interoperability, etc. Although patient selection can ensure the training set to have multi-modal data for model development, missing modality of clinical information can appear during test. This normally leads to performance degradation, which limits the use of multi-modal models in the clinic. To alleviate this problem, we propose a bidirectional distillation framework consisting of a multi-modal branch and a single-modal branch. The single-modal branch acquires the complete multi-modal knowledge from the multi-modal branch, while the multi-modal learns the robust features of WSI from the single-modal. We conduct experiments on a public dataset of Lymph Node Metastasis in Early Breast Cancer to validate the method. Our approach not only achieves state-of-the-art performance with an AUC of 0.861 on the test set without missing data, but also yields an AUC of 0.842 when the rate of missing modality is 80%. This shows the effectiveness of the approach in dealing with multi-modal data and missing modality. Such a model has the potential to improve treatment decision-making for early breast cancer patients who have axillary lymph node metastatic status. Shichuan Zhang, Sunyi Zheng, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
BIBM | 5 |
| 2023 | Task-Specific Fine-Tuning via Variational Information Bottleneck for Weakly-Supervised Pathology Whole Slide Image ClassificationabstractWhile Multiple Instance Learning (MIL) has shown promising results in digital Pathology Whole Slide Image (WSI) analysis, such a paradigm still faces performance and generalization problems due to high computational costs and limited supervision of Gigapixel WSIs. To deal with the computation problem, previous methods utilize a frozen model pretrained from ImageNet to obtain representations, however, it may lose key information owing to the large domain gap and hinder the generalization ability without image-level training-time augmentation. Though Self-supervised Learning (SSL) proposes viable representation learning schemes, the downstream task-specific features via partial label tuning are not explored. To alleviate this problem, we propose an efficient WSI fine-tuning framework motivated by the Information Bottleneck theory. The theory enables the framework to find the minimal sufficient statistics of WSI, thus supporting us to fine-tune the backbone into a task-specific representation only depending on WSI-level weak labels. The WSI-MIL problem is further analyzed to theoretically deduce our fine-tuning method. We evaluate the method on five pathological WSI datasets on various WSI heads. The experimental results show significant improvements in both accuracy and generalization compared with previous works. Source code will be available at https://github.com/invoker-LL/WSI-finetuning. Honglin Li 0001, Chenglu Zhu, Yuxuan Sun 0002, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, Lin Yang 0002 |
CVPR | 8 |
| 2023 | Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging SystemabstractWith the advancement of deep learning, computer-assisted clinical diagnosis, such as liquid-based cervical cytology, has attracted more attention. However, the fragile robustness of deep learning models has a non-negligible impact on their classification accuracy and reliability. To be more specific, various scanner parameters will be used depending on the pathologist’s preferences during the clinical diagnosis process (e.g., field source brightness, contrast, saturation, etc.), and this variation will lead to the unstable performance of the model. In this paper, we construct an evaluation pathway to assess the stability and consistency of deep learning models under various customized scanner parameters. Specifically, a multi-scanned dataset consists of 4200 whole slide images (WSIs) is generated by scanning 200 stained slices using various scanner parameters. Moreover, we conducted a large number of experiments to investigate the robustness of numerous models, including convolution-based and transformer-based models concerning various scanner parameter settings. Furthermore, we introduce several indicators to analyze the prediction accuracy, consistency and robustness of the model on the constructed dataset. The experimental results indicate that the deep learning models are sensitive to luminance-related scanner parameters. In addition, transformer-based models have better robustness than traditional convolutional neural networks. Our code has been made available1. Yuxuan Sun 0002, Chenglu Zhu, Honglin Li 0001, Pingyi Chen, Lin Yang 0002 |
ICASSP | 6 |
| 2023 | Exploring Unsupervised Cell Recognition with Prior Self-activation Maps
Pingyi Chen, Chenglu Zhu, Zhongyi Shui, Jiatong Cai, Sunyi Zheng, Shichuan Zhang, Lin Yang 0002 |
MICCAI (8) | 7 |
| 2023 | Segment Membranes and Nuclei from Histopathological Images via Nuclei Point-Level Supervision
Hansheng Li, Xiaoshuang Shi, Yuxin Kang, Qirong Bu, Hong Lv, Mingzhen Lin, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (6) | 13 |
| 2022 | Invariant Content Synergistic Learning for Domain Generalization on Medical Image SegmentationabstractAlthough deep convolution neural networks (DC-NNs) can achieve remarkable success on medical image segmentation, their performance might significantly deteriorate when confronting testing data with the new distribution. Recent studies suggest that one major cause of this issue is the strong inductive bias of DCNNs, which towards image styles (e.g., superficial texture) that are sensitive to change, instead of the invariant content (e.g., object shapes). Inspired by this, we propose a novel method, named Invariant Content Synergistic Learning (ICSL), to improve the generalization ability of DCNNs on unseen data by controlling the inductive bias. Specifically, ICSL first mixes the style of training instances to perturb the training distribution, so that more diverse domains or styles would be made available for training DCNNs. Then, based on the perturbed distribution, we carefully design a dual-branches invariant content synergistic learning strategy to prevent style-biased predictions and maintain the invariant content. Extensive experimental results demonstrate the superior performance of the proposed method over state-of-the-art domain generalization methods on two typical medical segmentation tasks. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Feihong Liu, Qingguo Yan, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 10 |
| 2022 | A Random Feature Augmentation for Domain Generalization in Medical Image SegmentationabstractDeep convolutional neural networks (DCNNs) significantly improve the performance of medical image segmentation. Nevertheless, medical images frequently experience distribution discrepancies, which fails to maintain their robustness when applying trained models to unseen clinical data. To address this problem, domain generalization methods were proposed to enhance the generalization ability of DCNNs. Feature space-based data augmentation methods have proven their effectiveness to improve domain generalization. However, existing methods still mainly rely on certain prior knowledge or assumption, which has limitations in enriching the diversity of source domain data. In this paper, we propose a random feature augmentation (RFA) method to diversify source domain data at the feature level without prior knowledge. Specifically, we explore the effectiveness of random convolution at the feature level for the first time and prove experimentallyt hat itc an adequately preserve domain-invariant information while perturbing domainspecific information. Furthermore, tocapture the same domain-invariant information from the augmented features of RFA, we present a domain-invariant consistent learning strategy to enable DCNNs to learn a more generalized representation. Our proposed method achieves state-of-the-art performance on two medical image segmentation tasks, including optic cup/disc segmentation on fundus images and prostate segmentation on MRI images. Yuxin Kang, Hansheng Li, Jiayu Luo, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 7 |
| 2022 | End-to-End Cell Recognition by Point Annotation
Zhongyi Shui, Shichuan Zhang, Chenglu Zhu, Bingchuan Wang, Pingyi Chen, Sunyi Zheng, Lin Yang 0002 |
MICCAI (4) | 7 |
| 2022 | Benchmarking the Robustness of Deep Neural Networks to Common Corruptions in Digital Pathology
Yuxuan Sun 0002, Honglin Li 0001, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
MICCAI (2) | 6 |
| 2022 | ChrSNet: Chromosome Straightening Using Self-attention Guided Networks
Sunyi Zheng, Jingxiong Li, Zhongyi Shui, Chenglu Zhu, Pingyi Chen, Lin Yang 0002 |
MICCAI (4) | 7 |
| 2022 | A Novel Encoding and Decoding Calibration Guiding Pathway for Pathological Image AnalysisabstractDiagnostic pathology is the foundation and gold standard for identifying carcinomas, and the accurate quantification of pathological images can provide objective clues for pathologists to make more convincing diagnosis. Recently, the encoder-decoder architectures (EDAs) of convolutional neural networks (CNNs) are widely used in the analysis of pathological images. Despite the rapid innovation of EDAs, we have conducted extensive experiments based on a variety of commonly used EDAs, and found them cannot handle the interference of complex background in pathological images, making the architectures unable to focus on the regions of interest (RoIs), thus making the quantitative results unreliable. Therefore, we proposed a pathway named GLobal Bank (GLB) to guide the encoder and the decoder to extract more features of RoIs rather than the complex background. Sufficient experiments have proved that the architecture remoulded by GLB can achieve significant performance improvement, and the quantitative results are more accurate. Hansheng Li, Yuxin Kang, Chunbao Wang 0002, Feihong Liu, Wenli Hui, Qirong Bo, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2022 | Harmonizing Pathological and Normal Pixels for Pseudo-Healthy SynthesisabstractSynthesizing a subject-specific pathology-free image from a pathological image is valuable for algorithm development and clinical practice. In recent years, several approaches based on the Generative Adversarial Network (GAN) have achieved promising results in pseudo-healthy synthesis. However, the discriminator (i.e., a classifier) in the GAN cannot accurately identify lesions and further hampers from generating admirable pseudo-healthy images. To address this problem, we present a new type of discriminator, the segmentor, to accurately locate the lesions and improve the visual quality of pseudo-healthy images. Then, we apply the generated images into medical image enhancement and utilize the enhanced results to cope with the low contrast problem existing in medical image segmentation. Furthermore, a reliable metric is proposed by utilizing two attributes of label noise to measure the health of synthetic images. Comprehensive experiments on the T2 modality of BraTS demonstrate that the proposed method substantially outperforms the state-of-the-art methods. The method achieves better performance than the existing methods with only 30% of the training data. The effectiveness of the proposed method is also demonstrated on the LiTS and the T1 modality of BraTS. The code and the pre-trained model of this study are publicly available at https://github.com/Au3C2/Generator-Versus-Segmentor. Yihong Zhuang, Liyan Sun, Yue Huang 0001, Xinghao Ding, Guisheng Wang, Lin Yang 0002, Yizhou Yu |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Boosting Boundary Representation for Gland Instance SegmentationabstractAccurate and automated gland instance segmentation on histology images can assist pathologists to analyze the malignancy degree of adenocarcinoma. Recently, deep-learning-based segmentation networks have been significantly developed to achieve this goal. However, the gland instances are generally proximate to each other and have indiscernible boundaries (i.e., homogeneous intensity values). Most of the existed networks do not define discriminative boundaries representation as context information, resulting in segmenting proximate instances incorrectly. In this paper, to improve the segmentation accuracy between proximate instances, we propose a Boundary Definition Module to boost boundaries feature representation by the guidance of the intra-and-extra glandular features. Moreover, we propose to use the Gumbel-Softmax distribution estimator to clarify the final prediction of boundaries further. Finally, we embed the Boundary Definition Module and Gumbel-Softmax distribution estimator into the gland instance network(FullNet) for performance verification. Experiments on the 2015 MICCAI Gland Segmentation Challenge dataset demonstrate that our proposed method achieves state-of-the-art performance. Yuxin Kang, Hansheng Li, Zhuoyue Wu, Feihong Liu, Dongqing Hu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 8 |
| 2021 | Robust Pathological Detector Training Method on Sparsely Annotated Datasets via Spatial CuesabstractComputer-aided diagnosis of pathological images usually requires detection and examination of all positive cells and lesions to make an accurate diagnosis. Therefore, there is an unprecedented demand for effective and reliable methods of training pathological detectors than ever. To train a reliable detector, the training dataset is required to fully annotate all positive instances, such a requirement is challenge and laborious, and is not guaranteed in most cases. However, sparse annotations will limit the training performance of detectors. Here, we propose a novel module named Collaborative Correction Sibling (CCS), which is embedded into the original object detection network to enhance the training performance on sparse annotations in a pioneering way. Specifically, instance-level annotations in the image space can be calibrated by positive instances’ spatial features provided by CCS. Extensive experiments have been conducted on both cellular-and-lesion-level detection tasks, compared with the state of the art methods, our CCS demonstrates the training effectiveness on pathological images. Hansheng Li, Yuxin Kang, Lingyu Hu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 7 |
| 2021 | Generalizing Nucleus Recognition Model in Multi-source Ki67 Immunohistochemistry Stained Images via Domain-Specific Pruning
Jiatong Cai, Chenglu Zhu, Honglin Li 0001, Shichuan Zhang, Lin Yang 0002 |
MICCAI (8) | 7 |
| 2021 | Automatic whole slide pathology image diagnosis framework via unit stochastic selection and attention fusion
Pingjun Chen, Yun Liang 0012, Xiaoshuang Shi, Lin Yang 0002, Paul D. Gader |
Neurocomputing | 4 |
| 2021 | Text-Guided Neural Network Training for Image Recognition in Natural Scenes and MedicineabstractConvolutional neural networks (CNNs) are widely recognized as the foundation for machine vision systems. The conventional rule of teaching CNNs to understand images requires training images with human annotated labels, without any additional instructions. In this article, we look into a new scope and explore the guidance from text for neural network training. We present two versions of attention mechanisms to facilitate interactions between visual and semantic information and encourage CNNs to effectively distill visual features by leveraging semantic features. In contrast to dedicated text-image joint embedding methods, our method realizes asynchronous training and inference behavior: a trained model can classify images, irrespective of the text availability. This characteristic substantially improves the model scalability to multiple (multimodal) vision tasks. We also apply the proposed method onto medical imaging, which learns from richer clinical knowledge and achieves attention-based interpretable decision-making. With comprehensive validation on two natural and two medical datasets, we demonstrate that our method can effectively make use of semantic knowledge to improve CNN performance. Our method performs substantial improvement on medical image datasets. Meanwhile, it achieves promising performance for multi-label image classification and caption-image retrieval as well as excellent performance for phrase-based and multi-object localization on public benchmarks. Zizhao Zhang 0002, Pingjun Chen, Xiaoshuang Shi, Lin Yang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | A Scalable Optimization Mechanism for Pairwise Based Discrete HashingabstractMaintaining the pairwise relationship among originally high-dimensional data into a low-dimensional binary space is a popular strategy to learn binary codes. One simple and intuitive method is to utilize two identical code matrices produced by hash functions to approximate a pairwise real label matrix. However, the resulting quartic problem in term of hash functions is difficult to directly solve due to the non-convex and non-smooth nature of the objective. In this paper, unlike previous optimization methods using various relaxation strategies, we aim to directly solve the original quartic problem using a novel alternative optimization mechanism to linearize the quartic problem by introducing a linear regression model. Additionally, we find that gradually learning each batch of binary codes in a sequential mode, i.e. batch by batch, is greatly beneficial to the convergence of binary code learning. Based on this significant discovery and the proposed strategy, we introduce a scalable symmetric discrete hashing algorithm that gradually and smoothly updates each batch of binary codes. To further improve the smoothness, we also propose a greedy symmetric discrete hashing algorithm to update each bit of batch binary codes. Moreover, we extend the proposed optimization mechanism to solve the non-convex optimization problems for binary code learning in many other pairwise based hashing algorithms. Extensive experiments on benchmark single-label and multi-label databases demonstrate the superior performance of the proposed mechanism over recent state-of-the-art methods on two kinds of retrieval tasks: similarity and ranking order. The source codes are available on https://github.com/xsshi2015/Scalable-Pairwise-based-Discrete-Hashing. Xiaoshuang Shi, Fuyong Xing, Zizhao Zhang 0002, Manish Sapkota, Zhenhua Guo 0001, Lin Yang 0002 |
IEEE Trans. Image Process. | 6 |
| 2021 | Lesion-Harvester: Iteratively Mining Unlabeled Lesions and Hard-Negative Examples at ScaleabstractThe acquisition of large-scale medical image data, necessary for training machine learning algorithms, is hampered by associated expert-driven annotation costs. Mining hospital archives can address this problem, but labels often incomplete or noisy, e.g., 50% of the lesions in DeepLesion are left unlabeled. Thus, effective label harvesting methods are critical. This is the goal of our work, where we introduce Lesion-Harvester-a powerful system to harvest missing annotations from lesion datasets at high precision. Accepting the need for some degree of expert labor, we use a small fully-labeled image subset to intelligently mine annotations from the remainder. To do this, we chain together a highly sensitive lesion proposal generator (LPG) and a very selective lesion proposal classifier (LPC). Using a new hard negative suppression loss, the resulting harvested and hard-negative proposals are then employed to iteratively finetune our LPG. While our framework is generic, we optimize our performance by proposing a new 3D contextual LPG and by using a global-local multi-view LPC. Experiments on DeepLesion demonstrate that Lesion-Harvester can discover an additional 9,805 lesions at a precision of 90%. We publicly release the harvested lesions, along with a new test set of completely annotated DeepLesion volumes. We also present a pseudo 3D IoU evaluation metric that corresponds much better to the real 3D IoU than current DeepLesion evaluation metrics. To quantify the downstream benefits of Lesion-Harvester we show that augmenting the DeepLesion annotations with our harvested lesions allows state-of-the-art detectors to boost their average precision by 7 to 10%. Jinzheng Cai, Adam P. Harrison, Youjing Zheng, Ke Yan 0006, Yuankai Huo, Jing Xiao 0006, Lin Yang 0002, Le Lu 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Loss-Based Attention for Deep Multiple Instance LearningabstractAlthough attention mechanisms have been widely used in deep learning for many tasks, they are rarely utilized to solve multiple instance learning (MIL) problems, where only a general category label is given for multiple instances contained in one bag. Additionally, previous deep MIL methods firstly utilize the attention mechanism to learn instance weights and then employ a fully connected layer to predict the bag label, so that the bag prediction is largely determined by the effectiveness of learned instance weights. To alleviate this issue, in this paper, we propose a novel loss based attention mechanism, which simultaneously learns instance weights and predictions, and bag predictions for deep multiple instance learning. Specifically, it calculates instance weights based on the loss function, e.g. softmax+cross-entropy, and shares the parameters with the fully connected layer, which is to predict instance and bag predictions. Additionally, a regularization term consisting of learned weights and cross-entropy functions is utilized to boost the recall of instances, and a consistency cost is used to smooth the training process of neural networks for boosting the model generalization performance. Extensive experiments on multiple types of benchmark databases demonstrate that the proposed attention mechanism is a general, effective and efficient framework, which can achieve superior bag and image classification performance over other state-of-the-art MIL methods, with obtaining higher instance precision and recall than previous attention mechanisms. Source codes are available on https://github.com/xsshi2015/Loss-Attention. Xiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Zizhao Zhang 0002, Lei Cui 0004, Lin Yang 0002 |
AAAI | 6 |
| 2020 | A Novel Loss Calibration Strategy for Object Detection Networks Training on Sparsely Annotated Pathological Datasets
Hansheng Li, Yuxin Kang, Xiaoshuang Shi, Mengdi Yan, Zixu Tong, Qirong Bu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (5) | 10 |
| 2020 | Rule-based automatic diagnosis of thyroid nodules from intraoperative frozen sections using deep learning
Pingjun Chen, Hai Su, Lin Yang 0002, Dingrong Zhong |
Artif. Intell. Medicine | 5 |
| 2020 | A deep learning-based framework for lung cancer survival analysis with biomarker interpretationabstractBACKGROUND: Lung cancer is the leading cause of cancer-related deaths in both men and women in the United States, and it has a much lower five-year survival rate than many other cancers. Accurate survival analysis is urgently needed for better disease diagnosis and treatment management. RESULTS: In this work, we propose a survival analysis system that takes advantage of recently emerging deep learning techniques. The proposed system consists of three major components. 1) The first component is an end-to-end cellular feature learning module using a deep neural network with global average pooling. The learned cellular representations encode high-level biologically relevant information without requiring individual cell segmentation, which is aggregated into patient-level feature vectors by using a locality-constrained linear coding (LLC)-based bag of words (BoW) encoding algorithm. 2) The second component is a Cox proportional hazards model with an elastic net penalty for robust feature selection and survival analysis. 3) The third commponent is a biomarker interpretation module that can help localize the image regions that contribute to the survival model's decision. Extensive experiments show that the proposed survival model has excellent predictive power for a public (i.e., The Cancer Genome Atlas) lung cancer dataset in terms of two commonly used metrics: log-rank test (p-value) of the Kaplan-Meier estimate and concordance index (c-index). CONCLUSIONS: In this work, we have proposed a segmentation-free survival analysis system that takes advantage of the recently emerging deep learning framework and well-studied survival analysis methods such as the Cox proportional hazards model. In addition, we provide an approach to visualize the discovered biomarkers, which can serve as concrete evidence supporting the survival model's decision. Lei Cui 0004, Hansheng Li, Wenli Hui, Lin Yang 0002, Yuxin Kang, Qirong Bo, Jun Feng 0003 |
BMC Bioinform. | 5 |
| 2020 | Anchor-Based Self-Ensembling for Semi-Supervised Deep Pairwise Hashing
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Yun Liang 0012, Lin Yang 0002 |
Int. J. Comput. Vis. | 5 |
| 2020 | Graph temporal ensembling based semi-supervised convolutional neural network with noisy labels for histopathology image analysis
Xiaoshuang Shi, Hai Su, Fuyong Xing, Yun Liang 0012, Gang Qu 0002, Lin Yang 0002 |
Medical Image Anal. | 6 |
| 2019 | Global Bank: A Guided Pathway of Encoding and Decoding for Pathological Image AnalysisabstractThe encoder-decoder architecture of convolutional neural networks (CNNs) is widely used in computer vision tasks and various analyses of medical images. However, extracting semantic features from regions of interest (RoIs) in pathological images remains a challenging task because RoIs of different morphologies and scales are embedded in a blurred background. Additionally, it is well known that the classic encoder-decoder architecture is vulnerable to interference from a blurred background and is thus not entirely suitable for precise analysis of pathological images. In this paper, we propose a pathway named global bank (GLB) to guide the encoder and decoder to focus more on the RoIs by providing the decoder with additional effective features of the RoIs. We extend the U-Net and feature pyramid network (FPN) with GLB and evaluate the resulting models on gland segmentation and cancer embolus detection tasks, respectively. Extensive experiments demonstrate that our proposal can significantly improve the performance of the encoder-decoder architecture. The U-Net with GLB achieves the best semantic segmentation performance on the 2015 MICCAI Gland Challenge dataset. Additionally, the FPN with GLB achieves improvements of 2% in average precision and 3.4% in recall on the embolus detection task. Hansheng Li, Jun Feng 0003, Baosheng Kang, Yuxin Kang, Feihong Liu, Wenli Hui, Qirong Bo, Chunbao Wang 0002, Lin Yang 0002, Lei Cui 0004 |
BIBM | 9 |
| 2019 | $S^{3}$ Net: Trained on a Small Sample Segmentation Network for Biomedical Image AnalysisabstractFully convolutional networks (FCNs) are powerful methods to extract hierarchies of features that have achieved remarkable success in various biomedical image analysis tasks. However, the successful training of FCN requires more than hundreds of pixel-level annotated training samples, which poses a challenge for biomedical image processing tasks. In this paper, we present S3Net, a network that makes more efficient use of available annotated samples on biomedical image segmentation. S3Net is essentially a deeply-supervised encoder-decoder network where the decoder has been redesigned to efficiently restore multi-level encoded feature resolution in a single step. We have conducted extensive experiments on the 2015 MICCAI Gland Challenge dataset. Compared with other methods, S3Net achieves a 0.781 dice-score using 11 images for training, higher than U-Net++ 18.6% and U-Net 12.9%, which verifies the performance of S3Net trained on a small sample. Further, with training on 85 images, S3Net achieves a 0.910 dice-score by using resnet50 as the encoder, which is higher than state-of-the-art semantic segmentation results by 4%. Mengdi Yan, Hansheng Li, Baosheng Kang, Jun Feng 0003, Yuxin Kang, Lin Yang 0002, Lei Cui 0004 |
BIBM | 7 |
| 2019 | Reducing Uncertainty in Undersampled MRI Reconstruction With Active AcquisitionabstractThe goal of MRI reconstruction is to restore a high fidelity image from partially observed measurements. This partial view naturally induces reconstruction uncertainty that can only be reduced by acquiring additional measurements. In this paper, we present a novel method for MRI reconstruction that, at inference time, dynamically selects the measurements to take and iteratively refines the prediction in order to best reduce the reconstruction error and, thus, its uncertainty. We validate our method on a large scale knee MRI dataset, as well as on ImageNet. Results show that (1) our system successfully outperforms active acquisition baselines; (2) our uncertainty estimates correlate with error maps; and (3) our ResNet-based architecture surpasses standard pixel-to-pixel models in the task of MRI reconstruction. The proposed method not only shows high-quality reconstructions but also paves the road towards more applicable solutions for accelerating MRI. Zizhao Zhang 0002, Adriana Romero, Matthew J. Muckley, Pascal Vincent, Lin Yang 0002, Michal Drozdzal |
CVPR | 5 |
| 2019 | Local and Global Consistency Regularized Mean Teacher for Semi-supervised Nuclei Classification
Hai Su, Xiaoshuang Shi, Jinzheng Cai, Lin Yang 0002 |
MICCAI (1) | 4 |
| 2019 | High throughput automatic muscle image segmentation using parallel frameworkabstractBACKGROUND: Fast and accurate automatic segmentation of skeletal muscle cell image is crucial for the diagnosis of muscle related diseases, which extremely reduces the labor-intensive manual annotation. Recently, several methods have been presented for automatic muscle cell segmentation. However, most methods exhibit high model complexity and time cost, and they are not adaptive to large-scale images such as whole-slide scanned specimens. METHODS: In this paper, we propose a novel distributed computing approach, which adopts both data and model parallel, for fast muscle cell segmentation. With a master-worker parallelism manner, the image data in the master is distributed onto multiple workers based on the Spark cloud computing platform. On each worker node, we first detect cell contours using a structured random forest (SRF) contour detector with fast parallel prediction and generate region candidates using a superpixel technique. Next, we propose a novel hierarchical tree based region selection algorithm for cell segmentation based on the conditional random field (CRF) algorithm. We divide the region selection algorithm into multiple sub-problems, which can be further parallelized using multi-core programming. RESULTS: We test the performance of the proposed method on a large-scale haematoxylin and eosin (H &E) stained skeletal muscle image dataset. Compared with the standalone implementation, the proposed method achieves more than 10 times speed improvement on very large-scale muscle images containing hundreds to thousands of cells. Meanwhile, our proposed method produces high-quality segmentation results compared with several state-of-the-art methods. CONCLUSIONS: This paper presents a parallel muscle image segmentation method with both data and model parallelism on multiple machines. The parallel strategy exhibits high compatibility to our muscle segmentation framework. The proposed method achieves high-throughput effective cell segmentation on large-scale muscle images. Lei Cui 0004, Jun Feng 0003, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 4 |
| 2019 | Towards pixel-to-pixel deep nucleus detection in microscopy imagesabstractBACKGROUND: Nucleus is a fundamental task in microscopy image analysis and supports many other quantitative studies such as object counting, segmentation, tracking, etc. Deep neural networks are emerging as a powerful tool for biomedical image computing; in particular, convolutional neural networks have been widely applied to nucleus/cell detection in microscopy images. However, almost all models are tailored for specific datasets and their applicability to other microscopy image data remains unknown. Some existing studies casually learn and evaluate deep neural networks on multiple microscopy datasets, but there are still several critical, open questions to be addressed. RESULTS: We analyze the applicability of deep models specifically for nucleus detection across a wide variety of microscopy image data. More specifically, we present a fully convolutional network-based regression model and extensively evaluate it on large-scale digital pathology and microscopy image datasets, which consist of 23 organs (or cancer diseases) and come from multiple institutions. We demonstrate that for a specific target dataset, training with images from the same types of organs might be usually necessary for nucleus detection. Although the images can be visually similar due to the same staining technique and imaging protocol, deep models learned with images from different organs might not deliver desirable results and would require model fine-tuning to be on a par with those trained with target data. We also observe that training with a mixture of target and other/non-target data does not always mean a higher accuracy of nucleus detection, and it might require proper data manipulation during model training to achieve good performance. CONCLUSIONS: We conduct a systematic case study on deep models for nucleus detection in a wide variety of microscopy images, aiming to address several important but previously understudied questions. We present and extensively evaluate an end-to-end, pixel-to-pixel fully convolutional regression network and report a few significant findings, some of which might have not been reported in previous studies. The model performance analysis and observations would be helpful to nucleus detection in microscopy images. Fuyong Xing, Yuanpu Xie, Xiaoshuang Shi, Pingjun Chen, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 6 |
| 2019 | Correction to: Towards pixel-to-pixel deep nucleus detection in microscopy imagesabstractFollowing publication of the original article [1], we have been notified of a few errors in the html version. Fuyong Xing, Yuanpu Xie, Xiaoshuang Shi, Pingjun Chen, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 6 |
| 2019 | Structured orthogonal matching pursuit for feature selection
Xiaoshuang Shi, Fuyong Xing, Zhenhua Guo 0001, Hai Su, Fujun Liu, Lin Yang 0002 |
Neurocomputing | 6 |
| 2019 | Towards cross-modal organ translation and segmentation: A cycle- and shape-consistent generative adversarial network
Jinzheng Cai, Zizhao Zhang 0002, Lei Cui 0004, Yefeng Zheng 0001, Lin Yang 0002 |
Medical Image Anal. | 5 |
| 2019 | Texture analysis for muscular dystrophy classification in MRI with improved class activation mapping
Jinzheng Cai, Fuyong Xing, Abhinandan Batra, Fujun Liu, Glenn A. Walter, Krista Vandenborne, Lin Yang 0002 |
Pattern Recognit. | 7 |
| 2019 | Deep Convolutional Hashing for Low-Dimensional Binary Embedding of Histopathological ImagesabstractCompact binary representations of histopa-thology images using hashing methods provide efficient approximate nearest neighbor search for direct visual query in large-scale databases. They can be utilized to measure the probability of the abnormality of the query image based on the retrieved similar cases, thereby providing support for medical diagnosis. They also allow for efficient managing of large-scale image databases because of a low storage requirement. However, the effectiveness of binary representations heavily relies on the visual descriptors that represent the semantic information in the histopathological images. Traditional approaches with hand-crafted visual descriptors might fail due to significant variations in image appearance. Recently, deep learning architectures provide promising solutions to address this problem using effective semantic representations. In this paper, we propose a deep convolutional hashing method that can be trained "point-wise" to simultaneously learn both semantic and binary representations of histopathological images. Specifically, we propose a convolutional neural network that introduces a latent binary encoding (LBE) layer for low-dimensional feature embedding to learn binary codes. We design a joint optimization objective function that encourages the network to learn discriminative representations from the label information, and reduce the gap between the real-valued low-dimensional embedded features and desired binary values. The binary encoding for new images can be obtained by forward propagating through the network and quantizing the output of the LBE layer. Experimental results on a large-scale histopathological image dataset demonstrate the effectiveness of the proposed method. Manish Sapkota, Xiaoshuang Shi, Fuyong Xing, Lin Yang 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Semi-supervised Deep Linear Discriminant Analysis for Histopathology Image Classification
Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 3 |
| 2018 | City-Wide Influenza Forecasting based on Multi-Source DataabstractSeasonal influenza epidemics which annually cause substantial diseases and deaths in high-risk population groups are a major public health concern around the world. Considering the hysteresis of traditional flu surveillance systems, this work aims to present a methodology capable of forecasting influenza activity of a city in China precisely 1 week ahead of the official publication. To that end, exogenous information collected from different sources were separately tested with historical influenza-like illness reports for the ability of detecting influenza activity, including climate surveillance, Internet users’ search activity, twitter and health inquiry on an online health consultation platform. Moreover, an ensemble model combining a time series analysis model and a tree boosting model based on those multisource data was applied to improve the accuracy and generalizability of influenza forecasting, in which a model fusion method based on the Kalman Filter was proposed. The validation experiments in this work were performed on the influenza-like illness reports collected from Chongqing city over 4 influenza seasons within 2014–2017. The results show that the proposed model outperformed other tested models by not only taking the periodic law of influenza into consideration but also incorporating information from diverse data sources. The mean absolute percentage error of the validation data set decreased to about 10%. This work provides a viable suggestion for improving the influenza activity forecasting of a city at its early stage. Baisong Li, Lin Yang 0002, Wenge Tang, Xiaowen Ruan, Shaofeng Lu, Xianxian Chen, Chaobo Shen, Jiaying Xu, Liang Xu 0010, Jing Xiao 0006 |
IEEE BigData | 6 |
| 2018 | Translating and Segmenting Multimodal Medical Volumes With Cycle- and Shape-Consistency Generative Adversarial NetworkabstractSynthesized medical images have several important applications, e.g., as an intermedium in cross-modality image registration and as supplementary training samples to boost the generalization capability of a classifier. Especially, synthesized computed tomography (CT) data can provide X-ray attenuation map for radiation therapy planning. In this work, we propose a generic cross-modality synthesis approach with the following targets: 1) synthesizing realistic looking 3D images using unpaired training data, 2) ensuring consistent anatomical structures, which could be changed by geometric distortion in cross-modality synthesis and 3) improving volume segmentation by using synthetic data for modalities with limited training samples. We show that these goals can be achieved with an end-to-end 3D convolutional neural network (CNN) composed of mutually-beneficial generators and segmentors for image synthesis and segmentation tasks. The generators are trained with an adversarial loss, a cycle-consistency loss, and also a shape-consistency loss, which is supervised by segmentors, to reduce the geometric distortion. From the segmentation view, the segmentors are boosted by synthetic data from generators in an online manner. Generators and segmentors prompt each other alternatively in an end-to-end training fashion. With extensive experiments on a dataset including a total of 4,496 CT and magnetic resonance imaging (MRI) cardiovascular volumes, we show both tasks are beneficial to each other and coupling these two tasks results in better performance than solving them exclusively. Zizhao Zhang 0002, Lin Yang 0002, Yefeng Zheng 0001 |
CVPR | 2 |
| 2018 | Iterative Attention Mining for Weakly Supervised Thoracic Disease Pattern Localization in Chest X-Rays
Jinzheng Cai, Le Lu 0001, Adam P. Harrison, Xiaoshuang Shi, Pingjun Chen, Lin Yang 0002 |
MICCAI (2) | 6 |
| 2018 | Accurate Weakly-Supervised Deep Lesion Segmentation Using Large-Scale Clinical Annotations: Slice-Propagated 3D Mask Generation from 2D RECIST
Jinzheng Cai, Youbao Tang, Le Lu 0001, Adam P. Harrison, Ke Yan 0006, Jing Xiao 0006, Lin Yang 0002, Ronald M. Summers |
MICCAI (4) | 7 |
| 2018 | Efficient and robust cell detection: A structured regression approach
Yuanpu Xie, Fuyong Xing, Xiaoshuang Shi, Xiangfei Kong, Hai Su, Lin Yang 0002 |
Medical Image Anal. | 6 |
| 2018 | Self-learning for face clustering
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Jinzheng Cai, Lin Yang 0002 |
Pattern Recognit. | 5 |
| 2018 | Pairwise based deep ranking hashing for histopathology image classification and retrieval
Xiaoshuang Shi, Manish Sapkota, Fuyong Xing, Fujun Liu, Lei Cui 0004, Lin Yang 0002 |
Pattern Recognit. | 6 |
| 2018 | Revisiting graph construction for fast image segmentation
Zizhao Zhang 0002, Fuyong Xing, Hanzi Wang, Yan Yan 0001, Xiaoshuang Shi, Lin Yang 0002 |
Pattern Recognit. | 7 |
| 2018 | AIIMDs: An Integrated Framework of Automatic Idiopathic Inflammatory Myopathy Diagnosis for MuscleabstractIdiopathic inflammatory myopathy (IIM) is a common skeletal muscle disease that relates to weakness and inflammation of muscle. Early diagnosis and prognosis of different types of IIMs will guide the effective treatment. Interpretation of digitized images of the cross-section muscle biopsy, which is currently done manually, provides the most reliable diagnostic information. With the increasing volume of images, the management and manual interpretation of the digitized muscle images suffer from low efficiency and high interobserver variabilities. In order to address these problems, we propose the first complete framework of automatic IIM diagnosis system for the management and interpretation of digitized skeletal muscle histopathology images. The proposed framework consists of several key components: (1) Automatic cell segmentation, perimysium annotation, and nuclei detection; (2) histogram-based feature extraction and quantification; (3) content-based image retrieval to search and retrieve similar cases in the database for comparative study; and (4) majority voting-based classification to provide decision support for computer-aided clinical diagnosis. Experiments show that the proposed diagnosis system provides efficient and robust interpretation of the digitized muscle image and computer-aided diagnosis of IIM. Manish Sapkota, Fujun Liu, Yuanpu Xie, Hai Su, Fuyong Xing, Lin Yang 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | Deep Learning in Microscopy Image Analysis: A SurveyabstractComputerized microscopy image analysis plays an important role in computer aided diagnosis and prognosis. Machine learning techniques have powered many aspects of medical investigation and clinical practice. Recently, deep learning is emerging as a leading machine learning tool in computer vision and has attracted considerable attention in biomedical image analysis. In this paper, we provide a snapshot of this fast-growing field, specifically for microscopy image analysis. We briefly introduce the popular deep neural networks and summarize current deep learning achievements in various tasks, such as detection, segmentation, and classification in microscopy image analysis. In particular, we explain the architectures and the principles of convolutional neural networks, fully convolutional networks, recurrent neural networks, stacked autoencoders, and deep belief networks, and interpret their formulations or modelings for specific tasks on various microscopy images. In addition, we discuss the open challenges and the potential trends of future research in microscopy image analysis using deep learning. Fuyong Xing, Yuanpu Xie, Hai Su, Fujun Liu, Lin Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Asymmetric Discrete Graph HashingabstractRecently, many graph based hashing methods have been emerged to tackle large-scale problems. However, there exists two major bottlenecks: (1) directly learning discrete hashing codes is an NP-hardoptimization problem; (2) the complexity of both storage and computational time to build a graph with n data points is O(n2). To address these two problems, in this paper, we propose a novel yetsimple supervised graph based hashing method, asymmetric discrete graph hashing, by preserving the asymmetric discrete constraint and building an asymmetric affinity matrix to learn compact binary codes.Specifically, we utilize two different instead of identical discrete matrices to better preserve the similarity of the graph with short binary codes. We generate the asymmetric affinity matrix using m (m << n) selected anchors to approximate the similarity among all training data so that computational time and storage requirement can be significantly improved. In addition, the proposed method jointly learns discrete binary codes and a low-dimensional projection matrix to further improve the retrieval accuracy. Extensive experiments on three benchmark large-scale databases demonstrate its superior performance over the recent state of the arts with lower training time costs. Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Manish Sapkota, Lin Yang 0002 |
AAAI | 5 |
| 2017 | MDNet: A Semantically and Visually Interpretable Medical Image Diagnosis NetworkabstractThe inability to interpret the model prediction in semantically and visually meaningful ways is a well-known shortcoming of most existing computer-aided diagnosis methods. In this paper, we propose MDNet to establish a direct multimodal mapping between medical images and diagnostic reports that can read images, generate diagnostic reports, retrieve images by symptom descriptions, and visualize attention, to provide justifications of the network diagnosis process. MDNet includes an image model and a language model. The image model is proposed to enhance multi-scale feature ensembles and utilization efficiency. The language model, integrated with our improved attention mechanism, aims to read and explore discriminative image feature descriptions from reports to learn a direct mapping from sentence words to image pixels. The overall network is trained end-to-end by using our developed optimization strategy. Based on a pathology bladder cancer images and its diagnostic reports (BCIDR) dataset, we conduct sufficient experiments to demonstrate that MDNet outperforms comparative baselines. The proposed image model obtains state-of-the-art performance on two CIFAR datasets as well. Zizhao Zhang 0002, Yuanpu Xie, Fuyong Xing, Mason McGough, Lin Yang 0002 |
CVPR | 5 |
| 2017 | Pancreas Segmentation in MRI Using Graph-Based Decision Fusion on Convolutional Neural Networks
Jinzheng Cai, Le Lu 0001, Yuanpu Xie, Fuyong Xing, Lin Yang 0002 |
MICCAI (3) | 5 |
| 2017 | Cell Encoding for Histopathology Image Classification
Xiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Hai Su, Lin Yang 0002 |
MICCAI (2) | 5 |
| 2017 | TandemNet: Distilling Knowledge from Medical Images Using Diagnostic Reports as Optional Semantic References
Zizhao Zhang 0002, Pingjun Chen, Manish Sapkota, Lin Yang 0002 |
MICCAI (3) | 4 |
| 2017 | Supervised graph hashing for histopathology image retrieval and classification
Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Yuanpu Xie, Hai Su, Lin Yang 0002 |
Medical Image Anal. | 6 |
| 2016 | SemiContour: A Semi-Supervised Learning Approach for Contour DetectionabstractSupervised contour detection methods usually require many labeled training images to obtain satisfactory performance. However, a large set of annotated data might be unavailable or extremely labor intensive. In this paper, we investigate the usage of semi-supervised learning (SSL) to obtain competitive detection accuracy with very limited training data (three labeled images). Specifically, we propose a semi-supervised structured ensemble learning approach for contour detection built on structured random forests (SRF). To allow SRF to be applicable to unlabeled data, we present an effective sparse representation approach to capture inherent structure in image patches by finding a compact and discriminative low-dimensional subspace representation in an unsupervised manner, enabling the incorporation of abundant unlabeled patches with their estimated structured labels to help SRF perform better node splitting. We re-examine the role of sparsity and propose a novel and fast sparse coding algorithm to boost the overall learning efficiency. To the best of our knowledge, this is the first attempt to apply SSL for contour detection. Extensive experiments on the BSDS500 segmentation dataset and the NYU Depth dataset demonstrate the superiority of the proposed method. Zizhao Zhang 0002, Fuyong Xing, Xiaoshuang Shi, Lin Yang 0002 |
CVPR | 4 |
| 2016 | Kernel-Based Supervised Discrete Hashing for Image Retrieval
Xiaoshuang Shi, Fuyong Xing, Jinzheng Cai, Zizhao Zhang 0002, Yuanpu Xie, Lin Yang 0002 |
ECCV (7) | 6 |
| 2016 | Pancreas Segmentation in MRI Using Graph-Based Decision Fusion on Convolutional Neural Networks
Jinzheng Cai, Le Lu 0001, Zizhao Zhang 0002, Fuyong Xing, Lin Yang 0002 |
MICCAI (2) | 5 |
| 2016 | Spatial Clockwork Recurrent Neural Network for Muscle Perimysium Segmentation
Yuanpu Xie, Zizhao Zhang 0002, Manish Sapkota, Lin Yang 0002 |
MICCAI (2) | 4 |
| 2016 | Transfer Shape Modeling Towards High-Throughput Microscopy Image Segmentation
Fuyong Xing, Xiaoshuang Shi, Zizhao Zhang 0002, Jinzheng Cai, Yuanpu Xie, Lin Yang 0002 |
MICCAI (3) | 6 |
| 2016 | Scalable histopathological image analysis via supervised hashing with multiple features
Menglin Jiang, Shaoting Zhang 0001, Junzhou Huang, Lin Yang 0002, Dimitris N. Metaxas |
Medical Image Anal. | 4 |
| 2016 | Two-Dimensional Whitening Reconstruction for Enhancing Robustness of Principal Component AnalysisabstractPrincipal component analysis (PCA) is widely applied in various areas, one of the typical applications is in face. Many versions of PCA have been developed for face recognition. However, most of these approaches are sensitive to grossly corrupted entries in a 2D matrix representing a face image. In this paper, we try to reduce the influence of grosses like variations in lighting, facial expressions and occlusions to improve the robustness of PCA. In order to achieve this goal, we present a simple but effective unsupervised preprocessing method, two-dimensional whitening reconstruction (TWR), which includes two stages: 1) A whitening process on a 2D face image matrix rather than a concatenated 1D vector; 2) 2D face image matrix reconstruction. TWR reduces the pixel redundancy of the internal image, meanwhile maintains important intrinsic features. In this way, negative effects introduced by gross-like variations are greatly reduced. Furthermore, the face image with TWR preprocessing could be approximate to a Gaussian signal, on which PCA is more effective. Experiments on benchmark face databases demonstrate that the proposed method could significantly improve the robustness of PCA methods on classification and clustering, especially for the faces with severe illumination changes. Xiaoshuang Shi, Zhenhua Guo 0001, Feiping Nie 0001, Lin Yang 0002, Jane You, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Robust Cell Detection of Histopathological Brain Tumor Images Using Sparse Reconstruction and Adaptive Dictionary SelectionabstractSuccessful diagnostic and prognostic stratification, treatment outcome prediction, and therapy planning depend on reproducible and accurate pathology analysis. Computer aided diagnosis (CAD) is a useful tool to help doctors make better decisions in cancer diagnosis and treatment. Accurate cell detection is often an essential prerequisite for subsequent cellular analysis. The major challenge of robust brain tumor nuclei/cell detection is to handle significant variations in cell appearance and to split touching cells. In this paper, we present an automatic cell detection framework using sparse reconstruction and adaptive dictionary learning. The main contributions of our method are: 1) A sparse reconstruction based approach to split touching cells; 2) An adaptive dictionary learning method used to handle cell appearance variations. The proposed method has been extensively tested on a data set with more than 2000 cells extracted from 32 whole slide scanned images. The automatic cell detection results are compared with the manually annotated ground truth and other state-of-the-art cell detection algorithms. The proposed method achieves the best cell detection accuracy with a F1 score = 0.96. Hai Su, Fuyong Xing, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2016 | An Automatic Learning-Based Framework for Robust Nucleus SegmentationabstractComputer-aided image analysis of histopathology specimens could potentially provide support for early detection and improved characterization of diseases such as brain tumor, pancreatic neuroendocrine tumor (NET), and breast cancer. Automated nucleus segmentation is a prerequisite for various quantitative analyses including automatic morphological feature computation. However, it remains to be a challenging problem due to the complex nature of histopathology images. In this paper, we propose a learning-based framework for robust and automatic nucleus segmentation with shape preservation. Given a nucleus image, it begins with a deep convolutional neural network (CNN) model to generate a probability map, on which an iterative region merging approach is performed for shape initializations. Next, a novel segmentation algorithm is exploited to separate individual nuclei combining a robust selection-based sparse shape model and a local repulsive deformable model. One of the significant benefits of the proposed framework is that it is applicable to different staining histopathology images. Due to the feature learning characteristic of the deep CNN and the high level shape prior modeling, the proposed method is general enough to perform well across multiple scenarios. We have tested the proposed algorithm on three large-scale pathology image datasets using a range of different tissue and stain preparations, and the comparative experiments with recent state of the arts demonstrate the superior performance of the proposed approach. Fuyong Xing, Yuanpu Xie, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2015 | Fine-grained histopathological image analysis via robust segmentation and large-scale retrievalabstractComputer-aided diagnosis of medical images requires thorough analysis of image details. For example, examining all cells enables fine-grained categorization of histopathological images. Traditional computational methods may have efficiency issues when performing such detailed analysis. In this paper, we propose a robust and scalable solution to achieve this. Specifically, a robust segmentation method is developed to delineate region-of-interests (e.g., cells) accurately, using hierarchical voting and repulsive active contour. A hashing-based large-scale retrieval approach is also designed to examine and classify them by comparing with a massive training database. We evaluate this proposed framework on a challenging and important clinical use case, i.e., differentiation of two types of lung cancers (the adenocarcinoma and the squamous carcinoma), using thousands of histopathological images extracted from hundreds of patients. Our method has achieved promising performance, i.e., 87.3% accuracy and 1.68 seconds by searching among half-million cells. Xiaofan Zhang 0002, Hai Su, Lin Yang 0002, Shaoting Zhang 0001 |
CVPR | 3 |
| 2015 | Within-class penalty based multi-class support vector machineabstractSupport vector machine (SVM) is a widely used maximum margin classifier, but the classification performance is largely affected by outliers. In this paper, we propose a novel multi-class SVM method to reduce the influence of outliers on the classification performance. Our proposed method includes an efficient optimization model via considering the within-class scatter and an optimization way. Specifically, the method is based on one assumption that penalizing the within-class scatter can reduce the number of misclassified outliers near the decision boundary, because data points of each class could be compacted by the within-class penalty. Experiments on benchmark databases demonstrate the effectiveness of the assumption and the proposed method. Xiaoshuang Shi, Zhenhua Guo 0001, Yujiu Yang 0001, Lin Yang 0002 |
ICIP | 4 |
| 2015 | Joint Kernel-Based Supervised Hashing for Scalable Histopathological Image Analysis
Menglin Jiang, Shaoting Zhang 0001, Junzhou Huang, Lin Yang 0002, Dimitris N. Metaxas |
MICCAI (3) | 4 |
| 2015 | Robust Muscle Cell Quantification Using Structured Edge Detection and Hierarchical Segmentation
Fujun Liu, Fuyong Xing, Zizhao Zhang 0002, Mason McGough, Lin Yang 0002 |
MICCAI (3) | 5 |
| 2015 | A Novel Cell Detection Method Using Deep Convolutional Neural Network and Maximum-Weight Independent Set
Fujun Liu, Lin Yang 0002 |
MICCAI (3) | 2 |
| 2015 | Robust Cell Detection and Segmentation in Histopathological Images Using Sparse Reconstruction and Stacked Denoising Autoencoders
Hai Su, Fuyong Xing, Xiangfei Kong, Yuanpu Xie, Shaoting Zhang 0001, Lin Yang 0002 |
MICCAI (3) | 6 |
| 2015 | Deep Voting: A Robust Approach Toward Nucleus Localization in Microscopy Images
Yuanpu Xie, Xiangfei Kong, Fuyong Xing, Fujun Liu, Hai Su, Lin Yang 0002 |
MICCAI (3) | 6 |
| 2015 | Beyond Classification: Structured Regression for Robust Cell Detection Using Convolutional Neural Network
Yuanpu Xie, Fuyong Xing, Xiangfei Kong, Hai Su, Lin Yang 0002 |
MICCAI (3) | 5 |
| 2015 | Fast Cell Segmentation Using Scalable Sparse Manifold Learning and Affine Transform-Approximated Active Contour
Fuyong Xing, Lin Yang 0002 |
MICCAI (3) | 2 |
| 2015 | Scalable analysis of Big pathology image data cohorts using efficient methods and high-performance computing strategiesabstractBACKGROUND: We describe a suite of tools and methods that form a core set of capabilities for researchers and clinical investigators to evaluate multiple analytical pipelines and quantify sensitivity and variability of the results while conducting large-scale studies in investigative pathology and oncology. The overarching objective of the current investigation is to address the challenges of large data sizes and high computational demands. RESULTS: The proposed tools and methods take advantage of state-of-the-art parallel machines and efficient content-based image searching strategies. The content based image retrieval (CBIR) algorithms can quickly detect and retrieve image patches similar to a query patch using a hierarchical analysis approach. The analysis component based on high performance computing can carry out consensus clustering on 500,000 data points using a large shared memory system. CONCLUSIONS: Our work demonstrates efficient CBIR algorithms and high performance computing can be leveraged for efficient analysis of large microscopy images to meet the challenges of clinically salient applications in pathology. These technologies enable researchers and clinical investigators to make more effective use of the rich informational content contained within digitized microscopy specimens. Tahsin M. Kurç, Xin Qi 0007, Daihou Wang, Fusheng Wang 0001, George Teodoro, Lee A. D. Cooper, Michael Nalisnik, Lin Yang 0002, Joel H. Saltz, David J. Foran |
BMC Bioinform. | 8 |
| 2015 | High-throughput histopathological image analysis via robust cell segmentation and hashing
Xiaofan Zhang 0002, Fuyong Xing, Hai Su, Lin Yang 0002, Shaoting Zhang 0001 |
Medical Image Anal. | 4 |
| 2014 | Mining Histopathological Images via Composite Hashing and Online Learning
Xiaofan Zhang 0002, Lin Yang 0002, Wei Liu 0005, Hai Su, Shaoting Zhang 0001 |
MICCAI (2) | 2 |
| 2014 | Parallel content-based sub-image retrieval using hierarchical searchingabstractMOTIVATION: The capacity to systematically search through large image collections and ensembles and detect regions exhibiting similar morphological characteristics is central to pathology diagnosis. Unfortunately, the primary methods used to search digitized, whole-slide histopathology specimens are slow and prone to inter- and intra-observer variability. The central objective of this research was to design, develop, and evaluate a content-based image retrieval system to assist doctors for quick and reliable content-based comparative search of similar prostate image patches. METHOD: Given a representative image patch (sub-image), the algorithm will return a ranked ensemble of image patches throughout the entire whole-slide histology section which exhibits the most similar morphologic characteristics. This is accomplished by first performing hierarchical searching based on a newly developed hierarchical annular histogram (HAH). The set of candidates is then further refined in the second stage of processing by computing a color histogram from eight equally divided segments within each square annular bin defined in the original HAH. A demand-driven master-worker parallelization approach is employed to speed up the searching procedure. Using this strategy, the query patch is broadcasted to all worker processes. Each worker process is dynamically assigned an image by the master process to search for and return a ranked list of similar patches in the image. RESULTS: The algorithm was tested using digitized hematoxylin and eosin (H&E) stained prostate cancer specimens. We have achieved an excellent image retrieval performance. The recall rate within the first 40 rank retrieved image patches is ∼90%. AVAILABILITY AND IMPLEMENTATION: Both the testing data and source code can be downloaded from http://pleiad.umdnj.edu/CBII/Bioinformatics/. Lin Yang 0002, Xin Qi 0007, Fuyong Xing, Tahsin M. Kurç, Joel H. Saltz, David J. Foran |
Bioinform. | 1 |
| 2014 | Content-based histopathology image retrieval using CometCloudabstractBACKGROUND: The development of digital imaging technology is creating extraordinary levels of accuracy that provide support for improved reliability in different aspects of the image analysis, such as content-based image retrieval, image segmentation, and classification. This has dramatically increased the volume and rate at which data are generated. Together these facts make querying and sharing non-trivial and render centralized solutions unfeasible. Moreover, in many cases this data is often distributed and must be shared across multiple institutions requiring decentralized solutions. In this context, a new generation of data/information driven applications must be developed to take advantage of the national advanced cyber-infrastructure (ACI) which enable investigators to seamlessly and securely interact with information/data which is distributed across geographically disparate resources. This paper presents the development and evaluation of a novel content-based image retrieval (CBIR) framework. The methods were tested extensively using both peripheral blood smears and renal glomeruli specimens. The datasets and performance were evaluated by two pathologists to determine the concordance. RESULTS: The CBIR algorithms that were developed can reliably retrieve the candidate image patches exhibiting intensity and morphological characteristics that are most similar to a given query image. The methods described in this paper are able to reliably discriminate among subtle staining differences and spatial pattern distributions. By integrating a newly developed dual-similarity relevance feedback module into the CBIR framework, the CBIR results were improved substantially. By aggregating the computational power of high performance computing (HPC) and cloud resources, we demonstrated that the method can be successfully executed in minutes on the Cloud compared to weeks using standard computers. CONCLUSIONS: In this paper, we present a set of newly developed CBIR algorithms and validate them using two different pathology applications, which are regularly evaluated in the practice of pathology. Comparative experimental results demonstrate excellent performance throughout the course of a set of systematic studies. Additionally, we present and evaluate a framework to enable the execution of these algorithms across distributed resources. We show how parallel searching of content-wise similar images in the dataset significantly reduces the overall computational time to ensure the practical utility of the proposed CBIR algorithms. Xin Qi 0007, Daihou Wang, Ivan Rodero, Javier Diaz Montes, Rebekah H. Gensure, Fuyong Xing, Lauri A. Goodell, Manish Parashar, David J. Foran, Lin Yang 0002 |
BMC Bioinform. | 11 |
| 2014 | Novel image markers for non-small cell lung cancer classification and survival predictionabstractBACKGROUND: Non-small cell lung cancer (NSCLC), the most common type of lung cancer, is one of serious diseases causing death for both men and women. Computer-aided diagnosis and survival prediction of NSCLC, is of great importance in providing assistance to diagnosis and personalize therapy planning for lung cancer patients. RESULTS: In this paper we have proposed an integrated framework for NSCLC computer-aided diagnosis and survival analysis using novel image markers. The entire biomedical imaging informatics framework consists of cell detection, segmentation, classification, discovery of image markers, and survival analysis. A robust seed detection-guided cell segmentation algorithm is proposed to accurately segment each individual cell in digital images. Based on cell segmentation results, a set of extensive cellular morphological features are extracted using efficient feature descriptors. Next, eight different classification techniques that can handle high-dimensional data have been evaluated and then compared for computer-aided diagnosis. The results show that the random forest and adaboost offer the best classification performance for NSCLC. Finally, a Cox proportional hazards model is fitted by component-wise likelihood based boosting. Significant image markers have been discovered using the bootstrap analysis and the survival prediction performance of the model is also evaluated. CONCLUSIONS: The proposed model have been applied to a lung cancer dataset that contains 122 cases with complete clinical information. The classification performance exhibits high correlations between the discovered image markers and the subtypes of NSCLC. The survival analysis demonstrates strong prediction power of the statistical model built from the discovered image markers. Fuyong Xing, Hai Su, Arnold J. Stromberg, Lin Yang 0002 |
BMC Bioinform. | 5 |
| 2014 | Automatic Myonuclear Detection in IsolatedSingle Muscle Fibers Using Robust EllipseFitting and Sparse RepresentationabstractAccurate and robust detection of myonuclei in isolated single muscle fibers is required to calculate myonuclear domain size. However, this task is challenging because: 1) shape and size variations of the nuclei, 2) overlapping nuclear clumps, and 3) multiple z-stack images with out-of-focus regions. In this paper, we have proposed a novel automatic detection algorithm to robustly quantify myonuclei in isolated single skeletal muscle fibers. The original z-stack images are first converted into one all-in-focus image using multi-focus image fusion. A sufficient number of ellipse fitting hypotheses are then generated from the myonuclei contour segments using heteroscedastic errors-in-variables (HEIV) regression. A set of representative training samples and a set of discriminative features are selected by a two-stage sparse model. The selected samples with representative features are utilized to train a classifier to select the best candidates. A modified inner geodesic distance based mean-shift clustering algorithm is used to produce the final nuclei detection results. The proposed method was extensively tested using 42 sets of z-stack images containing over 1,500 myonuclei. The method demonstrates excellent results that are better than current state-of-the-art approaches. Hai Su, Fuyong Xing, Jonah D. Lee, Charlotte A. Peterson, Lin Yang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2013 | An Integrated Framework for Automatic Ki-67 Scoring in Pancreatic Neuroendocrine Tumor
Fuyong Xing, Hai Su, Lin Yang 0002 |
MICCAI (1) | 3 |
| 2013 | Robust Selection-Based Sparse Shape Model for Lung Cancer Image Segmentation
Fuyong Xing, Lin Yang 0002 |
MICCAI (3) | 2 |
| 2013 | Robust Visual Tracking Using Local Sparse Appearance Model and K-SelectionabstractOnline learned tracking is widely used for its adaptive ability to handle appearance changes. However, it introduces potential drifting problems due to the accumulation of errors during the self-updating, especially for occluded scenarios. The recent literature demonstrates that appropriate combinations of trackers can help balance the stability and flexibility requirements. We have developed a robust tracking algorithm using a local sparse appearance model (SPT) and K-Selection. A static sparse dictionary and a dynamically updated online dictionary basis distribution are used to model the target appearance. A novel sparse representation-based voting map and a sparse constraint regularized mean shift are proposed to track the object robustly. Besides these contributions, we also introduce a new selection-based dictionary learning algorithm with a locally constrained sparse representation, called K-Selection. Based on a set of comprehensive experiments, our algorithm has demonstrated better performance than alternatives reported in the recent literature. Baiyang Liu, Junzhou Huang, Casimir A. Kulikowski, Lin Yang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Learning to translate with products of novices: a suite of open-ended challenge problems for teaching MTabstractMachine translation (MT) draws from several different disciplines, making it a complex subject to teach. There are excellent pedagogical texts, but problems in MT and current algorithms for solving them are best learned by doing. As a centerpiece of our MT course, we devised a series of open-ended challenges for students in which the goal was to improve performance on carefully constrained instances of four key MT tasks: alignment, decoding, evaluation, and reranking. Students brought a diverse set of techniques to the problems, including some novel solutions which performed remarkably well. A surprising and exciting outcome was that student solutions or their combinations fared competitively on some tasks, demonstrating that even newcomers to the field can help improve the state-of-the-art on hard NLP problems while simultaneously learning a great deal. The problems, baseline code, and results are freely available. Adam Lopez, Matt Post, Chris Callison-Burch, Jonathan Weese, Juri Ganitkevitch, Narges Ahmidi, Olivia Buzek, Leah Hanson, Beaniesh Jamil, Matthias A. Lee, Ya-Ting Lin, Henry Pao, Fatima Rivera, Leili Shahriyari, Debu Sinha, Adam R. Teichert, Stephen Wampler, Michael Weinberger, Daguang Xu, Lin Yang 0002, Shang Zhao 0002 |
Trans. Assoc. Comput. Linguistics | 20 |
| 2011 | Robust tracking using local sparse appearance model and K-selectionabstractOnline learned tracking is widely used for it's adaptive ability to handle appearance changes. However, it introduces potential drifting problems due to the accumulation of errors during the self-updating, especially for occluded scenarios. The recent literature demonstrates that appropriate combinations of trackers can help balance stability and flexibility requirements. We have developed a robust tracking algorithm using a local sparse appearance model (SPT). A static sparse dictionary and a dynamically online updated basis distribution model the target appearance. A novel sparse representation-based voting map and sparse constraint regularized mean-shift support the robust object tracking. Besides these contributions, we also introduce a new dictionary learning algorithm with a locally constrained sparse representation, called K-Selection. Based on a set of comprehensive experiments, our algorithm has demonstrated better performance than alternatives reported in the recent literature. Baiyang Liu, Junzhou Huang, Lin Yang 0002, Casimir A. Kulikowski |
CVPR | 3 |
| 2011 | ImageMiner: a software system for comparative analysis of tissue microarrays using content-based image retrieval, high-performance computing, and grid technologyabstractOBJECTIVE AND DESIGN: The design and implementation of ImageMiner, a software platform for performing comparative analysis of expression patterns in imaged microscopy specimens such as tissue microarrays (TMAs), is described. ImageMiner is a federated system of services that provides a reliable set of analytical and data management capabilities for investigative research applications in pathology. It provides a library of image processing methods, including automated registration, segmentation, feature extraction, and classification, all of which have been tailored, in these studies, to support TMA analysis. The system is designed to leverage high-performance computing machines so that investigators can rapidly analyze large ensembles of imaged TMA specimens. To support deployment in collaborative, multi-institutional projects, ImageMiner features grid-enabled, service-based components so that multiple instances of ImageMiner can be accessed remotely and federated. RESULTS: The experimental evaluation shows that: (1) ImageMiner is able to support reliable detection and feature extraction of tumor regions within imaged tissues; (2) images and analysis results managed in ImageMiner can be searched for and retrieved on the basis of image-based features, classification information, and any correlated clinical data, including any metadata that have been generated to describe the specified tissue and TMA; and (3) the system is able to reduce computation time of analyses by exploiting computing clusters, which facilitates analysis of larger sets of tissue samples. David J. Foran, Lin Yang 0002, Wenjin Chen, Lauri A. Goodell, Michael Reiss, Fusheng Wang 0001, Tahsin M. Kurç, Tony Pan, Ashish Sharma 0001, Joel H. Saltz |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Prediction Based Collaborative Trackers (PCT): A Robust and Accurate Approach Toward 3D Medical Object TrackingabstractRobust and fast 3D tracking of deformable objects, such as heart, is a challenging task because of the relatively low image contrast and speed requirement. Many existing 2D algorithms might not be directly applied on the 3D tracking problem. The 3D tracking performance is limited due to dramatically increased data size, landmarks ambiguity, signal drop-out or complex nonrigid deformation. In this paper, we present a robust, fast, and accurate 3D tracking algorithm: prediction based collaborative trackers (PCT). A novel one-step forward prediction is introduced to generate the motion prior using motion manifold learning. Collaborative trackers are introduced to achieve both temporal consistency and failure recovery. Compared with tracking by detection and 3D optical flow, PCT provides the best results. The new tracking algorithm is completely automatic and computationally efficient. It requires less than 1.5 s to process a 3D volume which contains millions of voxels. In order to demonstrate the generality of PCT, the tracker is fully tested on three large clinical datasets for three 3D heart tracking problems with two different imaging modalities: endocardium tracking of the left ventricle (67 sequences, 1134 3D volumetric echocardiography data), dense tracking in the myocardial regions between the epicardium and endocardium of the left ventricle (503 sequences, roughly 9000 3D volumetric echocardiography data), and whole heart four chambers tracking (20 sequences, 200 cardiac 3D volumetric CT data). Our datasets are much larger than most studies reported in the literature and we achieve very accurate tracking results compared with human experts' annotations and recent literature. Lin Yang 0002, Bogdan Georgescu, Yefeng Zheng 0001, Yang Wang 0001, Peter Meer, Dorin Comaniciu |
IEEE Trans. Medical Imaging | 1 |
| 2010 | Robust and Fast Collaborative Tracking with Two Stage Sparse Optimization
Baiyang Liu, Lin Yang 0002, Junzhou Huang, Peter Meer, Leiguang Gong, Casimir A. Kulikowski |
ECCV (4) | 2 |
| 2009 | A Parallel Point Matching Algorithm for Landmark Based Image Registration Using Multicore Platform
Lin Yang 0002, Leiguang Gong, John L. Nosher, David J. Foran |
Euro-Par | 1 |
| 2009 | Virtual Microscopy and Grid-Enabled Decision Support for Large-Scale Analysis of Imaged Pathology SpecimensabstractBreast cancer accounts for about 30% of all cancers and 15% of cancer deaths in women. Advances in computer-assisted analysis hold promise for classifying subtypes of disease and improving prognostic accuracy. We introduce a grid-enabled decision support system for performing automatic analysis of imaged breast tissue microarrays. To date, we have processed more than 1,00,000 digitized specimens (1200 x 1200 pixels each) on IBM's World Community Grid (WCG). As a part of the Help Defeat Cancer (HDC) project, we have analyzed that the data returned from WCG along with retrospective patient clinical profiles for a subset of 3744 breast tissue samples, and have reported the results in this paper. Texture-based features were extracted from the digitized specimens, and isometric feature mapping was applied to achieve nonlinear dimension reduction. Iterative prototyping and testing were performed to classify several major subtypes of breast cancer. Overall, the most reliable approach was gentle AdaBoost using an eight-node classification and regression tree as the weak learner. Using the proposed algorithm, a binary classification accuracy of 89% and the multiclass accuracy of 80% were achieved. Throughout the course of the experiments, only 30% of the dataset was used for training. Lin Yang 0002, Wenjin Chen, Peter Meer, Gratian Salaru, Lauri A. Goodell, Viktors Berstis, David J. Foran |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2009 | PathMiner: A Web-Based Tool for Computer-Assisted Diagnostics in PathologyabstractLarge-scale, multisite collaboration has become indispensable for a wide range of research and clinical activities that rely on the capacity of individuals to dynamically acquire, share, and assess images and correlated data. In this paper, we report the development of a Web-based system, PathMiner , for interactive telemedicine, intelligent archiving, and automated decision support in pathology. The PathMiner system supports network-based submission of queries and can automatically locate and retrieve digitized pathology specimens along with correlated molecular studies of cases from "ground-truth" databases that exhibit spectral and spatial profiles consistent with a given query image. The statistically most probable diagnosis is provided to the individual who is seeking decision support. To test the system under real-case scenarios, a pipeline infrastructure was developed and a network-based test laboratory was established at strategic sites at the University of Medicine and Dentistry of New Jersey-Robert Wood Johnson Medical School, Robert Wood Johnson University Hospital, the University of Pennsylvania School of Medicine, Hospital of the University of Pennsylvania, The Cancer Institute of New Jersey, and Rutgers University. The average five-class classification accuracy of the system was 93.18% based on a tenfold cross validation on a close dataset containing 3691 imaged specimens. We also conducted prospective performance studies with the PathMiner system in real applications in which the specimens exhibited large variations in staining characters compared with the training data. The average five-class classification accuracy in this open-set experiment was 87.22%. We also provide the comparative results with the previous literature and the PathMiner system shows superior performance. Lin Yang 0002, Oncel Tuzel, Wenjin Chen, Peter Meer, Gratian Salaru, Lauri A. Goodell, David J. Foran |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2008 | 3D ultrasound tracking of the left ventricle using one-step forward prediction and data fusion of collaborative trackersabstractTracking the left ventricle (LV) in 3D ultrasound data is a challenging task because of the poor image quality and speed requirements. Many previous algorithms applied standard 2D tracking methods to tackle the 3D problem. However, the performance is limited due to increased data size, landmarks ambiguity, signal drop-out or non-rigid deformation. In this paper we present a robust, fast and accurate 3D LV tracking algorithm. We propose a novel one-step forward prediction to generate the motion prior using motion manifold learning, and introduce two collaborative trackers to achieve both temporal consistency and failure recovery. Compared with tracking by detection and 3D optical flow, our algorithm provides the best results and sub-voxel accuracy. The new tracking algorithm is completely automatic and computationally efficient. It requires less than 1.5 seconds to process a 3D volume which contains 4,925,440 voxels. Lin Yang 0002, Bogdan Georgescu, Yefeng Zheng 0001, Peter Meer, Dorin Comaniciu |
CVPR | 1 |
| 2008 | Automatic Image Analysis of Histopathology Specimens Using Concave Vertex Graph
Lin Yang 0002, Oncel Tuzel, Peter Meer, David J. Foran |
MICCAI (1) | 1 |
| 2007 | Multiple Class Segmentation Using A Unified Framework over Mean-Shift PatchesabstractObject-based segmentation is a challenging topic. Most of the previous algorithms focused on segmenting a single or a small set of objects. In this paper, the multiple class object-based segmentation is achieved using the appearance and bag of keypoints models integrated over mean-shift patches. We also propose a novel affine invariant descriptor to model the spatial relationship of keypoints and apply the Elliptical Fourier Descriptor to describe the global shapes. The algorithm is computationally efficient and has been tested for three real datasets using less training samples. Our algorithm provides better results than other studies reported in the literature. Lin Yang 0002, Peter Meer, David J. Foran |
CVPR | 1 |
| 2007 | A Variational Framework for Partially Occluded Image Segmentation using Coarse to Fine Shape Alignment and Semi-Parametric Density ApproximationabstractIn this paper, we propose a variational framework which combines top-down and bottom-up information to address the challenge of partially occluded image segmentation. The algorithm applies shape priors and divides shape learning into shape mode clustering and non-rigid transformation estimation to handle intraclass and interclass coarse to fine variations. A semi-parametric density approximation using adaptive meanshift and L(2)E robust estimation is used to model the likelihood. A set of real images is used to show the good performance of the algorithm. Lin Yang 0002, David J. Foran |
ICIP (1) | 1 |
| 2007 | High Throughput Analysis of Breast Cancer Specimens on the Grid
Lin Yang 0002, Wenjin Chen, Peter Meer, Gratian Salaru, Michael D. Feldman, David J. Foran |
MICCAI (1) | 1 |
| 2007 | Classification of hematologic malignancies using texton signatures
Oncel Tuzel, Lin Yang 0002, Peter Meer, David J. Foran |
Pattern Anal. Appl. | 2 |
| 2005 | Unsupervised segmentation based on robust estimation and color active contour modelsabstractOne of the most commonly used clinical tests performed today is the routine evaluation of peripheral blood smears. In this paper, we investigate the design, development, and implementation of a robust color gradient vector flow (GVF) active contour model for performing segmentation, using a database of 1791 imaged cells. The algorithms developed for this research operate in Luv color space, and introduce a color gradient and L2E robust estimation into the traditional GVF snake. The accuracy of the new model was compared with the segmentation results using a mean-shift approach, the traditional color GVF snake, and several other commonly used segmentation strategies. The unsupervised robust color snake with L2E robust estimation was shown to provide results which were superior to the other unsupervised approaches, and was comparable with supervised segmentation, as judged by a panel of human experts. Lin Yang 0002, Peter Meer, David J. Foran |
IEEE Trans. Inf. Technol. Biomed. | 1 |