EDBT 2026 Demo / reviewers in the wild / expert
Shaohua Kevin Zhou
dblp:57/98 · also S. Kevin Zhou
· DBLP profile ↗
246ranked-venue papers
22as first author
131since 2021 · last 2026
0000-0002-6881-4444ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 163 · 15 first-author · 65 since 2021Applied, interdisciplinary, general and emerging computing · 132 · 4 first-author · 81 since 2021Artificial intelligence and machine learning · 77 · 13 first-author · 34 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical TextabstractArtificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limited to single-image, single-turn tasks, lacking multi-modal medical image integration and failing to capture the longitudinal and multi-modal interactive nature inherent to clinical practice. To address this gap, we introduce MedAtlas, a novel benchmark framework designed to evaluate large language models on realistic medical reasoning tasks. MedAtlas is characterized by four key features: multi-round visual question answering (VQA), Joint reasoning of multiple modalities of medical images, multi-task integration, and high clinical fidelity. It supports four core tasks: open-ended multi-round VQA, closed-ended multi-round VQA, multi-image joint reasoning, and comprehensive disease diagnosis. Each case is derived from real diagnostic workflows and incorporates temporal interactions between textual medical histories and multiple imaging modalities, including CT, MRI, PET, ultrasound, X-ray, etc., requiring models to perform deep integrative reasoning across images and clinical texts. MedAtlas provides expert-annotated gold standards for all tasks. Furthermore, we propose two novel evaluation metrics: Stage Chain Accuracy (SCA) and Error Propagation Suppression Coefficient (EPSC). Benchmark results with existing multi-modal models reveal substantial performance gaps in multi-stage clinical reasoning. MedAtlas establishes a challenging evaluation platform to advance the development of robust and trustworthy medical AI. Ronghao Xu, Zhen Huang 0007, Yangbo Wei, Xiaoqian Zhou, Zihang Jiang, Shaohua Kevin Zhou |
AAAI | 8 |
| 2026 | SIGHP: Scalable Information-Guided Hypergraph Partitioner
Huhao Guan, Zezhong Ding 0001, Ao Ke, Xike Xie, Shaohua Kevin Zhou |
KDD (1) | 5 |
| 2026 | BL-UDA: Towards Unsupervised Domain-Adaptive Surgical Instrument Segmentation with Source Box LabelsabstractRecent advances in unsupervised domain adaptation (UDA) by adapting the model from one domain to another unseen domain have shown considerable promise in improving surgical instrument segmentation performance across domains. However, existing UDA methods primarily rely on pixel-wise labels, which are always difficult to collect due to the labor-intensive annotation process. In this work, we aim to relax the dependence on pixel-level supervision and investigate a challenging UDA setting - source box annotations, where weak supervision and domain shifts coexist. To achieve this, we introduce a novel unsupervised domain adaptation framework, BL-UDA, which leverages bounding box annotations for surgical instrument segmentation across domains. By utilizing the Segment Anything Model (SAM) for pseudo label generation from box annotations, our method effectively bridges object-level and pixel-level domain adaptation. The proposed BL-UDA framework comprises a teacher-student network with entropy minimization for object detection and an entropy-based label selection strategy for generating box prompts to SAM, facilitating pixel-level domain adaptation. Extensive experiments on the EndoVis 2017 and 2018 datasets demonstrate the superiority of BL-UDA over existing UDA methods, significantly mitigating domain shifts and addressing weak supervision challenges with minimal annotation requirements. Ziyuan Zhao, Yifang Yin, Yichen Zhang 0002, Xulei Yang, Jun Cheng 0003, Roger Zimmermann, Cuntai Guan, Shaohua Kevin Zhou |
ICMR | 9 |
| 2026 | Equivariant Sampling for Improving Diffusion Model-based Image RestorationabstractRecent advances in generative models, especially diffusion models, have significantly improved image restoration (IR) performance. However, existing problem-agnostic diffusion model-based image restoration (DMIR) methods face challenges in fully leveraging diffusion priors, resulting in suboptimal performance. In this paper, we address the limitations of current problem-agnostic DMIR methods by analyzing their sampling process and providing effective solutions. We introduce EquS, a DMIR method that imposes equivariant information through dual sampling trajectories. To further boost EquS, we propose the Timestep-Aware Schedule (TAS) and introduce EquS+. TAS prioritizes deterministic steps to enhance certainty and sampling efficiency. Extensive experiments on benchmarks demonstrate that our method is compatible with previous problem-agnostic DMIR methods and significantly boosts their performance without increasing computational costs. Our code is available in https://github.com/FouierL/EquS. Chenxu Wu, Qingpeng Kong, Peiang Zhao, Wendi Yang, Fenghe Tang, Zihang Jiang, Shaohua Kevin Zhou |
WACV | 8 |
| 2026 | Unsupervised stain-aware pixel-adversarial transfer learning for virtual immunohistochemical staining
Qiuli Wang 0001, Yongxu Liu 0005, Yue Zhang 0042, Kaiyan Li 0005, Xianqi Wang 0002, Shaohua Kevin Zhou, Wei Chen 0090, Xiaohong Yao |
Knowl. Based Syst. | 8 |
| 2026 | Prediction model for pavement quality index (PQI) based on multi-source data fusion
Hansong Wu, Jinxi Zhang, Shaohua Kevin Zhou |
Knowl. Based Syst. | 5 |
| 2026 | PASS-Tr: PAtch-wise swin slice attention to leverage generalization of 2D large vision model to universal lesion detection
Jingsong Liu, Zhen Huang 0007, Xun Ma, Peter J. Schüffler, Nassir Navab, Shaohua Kevin Zhou |
Medical Image Anal. | 8 |
| 2026 | MoMBS: Mixed-order sampling improves training on heterogeneous-quality data for universal lesion detection
Jingsong Liu, Peter J. Schüffler, Hu Han 0001, Shaohua Kevin Zhou |
Medical Image Anal. | 5 |
| 2026 | CLIS: Causality-inspired Longitudinal Image Synthesis and its application to Alzheimer's disease characterization
Zhuowei Xu, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2026 | Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation
Fenghe Tang, Qingsong Yao, Chenxu Wu, Zihang Jiang, Shaohua Kevin Zhou |
Medical Image Anal. | 6 |
| 2026 | WSISum: WSI summarization via dual-level semantic reconstruction
Baizhi Wang, Kun Zhang 0040, Yunjie Gu, Haijing Luan, Taiyuan Hu, Zhidong Yang, Zihang Jiang, Rui Yan 0009, Shaohua Kevin Zhou |
Medical Image Anal. | 12 |
| 2026 | Pathway-Aware Multimodal Transformer (PAMT): Integrating Pathological Image and Gene Expression for Interpretable Cancer Survival AnalysisabstractIntegrating multimodal data of pathological image and gene expression for cancer survival analysis can achieve better results than using a single modality. However, existing multimodal learning methods ignore fine-grained interactions between both modalities, especially the interactions between biological pathways and pathological image patches. In this article, we propose a novel Pathway-Aware Multimodal Transformer (PAMT) framework for interpretable cancer survival analysis. Specifically, the PAMT learns fine-grained modality interaction through three stages: (1) In the intra-modal pathway-pathway / patch-patch interaction stage, we use the Transformer model to perform intra-modal information interaction; (2) In the inter-modal pathway-patch alignment stage, we introduce a novel label-free contrastive loss to aligns semantic information between different modalities so that the features of the two modalities are mapped to the same semantic space; and (3) In the inter-modal pathway-patch fusion stage, to model the medical prior knowledge of "genotype determines phenotype", we propose a pathway-to-patch cross fusion module to perform inter-modal information interaction under the guidance of pathway prior. In addition, the inter-modal cross fusion module of PAMT endows good interpretability, helping a pathologist to screen which pathway plays a key role, to locate where on whole slide image (WSI) are affected by the pathway, and to mine prognosis-relevant pathology image patterns. Experimental results based on three datasets of bladder urothelial carcinoma, lung squamous cell carcinoma, and lung adenocarcinoma demonstrate that the proposed framework significantly outperforms the state-of-the-art methods. Rui Yan 0009, Xueyuan Zhang, Zihang Jiang, Baizhi Wang, Xiuwu Bian, Shaohua Kevin Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | GSR: A Gaussian Splatting-Based Reconstruction Framework for EITabstractThis paper introduces 2D Gaussian Splatting (GS) to Electrical Impedance Tomography (EIT), marking its first application in this field. Initially developed for computer vision tasks such as scene reconstruction, GS enables continuous representation and efficient rendering of high-resolution images. Building on these capabilities, we propose a novel GS-based EIT reconstruction framework that models conductivity distributions as a set of Gaussian kernels. These kernels act as localized basis functions, dynamically adjusting their parameters (e.g., position, covariance, and amplitude) to enhance representation accuracy. To ensure regularization and physical constraints, we integrate a threshold-adjusted ReLU activation function to filter out insignificant components and a Sigmoid function to constrain conductivity values within a valid physical range. Experimental results on both simulated and real datasets demonstrate that our approach outperforms traditional model-driven methods and is competitive with conventional neural network-based methods in reconstruction quality. Furthermore, systematic ablation studies confirm the effectiveness of the key components of our framework. This work opens new possibilities for integrating advanced rendering techniques into EIT and inverse problem solving, bridging the gap between computer vision and biomedical imaging. Dong Liu 0007, Haoyuan Xia, Hongyan Xiang, Yukang Huang, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-Ray: Summary of the PENGWIN 2024 ChallengeabstractThe segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and efficiently delineating the bone fragments remains a significant challenge due to complex anatomy and imaging limitations. The PENGWIN challenge, organized as a MICCAI 2024 satellite event, aimed to advance automated fracture segmentation by benchmarking state-of-the-art algorithms on these complex tasks. A diverse dataset of 150 CT scans was collected from multiple clinical centers, and a large set of simulated X-ray images was generated using the DeepDRR method. Final submissions from 16 teams worldwide were evaluated under a rigorous multi-metric testing scheme. The top-performing CT algorithm achieved an average fragment-wise intersection over union (IoU) of 0.930, demonstrating satisfactory accuracy. However, in the X-ray task, the best algorithm achieved an IoU of 0.774, which is promising but not yet sufficient for intra-operative decision-making, reflecting the inherent challenges of fragment overlap in projection imaging. Beyond the quantitative evaluation, the challenge revealed methodological diversity in algorithm design. Variations in instance representation, such as primary-secondary classification versus boundary-core separation, led to differing segmentation strategies. Despite promising results, the challenge also exposed inherent uncertainties in fragment definition, particularly in cases of incomplete fractures. These findings suggest that interactive segmentation approaches, integrating human decision-making with task-relevant information, may be essential for improving model reliability and clinical applicability. Yudi Sang, Yanzhen Liu, Sutuke Yibulayimu, Yunning Wang, Benjamin Killeen, Mingxu Liu, Ping-Cheng Ku, Ole Johannsen, Karol Gotkowski, Maximilian Zenk, Klaus H. Maier-Hein, Fabian Isensee, Peiyan Yue, Yi Wang 0031, Zhaohong Pan, Xiaokun Liang, Daiqi Liu, Fuxin Fan, Artur Jurgas, Andrzej Skalski, Szymon Plotka, Rafal Litka, Yingchun Song, Mathias Unberath, Mehran Armand, Dan Ruan, Shaohua Kevin Zhou, Qiyong Cao, Chunpeng Zhao, Xinbao Wu, Yu Wang 0083 |
IEEE Trans. Medical Imaging | 32 |
| 2025 | GraphInsight: Unlocking Insights in Large Language Models for Graph Structure UnderstandingabstractAlthough Large Language Models (LLMs) have demonstrated potential in processing graphs, they struggle with comprehending graphical structure information through prompts of graph description sequences, especially as the graph size increases.We attribute this challenge to the uneven memory performance of LLMs across different positions in graph description sequences, known as "Positional bias".To address this, we propose GraphInsight, a novel framework aimed at improving LLMs' comprehension of both macro-and micro-level graphical information.GraphInsight is grounded in two key strategies: 1) placing critical graphical information in positions where LLMs exhibit stronger memory performance, and 2) investigating a lightweight external knowledge base for regions with weaker memory performance, inspired by retrieval-augmented generation (RAG).Moreover, GraphInsight explores integrating these two strategies into LLM agent processes for composite graph tasks that require multi-step reasoning.Extensive empirical studies on benchmarks with a wide range of evaluation tasks show that GraphInsight significantly outperforms all other graph description methods (e.g., prompting techniques and reordering strategies) in understanding graph structures of varying sizes. Yukun Cao, Zengyi Gao, Zezhong Ding 0001, Xike Xie, Shaohua Kevin Zhou |
ACL (1) | 6 |
| 2025 | Distilling Dataset Summarization via Constrained Diffusion Model for Shareable Synthetic Histopathology ImagesabstractThe rapid advancement of deep learning technologies significantly influences the field of histopathology. The primary challenges encountered include the necessity for large datasets and the protection of medical data privacy. Federated learning has emerged as a potential solution to these challenges by enabling local model training and subsequent uploading of model parameters for updates. Nevertheless, issues related to data distribution bias persist. Dataset distillation offers a novel approach to condensing extensive real datasets into smaller synthetic datasets. This method is not readily applicable in histopathology due to its reliance on human-unreadable representations, which are not suitable for pathologists' diagnostic processes, as well as its inadequate performance in downstream tasks. Some researchers suggest employing generative models to create samples for selection. Currently, challenges remain concerning the quality of the generated samples and the encapsulation of a maximally informative synthetic dataset. Our proposed method, DisDSCD(Distilling Dataset Summarization via Constrained Diffusion model), aims to address these issues. We develop a latent diffusion model and establish a constrained framework that integrates real images to regulate the generated images, thereby ensuring the quality of the synthesized outputs. Additionally, we select core and edge samples within clusters based on cosine distance in a specified ratio within the embedding space, collectively forming a summarized dataset. The performance of dataset summarization is evaluated by comparing it to the upper bound established using the complete real dataset in downstream tasks. Although the images included in the dataset summarization are synthetic and constitute merely 5 % of the real dataset, it attains comparable results to the full dataset in terms of accuracy (ACC) and area under the curve (AUC), and achieves 95% of the full dataset's performance with respect to the F1 score. Codes are available at https://anonymous.4open.science/r/DisDSCD. Wenqing Ye, Shaohua Kevin Zhou |
BIBM | 3 |
| 2025 | KANTrust: A Multi-Omics Framework for Uncertainty-Aware Disease SubtypingabstractThe integration of multi-omics data, including DNA methylation, mRNA expression, and miRNA profiles, is crucial for accurate disease subtyping and outcome prediction in complex disorders such as Alzheimer's disease and various cancers. However, the inherent heterogeneity and inconsistency among omics views present significant challenges for reliable data fusion. To address these issues, we propose KANTrust, a novel framework for trustworthy multi-omics classification that explicitly models both epistemic and aleatoric uncertainties. Our method combines a Kolmogorov-Arnold Network (KAN)enhanced robust representation module, a contrastive evidence consistency module, and an evidence-theoretic fusion module to achieve reliable multi-view integration. KANTrust adaptively highlights informative features within each omics modality, promotes semantic alignment across views, and quantifies uncertainty through a Dempster-Shafer framework. Experimental evaluations on four real-world biomedical datasets demonstrate that KANTrust consistently outperforms state-of-the-art methods in both binary and multi-class classification tasks. Code is available at https://github.com/wcj6/KANTrust. Chunjiang Wang, Rui Yan 0009, Kun Zhang 0040, Zihang Jiang, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
BIBM | 7 |
| 2025 | ICP: Immediate Compensation Pruning for Mid-to-high SparsityabstractThe increasing adoption of large-scale models under 7 billion parameters in both language and vision domains enables inference tasks on a single consumer-grade GPU but makes fine-tuning models of this scale, especially 7B models, challenging. This limits the applicability of pruning methods that require full fine-tuning. Meanwhile, pruning methods that do not require fine-tuning perform well at low sparsity levels (10%-50%) but struggle at mid-to-high sparsity levels (50%-70%), where the error behaves equivalently to that of semi-structured pruning. To address these issues, this paper introduces ICP, which finds a balance between full fine-tuning and zero fine-tuning. First, Sparsity Rearrange is used to reorganize the predefined sparsity levels, followed by Block-wise Compensate Pruning, which alternates pruning and compensation on the model’s backbone, fully utilizing inference results while avoiding full model fine-tuning. Experiments show that ICP improves performance at mid-to-high sparsity levels compared to baselines, with only a slight increase in pruning time and no additional peak memory overhead. Xueming Fu, Zihang Jiang, Shaohua Kevin Zhou |
CVPR | 4 |
| 2025 | AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIPabstractAnomaly detection (AD) identifies outliers for applications like defect and lesion detection. While CLIP shows promise for zero-shot AD tasks due to its strong generalization capabilities, its inherent Anomaly-Unawareness leads to limited discrimination between normal and abnormal features. To address this problem, we propose Anomaly-Aware CLIP (AA-CLIP), which enhances CLIP's anomaly discrimination ability in both text and visual spaces while preserving its generalization capability. AA-CLIP is achieved through a straightforward yet effective two-stage approach: it first creates anomaly-aware text anchors to differentiate normal and abnormal semantics clearly, then aligns patch-level visual features with these anchors for precise anomaly localization. This two-stage strategy, with the help of residual adapters, gradually adapts CLIP in a controlled manner, achieving effective AD while maintaining CLIP's class knowledge. Extensive experiments validate AA-CLIP as a resource-efficient solution for zero-shot AD tasks, achieving state-of-the-art results in industrial and medical applications. The code is available at https://github.com/Mwxinnn/AA-CLIP. Qingsong Yao, Fenghe Tang, Chenxu Wu, Yingtai Li, Rui Yan 0009, Zihang Jiang, Shaohua Kevin Zhou |
CVPR | 9 |
| 2025 | DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningabstractVision-Language (VL) alignment across image and text modalities is a challenging task due to the inherent semantic ambiguity of data with multiple possible meanings. Existing methods typically solve it by learning multiple sub-representation spaces to encode each input data as a set of embeddings, and constraining diversity between whole subspaces to capture diverse semantics for accurate VL alignment. Despite their promising outcomes, existing methods suffer two imperfections: 1) actually, specific semantics is mainly expressed by some local dimensions within the subspace. Ignoring this intrinsic property, existing diversity constraints imposed on the whole subspace may impair diverse embedding learning; 2) multiple embeddings are inevitably introduced, sacrificing computational and storage efficiency. In this paper, we propose a simple yet effective Diverse and Hybrid Set-embeddings learning framework (DH-Set), which is distinct from prior work in three aspects. DH-Set 1) devises a novel semantic importance dissecting method to focus on key local dimensions within each subspace; and thereby 2) not only imposes finer-grained diversity constraint to improve the accuracy of diverse embedding learning, 3) but also mixes key dimensions of all subspaces into the single hybrid embedding to boost inference efficiency. Extensive experiments on various benchmarks and model backbones show the superiority of DH-Set over state-of-the-art methods, achieving substantial 2.3%-14.7% rSum improvements while lowering computational and storage complexity. Kun Zhang 0040, Zhe Li 0028, Shaohua Kevin Zhou |
CVPR | 4 |
| 2025 | Thinning a Medical Image Segmentation Model via Dual-Level Multiscale FusionabstractMedical image segmentation plays a pivotal role in disease diagnosis and treatment planning, particularly in resource-constrained clinical settings where lightweight and generalizable models are urgently needed. However, existing lightweight models often compromise performance for efficiency and rarely adopt computationally expensive attention mechanisms, severely restricting their global contextual perception capabilities. Additionally, current architectures neglect the channel redundancy issue under the same convolutional kernels in medical imaging, which hinders effective feature extraction. To address these challenges, we propose LGMSNet, a novel lightweight framework based on local and global dual multiscale that achieves state-of-the-art performance with minimal computational overhead. LGMSNet employs heterogeneous intra-layer kernels to extract local high-frequency information while mitigating channel redundancy. In addition, the model integrates sparse transformer-convolutional hybrid branches to capture low-frequency global information. Extensive experiments across six public datasets demonstrate LGMSNet’s superiority over existing state-of-the-art methods. In particular, LGMSNet maintains exceptional performance in zero-shot generalization tests on four unseen datasets, underscoring its potential for real-world deployment in resource-limited medical scenarios. The whole project code is in https://github.com/cq-dong/LGMSNet. Chengqi Dong, Fenghe Tang, Rongge Mao, Xinpei Gao, Shaohua Kevin Zhou |
ECAI | 5 |
| 2025 | A Unified Framework for Few-Shot Medical Image Classification via Multi-agent Description Generation and Refined Contrastive Learning
Shenghao Chen, Zhen Huang 0007, Xiaoqian Zhou, Shaohua Kevin Zhou |
ICIC (28) | 4 |
| 2025 | Eliminating Ambiguities in One-Shot Medical Landmark Detection via Mask Drawing
Zhen Huang 0007, Xiaoqian Zhou, Shaohua Kevin Zhou |
ICIC (5) | 4 |
| 2025 | Self-Supervised Diffusion MRI Denoising via Iterative and Stable RefinementabstractMagnetic Resonance Imaging (MRI), including diffusion MRI (dMRI), serves as a ``microscope'' for anatomical structures and routinely mitigates the influence of low signal-to-noise ratio scans by compromising temporal or spatial resolution. However, these compromises fail to meet clinical demands for both efficiency and precision. Consequently, denoising is a vital preprocessing step, particularly for dMRI, where clean data is unavailable. In this paper, we introduce Di-Fusion, a fully self-supervised denoising method that leverages the latter diffusion steps and an adaptive sampling process. Unlike previous approaches, our single-stage framework achieves efficient and stable training without extra noise model training and offers adaptive and controllable results in the sampling process. Our thorough experiments on real and simulated data demonstrate that Di-Fusion achieves state-of-the-art performance in microstructure modeling, tractography tracking, and other downstream tasks. Code is available at https://github.com/FouierL/Di-Fusion. Chenxu Wu, Qingpeng Kong, Zihang Jiang, Shaohua Kevin Zhou |
ICLR | 4 |
| 2025 | Lego Sketch: A Scalable Memory-augmented Neural Network for Sketching Data StreamsabstractSketches, probabilistic structures for estimating item frequencies in infinite data streams with limited space, are widely used across various domains. Recent studies have shifted the focus from handcrafted sketches to neural sketches, leveraging memory-augmented neural networks (MANNs) to enhance the streaming compression capabilities and achieve better space-accuracy trade-offs. However, existing neural sketches struggle to scale across different data domains and space budgets due to inflexible MANN configurations. In this paper, we introduce a scalable MANN architecture that brings to life the Lego sketch, a novel sketch with superior scalability and accuracy. Much like assembling creations with modular Lego bricks, the Lego sketch dynamically coordinates multiple memory bricks to adapt to various space budgets and diverse data domains. Theoretical analysis and empirical studies demonstrate its scalability and superior space-accuracy trade-offs, outperforming existing handcrafted and neural sketches. Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou |
ICML | 5 |
| 2025 | Prototype-based Optimal Transport for Out-of-Distribution DetectionabstractDetecting Out-of-Distribution (OOD) inputs is crucial for improving the reliability of deep neural networks in the real-world deployment. In this paper, inspired by the inherent distribution shift between in-distribution (ID) and OOD data, we propose a novel method that leverages optimal transport to measure the distribution discrepancy between test inputs and ID prototypes. The resulting transport costs are used to quantify the individual contribution of each test input to the overall discrepancy, serving as a desirable measure for OOD detection. To address the issue that solely relying on the transport costs to ID prototypes is inadequate for identifying OOD inputs closer to ID data, we generate virtual outliers to approximate the OOD region via linear extrapolation. By combining the transport costs to ID prototypes with the costs to virtual outliers, the detection of OOD data near ID data is emphasized, thereby enhancing the distinction between ID and OOD inputs. Extensive evaluations demonstrate the superiority of our method over state-of-the-art methods. Ao Ke, Chuanwen Feng, Yukun Cao, Xike Xie, Shaohua Kevin Zhou, Lei Feng 0006 |
IJCAI | 6 |
| 2025 | MVP-CBM: Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image ClassificationabstractThe concept bottleneck model (CBM), as a technique improving interpretability via linking predictions to human-understandable concepts, makes high-risk and life-critical medical image classification credible. Typically, existing CBM methods associate the final layer of visual encoders with concepts to explain the model’s predictions. However, we empirically discover the phenomenon of concept preference variation, that is, the concepts are preferably associated with the features at different layers than those only at the final layer; yet a blind last-layer-based association neglects such a preference variation and thus weakens the accurate correspondences between features and concepts, impairing model interpretability. To address this issue, we propose a novel Multi-layer Visual Preference-enhanced Concept Bottleneck Model (MVP-CBM), which comprises two key novel modules: (1) intra-layer concept preference modeling, which captures the preferred association of different concepts with features at various visual layers, and (2) multi-layer concept sparse activation fusion, which sparsely aggregates concept activations from multiple layers to enhance performance. Thus, by explicitly modeling concept preferences, MVP-CBM can comprehensively leverage multi-layer visual information to provide a more nuanced and accurate explanation of model decisions. Extensive experiments on several public medical classification benchmarks demonstrate that MVP-CBM achieves state-of-the-art accuracy and interoperability, verifying its superiority. Code is available at https://github.com/wcj6/MVP-CBM. Chunjiang Wang, Kun Zhang 0040, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
IJCAI | 6 |
| 2025 | Slide-CLIP: A Simple and Effective Pruning Method for CLIPabstractPre-trained vision-language models have achieved great zero-shot performance in various downstream tasks. With the rapid development of vision-language models, many task-specific Contrastive Language-Image Pre-training (CLIP) models are proposed and utilized in diverse domains. However, their large model size hinders their utilization on platforms with limited hardware resources. Recently, several CLIP pruning methods have been proposed, we find they are effective but resources-consuming, costing a great amount of GPU hours at pruning stage. At the same time, we notice the emergence of fast large language model (LLM) pruning method such as SparseGPT which is simple and effective. All these lead to the question: Does there exist a simple but effective pruning method for CLIP? In this paper, we propose Slide-CLIP as an answer, which incorporates (i) applying a fast pruning method such as SparseGPT to obtain a specific layer with desired sparsity; (ii) applying Layer-wise Sliding Distillation (LSD) at the next layer of the pruned layer to minimize the mean squared error (MSE) of the next layer output before pruning and after pruning; (iii) iterating the pruning routine of (i) and (ii) in a slide manner until the last layer of model is reached; and (iv) applying a progressive process on top of (iii) at sparsity above 50% to minimize the performance drop. Extensive experiments on various CLIP models demonstrate the effectiveness of the proposed Slide-CLIP pruning method. The code is publicly available at https://github.com/bill426/Slide-CLIP Jiaxin Shi, Shaohua Kevin Zhou |
IJCNN | 3 |
| 2025 | Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuityabstractIn robotic surgery, instrument presence labels are typically recorded alongside video streams, offering a cost-effective alternative to manual annotations for segmentation tasks. Label-supervised surgical instrument segmentation (SIS), a weakly supervised segmentation setting where only instrument presence labels are available, remains underexplored due to its inherently ill-posed nature. Temporal information plays a vital role in capturing sequential dependencies, thereby enhancing representation learning even under incomplete supervision. This paper extends a two-stage label-supervised segmentation framework by leveraging the temporal characteristics of surgical videos from three perspectives. First, a temporal equivariance constraint is introduced to enforce pixel-level consistency across adjacent frames. Second, a class-aware semantic continuity constraint is applied to preserve coherence between global and local regions over time. Third, temporally-enhanced pseudo masks are generated from consecutive frames to suppress irrelevant regions and improve segmentation accuracy. We evaluate our method on two surgical video datasets: the Cholec80 cholecystectomy benchmark and a real-world robotic left lateral segmentectomy (RLLS) dataset. Instance-level instrument annotations, sampled at regular intervals and validated by an experienced clinician, provide a reliable basis for evaluation. Experimental results demonstrate that our method consistently achieves favorable performances over state-of-the-art methods. These findings highlight the effectiveness of incorporating temporal constraints into label-supervised frameworks, offering a promising strategy to reduce annotation costs and advance surgical video analysis. Qiyuan Wang 0001, Yanzhe Liu, Shang Zhao 0004, Shaohua Kevin Zhou |
IROS | 5 |
| 2025 | PATE: Enhancing Few-Shot Pathological Image Classification via Prompt-Based Text-Image Embedding Adaptation
Shenghao Chen, Zhen Huang 0007, Xiaoqian Zhou, Chunjiang Wang, Shaohua Kevin Zhou |
MICCAI (6) | 6 |
| 2025 | Dyna3DGR: 4D Cardiac Motion Tracking with Dynamic 3D Gaussian Representation
Xueming Fu, Yingtai Li, Zihang Jiang, Junhao Mei, Gaojun Teng, Shaohua Kevin Zhou |
MICCAI (2) | 9 |
| 2025 | More Performant and Scalable: Rethinking Contrastive Vision-Language Pre-training of Radiology in the LLM Era
Yingtai Li, Haoran Lai, Xiaoqian Zhou, Shuai Ming, Wei Wei 0006, Shaohua Kevin Zhou |
MICCAI (7) | 7 |
| 2025 | RadGS-Reg: Registering Spine CT with Biplanar X-Rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration
Xueming Fu, Luming Nong, Shaohua Kevin Zhou |
MICCAI (9) | 9 |
| 2025 | Pre-trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
Fenghe Tang, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou |
MICCAI (10) | 6 |
| 2025 | SimCroP: Radiograph Representation Learning with Similarity-Driven Cross-Granularity Pre-training
Rongsheng Wang 0003, Fenghe Tang, Qingsong Yao, Rui Yan 0009, Zhen Huang 0007, Haoran Lai, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou |
MICCAI (5) | 11 |
| 2025 | NeRF-Based CBCT Reconstruction Needs Normalization and Initialization
Zhuowei Xu, Dai Sun, Qingpeng Kong, Nassir Navab, Shaohua Kevin Zhou |
MICCAI (16) | 9 |
| 2025 | U-RWKV: Lightweight Medical Image Segmentation with Direction-Adaptive RWKV
Hongbo Ye, Fenghe Tang, Peiang Zhao, Zhen Huang 0007, Dexin Zhao, Minghao Bian, Shaohua Kevin Zhou |
MICCAI (11) | 7 |
| 2025 | Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationabstractIn clinical practice, medical image analysis often requires efficient execution on resource-constrained mobile devices. However, existing mobile models-primarily optimized for natural images-tend to perform poorly on medical tasks due to the significant information density gap between natural and medical domains. Combining computational efficiency with medical imaging-specific architectural advantages remains a challenge when developing lightweight, universal, and high-performing networks. To address this, we propose a mobile model called Mobile U-shaped Vision Transformer (Mobile U-ViT) tailored for medical image segmentation. Specifically, we employ the newly proposed ConvUtr as a hierarchical patch embedding, featuring a parameter-efficient large-kernel CNN with inverted bottleneck fusion. This design exhibits transformer-like representation learning capacity while being lighter and faster. To enable efficient local-global information exchange, we introduce a novel Large-kernel Local-Global-Local (LKLGL) block that effectively balances the low information density and high-level semantic discrepancy of medical images. Finally, we incorporate a shallow and lightweight transformer bottleneck for long-range modeling and employ a cascaded decoder with downsampled skip connections for dense prediction. Despite its reduced computational demands, our medical-optimized architecture achieves state-of-the-art performance across eight public 2D and 3D datasets covering diverse imaging modalities, including zero-shot testing on four unseen datasets. These results establish it as an efficient yet powerful and generalization solution for mobile medical image analysis. Code is available at: https://github.com/FengheTan9/Mobile-U-ViT. Fenghe Tang, Bingkun Nian, Jianrui Ding, Quan Quan, Chengqi Dong, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou |
ACM Multimedia | 9 |
| 2025 | LoCo: Training-Free Layout-to-Image Synthesis with Localized Constraints
Peiang Zhao, Ruiyang Jin, Shaohua Kevin Zhou |
ACM Multimedia | 4 |
| 2025 | Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceabstractLarge Language Models have excelled in various domains but face efficiency challenges due to the growing Key-Value (KV) cache required for long-sequence inference. Recent efforts aim to reduce KV cache size by evicting vast non-critical cache elements during runtime while preserving generation quality. However, these methods typically allocate compression budgets uniformly across all attention heads, ignoring the unique attention patterns of each head. In this paper, we establish a theoretical loss upper bound between pre- and post-eviction attention output, explaining the optimization target of prior cache eviction methods, while guiding the optimization of adaptive budget allocation. Base on this, we propose {\it Ada-KV}, the first head-wise adaptive budget allocation strategy. It offers plug-and-play benefits, enabling seamless integration with prior cache eviction methods. Extensive evaluations on 13 datasets from Ruler and 16 datasets from LongBench, all conducted under both question-aware and question-agnostic scenarios, demonstrate substantial quality improvements over existing methods. Our code is available at https://github.com/FFY0/AdaKV. Junlin Lv, Yukun Cao, Xike Xie, Shaohua Kevin Zhou |
NeurIPS | 5 |
| 2025 | MS-Glance: Bio-Inspired Non-Semantic Context Vectors and Their Applications in Supervising Image ReconstructionabstractNon-semantic context information is crucial for visual recognition, as the human visual perception system first uses global statistics to process scenes rapidly before identifying specific objects. However, while semantic information is increasingly incorporated into computer vision tasks such as image reconstruction, non-semantic information, such as global spatial structures, is often overlooked. To bridge the gap, we propose a biologically informed non-semantic context descriptor, MS-Glance, along with the Glance Index Measure for comparing two images. A Global Glance vector is formulated by randomly retrieving pixels based on a perception-driven rule from an image to form a vector representing non-semantic global context, while a local Glance vector is a flattened local image window, mimicking a zoom in observation. The Glance Index is defined as the inner product of two standardized sets of Glance vectors. We evaluate the effectiveness of incorporating Glance supervision in two reconstruction tasks: image fitting with implicit neural representation (INR) and undersampled MRI reconstruction. Extensive experimental results show that MS-Glance outperforms existing image restoration losses across both natural and medical images. The code is available at https://github.com/Z7Gao/MSGlance. Wendi Yang, Lei Xing 0001, Shaohua Kevin Zhou |
WACV | 5 |
| 2025 | Towards Accurate Unified Anomaly SegmentationabstractUnsupervised anomaly detection (UAD) from images strives to model normal data distributions, creating discriminative representations to distinguish and precisely localize anomalies. Despite recent advancements in the efficient and unified one-for-all scheme, challenges persist in accurately segmenting anomalies for further monitoring. Moreover, this problem is obscured by the widely-used AUROC metric under imbalanced UAD settings. This motivates us to emphasize the significance of precise segmentation of anomaly pixels using pAP and DSC as metrics. To address the unsolved segmentation task, we introduce the Unified Anomaly Segmentation (UniAS). UniAS presents a multi-level hybrid pipeline that progressively enhances normal information from coarse to fine, incorporating a novel multi-granularity gated CNN (MGG-CNN) into Transformer layers to explicitly aggregate local details from different granularities. UniAS achieves state-of-the-art anomaly segmentation performance, attaining 65.12/59.33 and 40.06/32.50 in pAP/DSC on the MVTec-AD and VisA datasets, respectively, surpassing previous methods significantly. The codes are shared at https://github.com/Mwxinnn/UniAS. Qingsong Yao, Zhelong Huang, Zihang Jiang, Shaohua Kevin Zhou |
WACV | 6 |
| 2025 | PostoMETRO: Pose Token Enhanced Mesh Transformer for Robust 3D Human Mesh RecoveryabstractWith the recent advancements in single-image-based 3D human pose and shape estimation (3DHPSE), there is a growing amount of works that can achieve good results on standard benchmarks but struggle to yield accurate hu-man mesh in extreme scenarios like occlusion. Previous works propose to leverage 2D poses to help 3D HPSE model improve performance under occlusion, but usually rely on manual design to integrate 2D poses and only aim for specific kinds of occlusion. In this paper, we present PostoMETRO (Pose token enhanced MEsh TRansfOrmer), which integrates 2D pose prior knowledge as tokens into transformers to improve model's performance under occlusion. Using a VQ- VAE-based pose tokenizer; we efficiently represent 2D poses as tokens andfeed them to transformers together with image tokens. Subsequently, these tokens are queried by vertex andJoint tokens to decode 3D coordinates of mesh vertices and human Joints. Our proposed 2D poses integration strategy is manual-design-free and suitable for various kinds of occlusion. Experiments on both standard and occlusion-specific benchmarks demonstrate the effectiveness of PostoMETRO. Code will be made available11https://github.com/PostoMETRO/PostoMETRO-Paper. Wendi Yang, Zihang Jiang, Shang Zhao 0004, Shaohua Kevin Zhou |
WACV | 4 |
| 2025 | 3DGR-CT: Sparse-view CT reconstruction with a 3D Gaussian representationabstractSparse-view computed tomography (CT) reduces radiation exposure by acquiring fewer projections, making it a valuable tool in clinical scenarios where low-dose radiation is essential. However, this often results in increased noise and artifacts due to limited data. In this paper we propose a novel 3D Gaussian representation (3DGR) based method for sparse-view CT reconstruction. Inspired by recent success in novel view synthesis driven by 3D Gaussian splatting, we leverage the efficiency and expressiveness of 3D Gaussian representation as an alternative to implicit neural representation. To unleash the potential of 3DGR for CT imaging scenario, we propose two key innovations: (i) FBP-image-guided Guassian initialization and (ii) efficient integration with a differentiable CT projector. Extensive experiments and ablations on diverse datasets demonstrate the proposed 3DGR-CT consistently outperforms state-of-the-art counterpart methods, achieving higher reconstruction accuracy with faster convergence. Furthermore, we showcase the potential of 3DGR-CT for real-time physical simulation, which holds important clinical applications while challenging for implicit neural representations. Code available at: https://github.com/SigmaLDC/3DGR-CT. Yingtai Li, Xueming Fu, Shang Zhao 0004, Ruiyang Jin, Shaohua Kevin Zhou |
Medical Image Anal. | 6 |
| 2025 | O-PRESS: Boosting OCT axial resolution with Prior guidance, Recurrence, and Equivariant Self-Supervision
Kaiyan Li 0005, Jingyuan Yang 0015, Wenxuan Liang, Xingde Li, Lulu Chen, Chan Wu, Xiao Zhang 0059, Zhiyan Xu, Yueling Wang, Lihui Meng, Yue Zhang 0042, Youxin Chen, Shaohua Kevin Zhou |
Medical Image Anal. | 14 |
| 2025 | MambaMIM: Pre-training Mamba with state space token interpolation and its application to medical image segmentation
Fenghe Tang, Bingkun Nian, Yingtai Li, Zihang Jiang, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou |
Medical Image Anal. | 7 |
| 2025 | ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training
Rongsheng Wang 0003, Qingsong Yao, Zihang Jiang, Haoran Lai, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
Medical Image Anal. | 7 |
| 2025 | LACOSTE: Exploiting stereo and temporal contexts for surgical instrument segmentation
Qiyuan Wang 0001, Shang Zhao 0004, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2025 | Skeleton2Mask: Skeleton-supervised airway segmentation
Mingyue Zhao, Xiuxiu Zhou, Li Fan 0002, Xiaolan Qiu, Shaohua Kevin Zhou |
Medical Image Anal. | 9 |
| 2025 | Capsule: An Out-of-Core Training Mechanism for Colossal GNNsabstractCutting-edge platforms of graph neural networks (GNNs), such as DGL and PyG, harness the parallel processing power of GPUs to extract structural information from graph data, achieving state-of-the-art (SOTA) performance in fields such as recommendation systems, knowledge graphs, and bioinformatics. Despite the computational advantages provided by GPUs, these GNN platforms struggle with scalability challenges due to the colossal graphical structures processed and the limited memory capacities of GPUs. In response, this work introduces Capsule, a new out-of-core mechanism for large-scale GNN training. Unlike existing out-of-core GNN systems, which use main or secondary memory as operative memory and use CPU kernels during non-backpropagation computation, Capsule uses GPU memory and GPU kernels. By substantially leveraging the parallelization capabilities of GPUs, Capsule significantly enhances GNN training efficiency. In addition, Capsule can be smoothly integrated to mainstream open-source GNN frameworks, DGL and PyG, in a play-and-plug manner. Through a prototype implementation and comprehensive experiments on real datasets, we demonstrate that Capsule can achieve up to a 12.02× improvement in runtime efficiency, while using only 22.24% of the main memory, compared to SOTA out-of-core GNN systems. Yongan Xiang, Zezhong Ding 0001, Shangyou Wang, Xike Xie, Shaohua Kevin Zhou |
Proc. ACM Manag. Data | 6 |
| 2025 | LEGO-GraphRAG: Modularizing Graph-based Retrieval-Augmented Generation for Design Space ExplorationabstractGraphRAG integrates (knowledge) graphs with large language models (LLMs) to improve reasoning accuracy and contextual relevance. Despite its promising applications and strong relevance to multiple research communities, such as databases and natural language processing, GraphRAG currently lacks modular workflow analysis, systematic solution frameworks, and insightful empirical studies. To bridge these gaps, we propose LEGO-GraphRAG , a modular framework that enables: 1 ) fine-grained decomposition of the GraphRAG workflow, 2 ) systematic classification of existing techniques and implemented GraphRAG instances, and 3 ) creation of new GraphRAG instances. Our framework facilitates comprehensive empirical studies of GraphRAG on large-scale real-world graphs and diverse query sets, revealing insights into balancing reasoning quality, runtime efficiency, and token or GPU cost, that are essential for building advanced GraphRAG systems. Yukun Cao, Zengyi Gao, Xike Xie, Shaohua Kevin Zhou, Jianliang Xu |
Proc. VLDB Endow. | 5 |
| 2025 | SRS: Siamese Reconstruction-Segmentation Network Based on Dynamic-Parameter ConvolutionabstractDynamic convolution demonstrates outstanding representation capabilities, which are crucial for natural image segmentation. However, it fails when applied to medical image segmentation (MIS) and infrared small target segmentation (IRSTS) due to limited data and limited fitting capacity. In this paper, we propose a new type of dynamic convolution called dynamic parameter convolution (DPConv) which shows superior fitting capacity, and it can efficiently leverage features from deep layers of encoder in reconstruction tasks to generate DPConv kernels that adapt to input variations. Moreover, we observe that DPConv, built upon deep features derived from reconstruction tasks, significantly enhances downstream segmentation performance. We refer to the segmentation network integrated with DPConv generated from reconstruction network as the siamese reconstruction-segmentation network (SRS). We conduct extensive experiments on seven datasets including five medical datasets and two infrared datasets, and the experimental results demonstrate that our method can show superior performance over several recently proposed methods. Furthermore, the zero-shot segmentation under unseen modality demonstrates the generalization of DPConv. The code is available at: https://github.com/fidshu/SRSNet. Bingkun Nian, Fenghe Tang, Jianrui Ding, Jie Yang 0002, Zhonglong Zheng, Shaohua Kevin Zhou, Wei Liu 0044 |
IEEE Trans. Image Process. | 6 |
| 2025 | IGU-Aug: Information-Guided Unsupervised Augmentation and Pixel-Wise Contrastive Learning for Medical Image AnalysisabstractContrastive learning (CL) is a form of self-supervised learning and has been widely used for various tasks. Different from widely studied instance-level contrastive learning, pixel-wise contrastive learning mainly helps with pixel-wise dense prediction tasks. The counterpart to an instance in instance-level CL is a pixel, along with its neighboring context, in pixel-wise CL. Aiming to build better feature representation, there is a vast literature about designing instance augmentation strategies for instance-level CL; but there is little similar work on pixel augmentation for pixel-wise CL with a pixel granularity. In this paper, we attempt to bridge this gap. We first classify a pixel into three categories, namely low-, medium-, and high-informative, based on the information quantity the pixel contains. We then adaptively design separate augmentation strategies for each category in terms of augmentation intensity and sampling ratio. Extensive experiments validate that our information-guided pixel augmentation strategy succeeds in encoding more discriminative representations and surpassing other competitive approaches in unsupervised local feature matching. Furthermore, our pretrained model improves the performance of both one-shot and fully supervised models. To the best of our knowledge, we are the first to propose a pixel augmentation method with a pixel granularity for enhancing unsupervised pixel-wise contrastive learning. Code is available at https://github.com/Curli-quan/IGU-Aug. Quan Quan, Qingsong Yao, Heqin Zhu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Unified Multi-Modal Image Synthesis for Missing Modality ImputationabstractMulti-modal medical images provide complementary soft-tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scanning time, image corruption and various imaging protocols often result in incomplete multi-modal images, thus limiting the usage of multi-modal data for clinical purposes. To address this issue, in this paper, we propose a novel unified multi-modal image synthesis method for missing modality imputation. Our method overall takes a generative adversarial architecture, which aims to synthesize missing modalities from any combination of available ones with a single model. To this end, we specifically design a Commonality- and Discrepancy-Sensitive Encoder for the generator to exploit both modality-invariant and specific information contained in input modalities. The incorporation of both types of information facilitates the generation of images with consistent anatomy and realistic details of the desired distribution. Besides, we propose a Dynamic Feature Unification Module to integrate information from a varying number of available modalities, which enables the network to be robust to random missing modalities. The module performs both hard integration and soft integration, ensuring the effectiveness of feature combination while avoiding information loss. Verified on two public multi-modal magnetic resonance datasets, the proposed method is effective in handling various synthesis tasks and shows superior performance compared to previous methods. Yue Zhang 0042, Chengtao Peng, Qiuli Wang 0001, Dan Song 0006, Kaiyan Li 0005, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 6 |
| 2024 | WeakPCSOD: Overcoming the Bias of Box Annotations for Weakly Supervised Point Cloud Salient Object DetectionabstractPoint cloud salient object detection (PCSOD) is a newly proposed task in 3D dense segmentation. However, the acquisition of accurate 3D dense annotations comes at a high cost, severely limiting the progress of PCSOD. To address this issue, we propose the first weakly supervised PCSOD (named WeakPCSOD) model, which relies solely on cheap 3D bounding box annotations. In WeakPCSOD, we extract noise-free supervision from coarse 3D bounding boxes while mitigating shape biases inherent in box annotations. To achieve this, we introduce a novel mask-to-box (M2B) transformation and a color consistency (CC) loss. The M2B transformation, from a shape perspective, disentangles predictions from labels, enabling the extraction of noiseless supervision from labels while preserving object shapes independently of the box bias. From an appearance perspective, we further introduce the CC loss to provide dense supervision, which mitigates the non-unique predictions stemming from weak supervision and substantially reduces prediction variability. Furthermore, we employ a self-training (ST) strategy to enhance performance by utilizing high-confidence pseudo labels. Notably, the M2B transformation, CC loss, and ST strategy are seamlessly integrated into any model and incur no computational costs for inference. Extensive experiments demonstrate the effectiveness of our WeakPCSOD model, even comparable to fully supervised models utilizing dense annotations. Jun Wei 0006, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
AAAI | 2 |
| 2024 | Taming Stable Diffusion for MRI Cross-Modality TranslationabstractIn this study, we explore using Stable Diffusion (SD) for unsupervised medical image-to-image translation. SD has shown remarkable performances in generating high-quality images and can be easily applied to generate custom contents by injecting standard plug-ins like LoRA, offering a promising solution to tackle the complexity caused by variations in imaging modalities, acquisition parameters, and body parts in medical imaging. However, We empirically find that existing pipelines designed for natural images fail to translate directly to medical images due to weak structural control and inappropriate color preservation. To address these issues, we propose a novel two-branch image translation pipeline. This pipeline decouples the generation of target image along the time axis and employs ControlNet to ensure precise structural preservation. Additionally, we customize SD to generate images of extreme brightness, a common feature in medical imaging. Our results on the BraTS dataset demonstrate that SD with task-specific plug-ins can generate high-quality medical images comparable to those generated by task-specific models. Since the development of these standard plug-ins can be easily done by clinicians without much knowledge of the underlying algorithm, such a mode holds the potential to significantly extend the use of medical image computing algorithms in the clinical environment. Yingtai Li, Shaohua Kevin Zhou |
BIBM | 5 |
| 2024 | CARZero: Cross-Attention Alignment for Radiology Zero-Shot ClassificationabstractThe advancement of Zero-Shot Learning in the medi-cal domain has been driven forward by using pretrained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely on cosine similarity for alignment, which may not fully capture the complex relationship between medical images and reports. To address this gap, we introduce a novel approach called Cross-Attention Alignment for Radiology Zero-Shot Classification (CARZero). Our approach innovatively leverages cross-attention mechanisms to process image and report features, creating a Similarity Representation that more accurately reflects the intricate relationships in medical semantics. This representation is then linearly projected to form an image-text similarity matrix for cross-modality alignment. Additionally, recognizing the pivotal role of prompt selection in zero-shot learning, CARZero in-corporates a Large Language Model-based prompt alignment strategy. This strategy standardizes diverse diagnostic expressions into a unified format for both training and inference phases, overcoming the challenges of manual prompt design. Our approach is simple yet effective, demonstrating state-of-the-art performance in zero-shot classification on five official chest radiograph diagnostic test sets, including remarkable results on datasets with long-tail distributions of rare diseases. This achievement is attributed to our new image-text alignment strategy, which effectively addresses the complex relationship between medical images and reports. Code and models are available at https://github.com/laihaoran/CARZero. Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang 0003, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
CVPR | 7 |
| 2024 | Out-of-Distribution Detection for Learning-Based Chest X-Ray DiagnosisabstractDeep learning has shown prominence in chest radiography interpretation, which is critical in evaluating various lung and chest diseases, such as pneumonia, emphysema, and tuberculosis. Deploying machine learning model, it is important to detect out-of-distribution (OOD) inputs, which are distinct from the training data. For example, a model trained on conventional pneumonia data may not be able to diagnose COVID-19 accurately, letting alone the vast body of possible abnormalities and variation in medical imaging. Thus, a trustworthy diagnosing model should be alerted with such OOD inputs to avoid the overconfidence with samples and therefore the risk of misdiagnosis and severe medical consequences. In this paper, we study the problem of OOD detection for deep learning-based chest X-ray imaging by proposing a distance-based OOD detection method. We show that our method (1) inherits the merits of distance-based methods in being applicable to different models and training losses; and (2) adopts probabilistic score functions with fewer statistics to mitigate the potential distribution assumption on training data. We conduct extensive experiments on benchmarks to demonstrate the superiority of our method. In particular, on challenging tasks, such as Chest X-ray 14 vs. COVID-19, our method achieves 12.62% lower FPR than the state-of-the-art solutions. Chuanwen Feng, Ao Ke, Xike Xie, Shaohua Kevin Zhou |
ICASSP | 5 |
| 2024 | Mayfly: a Neural Data Structure for Graph Stream SummarizationabstractA graph is a structure made up of vertices and edges used to represent complex relationships between entities, while a graph stream is a continuous flow of graph updates that convey evolving relationships between entities. The massive volume and high dynamism of graph streams promote research on data structures of graph summarization, which provides a concise and approximate view of graph streams with sub-linear space and linear construction time, enabling real-time graph analytics in various domains, such as social networking, financing, and cybersecurity.
In this work, we propose the Mayfly, the first neural data structure for summarizing graph streams. The Mayfly replaces handcrafted data structures with better accuracy and adaptivity.
To cater to practical applications, Mayfly incorporates two offline training phases.
During the larval phase, the Mayfly learns basic summarization abilities from automatically and synthetically constituted meta-tasks, and in the metamorphosis phase, it rapidly adapts to real graph streams via meta-tasks.
With specific configurations of information pathways, the Mayfly enables flexible support for miscellaneous graph queries, including edge, node, and connectivity queries.
Extensive empirical studies show that the Mayfly significantly outperforms its handcrafted competitors. Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou |
ICLR | 5 |
| 2024 | Prompt Learning with Extended Kalman Filter for Pre-trained Language Models
Xike Xie, Chao Wang 0003, Shaohua Kevin Zhou |
IJCAI | 4 |
| 2024 | Partial Optimal Transport Based Out-of-Distribution Detection for Open-Set Semi-Supervised Learning
Yilong Ren, Chuanwen Feng, Xike Xie, Shaohua Kevin Zhou |
IJCAI | 4 |
| 2024 | 3DGR-CAR: Coronary Artery Reconstruction from Ultra-sparse 2D X-Ray Views with a 3D Gaussians Representation
Xueming Fu, Yingtai Li, Fenghe Tang, Jun Li 0103, Mingyue Zhao, Gaojun Teng, Shaohua Kevin Zhou |
MICCAI (7) | 7 |
| 2024 | HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-training
Fenghe Tang, Ronghao Xu, Qingsong Yao, Xueming Fu, Quan Quan, Heqin Zhu, Zaiyi Liu, Shaohua Kevin Zhou |
MICCAI (11) | 8 |
| 2024 | SIX-Net: Spatial-Context Information miX-up for Electrode Landmark Detection
Heqin Zhu, Qingsong Yao, Yiyong Sun, Shaohua Kevin Zhou |
MICCAI (1) | 6 |
| 2024 | Hierarchical Multiple Instance Learning for COPD Grading with Relatively Specific Similarity
Mingyue Zhao, Mingzhu Liu, Jiejun Luo, Xiuxiu Zhou, Li Fan 0002, Shaohua Kevin Zhou |
MICCAI (1) | 12 |
| 2024 | See, Predict, Plan: Diffusion for Procedure Planning in Robotic Surgical Videos
Ziyuan Zhao, Fen Fang, Xulei Yang, Qianli Xu, Cuntai Guan, Shaohua Kevin Zhou |
MICCAI (6) | 6 |
| 2024 | FairMedFM: Fairness Benchmarking for Medical Imaging Foundation ModelsabstractThe advent of foundation models (FMs) in healthcare offers unprecedented opportunities to enhance medical diagnostics through automated classification and segmentation tasks. However, these models also raise significant concerns about their fairness, especially when applied to diverse and underrepresented populations in healthcare applications. Currently, there is a lack of comprehensive benchmarks, standardized pipelines, and easily adaptable libraries to evaluate and understand the fairness performance of FMs in medical imaging, leading to considerable challenges in formulating and implementing solutions that ensure equitable outcomes across diverse patient populations. To fill this gap, we introduce FairMedFM, a fairness benchmark for FM research in medical imaging. FairMedFM integrates with 17 popular medical imaging datasets, encompassing different modalities, dimensionalities, and sensitive attributes. It explores 20 widely used FMs, with various usages such as zero-shot learning, linear probing, parameter-efficient fine-tuning, and prompting in various downstream tasks -- classification and segmentation. Our exhaustive analysis evaluates the fairness performance over different evaluation metrics from multiple perspectives, revealing the existence of bias, varied utility-fairness trade-offs on different FMs, consistent disparities on the same datasets regardless FMs, and limited effectiveness of existing unfairness mitigation methods. Furthermore, FairMedFM provides an open-sourced codebase at https://github.com/FairMedFM/FairMedFM, supporting extendible functionalities and applications and inclusive for studies on FMs in medical imaging over the long term. Ruinan Jin, Yuan Zhong 0003, Qingsong Yao, Qi Dou 0001, Shaohua Kevin Zhou, Xiaoxiao Li 0001 |
NeurIPS | 6 |
| 2024 | Distortion-Disentangled Contrastive LearningabstractSelf-supervised learning is well known for its remarkable performance in representation learning and various downstream computer vision tasks. Recently, Positive-pair-Only Contrastive Learning (POCL) has achieved reliable performance without the need to construct positive-negative training sets. It reduces memory requirements by lessening the dependency on the batch size. The POCL method typically uses a single objective function to extract the distortion invariant representation (DIR) which describes the proximity of positive-pair representations affected by different distortions. This objective function implicitly enables the model to filter out or ignore the distortion variant representation (DVR) affected by different distortions. However, some recent studies have shown that proper use of DVR in contrastive can optimize the performance of models in some downstream domain-specific tasks. In addition, these POCL methods have been observed to be sensitive to augmentation strategies. To address these limitations, we propose a novel POCL framework named Distortion-Disentangled Contrastive Learning (DDCL) and a Distortion-Disentangled Loss (DDL). Our approach is the first to explicitly and adaptively disentangle and exploit the DVR inside the model and feature stream to improve the representation utilization efficiency, robustness and representation ability. Experiments demonstrate our framework’s superiority to Barlow Twins and Simsiam in terms of convergence, representation quality (including transferability and generalization), and robustness on several datasets. Jinfeng Wang 0008, Sifan Song, Jionglong Su, Shaohua Kevin Zhou |
WACV | 4 |
| 2024 | Two-dimensional materials for future information technology: status and prospectsabstractAbstract Over the past 70 years, the semiconductor industry has undergone transformative changes, largely driven by the miniaturization of devices and the integration of innovative structures and materials. Two-dimensional (2D) materials like transition metal dichalcogenides (TMDs) and graphene are pivotal in overcoming the limitations of silicon-based technologies, offering innovative approaches in transistor design and functionality, enabling atomic-thin channel transistors and monolithic 3D integration. We review the important progress in the application of 2D materials in future information technology, focusing in particular on microelectronics and optoelectronics. We comprehensively summarize the key advancements across material production, characterization metrology, electronic devices, optoelectronic devices, and heterogeneous integration on silicon. A strategic roadmap and key challenges for the transition of 2D materials from basic research to industrial development are outlined. To facilitate such a transition, key technologies and tools dedicated to 2D materials must be developed to meet industrial standards, and the employment of AI in material growth, characterizations, and circuit design will be essential. It is time for academia to actively engage with industry to drive the next 10 years of 2D material research. Hao Qiu 0001, Zhihao Yu, Tiange Zhao, Mingsheng Xu, Taotao Li, Wenzhong Bao, Yang Chai, Shula Chen, Hui-Ming Cheng, Daoxin Dai, Zengfeng Di, Zhuo Dong, Xidong Duan, Yuhan Feng, Jingshu Guo, Pengwen Guo, Yue Hao 0001, Jingyi Hu, Weida Hu, Zehua Hu, Ali Imran 0004, Ziqiang Kong, Bilu Liu, Chunsen Liu, Guanyu Liu, Kaihui Liu, Donglin Lu, Likuan Ma, Feng Miao, Zhenhua Ni, Anlian Pan, Haowen Shu, Quanyang Tao, Ziao Tian, Haomin Wang 0005, Yeliang Wang, Haidi Wu, Hongzhao Wu, Jiangbin Wu, Yanqing Wu, Longfei Xia, Baixu Xiang, Luwen Xing, Qihua Xiong, Jeffrey Xu, Yang Xu 0035, Yuekun Yang, Jincheng Zhang 0001, Tao Zhang 0090, Xinbo Zhang, Chunsong Zhao, Yuda Zhao, Ting Zheng, Peng Zhou 0021, Shaohua Kevin Zhou, Deren Yang |
Sci. China Inf. Sci. | 97 |
| 2024 | Which images to label for few-shot medical image analysis?
Quan Quan, Qingsong Yao, Heqin Zhu, Qiyuan Wang 0001, Shaohua Kevin Zhou |
Medical Image Anal. | 5 |
| 2024 | Cascaded Multi-path Shortcut Diffusion Model for Medical Image Translation
Yinchi Zhou, Huidong Xie, Nicha C. Dvornek, Shaohua Kevin Zhou, David L. Wilson, James S. Duncan, Chi Liu 0001, Bo Zhou 0009 |
Medical Image Anal. | 6 |
| 2024 | OS-SSVEP: One-shot SSVEP classification
Zhiwei Ji, Yijun Wang 0001, Shaohua Kevin Zhou |
Neural Networks | 4 |
| 2024 | Play like a Vertex: A Stackelberg Game Approach for Streaming Graph PartitioningabstractIn the realm of distributed systems tasked with managing and processing large-scale graph-structured data, optimizing graph partitioning stands as a pivotal challenge. The primary goal is to minimize communication overhead and runtime cost. However, alongside the computational complexity associated with optimal graph partitioning, a critical factor to consider is memory overhead. Real-world graphs often reach colossal sizes, making it impractical and economically unviable to load the entire graph into memory for partitioning. This is also a fundamental premise in distributed graph processing, where accommodating a graph with non-distributed systems is unattainable. Currently, existing streaming partitioning algorithms exhibit a skew-oblivious nature, yielding satisfactory partitioning results exclusively for specific graph types. In this paper, we propose a novel streaming partitioning algorithm, the Skewness-aware Vertex-cut Partitioner (S5P ), designed to leverage the skewness characteristics of real graphs for achieving high-quality partitioning. S5P offers high partitioning quality by segregating the graph's edge set into two subsets, head and tail sets. Following processing by a skewness-aware clustering algorithm, these two subsets subsequently undergo a Stackelberg graph game. Our extensive evaluations conducted on substantial real-world and synthetic graphs demonstrate that, in all instances, the partitioning quality of S5P surpasses that of existing streaming partitioning algorithms, operating within the same load balance constraints. For example, S5P can bring up to a 51% improvement in partitioning quality compared to the top partitioner among the baselines. Lastly, we showcase that the implementation of S5P results in up to an 81% reduction in communication cost and a 130% increase in runtime efficiency for distributed graph processing tasks on PowerGraph. Zezhong Ding 0001, Yongan Xiang, Shangyou Wang, Xike Xie, Shaohua Kevin Zhou |
Proc. ACM Manag. Data | 5 |
| 2024 | Learning to Sketch: A Neural Approach to Item Frequency Estimation in Streaming DataabstractRecently, there has been a trend of designing neural data structures to go beyond handcrafted data structures by leveraging patterns of data distributions for better accuracy and adaptivity. Sketches are widely used data structures in real-time web analysis, network monitoring, and self-driving to estimate item frequencies of data streams within limited space. However, existing sketches have not fully exploited the patterns of the data stream distributions, making it challenging to tightly couple them with neural networks that excel at memorizing pattern information. Starting from the premise, we envision a pure neural data structure as a base sketch, which we term the meta-sketch, to reinvent the base structure of conventional sketches. The meta-sketch learns basic sketching abilities from meta-tasks constituted with synthetic datasets following Zipf distributions in the pre-training phase and can be quickly adapted to real (skewed) distributions in the adaption phase. The meta-sketch not only surpasses its competitors in sketching conventional data streams but also holds good potential in supporting more complex streaming data, such as multimedia and graph stream scenarios. Extensive experiments demonstrate the superiority of the meta-sketch and offer insights into its working mechanism. Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Lung Nodule Segmentation and Uncertain Region Prediction With an Uncertainty-Aware Attention MechanismabstractRadiologists possess diverse training and clinical experiences, leading to variations in the segmentation annotations of lung nodules and resulting in segmentation uncertainty. Conventional methods typically select a single annotation as the learning target or attempt to learn a latent space comprising multiple annotations. However, these approaches fail to leverage the valuable information inherent in the consensus and disagreements among the multiple annotations. In this paper, we propose an Uncertainty-Aware Attention Mechanism (UAAM) that utilizes consensus and disagreements among multiple annotations to facilitate better segmentation. To this end, we introduce the Multi-Confidence Mask (MCM), which combines a Low-Confidence (LC) Mask and a High-Confidence (HC) Mask. The LC mask indicates regions with low segmentation confidence, where radiologists may have different segmentation choices. Following UAAM, we further design an Uncertainty-Guide Multi-Confidence Segmentation Network (UGMCS-Net), which contains three modules: a Feature Extracting Module that captures a general feature of a lung nodule, an Uncertainty-Aware Module that produces three features for the annotations' union, intersection, and annotation set, and an Intersection-Union Constraining Module that uses distances between the three features to balance the predictions of final segmentation and MCM. To comprehensively demonstrate the performance of our method, we propose a Complex-Nodule Validation on LIDC-IDRI, which tests UGMCS-Net's segmentation performance on lung nodules that are difficult to segment using common methods. Experimental results demonstrate that our method can significantly improve the segmentation performance on nodules that are difficult to segment using conventional methods. Qiuli Wang 0001, Yue Zhang 0042, Zhulin An, Chen Liu 0026, Xiaohong Zhang 0002, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Adversarial Medical Image With Hierarchical Feature HidingabstractDeep learning based methods for medical images can be easily compromised by adversarial examples (AEs), posing a great security flaw in clinical decision-making. It has been discovered that conventional adversarial attacks like PGD which optimize the classification logits, are easy to distinguish in the feature space, resulting in accurate reactive defenses. To better understand this phenomenon and reassess the reliability of the reactive defenses for medical AEs, we thoroughly investigate the characteristic of conventional medical AEs. Specifically, we first theoretically prove that conventional adversarial attacks change the outputs by continuously optimizing vulnerable features in a fixed direction, thereby leading to outlier representations in the feature space. Then, a stress test is conducted to reveal the vulnerability of medical images, by comparing with natural images. Interestingly, this vulnerability is a double-edged sword, which can be exploited to hide AEs. We then propose a simple-yet-effective hierarchical feature constraint (HFC), a novel add-on to conventional white-box attacks, which assists to hide the adversarial feature in the target feature distribution. The proposed method is evaluated on three medical datasets, both 2D and 3D, with different modalities. The experimental results demonstrate the superiority of HFC,i.e., it bypasses an array of state-of-the-art adversarial medical AE detectors more efficiently than competing adaptive attacks1, which reveals the deficiencies of medical reactive defense and allows to develop more robust defenses in future. Qingsong Yao, Zecheng He, Yuexiang Li, Yi Lin 0009, Kai Ma 0002, Yefeng Zheng 0001, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Hierarchically Contrastive Hard Sample Mining for Graph Self-Supervised PretrainingabstractContrastive learning has recently emerged as a powerful technique for graph self-supervised pretraining (GSP). By maximizing the mutual information (MI) between a positive sample pair, the network is forced to extract discriminative information from graphs to generate high-quality sample representations. However, we observe that, in the process of MI maximization (Infomax), the existing contrastive GSP algorithms suffer from at least one of the following problems: 1) treat all samples equally during optimization and 2) fall into a single contrasting pattern within the graph. Consequently, the vast number of well-categorized samples overwhelms the representation learning process, and limited information is accumulated, thus deteriorating the learning capability of the network. To solve these issues, in this article, by fusing the information from different views and conducting hard sample mining in a hierarchically contrastive manner, we propose a novel GSP algorithm called hierarchically contrastive hard sample mining (HCHSM). The hierarchical property of this algorithm is manifested in two aspects. First, according to the results of multilevel MI estimation in different views, the MI-based hard sample selection (MHSS) module keeps filtering the easy nodes and drives the network to focus more on hard nodes. Second, to collect more comprehensive information for hard sample learning, we introduce a hierarchically contrastive scheme to sequentially force the learned node representations to involve multilevel intrinsic graph features. In this way, as the contrastive granularity goes finer, the complementary information from different levels can be uniformly encoded to boost the discrimination of hard samples and enhance the quality of the learned graph embedding. Extensive experiments on seven benchmark datasets indicate that the HCHSM performs better than other competitors on node classification and node clustering tasks. The source code of HCHSM is available at https://github.com/WxTu/HCHSM. Wenxuan Tu, Shaohua Kevin Zhou, Xinwang Liu 0002, Chunpeng Ge 0001, Zhiping Cai, Yue Liu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Label-Free Nuclei Segmentation Using Intra-Image Self Similarity
Shaohua Kevin Zhou |
MICCAI (6) | 3 |
| 2023 | WeakPolyp: You only Look Bounding Box for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 4 |
| 2023 | FairAdaBN: Mitigating Unfairness with Adaptive Batch Normalization and Its Application to Dermatological Disease Classification
Shang Zhao 0004, Quan Quan, Qingsong Yao, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2023 | DiffULD: Diffusive Universal Lesion Detection
Peiang Zhao, Ruiyang Jin, Shaohua Kevin Zhou |
MICCAI (5) | 4 |
| 2023 | UOD: Universal One-Shot Detection of Anatomical Landmarks
Heqin Zhu, Quan Quan, Qingsong Yao, Zaiyi Liu, Shaohua Kevin Zhou |
MICCAI (1) | 5 |
| 2023 | Active CT Reconstruction with a Learned Sampling PolicyabstractComputed tomography (CT) is a widely-used imaging technology that assists clinical decision-making with high-quality human body representations. To reduce the radiation dose posed by CT, sparse-view (SV) CT is developed with preserved image quality. However, these methods are still stuck with a fixed uniform SV (USV) sampling strategy, which inhibits the possibility of acquiring a better image with an even reduced dose. In this paper, we explore this possibility via learning an active SV (ASV) sampling policy that optimizes the sampling positions for regions of interest (RoI)-specific, high-quality reconstruction. To this end, we design an sampling agent for the recommendation of ASV sampling positions based on on-the-fly reconstruction with obtained sinograms in a progressive fashion. With such a design, we achieve better performances on the NIH-AAPM dataset over popular USV sampling, especially when the number of views is small. Finally, such a design enables the RoI-aware reconstruction with improved local quality within the RoI that are clinically important. Experiments on the VerSe dataset demonstrate the ability of the proposed sampling policy, which is difficult to achieve with USV sampling. Ce Wang 0001, Kun Shang 0002, Haimiao Zhang, Shang Zhao 0004, Dong Liang 0001, Shaohua Kevin Zhou |
ACM Multimedia | 6 |
| 2023 | Unsupervised Polychromatic Neural Representation for CT Metal Artifact ReductionabstractEmerging neural reconstruction techniques based on tomography (e.g., NeRF, NeAT, and NeRP) have started showing unique capabilities in medical imaging. In this work, we present a novel Polychromatic neural representation (Polyner) to tackle the challenging problem of CT imaging when metallic implants exist within the human body. CT metal artifacts arise from the drastic variation of metal's attenuation coefficients at various energy levels of the X-ray spectrum, leading to a nonlinear metal effect in CT measurements. Recovering CT images from metal-affected measurements hence poses a complicated nonlinear inverse problem where empirical models adopted in previous metal artifact reduction (MAR) approaches lead to signal loss and strongly aliased reconstructions. Polyner instead models the MAR problem from a nonlinear inverse problem perspective. Specifically, we first derive a polychromatic forward model to accurately simulate the nonlinear CT acquisition process. Then, we incorporate our forward model into the implicit neural representation to accomplish reconstruction. Lastly, we adopt a regularizer to preserve the physical properties of the CT images across different energy levels while effectively constraining the solution space. Our Polyner is an unsupervised method and does not require any external training data. Experimenting with multiple datasets shows that our Polyner achieves comparable or better performance than supervised methods on in-domain datasets while demonstrating significant performance improvements on out-of-domain datasets. To the best of our knowledge, our Polyner is the first unsupervised MAR method that outperforms its supervised counterparts. The code for this work is available at: https://github.com/iwuqing/Polyner. Qing Wu 0001, Lixuan Chen, Ce Wang 0001, Hongjiang Wei, Shaohua Kevin Zhou, Jingyi Yu 0001, Yuyao Zhang 0005 |
NeurIPS | 5 |
| 2023 | Rethinking Semi-Supervised Medical Image Segmentation: A Variance-Reduction PerspectiveabstractFor medical image segmentation, contrastive learning is the dominant practice to improve the quality of visual representations by contrasting semantically similar and dissimilar pairs of samples. This is enabled by the observation that without accessing ground truth labels, negative examples with truly dissimilar anatomical features, if sampled, can significantly improve the performance. In reality, however, these samples may come from similar anatomical features and the models may struggle to distinguish the minority tail-class samples, making the tail classes more prone to misclassification, both of which typically lead to model collapse. In this paper, we propose $\texttt{ARCO}$, a semi-supervised contrastive learning (CL) framework with stratified group theory for medical image segmentation. In particular, we first propose building $\texttt{ARCO}$ through the concept of variance-reduced estimation, and show that certain variance-reduction techniques are particularly beneficial in pixel/voxel-level segmentation tasks with extremely limited labels. Furthermore, we theoretically prove these sampling techniques are universal in variance reduction. Finally, we experimentally validate our approaches on eight benchmarks, i.e., five 2D/3D medical and three semantic segmentation datasets, with different label settings, and our methods consistently outperform state-of-the-art semi-supervised methods. Additionally, we augment the CL frameworks with these sampling techniques and demonstrate significant gains over previous methods. We believe our work is an important step towards semi-supervised medical image segmentation by quantifying the limitation of current self-supervision objectives for accomplishing such challenging safety-critical tasks. Chenyu You, Weicheng Dai, Yifei Min, David A. Clifton, Shaohua Kevin Zhou, Lawrence H. Staib, James S. Duncan |
NeurIPS | 6 |
| 2023 | Transforming medical imaging with Transformers? A comparative review of key properties, current progresses, and future perspectives
Jun Li 0103, Junyu Chen 0002, Yucheng Tang, Ce Wang 0001, Bennett A. Landman, Shaohua Kevin Zhou |
Medical Image Anal. | 6 |
| 2023 | Histopathological bladder cancer gene mutation prediction with hierarchical deep multiple-instance learning
Rui Yan 0009, Xueyuan Zhang, Jintao Li 0001, Dingwei Ye, Shaohua Kevin Zhou |
Medical Image Anal. | 9 |
| 2023 | Radiology report generation with a learned knowledge base and multi-modal alignmentabstractIn clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a chest x-ray. Our approach, motivated by the observation that the descriptions in radiology reports are highly correlated with specific information of the x-ray images, features two distinct modules: (i) Learned knowledge base: To absorb the knowledge embedded in the radiology reports, we build a knowledge base that can automatically distill and restore medical knowledge from textual embedding without manual labor; (ii) Multi-modal alignment: to promote the semantic alignment among reports, disease labels, and images, we explicitly utilize textual embedding to guide the learning of the visual feature space. We evaluate the performance of the proposed model using metrics from both natural language generation and clinic efficacy on the public IU-Xray and MIMIC-CXR datasets. Our ablation study shows that each module contributes to improving the quality of generated reports. Furthermore, the assistance of both modules, our approach outperforms state-of-the-art methods over almost all the metrics. Code is available at https://github.com/LX-doctorAI1/M2KT. Shuxin Yang, Xian Wu 0001, Shen Ge, Zhuozhao Zheng, Shaohua Kevin Zhou, Li Xiao 0005 |
Medical Image Anal. | 5 |
| 2023 | FedFTN: Personalized federated learning with deep feature transformation network for multi-institutional low-count PET denoising
Bo Zhou 0009, Huidong Xie, Xiongchao Chen, Xueqi Guo, Zhicheng Feng, Shaohua Kevin Zhou, Axel Rominger, Kuangyu Shi, James S. Duncan, Chi Liu 0001 |
Medical Image Anal. | 8 |
| 2023 | Semi-Supervised CT Lesion Segmentation Using Uncertainty-Based Data Pairing and SwapMixabstractSemi-supervised learning (SSL) methods show their powerful performance to deal with the issue of data shortage in the field of medical image segmentation. However, existing SSL methods still suffer from the problem of unreliable predictions on unannotated data due to the lack of manual annotations for them. In this paper, we propose an unreliability-diluted consistency training (UDiCT) mechanism to dilute the unreliability in SSL by assembling reliable annotated data into unreliable unannotated data. Specifically, we first propose an uncertainty-based data pairing module to pair annotated data with unannotated data based on a complementary uncertainty pairing rule, which avoids two hard samples being paired off. Secondly, we develop SwapMix, a mixed sample data augmentation method, to integrate annotated data into unannotated data for training our model in a low-unreliability manner. Finally, UDiCT is trained by minimizing a supervised loss and an unreliability-diluted consistency loss, which makes our model robust to diverse backgrounds. Extensive experiments on three chest CT datasets show the effectiveness of our method for semi-supervised CT lesion segmentation. Pengchong Qiao, Guoli Song, Hu Han 0001, Yonghong Tian 0001, Yongsheng Liang 0001, Xi Li 0011, Shaohua Kevin Zhou, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Multi-Modal Tumor Segmentation With Deformable Aggregation and Uncertain Region InpaintingabstractMulti-modal tumor segmentation exploits complementary information from different modalities to help recognize tumor regions. Known multi-modal segmentation methods mainly have deficiencies in two aspects: First, the adopted multi-modal fusion strategies are built upon well-aligned input images, which are vulnerable to spatial misalignment between modalities (caused by respiratory motions, different scanning parameters, registration errors, etc). Second, the performance of known methods remains subject to the uncertainty of segmentation, which is particularly acute in tumor boundary regions. To tackle these issues, in this paper, we propose a novel multi-modal tumor segmentation method with deformable feature fusion and uncertain region refinement. Concretely, we introduce a deformable aggregation module, which integrates feature alignment and feature aggregation in an ensemble, to reduce inter-modality misalignment and make full use of cross-modal information. Moreover, we devise an uncertain region inpainting module to refine uncertain pixels using neighboring discriminative features. Experiments on two clinical multi-modal tumor datasets demonstrate that our method achieves promising tumor segmentation results and outperforms state-of-the-art methods. Yue Zhang 0042, Chengtao Peng, Ruofeng Tong 0001, Lanfen Lin, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 8 |
| 2023 | LE-UDA: Label-Efficient Unsupervised Domain Adaptation for Medical Image SegmentationabstractWhile deep learning methods hitherto have achieved considerable success in medical image segmentation, they are still hampered by two limitations: (i) reliance on large-scale well-labeled datasets, which are difficult to curate due to the expert-driven and time-consuming nature of pixel-level annotations in clinical practices, and (ii) failure to generalize from one domain to another, especially when the target domain is a different modality with severe domain shifts. Recent unsupervised domain adaptation (UDA) techniques leverage abundant labeled source data together with unlabeled target data to reduce the domain gap, but these methods degrade significantly with limited source annotations. In this study, we address this underexplored UDA problem, investigating a challenging but valuable realistic scenario, where the source domain not only exhibits domain shift w.r.t. the target domain but also suffers from label scarcity. In this regard, we propose a novel and generic framework called "Label-Efficient Unsupervised Domain Adaptation" (LE-UDA). In LE-UDA, we construct self-ensembling consistency for knowledge transfer between both domains, as well as a self-ensembling adversarial learning module to achieve better feature alignment for UDA. To assess the effectiveness of our method, we conduct extensive experiments on two different tasks for cross-modality segmentation between MRI and CT images. Experimental results demonstrate that the proposed LE-UDA can efficiently leverage limited source labels to improve cross-domain segmentation performance, outperforming state-of-the-art UDA approaches in the literature. Ziyuan Zhao, Fangcheng Zhou, Kaixin Xu, Zeng Zeng, Cuntai Guan, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 6 |
| 2022 | DeltaNet: Conditional Medical Report Generation for COVID-19 DiagnosisabstractFast screening and diagnosis are critical in COVID-19 patient treatment. In addition to the gold standard RT-PCR, radiological imaging like X-ray and CT also works as an important means in patient screening and follow-up. However, due to the excessive number of patients, writing reports becomes a heavy burden for radiologists. To reduce the workload of radiologists, we propose DeltaNet to generate medical reports automatically. Different from typical image captioning approaches that generate reports with an encoder and a decoder, DeltaNet applies a conditional generation process. In particular, given a medical image, DeltaNet employs three steps to generate a report: 1) first retrieving related medical reports, i.e., the historical reports from the same or similar patients; 2) then comparing retrieved images and current image to find the differences; 3) finally generating a new report to accommodate identified differences based on the conditional report. We evaluate DeltaNet on a COVID-19 dataset, where DeltaNet outperforms state-of-the-art approaches. Besides COVID-19, the proposed DeltaNet can be applied to other diseases as well. We validate its generalization capabilities on the public IU-Xray and MIMIC-CXR datasets for chest-related diseases. Xian Wu 0001, Shuxin Yang, Zhaopeng Qiu, Shen Ge, Yangtian Yan, Xingwang Wu, Yefeng Zheng 0001, Shaohua Kevin Zhou, Li Xiao 0005 |
COLING | 8 |
| 2022 | Which images to label for few-shot medical landmark detection?abstractThe success of deep learning methods relies on the availability of well-labeled large-scale datasets. However, for medical images, annotating such abundant training data often requires experienced radiologists and consumes their limited time. Few-shot learning is developed to alleviate this burden, which achieves competitive performances with only several labeled data. However, a crucial yet previously overlooked problem in few-shot learning is about the selection of template images for annotation before learning, which affects the final performance. We herein propose a novel Sample Choosing Policy (SCP) to select “the most worthy” images for annotation, in the context of few-shot medical landmark detection. SCP consists of three parts: 1) Self-supervised training for building a pre-trained deep model to extract features from radiological images, 2) Key Point Proposal for localizing informative patches, and 3) Representative Score Estimation for searching the most representative samples or templates. The advantage of SCP is demonstrated by various experiments on three widely-used public datasets. For one-shot medical landmark detection, its use reduces the mean radial errors on Cephalometric and HandXray datasets by 14.2% (from 3.595mm to 3.083mm) and 35.5% (4.114mm to 2.653mm), respectively. Quan Quan, Qingsong Yao, Jun Li 0103, Shaohua Kevin Zhou |
CVPR | 4 |
| 2022 | Weakly Supervised Object Localization Through Inter-class Feature Similarity and Intra-class Appearance Consistency
Jun Wei 0006, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
ECCV (30) | 3 |
| 2022 | SATr: Slice Attention with Transformer for Universal Lesion Detection
Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (3) | 4 |
| 2022 | Stabilize, Decompose, and Denoise: Self-supervised Fluoroscopy Denoising
Ruizhou Liu, Qiang Ma 0004, Yuanyuan Lyu, Jianji Wang 0003, Shaohua Kevin Zhou |
MICCAI (8) | 6 |
| 2022 | Learning Incrementally to Segment Multiple Organs in a CT Image
Pengbo Liu 0004, Mengsi Fan, Hongli Pan, Minmin Yin, Xiaohong Zhu, Dandan Du, Xiaoying Zhao, Li Xiao 0005, Lian Ding, Xingwang Wu, Shaohua Kevin Zhou |
MICCAI (4) | 12 |
| 2022 | Undersampled MRI Reconstruction with Side Information-Guided Normalisation
Xinwen Liu 0003, Jing Wang 0062, Cheng Peng 0008, Shekhar Chandra, Feng Liu 0005, Shaohua Kevin Zhou |
MICCAI (6) | 6 |
| 2022 | Towards Performant and Reliable Undersampled MR Reconstruction via Diffusion Model Sampling
Cheng Peng 0008, Shaohua Kevin Zhou, Vishal M. Patel, Rama Chellappa |
MICCAI (6) | 3 |
| 2022 | Rib Suppression in Digital Chest Tomosynthesis
Yihua Sun, Qingsong Yao, Yuanyuan Lyu, Jianji Wang 0003, Hongen Liao, Shaohua Kevin Zhou |
MICCAI (1) | 7 |
| 2022 | BoxPolyp: Boost Generalized Polyp Segmentation Using Extra Coarse Bounding Box Annotations
Jun Wei 0006, Yiwen Hu 0001, Guanbin Li, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 5 |
| 2022 | Meta-hallucinator: Towards Few-Shot Cross-Modality Cardiac Image Segmentation
Ziyuan Zhao, Fangcheng Zhou, Zeng Zeng, Cuntai Guan, Shaohua Kevin Zhou |
MICCAI (5) | 5 |
| 2022 | Aggregative Self-supervised Feature Learning from Limited Medical Images
Jiuwen Zhu, Yuexiang Li, Lian Ding, Shaohua Kevin Zhou |
MICCAI (8) | 4 |
| 2022 | GAN-based disentanglement learning for chest X-ray rib suppression
Luyi Han, Yuanyuan Lyu, Cheng Peng 0008, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2022 | Knowledge matters: Chest radiology report generation with general and specific knowledgeabstractAutomatic chest radiology report generation is critical in clinics which can relieve experienced radiologists from the heavy workload and remind inexperienced radiologists of misdiagnosis or missed diagnose. Existing approaches mainly formulate chest radiology report generation as an image captioning task and adopt the encoder-decoder framework. However, in the medical domain, such pure data-driven approaches suffer from the following problems: 1) visual and textual bias problem; 2) lack of expert knowledge. In this paper, we propose a knowledge-enhanced radiology report generation approach introduces two types of medical knowledge: 1) General knowledge, which is input independent and provides the broad knowledge for report generation; 2) Specific knowledge, which is input dependent and provides the fine-grained knowledge for chest X-ray report generation. To fully utilize both the general and specific knowledge, we also propose a knowledge-enhanced multi-head attention mechanism. By merging the visual features of the radiology image with general knowledge and specific knowledge, the proposed model can improve the quality of generated reports. The experimental results on the publicly available IU-Xray dataset show that the proposed knowledge-enhanced approach outperforms state-of-the-art methods in almost all metrics. And the results of MIMIC-CXR dataset show that the proposed knowledge-enhanced approach is on par with state-of-the-art methods. Ablation studies also demonstrate that both general and specific knowledge can help to improve the performance of chest radiology report generation. Shuxin Yang, Xian Wu 0001, Shen Ge, Shaohua Kevin Zhou, Li Xiao 0005 |
Medical Image Anal. | 4 |
| 2022 | DuDoDR-Net: Dual-domain data consistent recurrent network for simultaneous sparse view and metal artifact reduction in computed tomography
Bo Zhou 0009, Xiongchao Chen, Shaohua Kevin Zhou, James S. Duncan, Chi Liu 0001 |
Medical Image Anal. | 3 |
| 2022 | DuDoUFNet: Dual-Domain Under-to-Fully-Complete Progressive Restoration Network for Simultaneous Metal Artifact Reduction and Low-Dose CT ReconstructionabstractTo reduce the potential risk of radiation to the patient, low-dose computed tomography (LDCT) has been widely adopted in clinical practice for reconstructing cross-sectional images using sinograms with reduced x-ray flux. The LDCT image quality is often degraded by different levels of noise depending on the low-dose protocols. The image quality will be further degraded when the patient has metallic implants, where the image suffers from additional streak artifacts along with further amplified noise levels, thus affecting the medical diagnosis and other CT-related applications. Previous studies mainly focused either on denoising LDCT without considering metallic implants or full-dose CT metal artifact reduction (MAR). Directly applying previous LDCT or MAR approaches to the issue of simultaneous metal artifact reduction and low-dose CT (MARLD) may yield sub-optimal reconstruction results. In this work, we develop a dual-domain under-to-fully-complete progressive restoration network, called DuDoUFNet, for MARLD. Our DuDoUFNet aims to reconstruct images with substantially reduced noise and artifact by progressive sinogram to image domain restoration with a two-stage progressive restoration network design. Our experimental results demonstrate that our method can provide high-quality reconstruction, superior to previous LDCT and MAR methods under various low-dose and metal settings. Bo Zhou 0009, Xiongchao Chen, Huidong Xie, Shaohua Kevin Zhou, James S. Duncan, Chi Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | XraySyn: Realistic View Synthesis From a Single Radiograph Through CT PriorsabstractA radiograph visualizes the internal anatomy of a patient through the use of X-ray, which projects 3D information onto a 2D plane. Hence, radiograph analysis naturally requires physicians to relate their prior knowledge about 3D human anatomy to 2D radiographs. Synthesizing novel radiographic views in a small range can assist physicians in interpreting anatomy more reliably; however, radiograph view synthesis is heavily ill-posed, lacking in paired data, and lacking in differentiable operations to leverage learning-based approaches. To address these problems, we use Computed Tomography (CT) for radiograph simulation and design a differentiable projection algorithm, which enables us to achieve geometrically consistent transformations between the radiography and CT domains. Our method, XraySyn, can synthesize novel views on real radiographs through a combination of realistic simulation and finetuning on real radiographs. To the best of our knowledge, this is the first work on radiograph view synthesis. We show that by gaining an understanding of radiography in 3D space, our method can be applied to radiograph bone extraction and suppression without requiring groundtruth bone labels. Cheng Peng 0008, Haofu Liao, Gina Wong, Jiebo Luo 0001, Shaohua Kevin Zhou, Rama Chellappa |
AAAI | 5 |
| 2021 | Dual-GAN: Joint BVP and Noise Modeling for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) based physiological measurement has great application values in health monitoring, emotion analysis, etc. Existing methods mainly focus on how to enhance or extract the very weak blood volume pulse (BVP) signals from face videos, but seldom explicitly model the noises that dominate face video content. Thus, they may suffer from poor generalization ability in unseen scenarios. This paper proposes a novel adversarial learning approach for rPPG based physiological measurement by using Dual Generative Adversarial Networks (Dual-GAN) to model the BVP predictor and noise distribution jointly. The BVP-GAN aims to learn a noise-resistant mapping from input to ground-truth BVP, and the Noise-GAN aims to learn the noise distribution. The two GANs can promote each other’s capability, leading to improved feature disentanglement between BVP and noises. Besides, a plug-and-play block named ROI alignment and fusion (ROI-AF) block is proposed to alleviate the inconsistencies between different ROIs and exploit informative features from a wider receptive field in terms of ROIs. In comparison to state-of-the-art methods, our approach achieves better performance in heart rate, heart rate variability, and respiration frequency estimation from face videos. Hao Lu 0009, Hu Han 0001, Shaohua Kevin Zhou |
CVPR | 3 |
| 2021 | Shallow Feature Matters for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects by only utilizing image-level labels. Class activation maps (CAMs) are the commonly used features to achieve WSOL. However, previous CAM-based methods did not take full advantage of the shallow features, despite their importance for WSOL. Because shallow features are easily buried in background noise through conventional fusion. In this paper, we propose a simple but effective Shallow feature-aware Pseudo supervised Object Localization (SPOL) model for accurate WSOL, which makes the utmost of low-level features embedded in shallow layers. In practice, our SPOL model first generates the CAMs through a novel element-wise multiplication of shallow and deep feature maps, which filters the background noise and generates sharper boundaries robustly. Besides, we further propose a general class-agnostic segmentation model to achieve the accurate object mask, by only using the initial CAMs as the pseudo label without any extra annotation. Eventually, a bounding box extractor is applied to the object mask to locate the target. Experiments verify that our SPOL outperforms the state-of-the-art on both CUB- 200 and ImageNet-1K benchmarks, achieving 93.44% and 67.15% (i.e., 3.93% and 2.13% improvement) Top-5 localization accuracy, respectively. Jun Wei 0006, Qin Wang 0011, Zhen Li 0026, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui |
CVPR | 5 |
| 2021 | Decomposition-and-Fusion Network for HE-Stained Pathological Image Classification
Rui Yan 0009, Jintao Li 0001, Shaohua Kevin Zhou, Zhilong Lv, Xueyuan Zhang, Xiaosong Rao, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (3) | 3 |
| 2021 | Perceptual Quality Assessment of Chest Radiograph
Mengda Guan, Yuanyuan Lyu, Wanyue Cao, Xingwang Wu, Jingjing Lu, Shaohua Kevin Zhou |
MICCAI (7) | 6 |
| 2021 | Conditional Training with Bounding Map for Universal Lesion Detection
Hu Han 0001, Ying Chi, Shaohua Kevin Zhou |
MICCAI (5) | 5 |
| 2021 | Universal Undersampled MRI Reconstruction
Xinwen Liu 0003, Jing Wang 0062, Feng Liu 0005, Shaohua Kevin Zhou |
MICCAI (6) | 4 |
| 2021 | U-DuDoNet: Unpaired Dual-Domain Network for CT Metal Artifact Reduction
Yuanyuan Lyu, Jiajun Fu, Cheng Peng 0008, Shaohua Kevin Zhou |
MICCAI (6) | 4 |
| 2021 | DA-VSR: Domain Adaptable Volumetric Super-Resolution for Medical Images
Cheng Peng 0008, Shaohua Kevin Zhou, Rama Chellappa |
MICCAI (6) | 2 |
| 2021 | Improving Generalizability in Limited-Angle CT Reconstruction with Sinogram Extrapolation
Ce Wang 0001, Haimiao Zhang, Kun Shang 0002, Yuanyuan Lyu, Bin Dong 0001, Shaohua Kevin Zhou |
MICCAI (6) | 7 |
| 2021 | Shallow Attention Network for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Ruimao Zhang, Zhen Li 0026, Shaohua Kevin Zhou, Shuguang Cui |
MICCAI (1) | 5 |
| 2021 | A Hierarchical Feature Constraint to Camouflage Medical Adversarial Attacks
Qingsong Yao, Zecheng He, Yi Lin 0009, Kai Ma 0002, Yefeng Zheng 0001, Shaohua Kevin Zhou |
MICCAI (3) | 6 |
| 2021 | One-Shot Medical Landmark Detection
Qingsong Yao, Quan Quan, Li Xiao 0005, Shaohua Kevin Zhou |
MICCAI (2) | 4 |
| 2021 | You only Learn Once: Universal Anatomical Landmark Detection
Heqin Zhu, Qingsong Yao, Li Xiao 0005, Shaohua Kevin Zhou |
MICCAI (5) | 4 |
| 2021 | Few-shot Learning for Multi-Modality TasksabstractRecent deep learning methods rely on a large amount of labeled data to achieve high performance. These methods may be impractical in some scenarios, where manual data annotation is costly or the samples of certain categories are scarce (e.g., tumor lesions, endangered animals and rare individual activities). When only limited annotated samples are available, these methods usually suffer from the overfitting problem severely, which degrades the performance significantly. In contrast, humans can recognize the objects in the images rapidly and correctly with their prior knowledge after exposed to only a few annotated samples. To simulate the learning schema of humans and relieve the reliance on the large-scale annotation benchmarks, researchers start shifting towards the few-shot learning problem: they try to learn a model to correctly recognize novel categories with only a few annotated samples. Jie Chen 0001, Qixiang Ye, Xiaoshan Yang, Shaohua Kevin Zhou, Xiaopeng Hong, Li Zhang 0040 |
ACM Multimedia | 4 |
| 2021 | Marginal loss and exclusion loss for partially supervised multi-organ segmentation
Gonglei Shi, Li Xiao 0005, Yang Chen 0008, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2021 | Anatomy-guided multimodal registration by learning segmentation without ground truth: Application to intraprocedural CBCT/MR liver segmentation and registration
Bo Zhou 0009, Zachary Augenfeld, Julius Chapiro, Shaohua Kevin Zhou, Chi Liu 0001, James S. Duncan |
Medical Image Anal. | 4 |
| 2021 | Deep reinforcement learning in medical imaging: A literature review
Shaohua Kevin Zhou, T. Hoang Ngan Le, Khoa Luu, Hien Van Nguyen, Nicholas Ayache |
Medical Image Anal. | 1 |
| 2021 | A Review of Deep Learning in Medical Imaging: Imaging Traits, Technology Trends, Case Studies With Progress Highlights, and Future PromisesabstractSince its renaissance, deep learning has been widely used in various medical imaging tasks and has achieved remarkable success in many medical imaging applications, thereby propelling us into the so-called artificial intelligence (AI) era. It is known that the success of AI is mostly attributed to the availability of big data with annotations for a single task and the advances in high performance computing. However, medical imaging presents unique challenges that confront deep learning approaches. In this survey paper, we first present traits of medical imaging, highlight both clinical needs and technical challenges in medical imaging, and describe how emerging trends in deep learning are addressing these issues. We cover the topics of network architecture, sparse and noisy labels, federating learning, interpretability, uncertainty quantification, etc. Then, we present several case studies that are commonly found in clinical practice, including digital pathology and chest, brain, cardiovascular, and abdominal imaging. Rather than presenting an exhaustive literature survey, we instead describe some prominent research highlights related to these case study applications. We conclude with a discussion and presentation of promising future directions. Shaohua Kevin Zhou, Hayit Greenspan, Christos Davatzikos, James S. Duncan, Bram van Ginneken, Anant Madabhushi, Jerry L. Prince, Daniel Rueckert, Ronald M. Summers |
Proc. IEEE | 1 |
| 2021 | Semi-Supervised Natural Face De-OcclusionabstractOcclusions are often present in face images in the wild, e.g., under video surveillance and forensic scenarios. Existing face de-occlusion methods are limited as they require the knowledge of an occlusion mask. To overcome this limitation, we propose in this paper a new generative adversarial network (named OA-GAN) for natural face de-occlusion without an occlusion mask, enabled by learning in a semi-supervised fashion using (i) paired images with known masks of artificial occlusions and (ii) natural images without occlusion masks. The generator of our approach first predicts an occlusion mask, which is used for filtering the feature maps of the input image as a semantic cue for de-occlusion. The filtered feature maps are then used for face completion to recover a non-occluded face image. The initial occlusion mask prediction might not be accurate enough, but it gradually converges to the accurate one because of the adversarial loss we use to perceive which regions in a face image need to be recovered. The discriminator of our approach consists of an adversarial loss, distinguishing the recovered face images from natural face images, and an attribute preserving loss, ensuring that the face image after de-occlusion can retain the attributes of the input face image. Experimental evaluations on the widely used CelebA dataset and a dataset with natural occlusions we collected show that the proposed approach can outperform the state of the art methods in natural face de-occlusion. Jiancheng Cai, Hu Han 0001, Jiyun Cui, Jie Chen 0001, Li Liu 0002, Shaohua Kevin Zhou |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Deep Collocative Learning for Immunofixation Electrophoresis Image AnalysisabstractImmunofixation Electrophoresis (IFE) analysis is of great importance to the diagnosis of Multiple Myeloma, which is among the top-9 cancer killers in the United States, but has rarely been studied in the context of deep learning. Two possible reasons are: 1) the recognition of IFE patterns is dependent on the co-location of bands that forms a binary relation, different from the unary relation (visual features to label) that deep learning is good at modeling; 2) deep classification models may perform with high accuracy for IFE recognition but is not able to provide firm evidence (where the co-location patterns are) for its predictions, rendering difficulty for technicians to validate the results. We propose to address these issues with collocative learning, in which a collocative tensor has been constructed to transform the binary relations into unary relations that are compatible with conventional deep networks, and a location-label-free method that utilizes the Grad-CAM saliency map for evidence backtracking has been proposed for accurate localization. In addition, we have proposed Coached Attention Gates that can regulate the inference of the learning to be more consistent with human logic and thus support the evidence backtracking. The experimental results show that the proposed method has obtained a performance gain over its base model ResNet18 by 741.30% in IoU and also outperformed popular deep networks of DenseNet, CBAM, and Inception-v3. Xiaoyong Wei, Zhen-Qun Yang, Xulu Zhang, Ga Liao, Ailin Sheng, Shaohua Kevin Zhou, Yongkang Wu |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Label-Free Segmentation of COVID-19 Lesions in Lung CTabstractScarcity of annotated images hampers the building of automated solution for reliable COVID-19 diagnosis and evaluation from CT. To alleviate the burden of data annotation, we herein present a label-free approach for segmenting COVID-19 lesions in CT via voxel-level anomaly modeling that mines out the relevant knowledge from normal CT lung scans. Our modeling is inspired by the observation that the parts of tracheae and vessels, which lay in the high-intensity range where lesions belong to, exhibit strong patterns. To facilitate the learning of such patterns at a voxel level, we synthesize 'lesions' using a set of simple operations and insert the synthesized 'lesions' into normal CT lung scans to form training pairs, from which we learn a normalcy-recognizing network (NormNet) that recognizes normal tissues and separate them from possible COVID-19 lesions. Our experiments on three different public datasets validate the effectiveness of NormNet, which conspicuously outperforms a variety of unsupervised anomaly detection (UAD) methods. Qingsong Yao, Li Xiao 0005, Peihang Liu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Limited View Tomographic Reconstruction Using a Cascaded Residual Dense Spatial-Channel Attention Network With Projection Data Fidelity LayerabstractLimited view tomographic reconstruction aims to reconstruct a tomographic image from a limited number of projection views arising from sparse view or limited angle acquisitions that reduce radiation dose or shorten scanning time. However, such a reconstruction suffers from severe artifacts due to the incompleteness of sinogram. To derive quality reconstruction, previous methods use UNet-like neural architectures to directly predict the full view reconstruction from limited view data; but these methods leave the deep network architecture issue largely intact and cannot guarantee the consistency between the sinogram of the reconstructed image and the acquired sinogram, leading to a non-ideal reconstruction. In this work, we propose a cascaded residual dense spatial-channel attention network consisting of residual dense spatial-channel attention networks and projection data fidelity layers. We evaluate our methods on two datasets. Our experimental results on AAPM Low Dose CT Grand Challenge datasets demonstrate that our algorithm achieves a consistent and substantial improvement over the existing neural network methods on both limited angle reconstruction and sparse view reconstruction. In addition, our experimental results on Deep Lesion datasets demonstrate that our method is able to generate high-quality reconstruction for 8 major lesion types. Bo Zhou 0009, Shaohua Kevin Zhou, James S. Duncan, Chi Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | SAINT: Spatially Aware Interpolation NeTwork for Medical Slice SynthesisabstractDeep learning-based single image super-resolution (SISR) methods face various challenges when applied to 3D medical volumetric data (i.e., CT and MR images) due to the high memory cost and anisotropic resolution, which adversely affect their performance. Furthermore, mainstream SISR methods are designed to work over specific upsampling factors, which makes them ineffective in clinical practice. In this paper, we introduce a Spatially Aware Interpolation NeTwork (SAINT) for medical slice synthesis to alleviate the memory constraint that volumetric data poses. Compared to other super-resolution methods, SAINT utilizes voxel spacing information to provide desirable levels of details, and allows for the upsampling factor to be determined on the fly. Our evaluations based on 853 CT scans from four datasets that contain liver, colon, hepatic vessels, and kidneys show that SAINT consistently outperforms other SISR methods in terms of medical slice synthesis quality, while using only a single model to deal with different upsampling factors. Cheng Peng 0008, Wei-An Lin, Haofu Liao, Rama Chellappa, Shaohua Kevin Zhou |
CVPR | 5 |
| 2020 | DuDoRNet: Learning a Dual-Domain Recurrent Network for Fast MRI Reconstruction With Deep T1 PriorabstractMRI with multiple protocols is commonly used for diagnosis, but it suffers from a long acquisition time, which yields the image quality vulnerable to say motion artifacts. To accelerate, various methods have been proposed to reconstruct full images from under-sampled k-space data. However, these algorithms are inadequate for two main reasons. Firstly, aliasing artifacts generated in the image domain are structural and non-local, so that sole image domain restoration is insufficient. Secondly, though MRI comprises multiple protocols during one exam, almost all previous studies only employ the reconstruction of an individual protocol using a highly distorted undersampled image as input, leaving the use of fully-sampled short protocol (say T1) as complementary information highly underexplored. In this work, we address the above two limitations by proposing a Dual Domain Recurrent Network (DuDoRNet) with deep T1 prior embedded to simultaneously recover k-space and images for accelerating the acquisition of MRI with a long imaging protocol. Specifically, a Dilated Residual Dense Network (DRDNet) is customized for dual domain restorations from undersampled MRI data. Extensive experiments on different sampling patterns and acceleration rates demonstrate that our method consistently outperforms state-of-the-art methods, and can reconstruct high quality MRI. Bo Zhou 0009, Shaohua Kevin Zhou |
CVPR | 2 |
| 2020 | BCData: A Large-Scale Dataset and Benchmark for Cell Detection and Counting
Yao Ding 0006, Guoli Song, Lin Wang 0026, Ruizhe Geng, Yonghong Tian 0001, Yongsheng Liang 0001, Shaohua Kevin Zhou, Jie Chen 0001 |
MICCAI (5) | 11 |
| 2020 | Bounding Maps for Universal Lesion Detection
Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (4) | 3 |
| 2020 | Encoding Metal Mask Projection for Metal Artifact Reduction in Computed Tomography
Yuanyuan Lyu, Wei-An Lin, Haofu Liao, Jingjing Lu, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2020 | Dual-Level Selective Transfer Learning for Intrahepatic Cholangiocarcinoma Segmentation in Non-enhanced Abdominal CT
Wenzhe Wang, Qingyu Song 0004, Jiarong Zhou, Ruiwei Feng, Tingting Chen 0002, Wenhao Ge, Danny Ziyi Chen, Shaohua Kevin Zhou, Jian Wu 0001 |
MICCAI (1) | 8 |
| 2020 | Miss the Point: Targeted Adversarial Attack on Multiple Landmark Detection
Qingsong Yao, Zecheng He, Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (4) | 4 |
| 2020 | One-to-one Mapping for Unpaired Image-to-image TranslationabstractRecently image-to-image translation has attracted significant interests in the literature, starting from the successful use of the generative adversarial network (GAN), to the introduction of cyclic constraint, to extensions to multiple domains. However, in existing approaches, there is no guarantee that the mapping between two image domains is unique or one-to-one. Here we propose a self-inverse network learning approach for unpaired image-to-image translation. Building on top of CycleGAN, we learn a self-inverse function by simply augmenting the training samples by swapping inputs and outputs during training and with separated cycle consistency loss for each mapping direction. The outcome of such learning is a proven one-to-one mapping function. Our extensive experiments on a variety of datasets, including cross-modal medical image synthesis, object transfiguration, and semantic labeling, consistently demonstrate clear improvement over the CycleGAN method both qualitatively and quantitatively. Especially our proposed method reaches the state-of-the-art result on the cityscapes benchmark dataset for the label to photo un-paired directional image translation. Zengming Shen, Thomas S. Huang, Shaohua Kevin Zhou, Bogdan Georgescu, Xuqi Liu |
WACV | 4 |
| 2020 | Rubik's Cube+: A self-supervised feature learning framework for 3D medical image analysis
Jiuwen Zhu, Yuexiang Li, Kai Ma 0002, Shaohua Kevin Zhou, Yefeng Zheng 0001 |
Medical Image Anal. | 5 |
| 2020 | High-Resolution Chest X-Ray Bone Suppression Using Unpaired CT Structural PriorsabstractThere is clinical evidence that suppressing the bone structures in Chest X-rays (CXRs) improves diagnostic value, either for radiologists or computer-aided diagnosis. However, bone-free CXRs are not always accessible. We hereby propose a coarse-to-fine CXR bone suppression approach by using structural priors derived from unpaired computed tomography (CT) images. In the low-resolution stage, we use the digitally reconstructed radiograph (DRR) image that is computed from CT as a bridge to connect CT and CXR. We then perform CXR bone decomposition by leveraging the DRR bone decomposition model learned from unpaired CTs and domain adaptation between CXR and DRR. To further mitigate the domain differences between CXRs and DRRs and speed up the learning convergence, we perform all the aboved operations in Laplacian of Gaussian (LoG) domain. After obtaining the bone decomposition result in DRR, we upsample it to a high resolution, based on which the bone region in the original high-resolution CXR is cropped and processed to produce a high-resolution bone decomposition result. Finally, such a produced bone image is subtracted from the original high-resolution CXR to obtain the bone suppression result. We conduct experiments and clinical evaluations based on two benchmarking CXR databases to show that (i) the proposed method outperforms the state-of-the-art unsupervised CXR bone suppression approaches; (ii) the CXRs with bone suppression are instrumental to radiologists for reducing their false-negative rate of lung diseases from 15% to 8%; and (iii) state-of-the-art disease classification performances are achieved by learning a deep network that takes the original CXR and its bone-suppressed image as inputs. Hu Han 0001, Zeju Li, Jingjing Lu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 7 |
| 2020 | ADN: Artifact Disentanglement Network for Unsupervised Metal Artifact ReductionabstractCurrent deep neural network based approaches to computed tomography (CT) metal artifact reduction (MAR) are supervised methods that rely on synthesized metal artifacts for training. However, as synthesized data may not accurately simulate the underlying physical mechanisms of CT imaging, the supervised methods often generalize poorly to clinical applications. To address this problem, we propose, to the best of our knowledge, the first unsupervised learning approach to MAR. Specifically, we introduce a novel artifact disentanglement network that disentangles the metal artifacts from CT images in the latent space. It supports different forms of generations (artifact reduction, artifact transfer, and self-reconstruction, etc.) with specialized loss functions to obviate the need for supervision with synthesized data. Extensive experiments show that when applied to a synthesized dataset, our method addresses metal artifacts significantly better than the existing unsupervised models designed for natural image-to-image translation problems, and achieves comparable performance to existing supervised models for MAR. When applied to clinical datasets, our method demonstrates better generalization ability over the supervised models. The source code of this paper is publicly available at https:// github.com/liaohaofu/adn. Haofu Liao, Wei-An Lin, Shaohua Kevin Zhou, Jiebo Luo 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Multiview 2D/3D Rigid Registration via a Point-Of-Interest Network for Tracking and TriangulationabstractWe propose to tackle the problem of multiview 2D/3D rigid registration for intervention via a Point-Of-Interest Network for Tracking and Triangulation (POINT2). POINT2learns to establish 2D point-to-point correspondences between the pre- and intra-intervention images by tracking a set of random POIs. The 3D pose of the pre-intervention volume is then estimated through a triangulation layer. In POINT2, the unified framework of the POI tracker and the triangulation layer enables learning informative 2D features and estimating 3D pose jointly. In contrast to existing approaches, POINT2only requires a single forward-pass to achieve a reliable 2D/3D registration. As the POI tracker is shift-invariant, POINT2is more robust to the initial pose of the 3D pre-intervention image. Extensive experiments on a large-scale clinical cone-beam CT (CBCT) dataset show that the proposed POINT2method outperforms the existing learning-based method in terms of accuracy, robustness and running time. Furthermore, when used as an initial pose estimator, our method also improves the robustness and speed of the state-of-the-art optimization-based approaches by ten folds. Haofu Liao, Wei-An Lin, Jingdan Zhang, Jiebo Luo 0001, Shaohua Kevin Zhou |
CVPR | 6 |
| 2019 | DuDoNet: Dual Domain Network for CT Metal Artifact ReductionabstractComputed tomography (CT) is an imaging modality widely used for medical diagnosis and treatment. CT images are often corrupted by undesirable artifacts when metallic implants are carried by patients, which creates the problem of metal artifact reduction (MAR). Existing methods for reducing the artifacts due to metallic implants are inadequate for two main reasons. First, metal artifacts are structured and non-local so that simple image domain enhancement approaches would not suffice. Second, the MAR approaches which attempt to reduce metal artifacts in the X-ray projection (sinogram) domain inevitably lead to severe secondary artifact due to sinogram inconsistency. To overcome these difficulties, we propose an end-to-end trainable Dual Domain Network (DuDoNet) to simultaneously restore sinogram consistency and enhance CT images. The linkage between the sigogram and image domains is a novel Radon inversion layer that allows the gradients to back-propagate from the image domain to the sinogram domain during training. Extensive experiments show that our method achieves significant improvements over other single domain MAR approaches. To the best of our knowledge, it is the first end-to-end dual-domain network for MAR. Wei-An Lin, Haofu Liao, Cheng Peng 0008, Xiaohang Sun, Jingdan Zhang, Jiebo Luo 0001, Rama Chellappa, Shaohua Kevin Zhou |
CVPR | 8 |
| 2019 | 3D U2-Net: A 3D Universal U-Net for Multi-domain Medical Image Segmentation
Chao Huang 0009, Hu Han 0001, Qingsong Yao, Shankuan Zhu, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2019 | Encoding CT Anatomy Knowledge for Unpaired Chest X-ray Image Decomposition
Zeju Li, Hu Han 0001, Gonglei Shi, Jiannan Wang 0005, Shaohua Kevin Zhou |
MICCAI (6) | 6 |
| 2019 | Generative Mask Pyramid Network for CT/CBCT Metal Artifact Reduction with Joint Projection-Sinogram Correction
Haofu Liao, Wei-An Lin, Zhimin Huo, Levon Vogelsang, William J. Sehnert, Shaohua Kevin Zhou, Jiebo Luo 0001 |
MICCAI (6) | 6 |
| 2019 | Artifact Disentanglement Network for Unsupervised Metal Artifact Reduction
Haofu Liao, Wei-An Lin, Shaohua Kevin Zhou, Jiebo Luo 0001 |
MICCAI (6) | 4 |
| 2018 | A Deep Cascade Network for Unaligned Face Attribute ClassificationabstractHumans focus attention on different face regions when recognizing face attributes. Most existing face attribute classification methods use the whole image as input. Moreover, some of these methods rely on fiducial landmarks to provide defined face parts. In this paper, we propose a cascade network that simultaneously learns to localize face regions specific to attributes and performs attribute classification without alignment. First, a weakly-supervised face region localization network is designed to automatically detect regions (or parts) specific to attributes. Then multiple part-based networks and a whole-image-based network are separately constructed and combined together by the region switch layer and attribute relation layer for final attribute classification. A multi-net learning method and hint-based model compression is further proposed to get an effective localization model and a compact classification model, respectively. Our approach achieves significantly better performance than state-of-the-art methods on unaligned CelebA dataset, reducing the classification error by 30.9%. Hui Ding 0002, Shaohua Kevin Zhou, Rama Chellappa |
AAAI | 3 |
| 2018 | Face Completion with Semantic Knowledge and Collaborative Adversarial Learning
Haofu Liao, Gareth Funka-Lea, Yefeng Zheng 0001, Jiebo Luo 0001, Shaohua Kevin Zhou |
ACCV (1) | 5 |
| 2018 | Adversarial Sparse-View CBCT Artifact Reduction
Haofu Liao, Zhimin Huo, William J. Sehnert, Shaohua Kevin Zhou, Jiebo Luo 0001 |
MICCAI (1) | 4 |
| 2018 | More Knowledge Is Better: Cross-Modality Volume Completion and 3D+2D Segmentation for Intracardiac Echocardiography Contouring
Haofu Liao, Yucheng Tang, Gareth Funka-Lea, Jiebo Luo 0001, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2018 | 3D Anisotropic Hybrid Network: Transferring Convolutional Features from 2D Images to 3D Anisotropic Volumes
Siqi Liu 0001, Daguang Xu, Shaohua Kevin Zhou, Olivier Pauly, Sasa Grbic, Thomas Mertelmeier, Julia Wicklein, Anna K. Jerebko, Tom Weidong Cai, Dorin Comaniciu |
MICCAI (2) | 3 |
| 2018 | Less is More: Simultaneous View Classification and Landmark Detection for Abdominal Ultrasound Images
Zhoubing Xu, Yuankai Huo, Jin Hyeong Park, Bennett A. Landman, Andy Milkowski, Sasa Grbic, Shaohua Kevin Zhou |
MICCAI (2) | 7 |
| 2018 | Learning to Prune Filters in Convolutional Neural NetworksabstractMany state-of-the-art computer vision algorithms use large scale convolutional neural networks (CNNs) as basic building blocks. These CNNs are known for their huge number of parameters, high redundancy in weights, and tremendous computing resource consumptions. This paper presents a learning algorithm to simplify and speed up these CNNs. Specifically, we introduce a “try-and-learn” algorithm to train pruning agents that remove unnecessary CNN filters in a data-driven way. With the help of a novel reward function, our agents removes a significant number of filters in CNNs while maintaining performance at a desired level. Moreover, this method provides an easy control of the tradeoff between network performance and its scale. Performance of our algorithm is validated with comprehensive pruning experiments on several popular CNNs for visual recognition and semantic segmentation tasks. Qiangui Huang, Shaohua Kevin Zhou, Suya You, Ulrich Neumann |
WACV | 2 |
| 2017 | FaceNet2ExpNet: Regularizing a Deep Face Recognition Net for Expression RecognitionabstractRelatively small data sets available for expression recognition research make the training of deep networks very challenging. Although fine-tuning can partially alleviate the issue, the performance is still below acceptable levels as the deep features probably contain redundant information from the pretrained domain. In this paper, we present FaceNet2ExpNet, a novel idea to train an expression recognition network based on static images. We first propose a new distribution function to model the high-level neurons of the expression network. Based on this, a two-stage training algorithm is carefully designed. In the pre-training stage, we train the convolutional layers of the expression net, regularized by the face net; In the refining stage, we append fully-connected layers to the pre-trained convolutional layers and train the whole network jointly. Visualization results show that the model trained with our method captures improved high-level expression semantics. Evaluations on four public expression databases, CK+, Oulu- CASIA, TFD, and SFEW demonstrate that our method achieves better results than state-of-the-art. Hui Ding 0002, Shaohua Kevin Zhou, Rama Chellappa |
FG | 2 |
| 2017 | Supervised Action Classifier: Approaching Landmark Detection as Image Partitioning
Zhoubing Xu, Qiangui Huang, Jin Hyeong Park, Mingqing Chen, Daguang Xu, Dong Yang 0005, David Liu 0001, Shaohua Kevin Zhou |
MICCAI (3) | 8 |
| 2017 | Deep Image-to-Image Recurrent Network with Shape Basis Learning for Automatic Vertebra Labeling in Large-Scale 3D CT Volumes
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Zhoubing Xu, Mingqing Chen, Jin Hyeong Park, Sasa Grbic, Trac D. Tran, Sang (Peter) Chin, Dimitris N. Metaxas, Dorin Comaniciu |
MICCAI (3) | 4 |
| 2017 | Automatic Liver Segmentation Using an Adversarial Image-to-Image Network
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Bogdan Georgescu, Mingqing Chen, Sasa Grbic, Dimitris N. Metaxas, Dorin Comaniciu |
MICCAI (3) | 3 |
| 2016 | Iterative Multi-domain Regularized Deep Learning for Anatomical Structure Detection and Segmentation from Ultrasound Images
Hao Chen 0011, Yefeng Zheng 0001, Jin Hyeong Park, Pheng-Ann Heng, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2015 | Unsupervised Cross-Modal Synthesis of Subject-Specific ScansabstractRecently, cross-modal synthesis of subject-specific scans has been receiving significant attention from the medical imaging community. Though various synthesis approaches have been introduced in the recent past, most of them are either tailored to a specific application or proposed for the supervised setting, i.e., they assume the availability of training data from the same set of subjects in both source and target modalities. But, collecting multiple scans from each subject is undesirable. Hence, to address this issue, we propose a general unsupervised cross-modal medical image synthesis approach that works without paired training data. Given a source modality image of a subject, we first generate multiple target modality candidate values for each voxel independently using cross-modal nearest neighbor search. Then, we select the best candidate values jointly for all the voxels by simultaneously maximizing a global mutual information cost function and a local spatial consistency cost function. Finally, we use coupled sparse representation for further refinement of synthesized images. Our experiments on generating T1-MRI brain scans from T2-MRI and vice versa demonstrate that the synthesis capability of the proposed unsupervised approach is comparable to various state-of-the-art supervised approaches in the literature. Raviteja Vemulapalli, Hien Van Nguyen, Shaohua Kevin Zhou |
ICCV | 3 |
| 2015 | Cross-Domain Synthesis of Medical Images Using Efficient Location-Sensitive Deep Network
Hien Van Nguyen, Shaohua Kevin Zhou, Raviteja Vemulapalli |
MICCAI (1) | 2 |
| 2015 | Automatic Segmentation of Spinal Canals in CT Images via Iterative Topology RefinementabstractAccurate segmentation of the spinal canals in computed tomography (CT) images is an important task in many related studies. In this paper, we propose an automatic segmentation method and apply it to our highly challenging image cohort that is acquired from multiple clinical sites and from the CT channel of the PET-CT scans. To this end, we adapt the interactive random-walk solvers to be a fully automatic cascaded pipeline. The automatic segmentation pipeline is initialized with robust voxelwise classification using Haar-like features and probabilistic boosting tree. Then, the topology of the spinal canal is extracted from the tentative segmentation and further refined for the subsequent random-walk solver. In particular, the refined topology leads to improved seeding voxels or boundary conditions, which allow the subsequent random-walk solver to improve the segmentation result. Therefore, by iteratively refining the spinal canal topology and cascading the random-walk solvers, satisfactory segmentation results can be acquired within only a few iterations, even for cases with scoliosis, bone fractures and lesions. Our experiments validate the capability of the proposed method with promising segmentation performance, even though the resolution and the contrast of our dataset with 110 patient cases (90 for testing and 20 for training) are low and various bone pathologies occur frequently. Qian Wang 0001, Le Lu 0001, Dijia Wu, Noha Youssry El-Zehiry, Yefeng Zheng 0001, Dinggang Shen, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 7 |
| 2014 | Lung Segmentation from CT with Severe Pathologies Using Anatomical Constraints
Neil Birkbeck, Timo Kohlberger, Jingdan Zhang, Michal Sofka, Jens N. Kaftan, Dorin Comaniciu, Shaohua Kevin Zhou |
MICCAI (1) | 7 |
| 2014 | Segmentation of Multiple Knee Bones from CT for Orthopedic Knee Surgery Planning
Dijia Wu, Michal Sofka, Neil Birkbeck, Shaohua Kevin Zhou |
MICCAI (1) | 4 |
| 2014 | Discriminative anatomy detection: Classification vs regression
Shaohua Kevin Zhou |
Pattern Recognit. Lett. | 1 |
| 2014 | Automatic Detection and Measurement of Structures in Fetal Head Ultrasound Volumes Using Sequential Estimation and Integrated Detection Network (IDN)abstractRoutine ultrasound exam in the second and third trimesters of pregnancy involves manually measuring fetal head and brain structures in 2-D scans. The procedure requires a sonographer to find the standardized visualization planes with a probe and manually place measurement calipers on the structures of interest. The process is tedious, time consuming, and introduces user variability into the measurements. This paper proposes an automatic fetal head and brain (AFHB) system for automatically measuring anatomical structures from 3-D ultrasound volumes. The system searches the 3-D volume in a hierarchy of resolutions and by focusing on regions that are likely to be the measured anatomy. The output is a standardized visualization of the plane with correct orientation and centering as well as the biometric measurement of the anatomy. The system is based on a novel framework for detecting multiple structures in 3-D volumes. Since a joint model is difficult to obtain in most practical situations, the structures are detected in a sequence, one-by-one. The detection relies on Sequential Estimation techniques, frequently applied to visual tracking. The interdependence of structure poses and strong prior information embedded in our domain yields faster and more accurate results than detecting the objects individually. The posterior distribution of the structure pose is approximated at each step by sequential Monte Carlo. The samples are propagated within the sequence across multiple structures and hierarchical levels. The probabilistic model helps solve many challenges present in the ultrasound images of the fetus such as speckle noise, signal drop-out, shadows caused by bones, and appearance variations caused by the differences in the fetus gestational age. This is possible by discriminative learning on an extensive database of scans comprising more than two thousand volumes and more than thirteen thousand annotations. The average difference between ground truth and automatic measurements is below 2 mm with a running time of 6.9 s (GPU) or 14.7 s (CPU). The accuracy of the AFHB system is within inter-user variability and the running time is fast, which meets the requirements for clinical use. Michal Sofka, Jingdan Zhang, Sara Good, Shaohua Kevin Zhou, Dorin Comaniciu |
IEEE Trans. Medical Imaging | 4 |
| 2013 | Learning the Manifold of Quality Ultrasound Acquisition
Noha Youssry El-Zehiry, Michelle Yan, Sara Good, Tong Fang, Shaohua Kevin Zhou, Leo J. Grady |
MICCAI (1) | 5 |
| 2013 | Automatic Nuchal Translucency Measurement from Ultrasonography
Jin Hyeong Park, Michal Sofka, SunMi Lee, Shaohua Kevin Zhou |
MICCAI (3) | 5 |
| 2013 | Recognizing Interactive Group Activities Using Temporal Interaction Matrices and Their Riemannian Statistics
Rama Chellappa, Shaohua Kevin Zhou |
Int. J. Comput. Vis. | 3 |
| 2013 | Lymph node detection and segmentation in chest CT data using discriminative learning and a spatial prior
Johannes Feulner, Shaohua Kevin Zhou, Matthias Hammon, Joachim Hornegger, Dorin Comaniciu |
Medical Image Anal. | 2 |
| 2013 | Spine detection in CT and MR using iterated marginal space learning
B. Michael Kelm, Michael Wels, Shaohua Kevin Zhou, Sascha Seifert, Michael Sühling, Yefeng Zheng 0001, Dorin Comaniciu |
Medical Image Anal. | 3 |
| 2012 | A learning based deformable template matching method for automatic rib centerline extraction and labeling in CT imagesabstractThe automatic extraction and labeling of the rib centerlines is a useful yet challenging task in many clinical applications. In this paper, we propose a new approach integrating rib seed point detection and template matching to detect and identify each rib in chest CT scans. The bottom-up learning based detection exploits local image cues and top-down deformable template matching imposes global shape constraints. To adapt to the shape deformation of different rib cages whereas maintain high computational efficiency, we employ a Markov Random Field (MRF) based articulated rigid transformation method followed by Active Contour Model (ACM) deformation. Compared with traditional methods that each rib is individually detected, traced and labeled, the new approach is not only much more robust due to prior shape constraints of the whole rib cage, but removes tedious post-processing such as rib pairing and ordering steps because each rib is automatically labeled during the template matching. For experimental validation, we create an annotated database of 112 challenging volumes with ribs of various sizes, shapes, and pathologies such as metastases and fractures. The proposed approach shows orders of magnitude higher detection and labeling accuracy than state-of-the-art solutions and runs about 40 seconds for a complete rib cage on the average. Dijia Wu, David Liu 0001, Zoltan Puskas, Chao Lu 0011, Andreas Wimmer, Christian Tietjen, Grzegorz Soza, Shaohua Kevin Zhou |
CVPR | 8 |
| 2012 | Anatomical Landmark Detection Using Nearest Neighbor Matching and Submodular Optimization
David Liu 0001, Shaohua Kevin Zhou |
MICCAI (3) | 2 |
| 2012 | Precise Segmentation of Multiple Organs in CT Volumes Using Learning-Based Approach and Information Theory
Chao Lu 0011, Yefeng Zheng 0001, Neil Birkbeck, Jingdan Zhang, Timo Kohlberger, Christian Tietjen, Thomas Böttger, James S. Duncan, Shaohua Kevin Zhou |
MICCAI (2) | 9 |
| 2012 | Automatic Detection and Segmentation of Lymph Nodes From CT DataabstractLymph nodes are assessed routinely in clinical practice and their size is followed throughout radiation or chemotherapy to monitor the effectiveness of cancer treatment. This paper presents a robust learning-based method for automatic detection and segmentation of solid lymph nodes from CT data, with the following contributions. First, it presents a learning based approach to solid lymph node detection that relies on marginal space learning to achieve great speedup with virtually no loss in accuracy. Second, it presents a computationally efficient segmentation method for solid lymph nodes (LN). Third, it introduces two new sets of features that are effective for LN detection, one that self-aligns to high gradients and another set obtained from the segmentation result. The method is evaluated for axillary LN detection on 131 volumes containing 371 LN, yielding a 83.0% detection rate with 1.0 false positive per volume. It is further evaluated for pelvic and abdominal LN detection on 54 volumes containing 569 LN, yielding a 80.0% detection rate with 3.2 false positives per volume. The running time is 5-20 s per volume for axillary areas and 15-40 s for pelvic. An added benefit of the method is the capability to detect and segment conglomerated lymph nodes. Adrian Barbu, Michael Sühling, David Liu 0001, Shaohua Kevin Zhou, Dorin Comaniciu |
IEEE Trans. Medical Imaging | 5 |
| 2011 | Learning-based hypothesis fusion for robust catheter tracking in 2D X-ray fluoroscopyabstractCatheter tracking has become more and more important in recent interventional applications. It provides real time navigation for the physicians and can be used to control a motion compensated fluoro overlay reference image for other means of guidance, e.g. involving a 3D anatomical model. Tracking the coronary sinus (CS) catheter is effective to compensate respiratory and cardiac motion for 3D overlay navigation to assist positioning the ablation catheter in Atrial Fibrillation (Afib) treatments. During interventions, the CS catheter performs rapid motion and non-rigid deformation due to the beating heart and respiration. In this paper, we model the CS catheter as a set of electrodes. Novelly designed hypotheses generated by a number of learning-based detectors are fused. Robust hypothesis matching through a Bayesian framework is then used to select the best hypothesis for each frame. As a result, our tracking method achieves very high robustness against challenging scenarios such as low SNR, occlusion, foreshortening, non-rigid deformation, as well as the catheter moving in and out of ROI. Quantitative evaluation has been conducted on a database of 13221 frames from 1073 sequences. Our approach obtains 0.50mm median error and 0.76mm mean error. 97.8% of evaluated data have errors less than 2.00mm. The speed of our tracking algorithm reaches 5 frames-per-second on most data sets. Our approach is not limited to the catheters inside the CS but can be extended to track other types of catheters, such as ablation catheters or circumferential mapping catheters. Wen Wu 0004, Terrence Chen, Peng Wang 0005, Shaohua Kevin Zhou, Dorin Comaniciu, Adrian Barbu, Norbert Strobel |
CVPR | 4 |
| 2011 | Detection, Grading and Classification of Coronary Stenoses in Computed Tomography Angiography
B. Michael Kelm, Sushil Mittal, Yefeng Zheng 0001, Alexey Tsymbal, Dominik Bernhardt, Fernando Vega Higuera, Shaohua Kevin Zhou, Peter Meer, Dorin Comaniciu |
MICCAI (3) | 7 |
| 2011 | Automatic Multi-organ Segmentation Using Learning-Based Segmentation and Level Set Optimization
Timo Kohlberger, Michal Sofka, Jingdan Zhang, Neil Birkbeck, Jens Wetzl, Jens N. Kaftan, Jérôme Declerck, Shaohua Kevin Zhou |
MICCAI (3) | 8 |
| 2011 | Multi-stage Learning for Robust Lung Segmentation in Challenging CT Volumes
Michal Sofka, Jens Wetzl, Neil Birkbeck, Jingdan Zhang, Timo Kohlberger, Jens N. Kaftan, Jérôme Declerck, Shaohua Kevin Zhou |
MICCAI (3) | 8 |
| 2011 | Automatic Contrast Phase Estimation in CT Volumes
Michal Sofka, Dijia Wu, Michael Sühling, David Liu 0001, Christian Tietjen, Grzegorz Soza, Shaohua Kevin Zhou |
MICCAI (3) | 7 |
| 2011 | Efficient Detection of Native and Bypass Coronary Ostia in Cardiac CT Volumes: Anatomical vs. Pathological Structures
Yefeng Zheng 0001, Hüseyin Tek, Gareth Funka-Lea, Shaohua Kevin Zhou, Fernando Vega Higuera, Dorin Comaniciu |
MICCAI (3) | 4 |
| 2011 | Multi-part Left Atrium Modeling and Segmentation in C-Arm CT Volumes for Atrial Fibrillation Ablation
Yefeng Zheng 0001, Tianzhou Wang, Matthias John 0001, Shaohua Kevin Zhou, Jan M. Boese, Dorin Comaniciu |
MICCAI (3) | 4 |
| 2011 | A Probabilistic Model for Automatic Segmentation of the Esophagus in 3-D CT ScansabstractBeing able to segment the esophagus without user interaction from 3-D CT data is of high value to radiologists during oncological examinations of the mediastinum. The segmentation can serve as a guideline and prevent confusion with pathological tissue. However, limited contrast to surrounding structures and versatile shape and appearance make segmentation a challenging problem. This paper presents a multistep method. First, a detector that is trained to learn a discriminative model of the appearance is combined with an explicit model of the distribution of respiratory and esophageal air. In the next step, prior shape knowledge is incorporated using a Markov chain model. We follow a "detect and connect" approach to obtain the maximum a posteriori estimate of the approximate esophagus shape from hypothesis about the esophagus contour in axial image slices. Finally, the surface of this approximation is nonrigidly deformed to better fit the boundary of the organ. The method is compared to an alternative approach that uses a particle filter instead of a Markov chain to infer the approximate esophagus shape, to the performance of a human observer and also to state of the art methods, which are all semiautomatic. Cross-validation on 144 CT scans showed that the Markov chain based approach clearly outperforms the particle filter. It segments the esophagus with a mean error of 1.80 mm in less than 16 s on a standard PC. This is only 1 mm above the interobserver variability and can compete with the results of previously published semiautomatic methods. Johannes Feulner, Shaohua Kevin Zhou, Matthias Hammon, Sascha Seifert, Martin Huber 0001, Dorin Comaniciu, Joachim Hornegger, Alexander Cavallaro |
IEEE Trans. Medical Imaging | 2 |
| 2010 | Body landmark detection for a fully automatic AAA stent graft planning software systemabstractIn this paper, we present an approach to automate the planning of an endovascular stent graft for abdominal aortic aneurysms (AAAs), which are treated with bifurcated prosthesis (Y-stents) when located close to the iliac bifurcation. During the intervention, the folded Y-stent graft — consisting of several parts — is inserted via the iliac region and expanded inside the patient's body. The first step of the proposed approach is to detect different body landmarks with a statistical method. In the next step, these landmarks are used to calculate two vascular centerlines, which provide multiplanar reformatting (MPR) slices that are used for an automatic segmentation of the artery walls. The segmented artery walls provide the manufacturer specific measures to choose an adequate bifurcated prosthesis. In a final step, the expansion of the stent is simulated in the patient's data. Results for 50 abdominal aortic aneurysm cases provided by computed tomography angiography (CTA) acquisitions are successfully verified by a virtual stenting expert. Jan Egger, Shaohua Kevin Zhou, Stefan Großkopf, David Liu 0001, Christian Hopfgartner, Dominik Bernhardt, Christina Biermann, Christopher Nimsky, Bernd Freisleben |
CBMS | 2 |
| 2010 | Lymph node detection in 3-D chest CT using a spatial prior probabilityabstractLymph nodes have high clinical relevance but detection is challenging as they are hard to see due to low contrast and irregular shape. In this paper, a method for fully automatic mediastinal lymph node detection in 3-D computed tomography (CT) images of the chest area is proposed. Discriminative learning is used to detect lymph nodes based on their appearance. Because lymph nodes can easily be confused with other structures, it is vital to incorporate as much anatomical knowledge as possible to achieve good detection rates. Here, a learned prior of the spatial distribution is proposed to model this knowledge. As atlas matching is generally inaccurate in the chest area because of anatomical variations, this prior is not learned in the space of a single atlas, but in the space of multiple ones that are attached to anatomical structures. During test, the priors are weighted and merged according to spatial distances. Cross-validation on 54 CT datasets showed that the prior based detector yields a true positive rate of 52.3% for seven false positives per volume image, which is about two times better than without a spatial prior. Johannes Feulner, Shaohua Kevin Zhou, Martin Huber 0001, Joachim Hornegger, Dorin Comaniciu, Alexander Cavallaro |
CVPR | 2 |
| 2010 | Search strategies for multiple landmark detection by submodular maximizationabstractA fundamental issue in multiple landmark detection is the reduction of computational cost. This problem has previously been addressed mainly by reducing the complexity of each individual landmark detector. We address the problem by optimizing the search strategy of multiple landmarks. When the relative positions of landmarks are constrained, the search space can be reduced, thereby reducing the computation. The proposed method leverages the theory of submodular functions to provide a constant factor approximation guarantee of the optimal speed. Although the theory of submodular functions is well known, to the best of our knowledge, this is the first time it is applied to the landmark detection problem. We demonstrate our method by fast and accurate detection of human body landmarks including bones, organs, and vessels in 3D CT images from a diverse dataset of around 2000 volumes with pathological patients. We further provide different search space criteria and variations. David Liu 0001, Shaohua Kevin Zhou, Dominik Bernhardt, Dorin Comaniciu |
CVPR | 2 |
| 2010 | Multiple object detection by sequential monte carlo and Hierarchical Detection NetworkabstractIn this paper, we propose a novel framework for detecting multiple objects in 2D and 3D images. Since a joint multi-object model is difficult to obtain in most practical situations, we focus here on detecting the objects sequentially, one-by-one. The interdependence of object poses and strong prior information embedded in our domain of medical images results in better performance than detecting the objects individually. Our approach is based on Sequential Estimation techniques, frequently applied to visual tracking. Unlike in tracking, where the sequential order is naturally determined by the time sequence, the order of detection of multiple objects must be selected, leading to a Hierarchical Detection Network (HDN). We present an algorithm that optimally selects the order based on probability of states (object poses) within the ground truth region. The posterior distribution of the object pose is approximated at each step by sequential Monte Carlo. The samples are propagated within the sequence across multiple objects and hierarchical levels. We show on 2D ultrasound images of left atrium, that the automatically selected sequential order yields low mean detection error. We also quantitatively evaluate the hierarchical detection of fetal faces and three fetal brain structures in 3D ultrasound images. Michal Sofka, Jingdan Zhang, Shaohua Kevin Zhou, Dorin Comaniciu |
CVPR | 3 |
| 2010 | Automatic Detection and Segmentation of Axillary Lymph Nodes
Adrian Barbu, Michael Sühling, David Liu 0001, Shaohua Kevin Zhou, Dorin Comaniciu |
MICCAI (1) | 5 |
| 2010 | Model-Based Esophagus Segmentation from CT Scans Using a Spatial Probability Map
Johannes Feulner, Shaohua Kevin Zhou, Martin Huber 0001, Alexander Cavallaro, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 2 |
| 2010 | Cross-Modality Assessment and Planning for Pulmonary Trunk Treatment Using CT and MRI Imaging
Dime Vitanovski, Alexey Tsymbal, Razvan Ioan Ionasec, Bogdan Georgescu, Martin Huber 0001, Andrew Mayall Taylor, Silvia Schievano, Shaohua Kevin Zhou, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 8 |
| 2010 | Graph Based Interactive Detection of Curve Structures in 2D Fluoroscopy
Peng Wang 0005, Wei-shing Liao, Terrence Chen, Shaohua Kevin Zhou, Dorin Comaniciu |
MICCAI (3) | 4 |
| 2010 | Automatic Aorta Segmentation and Valve Landmark Detection in C-Arm CT: Application to Aortic Valve Implantation
Yefeng Zheng 0001, Matthias John 0001, Rui Liao, Jan M. Boese, Uwe Kirschstein, Bogdan Georgescu, Shaohua Kevin Zhou, Jörg Kempfert, Thomas Walther, Gernot Brockmann |
MICCAI (1) | 7 |
| 2010 | Shape regression machine and efficient segmentation of left ventricle endocardium from 2D B-mode echocardiogram
Shaohua Kevin Zhou |
Medical Image Anal. | 1 |
| 2010 | Robust Height Estimation of Moving Objects From Uncalibrated VideosabstractThis paper presents an approach for video metrology. From videos acquired by an uncalibrated stationary camera, we first recover the vanishing line and the vertical point of the scene based upon tracking moving objects that primarily lie on a ground plane. Using geometric properties of moving objects, a probabilistic model is constructed for simultaneously grouping trajectories and estimating vanishing points. Then we apply a single view mensuration algorithm to each of the frames to obtain height measurements. We finally fuse the multiframe measurements using the least median of squares (LMedS) as a robust cost function and the Robbins-Monro stochastic approximation (RMSA) technique. This method enables less human supervision, more flexibility and improved robustness. From the uncertainty analysis, we conclude that the method with auto-calibration is robust in practice. Results are shown based upon realistic tracking data from a variety of scenes. Jie Shao 0007, Shaohua Kevin Zhou, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2009 | Automatic fetal face detection from ultrasound volumes via learning 3D and 2D informationabstract3D ultrasound imaging has been increasingly used in clinics for fetal examination. However, manually searching for the optimal view of the fetal face in 3D ultrasound volumes is cumbersome and time-consuming even for expert physicians and sonographers. In this paper we propose a learning-based approach which combines both 3D and 2D information for automatic and fast fetal face detection from 3D ultrasound volumes. Our approach applies a new technique - constrained marginal space learning - for 3D face mesh detection, and combines a boosting-based 2D profile detection to refine 3D face pose. To enhance the rendering of the fetal face, an automatic carving algorithm is proposed to remove all obstructions in front of the face based on the detected face mesh. Experiments are performed on a challenging 3D ultrasound data set containing 1010 fetal volumes. The results show that our system not only achieves excellent detection accuracy but also runs very fast - it can detect the fetal face from the 3D data in 1 second on a dual-core 2.0 GHz computer. Shaolei Feng 0001, Shaohua Kevin Zhou, Sara Good, Dorin Comaniciu |
CVPR | 2 |
| 2009 | Learning multi-modal densities on Discriminative Temporal Interaction Manifold for group activity recognitionabstractWhile video-based activity analysis and recognition has received much attention, existing body of work mostly deals with single object/person case. Coordinated multi-object activities, or group activities, present in a variety of applications such as surveillance, sports, and biological monitoring records, etc., are the main focus of this paper. Unlike earlier attempts which model the complex spatial temporal constraints among multiple objects with a parametric Bayesian network, we propose a Discriminative Temporal Interaction Manifold (DTIM) framework as a data-driven strategy to characterize the group motion pattern without employing specific domain knowledge. In particular, we establish probability densities on the DTIM, whose element, the discriminative temporal interaction matrix, compactly describes the coordination and interaction among multiple objects in a group activity. For each class of group activity we learn a multi-modal density function on the DTIM. A Maximum a Posteriori (MAP) classifier on the manifold is then designed for recognizing new activities. Experiments on football play recognition demonstrate the effectiveness of the approach. Rama Chellappa, Shaohua Kevin Zhou |
CVPR | 3 |
| 2009 | Robust guidewire tracking in fluoroscopyabstractA guidewire is a medical device inserted into vessels during image guided interventions for balloon inflation. During interventions, the guidewire undergoes non-rigid deformation due to patients' breathing and cardiac motions, and such 3D motions are complicated when being projected onto the 2D fluoroscopy. Furthermore, in fluoroscopy there exist severe image artifacts and other wire-like structures. All these make robust guidewire tracking challenging. To address these challenges, this paper presents a probabilistic framework for robust guidewire tracking. We first introduce a semantic guidewire model that contains three parts, including a catheter tip, a guidewire tip and a guidewire body. Measurements of different parts are integrated into a Bayesian framework as measurements of a whole guidewire for robust guidewire tracking. Moreover, for each part, two types of measurements, one from learning-based detectors and the other from online appearance models, are applied and combined. A hierarchical and multi-resolution tracking scheme is then developed based on kernel-based measurement smoothing to track guidewires effectively and efficiently in a coarse-to-fine manner. The presented framework has been validated on a test set of 47 sequences, and achieves a mean tracking error of less than 2 pixels. This demonstrates the great potential of our method for clinical applications. Peng Wang 0005, Terrence Chen, Ying Zhu 0006, Wei Zhang 0018, Shaohua Kevin Zhou, Dorin Comaniciu |
CVPR | 5 |
| 2009 | Constrained marginal space learning for efficient 3D anatomical structure detection in medical imagesabstractRecently, we proposed marginal space learning (MSL) as a generic approach for automatic detection of 3D anatomical structures in many medical imaging modalities. To accurately localize a 3D object, we need to estimate nine parameters (three for position, three for orientation, and three for anisotropic scaling). Instead of uniformly searching the original nine-dimensional parameter space, only low-dimensional marginal spaces are uniformly searched in MSL, which significantly improves the speed. In many real applications, a strong correlation may exist among parameters in the same marginal spaces. For example, a large object may have large scaling values along all directions. In this paper, we propose constrained MSL to exploit this correlation for further speed-up. As another major contribution, we propose to use quaternions for 3D orientation representation and distance measurement to overcome the inherent drawbacks of Euler angles in the original MSL. The proposed method has been tested on three 3D anatomical structure detection problems in medical images, including liver detection in computed tomography (CT) volumes, and left ventricle detection in both CT and ultrasound volumes. Experiments on the largest datasets ever reported show that constrained MSL can improve the detection speed up to 14 times, while achieving comparable or better detection accuracy. It takes less than half a second to detect a 3D anatomical structure in a volume. Yefeng Zheng 0001, Bogdan Georgescu, Haibin Ling, Shaohua Kevin Zhou, Michael Scheuering, Dorin Comaniciu |
CVPR | 4 |
| 2009 | Automatic ovarian follicle quantification from 3D ultrasound data using global/local context with database guided segmentationabstractIn this paper, we present a novel probabilistic framework for automatic follicle quantification in 3D ultrasound data. The proposed framework robustly estimates size and location of each individual ovarian follicle by fusing the information from both global and local context. Follicle candidates at detected locations are then segmented by a novel database guided segmentation method. To efficiently search hypothesis in a high dimensional space for multiple object detection, a clustered marginal space learning approach is introduced. Extensive evaluations conducted on 501 volumes containing 8108 follicles showed that our method is able to detect and segment ovarian follicles with high robustness and accuracy. It is also much faster than the current ultrasound manual workflow. The proposed method is able to streamline the clinical workflow and improve the accuracy of existing follicular measurements. Terrence Chen, Wei Zhang 0018, Sara Good, Shaohua Kevin Zhou, Dorin Comaniciu |
ICCV | 4 |
| 2009 | Context ranking machine and its application to rigid localization of deformable objectsabstractIn this paper, we exploit the context information embodied in an image to develop a machine learning method called context ranking machine (CRM). Specifically, we leverage two kinds of context information: identity context and metric context. The identity context of an image patch refers to its origin (e.g., from which image it is cropped), and the metric context refers to its distance to the exact surrounding box of the target object inside the image. We use these context information in two ways. First, for object localization, instead of learning classifiers to separate the whole negative pool from all positives, we separate each positive from its own negatives sharing the same identity context. Second, we rank image patches according to their resemblance to the ground truth by establishing a connection between appearance based features and metric properties of the image. The CRM learns an image-based ranking algorithm via boosting and achieves an improved localization accuracy. We performed tests on echocardiogram images to localize heart chambers, and face images for eye band localization. Birkan Tunç, Shaohua Kevin Zhou, Jin Hyeong Park, Muhittin Gökmen |
ICIP | 2 |
| 2009 | Fast Automatic Segmentation of the Esophagus from 3D CT Data Using a Probabilistic Model
Johannes Feulner, Shaohua Kevin Zhou, Alexander Cavallaro, Sascha Seifert, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 2 |
| 2009 | Coronary Tree Extraction Using Motion Layer Separation
Wei Zhang 0018, Haibin Ling, Simone Prummer, Shaohua Kevin Zhou, Martin Ostermeier, Dorin Comaniciu |
MICCAI (1) | 4 |
| 2009 | Variational Graph Embedding for Globally and Locally Consistent Feature Extraction
Shuang-Hong Yang, Hongyuan Zha, Shaohua Kevin Zhou, Bao-Gang Hu |
ECML/PKDD (2) | 3 |
| 2009 | Appearance Modeling Using a Geometric TransformabstractA general transform, called the geometric transform (GeT), that models the appearance inside a closed contour is proposed. The proposed GeT is a functional of an image intensity function and a region indicator function derived from a closed contour. It can be designed to combine the shape and appearance information at different resolutions and to generate models invariant to deformation, articulation, or occlusion. By choosing appropriate functionals and region indicator functions, the GeT unifies Radon transform, trace transform, and a class of image warpings. By varying the region indicator and the types of features used for appearance modeling, five novel types of GeTs are introduced and applied to fingerprinting the appearance inside a contour. They include the GeTs based on a level set, shape matching, feature curves, and the GeT invariant to occlusion, and a multiresolution GeT (MRGeT). Applications of GeT to pedestrian identity recognition, human body part segmentation, and image synthesis are illustrated. The proposed approach produces promising results when applied to fingerprinting the appearance of a human and body parts despite the presence of nonrigid deformations and articulated motion. Jian Li 0022, Shaohua Kevin Zhou, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2008 | Hierarchical, learning-based automatic liver segmentationabstractIn this paper we present a hierarchical, learning-based approach for automatic and accurate liver segmentation from 3D CT volumes. We target CT volumes that come from largely diverse sources (e.g., diseased in six different organs) and are generated by different scanning protocols (e.g., contrast and non-contrast, various resolution and position). Three key ingredients are combined to solve the segmentation problem. First, a hierarchical framework is used to efficiently and effectively monitor the accuracy propagation in a coarse-to-fine fashion. Second, two new learning techniques, marginal space learning and steerable features, are applied for robust boundary inference. This enables handling of highly heterogeneous texture pattern. Third, a novel shape space initialization is proposed to improve traditional methods that are limited to similarity transformation. The proposed approach is tested on a challenging dataset containing 174 volumes. Our approach not only produces excellent segmentation accuracy, but also runs about fifty times faster than state-of-the-art solutions [7, 9]. Haibin Ling, Shaohua Kevin Zhou, Yefeng Zheng 0001, Bogdan Georgescu, Michael Sühling, Dorin Comaniciu |
CVPR | 2 |
| 2008 | Conditional density learning via regression with application to deformable shape segmentationabstractMany vision problems can be cast as optimizing the conditional probability density function p(C\I) where I is an image and C is a vector of model parameters describing the image. Ideally, the density function p(C\I) would be smooth and unimodal allowing local optimization techniques, such as gradient descent or simplex, to converge to an optimal solution quickly, while preserving significant nonlinearities of the model. We propose to learn a conditional probability density satisfying these desired properties for the given training data set. To do this, we formulate a novel regression problem that finds a function approximating the target density. Learning the regressor is challenging due to the high dimensionality of model parameters, C, and the complexity of relating the image and the model. Our approach makes two contributions. First, we take a multilevel refinement approach by learning a series of density functions, each of which guides the solution of optimization algorithms increasingly converging to the correct solution. Second, we propose a new data sampling algorithm that takes into account the gradient information of the target function. We have applied this learning approach to deformable shape segmentation and have achieved better accuracy than the previous methods. Jingdan Zhang, Shaohua Kevin Zhou, Dorin Comaniciu, Leonard McMillan |
CVPR | 2 |
| 2008 | Discriminative Learning for Deformable Shape Segmentation: A Comparative Study
Jingdan Zhang, Shaohua Kevin Zhou, Dorin Comaniciu, Leonard McMillan |
ECCV (1) | 2 |
| 2008 | Automatic Mitral Valve Inflow Measurements from Doppler Echocardiography
Jin Hyeong Park, Shaohua Kevin Zhou, John Jackson, Dorin Comaniciu |
MICCAI (1) | 2 |
| 2008 | AutoGate: Fast and Automatic Doppler Gate Localization in B-Mode Echocardiogram
Jin Hyeong Park, Shaohua Kevin Zhou, Costas Simopoulos, Dorin Comaniciu |
MICCAI (2) | 2 |
| 2007 | Joint Real-time Object Detection and Pose Estimation Using Probabilistic Boosting NetworkabstractIn this paper, we present a learning procedure called probabilistic boosting network (PBN) for joint real-time object detection and pose estimation. Grounded on the law of total probability, PBN integrates evidence from two building blocks, namely a multiclass boosting classifier for pose estimation and a boosted detection cascade for object detection. By inferring the pose parameter, we avoid the exhaustive scanning for the pose, which hampers real time requirement. In addition, we only need one integral image/volume with no need of image/volume rotation. We implement PBN using a graph-structured network that alternates the two tasks of foreground/background discrimination and pose estimation for rejecting negatives as quickly as possible. Compared with previous approaches, we gain accuracy in object localization and pose estimation while noticeably reducing the computation. We invoke PBN to detect the left ventricle from a 3D ultrasound volume, processing about 10 volumes per second, and the left atrium from 2D images in real time. Jingdan Zhang, Shaohua Kevin Zhou, Leonard McMillan, Dorin Comaniciu |
CVPR | 2 |
| 2007 | A boosting regression approach to medical anatomy detectionabstractThe state-of-the-art object detection algorithm learns a binary classifier to differentiate the foreground object from the background. Since the detection algorithm exhaustively scans the input image for object instances by testing the classifier, its computational complexity linearly depends on the image size and, if say orientation and scale are scanned, the number of configurations in orientation and scale. We argue that exhaustive scanning is unnecessary when detecting medical anatomy because a medical image offers strong contextual information. We then present an approach to effectively leveraging the medical context, leading to a solution that needs only one scan in theory or several sparse scans in practice and only one integral image even when the rotation is considered. The core is to learn a regression function, based on an annotated database, that maps the appearance observed in a scan window to a displacement vector, which measures the difference between the configuration being scanned and that of the target object. To achieve the learning task, we propose an image-based boosting ridge regression algorithm, which exhibits good generalization capability and training efficiency. Coupled with a binary classifier as a confidence scorer, the regression approach becomes an effective tool for detecting left ventricle in echocardiogram, achieving improved accuracy over the state-of-the-art object detection algorithm with significantly less computation. Shaohua Kevin Zhou, Jinghao Zhou, Dorin Comaniciu |
CVPR | 1 |
| 2007 | Automatic Cardiac View Classification of EchocardiogramabstractWe propose a fully automatic system for cardiac view classification of echocardiogram. Given an echo study video sequence, the system outputs a view label among the pre-defined standard views. The system is built based on a machine learning approach that extracts knowledge from an annotated database. It characterizes three features: 1) integrating local and global evidence, 2) utilizing view specific knowledge, and 3) employing a multi-class Logit-boost algorithm. In our prototype system, we classify four standard cardiac views: apical four chamber and apical two chamber, parasternal long axis and parasternal short axis (at mid cavity). We achieve a classification accuracy over 96% both of training and test data sets and the system runs in a second in the environment of Pentium 4 PC with 3.4 GHz CPU and 1.5 G RAM. Jin Hyeong Park, Shaohua Kevin Zhou, Costas Simopoulos, Joanne Otsuki, Dorin Comaniciu |
ICCV | 2 |
| 2007 | Robust Visual Tracking Using the Time-Reversibility ConstraintabstractVisual tracking is a very important front-end to many vision applications. We present a new framework for robust visual tracking in this paper. Instead of just looking forward in the time domain, we incorporate both forward and backward processing of video frames using a novel time-reversibility constraint. This leads to a new minimization criterion that combines the forward and backward similarity functions and the distances of the state vectors between the forward and backward states of the tracker. The new framework reduces the possibility of the tracker getting stuck in local minima and significantly improves the tracking robustness and accuracy. Our approach is general enough to be incorporated into most of the current tracking algorithms. We illustrate the improvements due to the proposed approach for the popular KLT tracker and a search based tracker. The experimental results show that the improved KLT tracker significantly outperforms the original KLT tracker. The time-reversibility constraint used for tracking can be incorporated to improve the performance of optical flow, mean shift tracking and other algorithms. Hao Wu 0014, Rama Chellappa, Aswin C. Sankaranarayanan, Shaohua Kevin Zhou |
ICCV | 4 |
| 2007 | A probabilistic, hierarchical, and discriminant framework for rapid and accurate detection of deformable anatomic structureabstractWe propose a probabilistic, hierarchical, and discriminant (PHD) framework for fast and accurate detection of deformable anatomic structures from medical images. The PHD framework has three characteristics. First, it integrates distinctive primitives of the anatomic structures at global, segmental, and landmark levels in a probabilistic manner. Second, since the configuration of the anatomic structures lies in a high-dimensional parameter space, it seeks the best configuration via a hierarchical evaluation of the detection probability that quickly prunes the search space. Finally, to separate the primitive from the background, it adopts a discriminative boosting learning implementation. We apply the PHD framework for accurately detecting various deformable anatomic structures from M- mode and Doppler echocardiograms in about a second. Shaohua Kevin Zhou, Gustavo Carneiro 0001, John Jackson, M. Brendel, Costas Simopoulos, Joanne Otsuki, Dorin Comaniciu |
ICCV | 1 |
| 2007 | Probabilistic Visual Tracking via Robust Template Matching and Incremental Subspace UpdateabstractIn this paper, we present a probabilistic algorithm for visual tracking that incorporates robust template matching and incremental sub-space update. There are two template matching methods used in the tracker: one is robust to small perturbation and the other to background clutter. Each method yields a probability of matching. Further, the templates are modeled using mixed probabilities and updated once the templates in the library cannot capture the variation of object appearance. We also model the tracking history using a nonlinear subspace that is described by probabilistic kernel principal components analysis, which provides a third probability. The most-recent tracking result is added to the nonlinear subspace incrementally. This update is performed efficiently by augmenting the kernel Gram matrix with one row and one column. The product of the three probabilities is defined as the observation likelihood used in a particle filter to derive the tracking result. Experimental results demonstrate the efficiency and effectiveness of the proposed algorithm. Xue Mei, Shaohua Kevin Zhou, Fatih Porikli |
ICME | 2 |
| 2007 | Appearance Characterization of Linear Lambertian Objects, Generalized Photometric Stereo, and Illumination-Invariant Face RecognitionabstractTraditional photometric stereo algorithms employ a Lambertian reflectance model with a varying albedo field and involve the appearance of only one object. In this paper, we generalize photometric stereo algorithms to handle all appearances of all objects in a class, in particular the human face class, by making use of the linear Lambertian property. A linear Lambertian object is one which is linearly spanned by a set of basis objects and has a Lambertian surface. The linear property leads to a rank constraint and, consequently, a factorization of an observation matrix that consists of exemplar images of different objects (e.g., faces of different subjects) under different, unknown illuminations. Integrability and symmetry constraints are used to fully recover the subspace bases using a novel linearized algorithm that takes the varying albedo field into account. The effectiveness of the linear Lambertian property is further investigated by using it for the problem of illumination-invariant face recognition using just one image. Attached shadows are incorporated in the model by a careful treatment of the inherent nonlinearity in Lambert's law. This enables us to extend our algorithm to perform face recognition in the presence of multiple illumination sources. Experimental results using standard data sets are presented. Shaohua Kevin Zhou, Gaurav Aggarwal, Rama Chellappa, David Jacobs 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | BoostMotion: Boosting a Discriminative Similarity Function for Motion EstimationabstractMotion estimation for applications where appearance undergoes complex changes is challenging due to lack of an appropriate similarity function. In this paper, we propose to learn a discriminative similarity function based on an annotated database that exemplifies the appearance variations. We invoke the LogitBoost algorithm to selectively combine weak learners into one strong similarity function. The weak learners based on local rectangle features are constructed as nonparametric 2D piecewise constant functions, using the feature responses from both images, to strengthen the modeling power and accommodate fast evaluation. Because the negatives possess a location parameter measuring their closeness to the positives, we present a locationsensitive cascade training procedure, which bootstraps negatives for later stages of the cascade from the regions closer to the positives. This allows viewing a large number of negatives and steering the training process to yield lower training and test errors. In experiments of estimating the motion for the endocardial wall of the left ventricle in echocardiography, we compare the learned similarity function with conventional ones and obtain improved performances. We also contrast the proposed method with a learning-based detection algorithm to demonstrate the importance of temporal information in motion estimation. Finally, we insert the learned similarity function into a simple contour tracking algorithm and find that it reduces drifting. Shaohua Kevin Zhou, Bogdan Georgescu, Dorin Comaniciu, Jie Shao 0007 |
CVPR (2) | 1 |
| 2006 | Image-Based Multiclass Boosting and Echocardiographic View ClassificationabstractWe tackle the problem of automatically classifying cardiac view for an echocardiographic sequence as a multiclass object detection. As a solution, we present an imagebased multiclass boosting procedure. In contrast with conventional approaches for multiple object detection that train multiple binary classifiers, one per object, we learn only one multiclass classifier using the LogitBoosting algorithm. To utilize the fact that, in the midst of boosting, one class is fully separated from the remaining classes, we propose to learn a tree structure that focuses on the remaining classes to improve learning efficiency. Further, we accommodate the large number of background images using a cascade of boosted multiclass classifiers, which is able to simultaneously detect and classify multiple objects while rejecting the background class quickly. Our experiments on echocardiographic view classification demonstrate promising performances of image-based multiclass boosting. Shaohua Kevin Zhou, Jin Hyeong Park, Bogdan Georgescu, Dorin Comaniciu, Costas Simopoulos, Joanne Otsuki |
CVPR (2) | 1 |
| 2006 | Example Based Non-rigid Shape Detection
Yefeng Zheng 0001, Xiang Sean Zhou, Bogdan Georgescu, Shaohua Kevin Zhou, Dorin Comaniciu |
ECCV (4) | 4 |
| 2006 | Integrated Detection, Tracking and Recognition for IR Video-Based Vehicle ClassificationabstractWe present an approach for vehicle classification in IR video sequences by integrating detection, tracking and recognition. The method has two steps. First, the moving target is automatically detected using a detection algorithm. Next, we perform simultaneous tracking and recognition using an appearance-model based particle filter. The tracking result is evaluated at each frame. Low confidence in tracking performance initiates a new cycle of detection, tracking and classification. We demonstrate the robustness of the proposed method using outdoor IR video sequences Xue Mei, Shaohua Kevin Zhou, Hao Wu 0014 |
ICASSP (5) | 2 |
| 2006 | Pairwise Active Appearance Model and Its Application to Echocardiography Tracking
Shaohua Kevin Zhou, Jie Shao 0007, Bogdan Georgescu, Dorin Comaniciu |
MICCAI (1) | 1 |
| 2006 | From Sample Similarity to Ensemble Similarity: Probabilistic Distance Measures in Reproducing Kernel Hilbert SpaceabstractThis paper addresses the problem of characterizing ensemble similarity from sample similarity in a principled manner. Using reproducing kernel as a characterization of sample similarity, we suggest a probabilistic distance measure in the reproducing kernel Hilbert space (RKHS) as the ensemble similarity. Assuming normality in the RKHS, we derive analytic expressions for probabilistic distance measures that are commonly used in many applications, such as Chernoff distance (or the Bhattacharyya distance as its special case), Kullback-Leibler divergence, etc. Since the reproducing kernel implicitly embeds a nonlinear mapping, our approach presents a new way to study these distances whose feasibility and efficiency is demonstrated using experiments with synthetic and real examples. Further, we extend the ensemble similarity to the reproducing kernel for ensemble and study the ensemble similarity for more general data representations. Shaohua Kevin Zhou, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Video Background Retrieval using Mosaic ImagesabstractContent-based video retrieval is one of the most active and exciting research areas in the field of multimedia technology. In this paper, we present an approach for video background retrieval using mosaic images and a support vector machine (SVM). The video is captured by a moving camera and a portion of the scene is visible at any time. The Kanade-Lucas-Tomasi (KLT) feature tracker is used to get the correspondences between consecutive images and the homography is calculated using these correspondences. We use the homography to construct the mosaic background image and a mixture of Gaussian (MoG) background subtraction algorithm to remove the moving objects in the scene. An SVM is then applied to classify the mosaic background image. The experimental results show the efficiency and effectiveness of the proposed approach. Xue Mei, Mahesh Ramachandran, Shaohua Kevin Zhou |
ICASSP (2) | 3 |
| 2005 | A method for converting a smiling face to a neutral face with applications to face recognitionabstractThe human face displays a variety of expressions, like smile, sorrow, surprise, etc. All these expressions constitute nonrigid motions of various features of the face. These expressions lead to a significant change in the appearance of a facial image which leads to a drop in the recognition accuracy of a face-recognition system trained with neutral faces. There are other factors like pose and illumination which also lead to performance drops. Researchers have proposed methods to tackle the effects of pose and illumination; however, there has been little work on how to tackle expressions. We attempt to address the issue of expression invariant face-recognition. We present preprocessing steps for converting a smiling face to a neutral face. We expect that this would in turn make the vector in the feature space to be closer to the correct vector in the gallery, in an appearance-based face recognition. This conjecture is supported by our recognition results which demonstrate that the accuracy goes up if we include the expression-normalization block. Mahesh Ramachandran, Shaohua Kevin Zhou, Divya Jhalani, Rama Chellappa |
ICASSP (2) | 2 |
| 2005 | Tracking Algorithm Using Background-Foreground Motion Models and Multiple CuesabstractWe present a stochastic tracking algorithm for surveillance videos where targets are dim and of low resolution. Our tracker is mainly based on the particle filter algorithm. Two important novel features of the tracker include: a motion model consisting of both background and foreground motion parameters; multiple cues are adaptively integrated in a system observation model when estimating the likelihood functions. Based on these features, the accuracy and robustness of the tracker has been improved, which is very important for surveillance problems. We present the results of applying the proposed algorithm to many videos. Jie Shao 0007, Shaohua Kevin Zhou, Rama Chellappa |
ICASSP (2) | 2 |
| 2005 | Appearance Modeling Under Geometric ContextabstractWe propose a unified framework based on a general definition of geometric transform (GeT) for modeling appearance. GeT represents the appearance by applying designed functionals over certain geometric sets. We show that image warping, Radon transform, trace transform, etc. are special cases of our definition. Moreover, three different types of GeTs are designed to handle deformation, articulation and occlusion and applied to fingerprinting the appearance inside a contour. They include the contour-driven GeT, the feature curve based GeT and selecting functionals to model the appearance inside the convex hull of the contour. A multi-resolution representation that combines both shape and appearance information is also proposed. We apply our approach to image synthesis and object recognition. The proposed approach produces promising results when applied to fingerprinting the appearance of human and body parts despite the challenges due to articulated motion and deformations. Jian Li 0022, Shaohua Kevin Zhou, Rama Chellappa |
ICCV | 2 |
| 2005 | Image Based Regression Using Boosting MethodabstractWe present a general algorithm of image based regression that is applicable to many vision problems. The proposed regressor that targets a multiple-output setting is learned using boosting method. We formulate a multiple-output regression problem in such a way that overfitting is decreased and an analytic solution is admitted. Because we represent the image via a set of highly redundant Haar-like features that can be evaluated very quickly and select relevant features through boosting to absorb the knowledge of the training data, during testing we require no storage of the training data and evaluate the regression function almost in no time. We also propose an efficient training algorithm that breaks the computational bottleneck in the greedy feature selection process. We validate the efficiency of the proposed regressor using three challenging tasks of age estimation, tumor detection, and endocardial wall localization and achieve the best performance with a dramatic speed, e.g., more than 1000 times faster than conventional data-driven techniques such as support vector regressor in the experiment of endocardial wall localization. Shaohua Kevin Zhou, Bogdan Georgescu, Xiang Sean Zhou, Dorin Comaniciu |
ICCV | 1 |
| 2004 | Probabilistic Identity Characterization for Face Recognition
Shaohua Kevin Zhou, Rama Chellappa |
CVPR (2) | 1 |
| 2004 | Characterization of Human Faces under Illumination Variations Using Rank, Integrability, and Symmetry Constraints
Shaohua Kevin Zhou, Rama Chellappa, David Jacobs 0001 |
ECCV (1) | 1 |
| 2004 | Probabilistic face recognition from compressed imageryabstractThe effects of image and video compression on face recognition in the still-to-video setting are studied in this paper. We use the probabilistic framework described in (S. Zhou et al., Computer Vision and Image Understanding, vol.91, p.214-245, 2003), which solves tracking and recognition problems simultaneously via sequential importance sampling (SIS). To account for the illumination and pose variations in test sequences, intrapersonal space (IPS) is constructed from exemplary views and used to calculate the likelihood density. Both the gallery images and probe videos are compressed and several experiments are run to study their effects on the recognition rate. Some useful conclusions are drawn from the analysis of the experimental results, which are helpful for future research on the interaction between recognition and compression. Meanwhile, the experiments also demonstrate the robustness of the proposed methods. Jian Li 0022, Shaohua Kevin Zhou |
ICASSP (5) | 2 |
| 2004 | Appearance-based tracking and recognition using the 3D trilinear tensorabstractThe paper presents an appearance-based adaptive algorithm for simultaneous tracking and recognition by generalizing the transformation model to 3D perspective transformation. A trilinear tensor operator is used to represent the 3D geometrical structure. The tensor is estimated by predicting the corresponding points using the existing affine-transformation based algorithm. The estimated tensor is used to synthesize novel views to update the appearance templates. Some experimental results using airborne video are presented. Jie Shao 0007, Shaohua Kevin Zhou, Rama Chellappa |
ICASSP (3) | 2 |
| 2004 | Robust two-camera tracking using homographyabstractThe paper introduces a two view tracking method which uses the homography relation between the two views to handle occlusions. An adaptive appearance-based model is incorporated in a particle filter to realize robust visual tracking. Occlusion is detected using robust statistics. When there is occlusion in one view, the homography from this view to other views is estimated from previous tracking results and used to infer the correct transformation for the occluded view. Experimental results show the robustness of the two view tracker. Zhanfeng Yue, Shaohua Kevin Zhou, Rama Chellappa |
ICASSP (3) | 2 |
| 2004 | Simultaneous background and foreground modeling for tracking in surveillance videoabstractWe present a stochastic tracking algorithm for surveillance video where targets are dim and at low resolution. The algorithm builds motion models for both background and foreground by integrating motion and intensity information. Some other merits of the algorithm include adaptive selection of feature points for scene description and defining proper cost functions for displacement estimation. The experimental results show tracking robustness and precision in a challenging video sequences. Jie Shao 0007, Shaohua Kevin Zhou, Rama Chellappa |
ICIP | 2 |
| 2004 | Visual tracking and recognition using appearance-adaptive models in particle filtersabstractWe present an approach that incorporates appearance-adaptive models in a particle filter to realize robust visual tracking and recognition algorithms. Tracking needs modeling interframe motion and appearance changes, whereas recognition needs modeling appearance changes between frames and gallery images. In conventional tracking algorithms, the appearance model is either fixed or rapidly changing, and the motion model is simply a random walk with fixed noise variance. Also, the number of particles is typically fixed. All these factors make the visual tracker unstable. To stabilize the tracker, we propose the following modifications: an observation model arising from an adaptive appearance model, an adaptive velocity motion model with adaptive noise variance, and an adaptive number of particles. The adaptive-velocity model is derived using a first-order linear predictor based on the appearance difference between the incoming observation and the previous particle configuration. Occlusion analysis is implemented using robust statistics. Experimental results on tracking visual objects in long outdoor and indoor video sequences demonstrate the effectiveness and robustness of our tracking algorithm. We then perform simultaneous tracking and recognition by embedding them in a particle filter. For recognition purposes, we model the appearance changes between frames and gallery images by constructing the intra- and extrapersonal spaces. Accurate recognition is achieved when confronted by pose and view variations. Shaohua Kevin Zhou, Rama Chellappa, Baback Moghaddam |
IEEE Trans. Image Process. | 1 |
| 2003 | A comparison of subspace analysis for face recognitionabstractWe report the results of a comparative study on subspace analysis methods for face recognition. In particular, we have studied four different subspace representations and their 'kernelized' versions if available. They include both unsupervised methods such as principal component analysis (PCA) and independent component analysis (ICA), and supervised methods such as Fisher discriminant analysis (FDA) and probabilistic PCA (PPCA) used in a discriminative manner. The 'kernelized' versions of these methods provide subspaces of high-dimensional feature spaces induced by non-linear mappings. To test the effectiveness of these subspace representations, we experiment on two databases with three typical variations of face images, i.e, pose, illumination and facial expression changes. The comparison of these methods applied to different variations in face images offers a comprehensive view of all the subspace methods currently used in face recognition. Jian Li 0022, Shaohua Kevin Zhou, Chandra Shekhar 0002 |
ICASSP (3) | 2 |
| 2003 | Simultaneous tracking and recognition of human faces from videoabstractThe paper investigates the interaction between tracking and recognition of human faces from video under a framework proposed earlier (Shaohua Zhou et al., Proc. 5th Int. Conf. on Face and Gesture Recog., 2002; Shaohua Zhou and Chellappa, R., Proc. European Conf. on Computer Vision, 2002), where a time series model is used to resolve the uncertainties in both tracking and recognition. However, our earlier efforts employed only a simple likelihood measurement in the form of a Laplacian density to deal with appearance changes between frames and between the observation and gallery images, yielding poor accuracies in both tracking and recognition when confronted by pose and illumination variations. The interaction between tracking and recognition was not well understood. We address the interdependence between tracking and recognition using a series of experiments and quantify the interacting nature of tracking and recognition. Shaohua Kevin Zhou, Rama Chellappa |
ICASSP (3) | 1 |
| 2003 | A comparison of subspace analysis for face recognitionabstractWe report the results of a comparative study on subspace analysis methods for face recognition. In particular, we have studied four different subspace representations and their 'kernelized' versions if available. They include both unsupervised methods such as principal component analysis (PCA) and independent component analysis (ICA), and supervised methods such as Fisher discriminant analysis (FDA) and probabilistic PCA (PPCA) used in a discriminative manner. The 'kernelized' versions of these methods provide subspaces of high-dimensional feature spaces induced by non-linear mappings. To test the effectiveness of these subspace representations, we experiment on two databases with three typical variations of face images, i.e., pose, illumination and facial expression changes. The comparison of these methods applied to different variations in face images offers a comprehensive view of all the subspace methods currently used in face recognition. Jian Li 0022, Shaohua Kevin Zhou, Chandra Shekhar 0002 |
ICME | 2 |
| 2003 | Simultaneous tracking and recognition of human faces from videoabstractThis paper investigates the interaction between tracking and recognition of human faces from video under the framework proposed earlier [(Shaohua Zhou et al., 2002), (Shaohua Zhou and R. Chellapa, 2002)], where a time series model is used to resolve the uncertainties in both tracking and recognition. However, our earlier efforts employed only a simple likelihood measurement in the form of a Laplacian density to deal with appearance changes between the frames and between the observation and the gallery images, yielding poor accuracies in both tracking and recognition when confronted by pose and illumination variations. The interaction between tracking and recognition was not well understood. We address the interdependence between tracking and recognition using a series of experiments and quantify the interacting nature of tracking and recognition. Shaohua Kevin Zhou, Rama Chellappa |
ICME | 1 |
| 2003 | Adaptive visual tracking and recognition using particle filtersabstractThis paper presents an improved method for simultaneous tracking and recognition of human faces from video, where a time series model is used to resolve the uncertainties in tracking and recognition. The improvements mainly arise from three aspects: (i) modeling the inter-frame appearance changes within the video sequence using an adaptive appearance model and an adaptive-velocity motion model; (ii) modeling the appearance changes between the video frames and gallery images by constructing intra- and extra-personal spaces; and (iii) utilization of the fact that the gallery images are in frontal views. By embedding them in a particle filter, we are able to achieve a stabilized tracker and an accurate recognizer when confronted by pose and illumination variations. Shaohua Kevin Zhou, Rama Chellappa, Baback Moghaddam |
ICME | 1 |
| 2003 | Probabilistic recognition of human faces from video
Shaohua Kevin Zhou, Volker Krüger, Rama Chellappa |
Comput. Vis. Image Underst. | 1 |
| 2002 | Exemplar-Based Face Recognition from Video
Volker Krüger, Shaohua Kevin Zhou |
ECCV (4) | 2 |
| 2002 | Probabilistic Human Recognition from Video
Shaohua Kevin Zhou, Rama Chellappa |
ECCV (3) | 1 |
| 2002 | Synthesis of visual textures and perceptual asymmetry of visual texture perceptionabstractA new simple scheme is proposed to synthesize visual textures with positional jitter and/or orientation randomness that find extensive use in controlled psychophysical experiments for modelling (human) preattentive vision. Spectral properties of synthesized textures are then analyzed with the goal of creating a framework for computing signal-to-noise ratio from the spectral that offers the potential to interpret quantitatively certain perceptual phenomena, such as perceptual asymmetry. Shaohua Kevin Zhou, Y. V. Venkatesh |
ICARCV | 1 |
| 2002 | Bayesian methods for face recognition from videoabstractFace recognition (FR) from video necessitates simultaneously solving two asks, recognition and tracking. To accommodate the video, a time series state space model is introduced in a Bayesian approach. Given this model, the goal reduces to estimating the posterior distribution of the state vector given the observations up to the present. The Sequential Importance Sampling (SIS) technique is invoked to generate a numerical solution to this model. However, the ultimate goal is to estimate the posterior distribution of the identity of humans for recognition purposes. Presented here are two methods to approximate the above distribution under different experimental scenarios. Rama Chellappa, Shaohua Kevin Zhou, Baoxin Li |
ICASSP | 2 |
| 2002 | Probabilistic recognition of human faces from videoabstractMost present face recognition approaches recognize faces based on still images. We present a novel approach to recognize faces in video. In that scenario, the face gallery may consist of still images or may be derived from a videos. For evidence integration we use classical Bayesian propagation over time and compute the posterior distribution using sequential importance sampling. The probabilistic approach allows us to handle uncertainties in a systematic manner. Experimental results using videos collected by NIST/USF and CMU illustrate the effectiveness of this approach in both still-to-video and video-to-video scenarios with appropriate model choices. Shaohua Kevin Zhou, Volker Krüger, Rama Chellappa |
ICIP (1) | 1 |