Yan Wang 0033

dblp:59/2227-33 · DBLP profile ↗
← Back
69ranked-venue papers
11as first author
46since 2021 · last 2026
0000-0002-7865-9580ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 24 · 3 first-author · 15 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Physics-Aware Accelerated Unrolling Model for Sparse-View CT Reconstruction
abstract
Deep unrolling models (DUMs) have shown great poten-tial in sparse-view CT reconstruction by combining itera-tive optimization and deep learning. However, most DUMsinsufficiently account for physical degradation from sparse-view imaging, leading to slow convergence and persistentartifacts. To address this, we propose PAUM, a Physics-Aware Accelerated Unrolling Model explicitly incorporatingCT imaging physics into the iterative reconstruction. PAUMfirst introduces a Dual-Domain Physics-Aware Extrapolation(DDPE) module. By modeling dual-domain degradations, itperforms row-wise extrapolation in the sinogram domain toimprove missing view recovery, and pixel-wise extrapolationin the image domain to address spatially variant degradationfrom incomplete backprojection. This physics-aware extrap-olation aligns optimization dynamics with underlying physi-cal imaging degradation, significantly enhances structural up-dates, thereby accelerating convergence. Subsequently, wedevelop a lightweight Block-Attention Deformable Regu-larization Network (BDRN), leveraging deformable convo-lutions and block-wise attention to model spatially variantand structured artifact physical characteristics. This enablesspatially adaptive regularization on extrapolated results, ef-fectively improving reconstruction quality. Extensive exper-iments demonstrate PAUM achieves over 1dB improvementcompared to SOTA methods, while reducing iteration countby 50%.
Shaojie Guo, Yingying Fang, Junkang Zhang, Yan Wang 0033
AAAI4
2026 Scribble-Supervised Multi-Organ Segmentation via Epistemic-Driven Hardness-Adaptive Focusing
abstract
Scribble supervision reduces annotation costs in multi-organ segmentation. However, its sparsity results in insufficient supervision for most regions and inadequate feature learning in hard areas (e.g., organ boundaries). These hard areas cause model confirmation bias and high epistemic uncertainty, which existing methods fail to address. To overcome these core challenges, we propose an epistemic-driven hardness-adaptive focusing framework. This framework establishes a self-improving loop: quantified epistemic uncertainty guides hard sample generation, while hard sample learning and feature alignment jointly reduce epistemic uncertainty. Specifically, we first propose a phase-adaptive hardness-aware loss function to quantify epistemic uncertainty and generate dynamic hardness maps during training. Based on these maps, we employ a distribution-divergence-aware copy-paste operation to create hard samples, which are progressively incorporated into learning to reduce epistemic uncertainty. Furthermore, we introduce feature distribution alignment to mitigate bias and epistemic uncertainty by aligning organ-specific hard regions with global features. Extensive experiments on multi-organ CT and ultrasound datasets demonstrate the competitiveness and effectiveness of our method. The framework's generalizability and robustness are further validated under cross-dataset and noise-corrupted scenarios. This work offers a practical solution for clinical applications where annotation efficiency is critical.
Xiaoxiang Han 0001, Yiman Liu, Jiang Shang, Haobo Chen, Xiaohong Liu 0001, Zhen Qiu 0001, Yan Wang 0033, Qi Zhang 0003
IEEE Trans. Medical Imaging7
2025 MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image Segmentation
abstract
Medical Hyperspectral Imaging (MHSI) offers potential for computational pathology and precision medicine. However, existing CNN and Transformer struggle to balance segmentation accuracy and speed due to high spatial-spectral dimensionality. In this study, we leverage Mamba’s global context modeling to propose a dual-stream architecture for joint spatial-spectral feature extraction. To address the limitation of Mamba’s unidirectional aggregation, we introduce a recurrent spectral sequence representation to capture low-redundancy global spectral features. Experiments on a public Multi-Dimensional Choledoch dataset and a private Cervical Cancer dataset show that our method outperforms state-of-the-art approaches in segmentation accuracy while minimizing resource usage and achieving the fastest inference speed. Our code will be available at https://github.com/DeepMed-Lab-ECNU/MDN.
Shijie Lin, Boxiang Yun, Wei Shen 0002, Qingli Li, Anqiang Yang, Yan Wang 0033
ICASSP6
2025 Self-Prompting Driven SAM2 for 3D Medical Image Segmentation
abstract
The latest advancement in large foundational model, SAM2, has demonstrated significant potential in 3D medical image segmentation due to their capability to effectively segment video streams. However, its application in medical image segmentation presents challenges, requiring extensive training on medical images or high-quality prompts provided by experts to achieve optimal performance. To address the aforementioned limitations, we propose SAM2-SP, which adopts Low-Rank Adaption for parameter-efficient fine-tuning and introduces a novel dynamic self-prompting strategy that generates most confident prompt templates from voxel features, enabling SAM2 to achieve domain adaptation in medical image segmentation without reliance on expert-level prompts. Extensive experiments show that SAM2-SP achieves state-of-the-art performance on the public Synapse dataset and the private EDC dataset, and even outperforms the compared task-specific segmentation approaches, the vanilla SAM and other SAM-based approaches.
Sheng Wei 0004, Song Qiu, Mei Zhou, He Zhang 0023, Yan Wang 0033, Qingli Li
ICASSP5
2025 Advancing Stain Transfer for Multi-Biomarkers: A Human Annotation-Free Method Based on Auxiliary Task Supervision
abstract
Histopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemical (IHC) staining. Though IHC provides more crucial molecular information for diagnosis, it is more costly than H&E staining. Stain transfer technology seeks to efficiently generate virtual IHC images from H&E images. While current deep learning-based methods have made progress, they still struggle to maintain pathological and structural consistency across biomarkers without pixel-level aligned reference. To address the problem, we propose an Auxiliary Task supervision-based Stain Transfer method for multi-biomarkers (ATST-Net), which pioneeringly employs human annotation-free masks as ground truth (GT). ATST-Net ensures pathological consistency, structural preservation and style transfer. It automatically annotates H&E masks in a cost-effective manner by utilizing consecutive IHC sections. Multiple auxiliary tasks provide diverse supervisory information on the location and intensity of biomarker expression, ensuring model accuracy and interpretability. We design a pretrained model-based generator to extract deep feature in H&E images, improving generalization performance. Extensive experiments demonstrate the effectiveness of ATST-Net's components. Compared to existing methods, ATST-Net achieves state-of-the-art (SOTA) accuracy on datasets with multiple biomarkers and intensity levels, while also reflecting high practical value. Code is available at https://github.com/SikangSHU/ATST-Net.
Haofei Song, Yingjiao Deng, Jiansheng Wang, Yan Wang 0033, Qingli Li
IJCAI5
2025 Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation
Boxiang Yun, Qingli Li, Yan Wang 0033
MICCAI (6)4
2025 Deep Association Multimodal Learning for Zero-Shot Spatial Transcriptomics Prediction
Yijing Zhou, Yadong Lu, Qingli Li, Xinxing Li, Yan Wang 0033
MICCAI (6)5
2025 Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification
abstract
Accurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume. Existing WSI classification methods encounter critical limitations: (1) inadequate patch selection strategies that either compromise performance or fail to sufficiently reduce computational demands, and (2) the loss of global histological context resulting from patch-level processing approaches. To address these challenges, we propose an Efficient multimodal framework for Patient-level placental disease Diagnosis, named EmmPD. Our approach introduces a two-stage patch selection module that combines parameter-free and learnable compression strategies, optimally balancing computational efficiency with critical feature preservation. Additionally, we develop a hybrid multimodal fusion module that leverages adaptive graph learning to enhance pathological feature representation and incorporates textual medical reports to enrich global contextual understanding. Extensive experiments conducted on both a self-constructed patient-level Placental dataset and two public datasets demonstrating that our method achieves state-of-the-art diagnostic performance. The code is available at https://github.com/ECNU-MultiDimLab/EmmPD.
Zixuan Gao, Siyuan Yang 0001, Shulin Peng, Xiang Tao, Yan Wang 0033, Qingli Li
ACM Multimedia8
2025 WFANet-DDCL: Wavelet-Based Frequency Attention Network and Dual Domain Consistency Learning for 7T MRI Synthesis From 3T MRI
abstract
Ultra-high field magnetic resonance imaging (MRI), such as 7-Tesla (7T) MRI, provides significantly enhanced tissue contrast and anatomical details compared to 3T MRI. However, 7T MRI scanners are more costly and less accessible in clinical settings than 3T scanners. In this paper, we propose a wavelet-based frequency attention network (WFANet) and a semi-supervised method named dual domain consistency learning (DDCL), and combine them to form a WFANet-DDCL framework for 7T MRI synthesis. WFANet leverages the frequency sensitivity of the proposed wavelet-based frequency attention encoder (WFAE) along with the large receptive field of dilated convolution. WFAE is proposed as an independent module to capture multi-scale frequency attention via the proposed wavelet-based frequency attention (WFA) mechanism. WFAE can be integrated into any backbone network as a plug-and-play component and improve network performance. To tackle the challenge of limited paired data for network training, DDCL is proposed to take advantage of both paired and unpaired data. Frequency domain perturbation is proposed and combined with Gaussian noise to regularize the supervised learning process in dual domains, better avoiding overfitting. Extensive experimental results demonstrate that WFANet-DDCL can achieve comparable performance to state-of-the-art supervised methods even using 66% of all paired data.
Song Qiu, Mei Zhou, Weijie Le, Qingli Li, Yan Wang 0033
IEEE Trans. Circuits Syst. Video Technol.6
2025 S4R: Separated Self-Supervised Spectral Regression for Hyperspectral Histopathology Image Diagnosis
abstract
Hyperspectral images (HSIs) offer great potential for computational pathology. But, limited by the lack of adequate annotated data and the high spectral redundancy of HSIs, traditional supervised learning techniques are usually bottlenecked. To exploit the structural properties of HSIs and learn representations with good transferability, we propose Separated Self-Supervised Spectral Regression (S4R). Concretely, we find one spectral band can be represented by a linear combination of the remaining bands. Regressing the distribution of the linear coefficients learns the inherent properties of HSIs and pathological information about the tissue. Besides, reconstructing the missing band, especially the tissue boundaries makes the model learn pathology details that are critical to downstream tasks. Coupling these two pretext tasks makes the self-supervised model understand spectral structures of HSIs w.r.t. pathological semantics and spatial micro details. Furthermore, we design two brand-new architectures to avoid the interference of extraneous signal based on S4R: S4R-CLS and S4R-SEG for HSI classification and segmentation, respectively. Two downstream tasks are incorporated into a unified framework, which first encodes different bands from HSIs via a depthwise separable encoder, and then selectively aggregates band features to generate final predictions. In S4R-SEG, we propose to pick the best matching bands with the guidance of a classification paradigm. Extensive experiments show S4R performs much better than competitors on both tasks. Theoretical analysis and clinical discussion also indicate the great potential for further medical applications. The code and pre-trained checkpoints are available at https://github.com/DeepMed-Lab-ECNU/S4R.
Yan Wang 0033, Xingran Xie, Benyan Zhang, Chunhua Zhou, Duowu Zou, Le Lu 0001, Qingli Li
IEEE Trans. Image Process.1
2025 Clinical Stage Prompt Induced Multi-Modal Prognosis
abstract
Histology analysis of the tumor micro-environment integrated with genomic assays is widely regarded as the cornerstone for cancer analysis and survival prediction. This paper jointly incorporates genomics and Whole Slide Images (WSIs), and focuses on addressing the primary challenges involved in multi-modality prognosis analysis: 1) the high-order relevance is difficult to be modeled from dimensional imbalanced gigapixel WSIs and tens of thousands of genetic sequences, and 2) the lack of medical expertise and clinical knowledge hampers the effectiveness of prognosis-oriented multi-modal fusion. Due to the nature of the prognosis task, statistical priors and clinical knowledge are essential factors to provide the likelihood of survival over time, which, however, has been under-studied. To this end, we propose a prognosis-oriented image-omics fusion framework, dubbed Clinical Stage Prompt induced Multimodal Prognosis (CiMP). Concretely, we leverage the capabilities of the advanced LLM to generate descriptions derived from structured clinical records and utilize the generated clinical staging prompts to inquire critical prognosis-related information from each modality intentionally. In addition, we propose a Group Multi-Head Self-Attention module to capture structured group-specific features within cohorts of genomic data. Experimental results on five TCGA datasets show the superiority of our proposed method, achieving state-of-the-art performance compared to previous multi-modal prognostic models. Furthermore, the clinical interpretability and discussion also highlight the immense potential for further medical applications. Our code will be released at https://github.com/DeepMed-Lab-ECNU/CiMP/.
Xingran Xie, Qingli Li, Xinxing Li, Yan Wang 0033
IEEE Trans. Medical Imaging5
2025 Debiasing Medical Knowledge for Prompting Universal Model in CT Image Segmentation
abstract
With the assistance of large language models, which offer universal medical prior knowledge via text prompts, state-of-the-art Universal Models (UM) have demonstrated considerable potential in the field of medical image segmentation. Semantically detailed text prompts, on the one hand, indicate comprehensive knowledge; on the other hand, they bring biases that may not be applicable to specific cases involving heterogeneous organs or rare cancers. To this end, we propose a Debiased Universal Model (DUM) to consider instance-level context information and remove knowledge biases in text prompts from the causal perspective. We are the first to discover and mitigate the bias introduced by universal knowledge. Specifically, we propose to extract organ-level text prompts via language models and instance-level context prompts from the visual features of each image. We aim to highlight more on factual instance-level information and mitigate organ-level's knowledge bias. This process can be derived and theoretically supported by a causal graph, and instantiated by designing a standard UM (SUM) and a biased UM. The debiased output is finally obtained by subtracting the likelihood distribution output by biased UM from that of the SUM. Experiments on three large-scale multi-center external datasets and MSD internal tumor datasets show that our method enhances the model's generalization ability in handling diverse medical scenarios and reducing the potential biases, even with an improvement of 4.16% compared with popular universal model on the AbdomenAtlas dataset, showing the strong generalizability. The code is publicly available at https://github.com/DeepMed-Lab-ECNU/DUM.
Boxiang Yun, Shitian Zhao, Qingli Li, Alex Chichung Kot, Yan Wang 0033
IEEE Trans. Medical Imaging5
2024 Partial Label Learning with a Partner
abstract
In partial label learning (PLL), each instance is associated with a set of candidate labels among which only one is ground-truth. The majority of the existing works focuses on constructing robust classifiers to estimate the labeling confidence of candidate labels in order to identify the correct one. However, these methods usually struggle to rectify mislabeled samples. To help existing PLL methods identify and rectify mislabeled samples, in this paper, we introduce a novel partner classifier and propose a novel ``mutual supervision'' paradigm. Specifically, we instantiate the partner classifier predicated on the implicit fact that non-candidate labels of a sample should not be assigned to it, which is inherently accurate and has not been fully investigated in PLL. Furthermore, a novel collaborative term is formulated to link the base classifier and the partner one. During each stage of mutual supervision, both classifiers will blur each other's predictions through a blurring mechanism to prevent overconfidence in a specific label. Extensive experiments demonstrate that the performance and disambiguation ability of several well-established stand-alone and deep-learning based PLL approaches can be significantly improved by coupling with this learning paradigm.
Chongjie Si, Zekun Jiang, Yan Wang 0033, Xiaokang Yang 0001, Wei Shen 0002
AAAI4
2024 A Semi-Supervised Approach with Error Reflection for Echocardiography Segmentation
abstract
Segmenting internal structure from echocardiography is essential for the diagnosis and treatment of various heart diseases. Semi-supervised learning shows its ability in alleviating annotations scarcity. While existing semi-supervised methods have been successful in image segmentation across various medical imaging modalities, few have attempted to design methods specifically addressing the challenges posed by the poor contrast, blurred edge details and noise of echocardiography. These characteristics pose challenges to the generation of high-quality pseudo-labels in semi-supervised segmentation based on Mean Teacher. Inspired by human reflection on erroneous practices, we devise an error reflection strategy for echocardiography semi-supervised segmentation architecture. The process triggers the model to reflect on inaccuracies in unlabeled image segmentation, thereby enhancing the robustness of pseudo-label generation. Specifically, the strategy is divided into two steps. The first step is called reconstruction reflection. The network is tasked with reconstructing authentic proxy images from the semantic masks of unlabeled images and their auxiliary sketches, while maximizing the structural similarity between the original inputs and the proxies. The second step is called guidance correction. Reconstruction error maps decouple unreliable segmentation regions. Then, reliable data that are more likely to occur near high-density areas are leveraged to guide the optimization of unreliable data potentially located around decision boundaries. Additionally, we introduce an effective data augmentation strategy, termed as multi-scale mixing up strategy, to minimize the empirical distribution gap between labeled and unlabeled images and perceive diverse scales of cardiac anatomical structures. Extensive experiments on a public echocardiography dataset CAMUS, and a private clinical echocardiography dataset demonstrate the competitiveness of the proposed method.
Xiaoxiang Han 0001, Yiman Liu, Jiang Shang, Qingli Li, Menghan Hu, Qi Zhang 0003, Yan Wang 0033
BIBM9
2024 Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
abstract
The Segment Anything Model (SAM) has garnered significant attention for its versatile segmentation abilities and intuitive prompt-based interface. However, its application in medical imaging presents challenges, requiring either substantial training costs and extensive medical datasets for full model fine-tuning or high-quality prompts for optimal performance. This paper introduces H-SAM: a prompt-free adaptation of SAM tailored for efficient fine-tuning of medical images via a two-stage hierarchical decoding procedure. In the initial stage, H-SAM employs SAM's original decoder to generate a prior probabilistic mask, guiding a more intricate decoding process in the second stage. Specifically, we propose two key designs: 1) A class-balanced, mask-guided self-attention mechanism addressing the unbalanced label distribution, enhancing image embedding; 2) A learnable mask cross-attention mechanism spatially modulating the interplay among different image regions based on the prior mask. Moreover, the inclusion of a hierarchical pixel decoder in H-SAM enhances its proficiency in capturing fine-grained and localized details. This approach enables SAM to effectively integrate learned medical priors, facilitating enhanced adaptation for medical image segmentation with limited samples. Our H-SAM demonstrates a 4.78% improvement in average Dice compared to existing prompt-free SAM variants for multi-organ segmentation using only 10% of 2D slices. Notably, without using any unlabeled data, H-SAM even outperforms state-of-the-art semisupervised models relying on extensive unlabeled training data across various medical datasets. Our code is available at https://github.com/Cccccczh404/H-SAM.
Zhiheng Cheng, Qingyue Wei, Hongru Zhu, Yan Wang 0033, Liangqiong Qu, Wei Shao 0008, Yuyin Zhou
CVPR4
2024 AdaRevD: Adaptive Patch Exiting Reversible Decoder Pushes the Limit of Image Deblurring
abstract
Despite the recent progress in enhancing the efficacy of image deblurring, the limited decoding capability constrains the upper limit of State-Of- The-Art (SOTA) methods. This paper proposes a pioneering work, Adaptive Patch Ex-iting Reversible Decoder (AdaRevD), to explore their in-sufficient decoding capability. By inheriting the weights of the well-trained encoder, we refactor a reversible de-coder which scales up the single-decoder training to multi-decoder training while remaining GPU memory-friendly. Meanwhile, we show that our reversible structure gradually disentangles high-level degradation degree and low-level blur pattern (residual of the blur image and its sharp counterpart) from compact degradation representation. Besides, due to the spatially-variant motion blur kernels, different blur patches have various deblurring difficulties. We further introduce a classifier to learn the degradation degree of image patches, enabling them to exit at different sub-decoders for speedup. Experiments show that our AdaRevD pushes the limit of image deblurring, e.g., achieving 34.60 dB in PSNR on GoPro dataset.
Xintian Mao, Qingli Li, Yan Wang 0033
CVPR3
2024 Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models
abstract
While Multi-modal Language Models (MLMs) demonstrate impressive multimodal ability, they still struggle on providing factual and precise responses for tasks like visual question answering (VQA). In this paper, we address this challenge from the perspective of contextual information. We propose Causal Context Generation, Causal-CoG, which is a prompting strategy that engages contextual information to enhance precise VQA during inference. Specifically, we prompt MLMs to generate contexts, i.e, text description of an image, and engage the generated contexts for question answering. Moreover, we investigate the ad-vantage of contexts on VQA from a causality perspective, introducing causality filtering to select samples for which contextual information is helpful. To show the effectiveness of Causal-CoG, we run extensive experiments on 10 multimodal benchmarks and show consistent improvements, e.g., +6.30% on POPE, +13.69% on Vizwiz and +6.43% on VQAv2 compared to direct decoding, surpassing existing methods. We hope Casual-CoG inspires explorations of context knowledge in multimodal models, and serves as a plug-and-play strategy for MLM decoding.11Code is released zhaoshitian/Causal-CoG
Shitian Zhao, Zhuowan Li, Yadong Lu, Alan L. Yuille, Yan Wang 0033
CVPR5
2024 Spatially-Variant Degradation Model for Dataset-Free Super-Resolution
Shaojie Guo, Haofei Song, Qingli Li, Yan Wang 0033
ECCV (25)4
2024 Medical Image Classification Attack Based on Texture Manipulation
Yunrui Gu, Cong Kong, Zhao-Xia Yin, Yan Wang 0033, Qingli Li
ICPR (12)4
2024 Multi-stage Multi-granularity Focus-Tuned Learning Paradigm for Medical HSI Segmentation
Haichuan Dong, Runjie Zhou, Boxiang Yun, Benyan Zhang, Qingli Li, Yan Wang 0033
MICCAI (8)7
2024 EndoGSLAM: Real-Time Dense Reconstruction and Tracking in Endoscopic Surgeries Using Gaussian Splatting
Kailing Wang, Chen Yang 0023, Yuehao Wang, Sikuang Li, Yan Wang 0033, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (6)5
2024 Prompting Whole Slide Image Based Genetic Biomarker Prediction
Boxiang Yun, Xingran Xie, Qingli Li, Xinxing Li, Yan Wang 0033
MICCAI (4)6
2024 LoFormer: Local Frequency Transformer for Image Deblurring
Xintian Mao, Jiansheng Wang, Xingran Xie, Qingli Li, Yan Wang 0033
ACM Multimedia5
2024 GeNSeg-Net: A General Segmentation Framework for Any Nucleus in Immunohistochemistry Images
Haofei Song, Jiansheng Wang, Yan Wang 0033, Qingli Li
ACM Multimedia5
2024 TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers
abstract
Medical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively.
Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou
Medical Image Anal.10
2024 SpecTr: Spectral Transformer for Microscopic Hyperspectral Pathology Image Segmentation
abstract
Hyperspectral imaging (HSI) unlocks the huge potential to a wide variety of applications relying on high-precision pathology image segmentation, such as computational pathology. It can acquire biochemical properties even invisible to naked eyes from histological specimens. Since 1) spectra contain discriminative and continuous patterns for differentiating tissues/cells, and 2) the discriminability of spectra relies on both fine-grained relations in the high-resolution spectrum and coarse relations in the low-resolution spectrum, the key to achieving high-precision hyperspectral pathology image segmentation is to felicitously model the intra- and inter-scale context especially for spectra. In this paper, we propose a spectral transformer (SpecTr) for hyperspectral pathology image segmentation, which first captures global context for intra-scale spectral features, and subsequently extract coarse and fine-grained discriminative spectral information from inter-scale features, respectively. To learn intra-scale spectral context, we propose a Spectral Attentive Module (SAM). Unlike the existing Transformer model that is designed for modalities such as natural images, our proposed SAM is efficient in capturing sparse and pivotal spectral context while avoiding the heterogeneous underlying distributions and noises of different bands. Besides, to reduce the computational complexity of the HSI segmentation model, we further propose a global-local attention module to effectively learn a condensed spectral feature. Experiments show that HSIs can become a more powerful image modality for understanding microscopic pathology images than RGB images, and the proposed SpecTr outperforms other competing methods for hyperspectral pathology image segmentation, with an improvement of 3% compared with the popular 3D-nnUNet and other transformer-based methods. Our code is available at https://github.com/DeepMed-Lab-ECNU/SpecTr.
Boxiang Yun, Bai Ying Lei, Jieneng Chen, Song Qiu, Wei Shen 0002, Qingli Li, Yan Wang 0033
IEEE Trans. Circuits Syst. Video Technol.8
2024 I³Net: Inter-Intra-Slice Interpolation Network for Medical Slice Synthesis
abstract
Medical imaging is limited by acquisition time and scanning equipment. CT and MR volumes, reconstructed with thicker slices, are anisotropic with high in-plane resolution and low through-plane resolution. We reveal an intriguing phenomenon that due to the mentioned nature of data, performing slice-wise interpolation from the axial view can yield greater benefits than performing super-resolution from other views. Based on this observation, we propose an Inter-Intra-slice Interpolation Network ( [Formula: see text]Net), which fully explores information from high in-plane resolution and compensates for low through-plane resolution. The through-plane branch supplements the limited information contained in low through-plane resolution from high in-plane resolution and enables continual and diverse feature learning. In-plane branch transforms features to the frequency domain and enforces an equal learning opportunity for all frequency bands in a global context learning paradigm. We further propose a cross-view block to take advantage of the information from all three views online. Extensive experiments on two public datasets demonstrate the effectiveness of [Formula: see text]Net, and noticeably outperforms state-of-the-art super-resolution, video frame interpolation and slice interpolation methods by a large margin. We achieve 43.90dB in PSNR, with at least 1.14dB improvement under the upscale factor of ×2 on MSD dataset with faster inference. Code is available at https://github.com/DeepMed-Lab-ECNU/Medical-Image-Reconstruction.
Haofei Song, Xintian Mao, Qingli Li, Yan Wang 0033
IEEE Trans. Medical Imaging5
2023 Intriguing Findings of Frequency Selection for Image Deblurring
abstract
Blur was naturally analyzed in the frequency domain, by estimating the latent sharp image and the blur kernel given a blurry image. Recent progress on image deblurring always designs end-to-end architectures and aims at learning the difference between blurry and sharp image pairs from pixel-level, which inevitably overlooks the importance of blur kernels. This paper reveals an intriguing phenomenon that simply applying ReLU operation on the frequency domain of a blur image followed by inverse Fourier transform, i.e., frequency selection, provides faithful information about the blur pattern (e.g., the blur direction and blur level, implicitly shows the kernel pattern). Based on this observation, we attempt to leverage kernel-level information for image deblurring networks by inserting Fourier transform, ReLU operation, and inverse Fourier transform to the standard ResBlock. 1 × 1 convolution is further added to let the network modulate flexible thresholds for frequency selection. We term our newly built block as Res FFT-ReLU Block, which takes advantages of both kernel-level and pixel-level features via learning frequency-spatial dual-domain representations. Extensive experiments are conducted to acquire a thorough analysis on the insights of the method. Moreover, after plugging the proposed block into NAFNet, we can achieve 33.85 dB in PSNR on GoPro dataset. Our method noticeably improves backbone architectures without introducing many parameters, while maintaining low computational complexity. Code is available at https://github.com/DeepMed-Lab/DeepRFT-AAAI2023.
Xintian Mao, Fengze Liu, Qingli Li, Wei Shen 0002, Yan Wang 0033
AAAI6
2023 Bidirectional Copy-Paste for Semi-Supervised Medical Image Segmentation
abstract
In semi-supervised medical image segmentation, there exist empirical mismatch problems between labeled and un-labeled data distribution. The knowledge learned from the labeled data may be largely discarded if treating labeled and unlabeled data separately or in an inconsistent manner. We propose a straightforward method for alleviating the problem-copy-pasting labeled and unlabeled data bidirectionally, in a simple Mean Teacher architecture. The method encourages unlabeled data to learn comprehensive common semantics from the labeled data in both inward and outward directions. More importantly, the consistent learning procedure for labeled and unlabeled data can largely reduce the empirical distribution gap. In detail, we copy-paste a random crop from a labeled image (foreground) onto an unlabeled image (background) and an unlabeled image (foreground) onto a labeled image (background), respectively. The two mixed images are fed into a Student network and supervised by the mixed supervisory signals of pseudo-labels and ground-truth. We reveal that the simple mechanism of copy-pasting bidirectionally between labeled and unlabeled data is good enough and the experiments show solid gains (e.g., over 21% Dice improvement on ACDC dataset with 5% labeled data) compared with other state-of-the-arts on various semi-supervised medical image segmentation datasets. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/BCP.
Yunhao Bai, Duowen Chen 0002, Qingli Li, Wei Shen 0002, Yan Wang 0033
CVPR5
2023 MagicNet: Semi-Supervised Multi-Organ Segmentation via Magic-Cube Partition and Recovery
abstract
We propose a novel teacher-student model for semi-supervised multi-organ segmentation. In teacher-student model, data augmentation is usually adopted on unlabeled data to regularize the consistent training between teacher and student. We start from a key perspective that fixed relative locationsand variable sizes of different organs can provide distribution information where a multi-organ CT scan is drawn. Thus, we treat the prior anatomy as a strong tool to guide the data augmentation and reduce the mismatch between labeled and unlabeled images for semi-supervised learning. More specifically, we propose a data augmentation strategy based on partition-and-recovery N3cubes cross-and within-labeled and unlabeled images. Our strategy encourages unlabeled images to learn organ semantics in relative locations from the labeled images (cross-branch) and enhances the learning ability for small organs (within-branch). For within-branch, we further propose to refine the quality of pseudo labels by blending the learned representations from small cubes to incorporate local attributes. Our method is termed as MagicNet, since it treats the CT volume as a magic-cube and N3-cube partition-and-recovery process matches with the rule of playing a magic-cube. Extensive experiments on two public CT multi-organ datasets demonstrate the effectiveness of MagicNet, and noticeably outperforms state-of-the-art semi-supervised medical image segmentation approaches, with + 7% DSC improvement on MACT dataset with 10% labeled images. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/MagicNet.
Duowen Chen 0002, Yunhao Bai, Wei Shen 0002, Qingli Li, Lequan Yu, Yan Wang 0033
CVPR6
2023 Class Balanced Adaptive Pseudo Labeling for Federated Semi-Supervised Learning
abstract
This paper focuses on federated semi-supervised learning (FSSL), assuming that few clients have fully labeled data (labeled clients) and the training datasets in other clients are fully unlabeled (unlabeled clients). Existing methods attempt to deal with the challenges caused by not independent and identically distributed data (Non-IID) setting. Though methods such as sub-consensus models have been proposed, they usually adopt standard pseudo labeling or consistency regularization on unlabeled clients which can be easily influenced by imbalanced class distribution. Thus, problems in FSSL are still yet to be solved. To seek for a fundamental solution to this problem, we present Class Balanced Adaptive Pseudo Labeling (CBAFed), to study FSSL from the perspective of pseudo labeling. In CBAFed, the first key element is a fixed pseudo labeling strategy to handle the catastrophic forgetting problem, where we keep a fixed set by letting pass information of unlabeled data at the beginning of the unlabeled client training in each communication round. The second key element is that we design class balanced adaptive thresholds via considering the empirical distribution of all training data in local clients, to encourage a balanced training process. To make the model reach a better optimum, we further propose a residual weight connection in local supervised training and global model aggregation. Extensive experiments on five datasets demonstrate the superiority of CBAFed. Code will be available at https://github.com/minglllli/CBAFed.
Qingli Li, Yan Wang 0033
CVPR3
2023 A Generative Data Augmentation Trained by Low-quality Annotations for Cholangiocarcinoma Hyperspectral Image Segmentation
abstract
Microscopic hyperspectral imaging technology combined with deep learning method emerges medical field recently as a multiplexed imaging technology. With the semantic segmentation of hyperspectral histopathological image of pathological tissue, doctors can quickly locate suspicious areas, diagnose and arrange treatment accurately and rapidly, reducing the workload of them. Cholangiocarcinoma is a rare and devastating disease with few hyperspectral histopathological data. Moreover, achieving high-quality annotations of hyperspectral histopathological image is challenging and costs time for pathologists, so generally, rough labels are annotated, but directly using the low-quality labels will reduce the performance of segmentation networks. So how to fully utilize few high-quality annotations and dozens of low-quality labels to enhance the segmentation performance of cholangiocarcinoma hyperspectral image remains to be resolved. In this paper, we proposed a two-stage hyperspectral segmentation deep learning framework based on Labels-to-Photo translation and Swin-Spec Transformer(L2P-SST). In stage-I, the OASIS generative network and the Swin-Spec Transformer discriminative network are used for adversarial training, and a spectral perceptual loss function is proposed to generate highquality hyperspectral images; in stage-II, parameters of the generative network is fixed and the generated hyperspectral images are used as data augmentation in the training of Swin-Spec Transformer segmentation network. The proposed framework achieved 76.16% mIoU(mean Intersection over Union), 85.80% mDice(mean Dice), 90.96% Accuracy and 71.65% Kappa coefficient in the semantic segmentation task of the Multidimensional Choledoch Database. Compared with other methods, the results demonstrate our framework provides a competitive segmentation performance.
Kaijie Dai, Zehao Zhou, Song Qiu, Yan Wang 0033, Mei Zhou, Mingshuai Li, Qingli Li
IJCNN4
2023 Gene-Induced Multimodal Pre-training for Image-Omic Classification
Xingran Xie, Renjie Wan, Qingli Li, Yan Wang 0033
MICCAI (6)5
2023 Deep Mutual Distillation for Semi-supervised Medical Image Segmentation
Yushan Xie, Yuejia Yin, Qingli Li, Yan Wang 0033
MICCAI (3)4
2023 Factor Space and Spectrum for Medical Hyperspectral Image Segmentation
Boxiang Yun, Qingli Li, Lubov B. Mitrofanova, Chunhua Zhou, Yan Wang 0033
MICCAI (4)5
2023 Exploring Hyperspectral Histopathology Image Segmentation from a Deformable Perspective
abstract
Hyperspectral images (HSIs) offer great potential for computational pathology. However, limited by the spectral redundancy and the lack of spectral prior in popular 2D networks, previous HSI based techniques do not perform well. To address these problems, we propose to segment HSIs from a deformable perspective, which processes different spectral bands independently and fuses spatiospectral features of interest via deformable attention mechanisms. In addition, we propose Deformable Self-Supervised Spectral Regression (DF-S3R), which introduces two self-supervised pre-text tasks based on the low rank prior of HSIs enabling the network learning with spectrum-related features. During pre-training, DF-S3R learns both spectral structures and spatial morphology, and the jointly pre-trained architectures help alleviate the transfer risk to downstream fine-tuning. Compared to previous works, experiments show that our deformable architecture and pre-training method perform much better than other competitive methods on pathological semantic segmentation tasks, and the visualizations indicate that our method can trace the critical spectral characteristics from subtle spectral disparities. Code will be released at https://github.com/Ayakax/DFS3R.
Xingran Xie, Boxiang Yun, Qingli Li, Yan Wang 0033
ACM Multimedia5
2023 Uni-Dual: A Generic Unified Dual-Task Medical Self-Supervised Learning Framework
abstract
RGB images and medical hyperspectral images (MHSIs) are two widely-used modalities in computational pathology. The former is cheap, easy and fast to obtain while lacking pathological information such as physiochemical state. The latter is an emerging modality which captures electromagnetic radiation matter interaction but suffers from problems such as high time cost and low spatial resolution. In this paper, we bring forward a unified dual-task multi-modality self-supervised learning (SSL) framework, called Uni-Dual, which takes the most use of both paired and unpaired RGB-MHSIs. Concretely, we design a unified SSL paradigm for RGB images and MHSIs. Two tasks are proposed: (1) a discrimination learning task which learns high-level semantics via mining the cross-correlation across unpaired RGB-MHSIs, (2) a reconstruction learning task which models low-level stochastic variations via furthering the interaction across RGB-MHSI pairs. Our Uni-Dual enjoys the following benefits: (1) A unified model which can be easily transferred to different downstream tasks on various modality combinations. (2) We consider multi-constituent and structured information learning from MHSIs and RGB images for low-cost high-precision clinical purposes. Experiments conducted on various downstream tasks with different modalities show the proposed Uni-Dual substantially outperforms other competitive SSL methods.
Boxiang Yun, Xingran Xie, Qingli Li, Yan Wang 0033
ACM Multimedia4
2022 ContrastMask: Contrastive Learning to Segment Every Thing
abstract
Partially-supervised instance segmentation is a task which requests segmenting objects from novel categories via learning on limited base categories with annotated masks thus eliminating demands of heavy annotation burden. The key to addressing this task is to build an effective class-agnostic mask segmentation model. Unlike previous methods that learn such models only on base categories, in this paper, we propose a new method, named ContrastMask, which learns a mask segmentation model on both base and novel categories under a unified pixel-level contrastive learning framework. In this framework, annotated masks of base categories and pseudo masks of novel categories serve as a prior for contrastive learning, where features from the mask regions (foreground) are pulled together, and are contrasted against those from the background, and vice versa. Through this framework, feature discrimination between foreground and background is largely improved, facilitating learning of the class-agnostic mask segmentation model. Exhaustive experiments on the COCO dataset demonstrate the superiority of our method, which outperforms previous state-of-the-arts.
Kai Zhao 0012, Shouhong Ding, Yan Wang 0033, Wei Shen 0002
CVPR5
2022 S3R: Self-supervised Spectral Regression for Hyperspectral Histopathology Image Classification
Xingran Xie, Yan Wang 0033, Qingli Li
MICCAI (2)2
2022 External Attention Assisted Multi-Phase Splenic Vascular Injury Segmentation With Limited Data
abstract
The spleen is one of the most commonly injured solid organs in blunt abdominal trauma. The development of automatic segmentation systems from multi-phase CT for splenic vascular injury can augment severity grading for improving clinical decision support and outcome prediction. However, accurate segmentation of splenic vascular injury is challenging for the following reasons: 1) Splenic vascular injury can be highly variant in shape, texture, size, and overall appearance; and 2) Data acquisition is a complex and expensive procedure that requires intensive efforts from both data scientists and radiologists, which makes large-scale well-annotated datasets hard to acquire in general. In light of these challenges, we hereby design a novel framework for multi-phase splenic vascular injury segmentation, especially with limited data. On the one hand, we propose to leverage external data to mine pseudo splenic masks as the spatial attention, dubbed external attention, for guiding the segmentation of splenic vascular injury. On the other hand, we develop a synthetic phase augmentation module, which builds upon generative adversarial networks, for populating the internal data by fully leveraging the relation between different phases. By jointly enforcing external attention and populating internal data representation during training, our proposed method outperforms other competing methods and substantially improves the popular DeepLab-v3+ baseline by more than 7% in terms of average DSC, which confirms its effectiveness.
Yuyin Zhou, David Dreizin, Yan Wang 0033, Fengze Liu, Wei Shen 0002, Alan L. Yuille
IEEE Trans. Medical Imaging3
2021 Proposal Learning for Semi-Supervised Object Detection
abstract
In this paper, we focus on semi-supervised object detection to boost performance of proposal-based object detectors (a.k.a. two-stage object detectors) by training on both labeled and unlabeled data. However, it is non-trivial to train object detectors on unlabeled data due to the un-availability of ground truth labels. To address this problem, we present a proposal learning approach to learn proposal features and predictions from both labeled and unlabeled data. The approach consists of a self-supervised proposal learning module and a consistency-based proposal learning module. In the self-supervised proposal learning module, we present a proposal location loss and a contrastive loss to learn context-aware and noise-robust proposal features respectively. In the consistency-based proposal learning module, we apply consistency losses to both bounding box classification and regression predictions of proposals to learn noise-robust proposal features and predictions. Our approach enjoys the following benefits: 1) encouraging more context information to be delivered in the proposals learning procedure; 2) noisy proposal features and enforcing consistency to allow noise-robust object detection; 3) building a general and high-performance semi-supervised object detection framework, which can be easily adapted to proposal-based object detectors with different backbone architectures. Experiments are conducted on the COCO dataset with all available labeled and unlabeled data. Results demonstrate that our approach consistently improves the performance of fully-supervised baselines. In particular, after combining with data distillation [39], our approach improves AP by about 2.0% and 0.9% on average compared to fully-supervised baselines and data distillation baselines respectively.
Peng Tang 0005, Chetan Ramaiah, Yan Wang 0033, Ran Xu 0001, Caiming Xiong
WACV3
2021 Deep Differentiable Random Forests for Age Estimation
abstract
Age estimation from facial images is typically cast as a label distribution learning or regression problem, since aging is a gradual progress. Its main challenge is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging. In this paper, we propose two Deep Differentiable Random Forests methods, Deep Label Distribution Learning Forest (DLDLF) and Deep Regression Forest (DRF), for age estimation. Both of them connect split nodes to the top layer of convolutional neural networks (CNNs) and deal with inhomogeneous data by jointly learning input-dependent data partitions at the split nodes and age distributions at the leaf nodes. This joint learning follows an alternating strategy: (1) Fixing the leaf nodes and optimizing the split nodes and the CNN parameters by Back-propagation; (2) Fixing the split nodes and optimizing the leaf nodes by Variational Bounding. Two Deterministic Annealing processes are introduced into the learning of the split and leaf nodes, respectively, to avoid poor local optima and obtain better estimates of tree parameters free of initial values. Experimental results show that DLDLF and DRF achieve state-of-the-art performance on three age estimation datasets.
Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 A 76-81-GHz Four-Channel Digitally Controlled CMOS Receiver for Automotive Radars
abstract
This paper presents a fully-integrated 76-81 GHz four-channel digitally controlled receiver (RX) in 65-nm CMOS, which utilizes a four-channel RX front-end with high linearity, an LO distribution network, and a reconfigurable 4-channel analog baseband (ABB). It achieves high integration level and is capable of flexible reconfigurability automotive radars with wide and flattened temperature characteristic. Each RX front-end consists of a low noise amplifier (LNA) employing a neutralized common source (CS) with configurable high-linearity and an active mixer with linearity-priority design. Additionally, a bias circuit with positive temperature coefficient is used for temperature compensation of the millimeter wave (mm-Wave) amplifiers and mixer. The RX chain offers tunable gain and bandwidth (BW) with digital control and bandwidth tuning for the ABB. The measured results show that the RX achieves an input 1-dB compression point (Pin,1dB) of -7 dBm and a noise figure of 11-26 dB across a temperature range from -45 to +125°C, both over the range of 76-81 GHz. The RX has a tunable gain from 18 to 66 dB over the 3-dB programmable IF BW from 0.1 to 10 MHz. The power consumption of the entire receiver is 0.52 W.
Dongfang Pan, Zongming Duan, Yan Wang 0023, Yan Wang 0033, Liguo Sun, Ping Gui, Lin Cheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 Identification of Melanoma From Hyperspectral Pathology Image Using 3D Convolutional Networks
abstract
Skin biopsy histopathological analysis is one of the primary methods used for pathologists to assess the presence and deterioration of melanoma in clinical. A comprehensive and reliable pathological analysis is the result of correctly segmented melanoma and its interaction with benign tissues, and therefore providing accurate therapy. In this study, we applied the deep convolution network on the hyperspectral pathology images to perform the segmentation of melanoma. To make the best use of spectral properties of three dimensional hyperspectral data, we proposed a 3D fully convolutional network named Hyper-net to segment melanoma from hyperspectral pathology images. In order to enhance the sensitivity of the model, we made a specific modification to the loss function with caution of false negative in diagnosis. The performance of Hyper-net surpassed the 2D model with the accuracy over 92%. The false negative rate decreased by nearly 66% using Hyper-net with the modified loss function. These findings demonstrated the ability of the Hyper-net for assisting pathologists in diagnosis of melanoma based on hyperspectral pathology images.
Qian Wang 0046, Li Sun 0012, Yan Wang 0033, Mei Zhou, Menghan Hu, Ying Wen 0003, Qingli Li
IEEE Trans. Medical Imaging3
2021 Learning Inductive Attention Guidance for Partially Supervised Pancreatic Ductal Adenocarcinoma Prediction
abstract
Pancreatic ductal adenocarcinoma (PDAC) is the third most common cause of cancer death in the United States. Predicting tumors like PDACs (including both classification and segmentation) from medical images by deep learning is becoming a growing trend, but usually a large number of annotated data are required for training, which is very labor-intensive and time-consuming. In this paper, we consider a partially supervised setting, where cheap image-level annotations are provided for all the training data, and the costly per-voxel annotations are only available for a subset of them. We propose an Inductive Attention Guidance Network (IAG-Net) to jointly learn a global image-level classifier for normal/PDAC classification and a local voxel-level classifier for semi-supervised PDAC segmentation. We instantiate both the global and the local classifiers by multiple instance learning (MIL), where the attention guidance, indicating roughly where the PDAC regions are, is the key to bridging them: For global MIL based normal/PDAC classification, attention serves as a weight for each instance (voxel) during MIL pooling, which eliminates the distraction from the background; For local MIL based semi-supervised PDAC segmentation, the attention guidance is inductive, which not only provides bag-level pseudo-labels to training data without per-voxel annotations for MIL training, but also acts as a proxy of an instance-level classifier. Experimental results show that our IAG-Net boosts PDAC segmentation accuracy by more than 5% compared with the state-of-the-arts.
Yan Wang 0033, Peng Tang 0005, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
IEEE Trans. Medical Imaging1
2021 Inter-Slice Context Residual Learning for 3D Medical Image Segmentation
abstract
Automated and accurate 3D medical image segmentation plays an essential role in assisting medical professionals to evaluate disease progresses and make fast therapeutic schedules. Although deep convolutional neural networks (DCNNs) have widely applied to this task, the accuracy of these models still need to be further improved mainly due to their limited ability to 3D context perception. In this paper, we propose the 3D context residual network (ConResNet) for the accurate segmentation of 3D medical images. This model consists of an encoder, a segmentation decoder, and a context residual decoder. We design the context residual module and use it to bridge both decoders at each scale. Each context residual module contains both context residual mapping and context attention mapping, the formal aims to explicitly learn the inter-slice context information and the latter uses such context as a kind of attention to boost the segmentation accuracy. We evaluated this model on the MICCAI 2018 Brain Tumor Segmentation (BraTS) dataset and NIH Pancreas Segmentation (Pancreas-CT) dataset. Our results not only demonstrate the effectiveness of the proposed 3D context residual learning scheme but also indicate that the proposed ConResNet is more accurate than six top-ranking methods in brain tumor segmentation and seven top-ranking methods in pancreas segmentation.
Yutong Xie 0001, Yan Wang 0033, Yong Xia 0001
IEEE Trans. Medical Imaging3
2020 Deep Distance Transform for Tubular Structure Segmentation in CT Scans
abstract
Tubular structure segmentation in medical images, e.g., segmenting vessels in CT scans, serves as a vital step in the use of computers to aid in screening early stages of related diseases. But automatic tubular structure segmentation in CT scans is a challenging problem, due to issues such as poor contrast, noise and complicated background. A tubular structure usually has a cylinder-like shape which can be well represented by its skeleton and cross-sectional radii (scales). Inspired by this, we propose a geometry-aware tubular structure segmentation method, Deep Distance Transform (DDT), which combines intuitions from the classical distance transform for skeletonization and modern deep segmentation networks. DDT first learns a multi-task network to predict a segmentation mask for a tubular structure and a distance map. Each value in the map represents the distance from each tubular structure voxel to the tubular structure surface. Then the segmentation mask is refined by leveraging the shape prior reconstructed from the distance map. We apply our DDT on six medical image datasets. Results show that (1) DDT can boost tubular structure segmentation performance significantly (e.g., over 13% DSC improvement for pancreatic duct segmentation), and (2) DDT additionally provides a geometrical measurement for a tubular structure, which is important for clinical diagnosis (e.g., the cross-sectional scale of a pancreatic duct can be an indicator for pancreatic cancer).
Yan Wang 0033, Fengze Liu, Jieneng Chen, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
CVPR1
2020 Domain Adaptive Relational Reasoning for 3D Multi-organ Segmentation
Shuhao Fu, Yongyi Lu, Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
MICCAI (1)3
2020 Robust Face Detection via Learning Small Faces on Hard Images
abstract
Recent anchor-based deep face detectors have achieved promising performance, but they are still struggling to detect hard faces, such as small, blurred and partially occluded faces. One reason is that they treat all images and faces equally, and ignore the imbalance between easy images and hard images; however large amounts of training images only contain easy faces, which are less helpful to learn robust detectors for hard faces. In this paper, we propose that the robustness of a face detector against hard faces can be improved by learning small faces on hard images. Our intuitions are (1) hard images are the images which contain at least one hard face, thus they facilitate training robust face detectors; (2) most hard faces are small faces and other types of hard faces can be easily shrunk to small faces. To this end, we build an anchor-based deep face detector, which only outputs a single high-resolution feature map with small anchors, to specifically learn small faces and train it by a novel hard image mining strategy which automatically adjusts training weights on images according to their difficulties. Extensive experiments have been conducted on WIDER FACE, FDDB, Pascal Faces, and AFW datasets and our method achieves APs of 95.7, 94.9 and 89.7 on easy, medium and hard WIDER FACE val dataset respectively, which verify the effectiveness of our methods, especially on detecting hard faces. Our detector is also lightweight and enjoys a fast inference speed. Code and model are available at https://github.com/bairdzhang/smallhardface.
Zhishuai Zhang, Wei Shen 0002, Siyuan Qiao, Yan Wang 0033, Bo Wang 0044, Alan L. Yuille
WACV4
2020 Recurrent Saliency Transformation Network for Tiny Target Segmentation in Abdominal CT Scans
abstract
We aim at segmenting a wide variety of organs, including tiny targets (e.g., adrenal gland), and neoplasms (e.g., pancreatic cyst), from abdominal CT scans. This is a challenging task in two aspects. First, some organs (e.g., the pancreas), are highly variable in both anatomy and geometry, and thus very difficult to depict. Second, the neoplasms often vary a lot in its size, shape, as well as its location within the organ. Third, the targets (organs and neoplasms) can be considerably small compared to the human body, and so standard deep networks for segmentation are often less sensitive to these targets and thus predict less accurately especially around their boundaries. In this paper, we present an end-to-end framework named recurrent saliency transformation network (RSTN) for segmenting tiny and/or variable targets. The RSTN is a coarse-to-fine approach that uses prediction from the first (coarse) stage to shrink the input region for the second (fine) stage. A saliency transformation module is inserted between these two stages so that 1) the coarse-scaled segmentation mask can be transferred as spatial weights and applied to the fine stage and 2) the gradients can be back-propagated from the loss layer to the entire network so that the two stages are optimized in a joint manner. In the testing stage, we perform segmentation iteratively to improve accuracy. In this extended journal paper, we allow a gradual optimization to improve the stability of the RSTN, and introduce a hierarchical version named H-RSTN to segment tiny and variable neoplasms such as pancreatic cysts. Experiments are performed on several CT datasets including a public pancreas segmentation dataset, our own multi-organ dataset, and a cystic pancreas dataset. In all these cases, the RSTN outperforms the baseline (a stage-wise coarse-to-fine approach) significantly. Confirmed by the radiologists in our team, these promising segmentation results can help early diagnosis of pancreatic cancer. The code and pre-trained models of our project were made available at https://github.com/198808xc/OrganSegRSTN.
Lingxi Xie, Qihang Yu, Yuyin Zhou, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille
IEEE Trans. Medical Imaging4
2019 A Digitally Controlled CMOS Receiver with -14 dBm P1dB for 77 GHz Automotive Radar
abstract
This paper presents a 77 GHz Receiver (Rx) for automotive radar with high linearity, low noise figure (NF) and reconfigurable gain and bandwidth of the analog baseband (ABB), implemented in TSMC 65 nm CMOS general purpose (GP) technology. The Rx includes a low-noise amplifier (LNA), sharing the transconductance stage with a double-balanced mixer, an LO chain with a frequency doubler and an analog baseband. The entire Rx consumes 123 mW. Measurement results show that the input 1 dB compression point (P1dB) is -14 dBm at 81 GHz, a conversion gain (CG) of 20.5-62.5 dB is digitally tunable at 77 GHz, and an NF of 11.2-18.3 dB is achieved over the frequency range of 76-81 GHz.
Dongfang Pan, Zongming Duan, Yan Wang 0023, Yan Wang 0033, Ping Gui, Liguo Sun
ISCAS6
2019 Hyper-Pairing Network for Multi-phase Pancreatic Ductal Adenocarcinoma Segmentation
Yuyin Zhou, Yingwei Li 0002, Zhishuai Zhang, Yan Wang 0033, Angtian Wang, Elliot K. Fishman, Alan L. Yuille, Seyoun Park
MICCAI (2)4
2019 Semi-Supervised 3D Abdominal Multi-Organ Segmentation Via Deep Multi-Planar Co-Training
abstract
In multi-organ segmentation of abdominal CT scans, most existing fully supervised deep learning algorithms require lots of voxel-wise annotations, which are usually difficult, expensive, and slow to obtain. In comparison, massive unlabeled 3D CT volumes are usually easily accessible. Current mainstream works to address semi-supervised biomedical image segmentation problem are mostly graph-based. By contrast, deep network based semi-supervised learning methods have not drawn much attention in this field. In this work, we propose Deep Multi-Planar Co-Training (DMPCT), whose contributions can be divided into two folds: 1) The deep model is learned in a co-training style which can mine consensus information from multiple planes like the sagittal, coronal, and axial planes; 2) Multi-planar fusion is applied to generate more reliable pseudo-labels, which alleviates the errors occurring in the pseudo-labels and thus can help to train better segmentation networks. Experiments are done on our newly collected large dataset with 100 unlabeled cases as well as 210 labeled cases where 16 anatomical structures are manually annotated by four radiologists and confirmed by a senior expert. The results suggest that DMPCT significantly outperforms the fully supervised method by more than 4% especially when only a small set of annotations is used.
Yuyin Zhou, Yan Wang 0033, Peng Tang 0005, Song Bai 0001, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
WACV2
2019 Abdominal multi-organ segmentation with organ-attention networks and statistical fusion
Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Seyoun Park, Elliot K. Fishman, Alan L. Yuille
Medical Image Anal.1
2018 Deep Regression Forests for Age Estimation
abstract
Age estimation from facial images is typically cast as a nonlinear regression problem. The main challenge of this problem is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging patterns. In this paper, we propose Deep Regression Forests (DRFs), an end-to-end model, for age estimation. DRFs connect the split nodes to a fully connected layer of a convolutional neural network (CNN) and deal with inhomogeneous data by jointly learning input-dependant data partitions at the split nodes and data abstractions at the leaf nodes. This joint learning follows an alternating strategy: First, by fixing the leaf nodes, the split nodes as well as the CNN parameters are optimized by Back-propagation; Then, by fixing the split nodes, the leaf nodes are optimized by iterating a step-size free update rule derived from Variational Bounding. We verify the proposed DRFs on three standard age estimation benchmarks and achieve state-of-the-art results on all of them.
Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille
CVPR3
2018 Recurrent Saliency Transformation Network: Incorporating Multi-Stage Visual Cues for Small Organ Segmentation
abstract
We aim at segmenting small organs (e.g., the pancreas) from abdominal CT scans. As the target often occupies a relatively small region in the input image, deep neural networks can be easily confused by the complex and variable background. To alleviate this, researchers proposed a coarse-to-fine approach [46], which used prediction from the first (coarse) stage to indicate a smaller input region for the second (fine) stage. Despite its effectiveness, this algorithm dealt with two stages individually, which lacked optimizing a global energy function, and limited its ability to incorporate multi-stage visual cues. Missing contextual information led to unsatisfying convergence in iterations, and that the fine stage sometimes produced even lower segmentation accuracy than the coarse stage. This paper presents a Recurrent Saliency Transformation Network. The key innovation is a saliency transformation module, which repeatedly converts the segmentation probability map from the previous iteration as spatial weights and applies these weights to the current iteration. This brings us two-fold benefits. In training, it allows joint optimization over the deep networks dealing with different input scales. In testing, it propagates multi-stage visual information throughout iterations to improve segmentation accuracy. Experiments in the NIH pancreas segmentation dataset demonstrate the state-of-the-art accuracy, which outperforms the previous best by an average of over 2%. Much higher accuracies are also reported on several small organs in a larger dataset collected by ourselves. In addition, our approach enjoys better convergence properties, making it more efficient and reliable in practice.
Qihang Yu, Lingxi Xie, Yan Wang 0033, Yuyin Zhou, Elliot K. Fishman, Alan L. Yuille
CVPR3
2018 Multi-scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang 0033, Lingxi Xie, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Alan L. Yuille
ECCV (13)1
2018 Training Multi-organ Segmentation Networks with Sample Selection by Relaxed Upper Confident Bound
Yan Wang 0033, Yuyin Zhou, Peng Tang 0005, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
MICCAI (4)1
2017 Multi-stage Multi-recursive-input Fully Convolutional Networks for Neuronal Boundary Detection
abstract
In the field of connectomics, neuroscientists seek to identify cortical connectivity comprehensively. Neuronal boundary detection from the Electron Microscopy (EM) images is often done to assist the automatic reconstruction of neuronal circuit. But the segmentation of EM images is a challenging problem, as it requires the detector to be able to detect both filament-like thin and blob-like thick membrane, while suppressing the ambiguous intracellular structure. In this paper, we propose multi-stage multi-recursiveinput fully convolutional networks to address this problem. The multiple recursive inputs for one stage, i.e., the multiple side outputs with different receptive field sizes learned from the lower stage, provide multi-scale contextual boundary information for the consecutive learning. This design is biologically-plausible, as it likes a human visual system to compare different possible segmentation solutions to address the ambiguous boundary issue. Our multi-stage networks are trained end-to-end. It achieves promising results on two public available EM segmentation datasets, the mouse piriform cortex dataset and the ISBI 2012 EM dataset.
Wei Shen 0002, Bin Wang 0027, Yuan Jiang 0002, Yan Wang 0033, Alan L. Yuille
ICCV4
2017 SORT: Second-Order Response Transform for Visual Recognition
abstract
In this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct advantage of SORT is to facilitate cross-branch response propagation, so that each branch can update its weights based on the current status of the other branch. Moreover, SORT augments the family of transform operations and increases the nonlinearity of the network, making it possible to learn flexible functions to fit the complicated distribution of feature space. SORT can be applied to a wide range of network architectures, including a branched variant of a chain-styled network and a residual network, with very light-weighted modifications. We observe consistent accuracy gain on both small (CIFAR10, CIFAR100 and SVHN) and big (ILSVRC2012) datasets. In addition, SORT is very efficient, as the extra computation overhead is less than 5%.
Yan Wang 0033, Lingxi Xie, Chenxi Liu 0001, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Qi Tian 0001, Alan L. Yuille
ICCV1
2017 A Fixed-Point Model for Pancreas Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Wei Shen 0002, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille
MICCAI (1)4
2017 DeepSkeleton: Learning Multi-Task Scale-Associated Deep Side Outputs for Object Skeleton Extraction in Natural Images
abstract
Object skeletons are useful for object representation and object detection. They are complementary to the object contour, and provide extra information, such as how object scale (thickness) varies among object parts. But object skeleton extraction from natural images is very challenging, because it requires the extractor to be able to capture both local and non-local image context in order to determine the scale of each skeleton pixel. In this paper, we present a novel fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the different layers in the network and the skeleton scales they can capture, we introduce two scale-associated side outputs to each stage of the network. The network is trained by multi-task learning, where one task is skeleton localization to classify whether a pixel is a skeleton pixel or not, and the other is skeleton scale prediction to regress the scale of each skeleton pixel. Supervision is imposed at different stages by guiding the scale-associated side outputs toward the ground-truth skeletons at the appropriate scales. The responses of the multiple scale-associated side outputs are then fused in a scale-specific way to detect skeleton pixels using multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. In addition, the usefulness of the obtained skeletons and scales (thickness) are verified on two object detection applications: foreground object segmentation and object proposal detection.
Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Xiang Bai, Alan L. Yuille
IEEE Trans. Image Process.4
2016 Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side Outputs
abstract
Object skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be able to capture both local and global image context to determine the intrinsic scale of each skeleton pixel. Existing methods rely on per-pixel based multi-scale feature computation, which results in difficult modeling and high time consumption. In this paper, we present a fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the sequential stages in the network and the skeleton scales they can capture, we introduce a scale-associated side output to each stage. We impose supervision to different stages by guiding the scale-associated side outputs toward groundtruth skeletons of different scales. The responses of the multiple scaleassociated side outputs are then fused in a scale-specific way to localize skeleton pixels with multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors.
Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Zhijiang Zhang, Xiang Bai
CVPR4
2016 On Branded Handbag Recognition
abstract
Manufacturing branded handbags is a big business in the fashion world. Shoppers’ feedback showing photos of their purchased handbags in social networks or blogs is important for branding purposes. In this paper, we deal with handbag recognition. It is a challenging problem due to the inter-class style similarity and the intra-class color variation. We focus on developing discriminative representations of handbag style and color. For handbag style representation, two supervised mid-level patch selection procedures are proposed to select discriminative patches, regarding individual classes and pairwise classes. We also propose a low-level complementary feature, extracted from texture-enhanced mid-level patches, to capture the fine details of the mid-level patches. For handbag color representation, we propose to extract dominant color features to handle the illumination changes. The performance of our proposed method is evaluated on a newly built branded handbag dataset. The results show that our method performs favorably in recognizing handbags, with around$10\%$improvement in accuracy when compared with the existing fine-grained or generic object recognition methods.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Multim.1
2015 DeepContour: A deep convolutional feature learned by positive-sharing loss for contour detection
abstract
Contour detection serves as the basis of a variety of computer vision tasks such as image segmentation and object recognition. The mainstream works to address this problem focus on designing engineered gradient features. In this work, we show that contour detection accuracy can be improved by instead making the use of the deep features learned from convolutional neural networks (CNNs). While rather than using the networks as a blackbox feature extractor, we customize the training strategy by partitioning contour (positive) data into subclasses and fitting each subclass by different model parameters. A new loss function, named positive-sharing loss, in which each subclass shares the loss for the whole positive class, is proposed to learn the parameters. Compared to the sofmax loss function, the proposed one, introduces an extra regularizer to emphasizes the losses for the positive and negative classes, which facilitates to explore more discriminative features. Our experimental results demonstrate that learned deep features can achieve top performance on Berkeley Segmentation Dataset and Benchmark (BSDS500) and obtain competitive cross dataset generalization result on the NYUD dataset.
Wei Shen 0002, Xinggang Wang, Yan Wang 0033, Xiang Bai, Zhijiang Zhang
CVPR3
2015 Joint learning for image-based handbag recommendation
abstract
Fashion recommendation helps shoppers to find desirable fashion items, which facilitates online interaction and product promotion. In this paper, we propose a method to recommend handbags to each shopper, based on the handbag images the shopper has clicked. This is performed by Joint learning of attribute Projection and One-class SVM classification (JPO) based on the images of the shopper's preferred handbags. More specifically, for the handbag images clicked by each shopper, we project the original image feature space into an attribute space which is more compact. The projection matrix is learned jointly with a one-class SVM to yield a shopper-specific one-class classifier. The results show that the proposed JPO handbag recommendation performs favorably based on initial subject testing.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
ICME1
2015 DeepBag: Recognizing Handbag Models
abstract
In this paper, we address the problem of branded handbag recognition. It is a challenging problem due to the non-rigid deformation, illumination changes, and inter-class similarity. We propose a novel framework based on deep convolutional neural network (CNN). Concretely, we propose a new CNN model, called feature selective joint classification - regression CNN (FSCR-CNN). Its advantages lie in two folds: 1) it alleviates the illumination changes by a feature selection strategy to focus on the color- nondiscriminative features in the network learning, and 2) rather than only targeting on the hard label (i.e., the handbag model), it also incorporates a soft label (i.e., a distribution measuring the similarity between the ground truth model and all the models to be trained) to construct the loss function for training CNN, which leads to a better classifier for handbags with large inter-class similarity. We evaluate the performance of our framework on a newly built branded handbag dataset. The results show that it performs favorably for recognizing handbags with 94.48% in accuracy. We also apply the proposed FSCR-CNN model in recognizing other fine-grained objects with state-of-the-art CNN architectures, which is able to achieve over 5% improvement in accuracy.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
IEEE Trans. Multim.1
2014 Complementary feature extraction for branded handbag recognition
abstract
Fine-grained object recognition aims at recognizing objects belonging to the same basic-level class such as dog, bird or fish, which is a challenging problem in computer vision. In this paper, we consider the problem of recognizing handbags that belong to a specific brand. In order to identify the subtle differences among handbags, we propose to enhance the handbag local structure pattern by using the Hölder exponent, and extract the feature from the enhanced handbag image to complement the feature extracted directly from the original handbag image. We term such two types of features as the complementary and original features. These features will then be fused by using Multiple Kernel Learning (MKL) for branded handbag recognition. We conduct the experiments on a newly built branded handbag dataset, the results of which demonstrate the effectiveness of the proposed complementary feature in recognizing the handbags.
Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot
ICIP1
2013 Shape clustering: Common structure discovery
Wei Shen 0002, Yan Wang 0033, Xiang Bai, Longin Jan Latecki
Pattern Recognit.2