Jiancheng Yang

dblp:09/7763 · DBLP profile ↗
← Back
42ranked-venue papers
13as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 10 first-author · 21 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AortaDiff: A Unified Multitask Diffusion Framework for Contrast-Free AAA Imaging
abstract
While contrast-enhanced CT (CECT) is standard for assessing abdominal aortic aneurysms (AAA) , the required iodinated contrast agents pose significant risks, including patient allergies and environmental harm. To reduce contrast agent use, methods have focused on generating synthetic CECT from non-contrast CT (NCCT) scans. However, most adopt a multi-stage pipeline that first generates images and then performs segmentation, which leads to error accumulation and fails to leverage shared information. To address this, we propose a unified framework that generates synthetic CECT images from NCCT scans while simultaneously segmenting the aortic lumen and thrombus. Our approach integrates conditional diffusion models (CDM) with multi-task learning, enabling end-to-end joint optimization of image synthesis and anatomical segmentation. Unlike previous multitask diffusion models, our approach requires no initial predictions (e.g., a coarse segmentation mask), shares both encoder and decoder parameters across tasks, and employs a semi-supervised training strategy to learn from scans with missing segmentation labels, a common constraint in clinical data. Evaluated on a cohort of 264 patients, our method consistently outperformed state-of-the-art single-task and multi-stage models. For image synthesis, it achieved a PSNR of 25.61 dB, compared to 23.80 dB from a single-task CDM. For segmentation, it improved the lumen Dice score to 0.89 from 0.87 and the challenging thrombus Dice score to 0.53 from 0.48 (nnU-Net). These segmentation enhancements led to more accurate clinical measurements, reducing the lumen diameter MAE to 4.19 mm from 5.78 mm and the thrombus area error to 33.85% from 41.45%. Code is at https://github.com/yuxuanou623/AortaDiff.
Yuxuan Ou, Ning Bi, Jiazheng Pan, Jiancheng Yang, Boliang Yu, Usama Zidan, Regent Lee, Vicente Grau
WACV4
2026 Template-guided reconstruction of pulmonary segments with neural implicit functions
abstract
High-quality 3D reconstruction of pulmonary segments plays a crucial role in segmentectomy and surgical planning for the treatment of lung cancer. Due to the resolution requirement of the target reconstruction, conventional deep learning-based methods often suffer from computational resource constraints or limited granularity. Conversely, implicit modeling is favored due to its computational efficiency and continuous representation at any resolution. We propose a neural implicit function-based method to learn a 3D surface to achieve anatomy-aware, precise pulmonary segment reconstruction, represented as a shape by deforming a learnable template. Additionally, we introduce two clinically relevant evaluation metrics to comprehensively assess the quality of the reconstruction. Furthermore, to address the lack of publicly available shape datasets for benchmarking reconstruction algorithms, we developed a shape dataset named Lung3D, which includes the 3D models of 800 labeled pulmonary segments and their corresponding airways, arteries, veins, and intersegmental veins. We demonstrate that the proposed approach outperforms existing methods, providing a new perspective for pulmonary segment reconstruction. Code and data will be available at https://github.com/HINTLab/ImPulSe.
Kangxian Xie, Kaiming Kuang, Li Zhang 0085, Hongwei Li 0004, Mingchen Gao, Jiancheng Yang
Medical Image Anal.7
2025 LeFusion: Controllable Pathology Synthesis via Lesion-Focused Diffusion Models
abstract
Patient data from real-world clinical practice often suffers from data scarcity and long-tail imbalances, leading to biased outcomes or algorithmic unfairness. This study addresses these challenges by generating lesion-containing image-segmentation pairs from lesion-free images. Previous efforts in medical imaging synthesis have struggled with separating lesion information from background, resulting in low-quality backgrounds and limited control over the synthetic output. Inspired by diffusion-based image inpainting, we propose LeFusion, a lesion-focused diffusion model. By redesigning the diffusion learning objectives to focus on lesion areas, we simplify the learning process and improve control over the output while preserving high-fidelity backgrounds by integrating forward-diffused background contexts into the reverse diffusion process. Additionally, we tackle two major challenges in lesion texture synthesis: 1) multi-peak and 2) multi-class lesions. We introduce two effective strategies: histogram-based texture control and multi-channel decomposition, enabling the controlled generation of high-quality lesions in difficult scenarios. Furthermore, we incorporate lesion mask diffusion, allowing control over lesion size, location, and boundary, thus increasing lesion diversity. Validated on 3D cardiac lesion MRI and lung nodule CT datasets, LeFusion-generated data significantly improves the performance of state-of-the-art segmentation models, including nnUNet and SwinUNETR.
Yuhe Liu, Jiancheng Yang, Shouhong Wan, Pascal Fua
ICLR3
2025 Pairwise-Constrained Implicit Functions for 3D Human Heart Modeling
Hieu Le 0001, Nicolas Talabot, Jiancheng Yang, Pascal Fua
MICCAI (16)4
2025 DiffAtlas: GenAI-Fying Atlas Segmentation via Image-Mask Diffusion
Yuhe Liu, Jiancheng Yang, Weidong Guo, Pascal Fua
MICCAI (16)3
2025 Efficient anatomical labeling of pulmonary tree structures via deep point-graph representation-based implicit fields
abstract
Pulmonary diseases rank prominently among the principal causes of death worldwide. Curing them will require, among other things, a better understanding of the complex 3D tree-shaped structures within the pulmonary system, such as airways, arteries, and veins. Traditional approaches using high-resolution image stacks and standard CNNs on dense voxel grids face challenges in computational efficiency, limited resolution, local context, and inadequate preservation of shape topology. Our method addresses these issues by shifting from dense voxel to sparse point representation, offering better memory efficiency and global context utilization. However, the inherent sparsity in point representation can lead to a loss of crucial connectivity in tree-shaped structures. To mitigate this, we introduce graph learning on skeletonized structures, incorporating differentiable feature fusion for improved topology and long-distance context capture. Furthermore, we employ an implicit function for efficient conversion of sparse representations into dense reconstructions end-to-end. The proposed method not only delivers state-of-the-art performance in labeling accuracy, both overall and at key locations, but also enables efficient inference and the generation of closed surface shapes. Addressing data scarcity in this field, we have also curated a comprehensive dataset to validate our approach. Data and code are available at https://github.com/M3DV/pulmonary-tree-labeling.
Kangxian Xie, Jiancheng Yang, Donglai Wei 0001, Ziqiao Weng, Pascal Fua
Medical Image Anal.2
2025 SimTA++: Simple attention neural network for clinical asynchronous time series
Kaiming Kuang, Baoyu Jing, Bo Du 0001, Jiancheng Yang
Neural Networks7
2025 Exploring Causal Information Bottleneck for Adversarial Defense
abstract
Information bottleneck (IB) is a promising defense solution against adversarial attacks on deep neural networks. However, these methods often suffer from spurious correlations. A correlation exists between the prediction and the non-robust features, yet it does not reflect the causal relationship well. Such spurious correlations induce the neural networks to learn fragile and incomprehensible (non-robust) features. This issue limits its potential for further improving adversarial robustness. This paper addresses this issue by incorporating causal inference into the IB-based defense framework. Specifically, we propose a novel defense method that use the instrumental variables to enhance the adversarial robustness. Our proposed method divides the features into two parts for causal effect estimation: robust and non-robust features. The robust features relate to understanding semantic information, and the non-robust features link to the vulnerable style information. By employing this framework, the IB method can mitigate the influence of non-robust features and extract the robust features linking to the semantic information of objects. We conduct a thorough analysis of the effectiveness of our proposed method. Notably, the experiments on MNIST, FashionMNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that our method significantly boosts the adversarial robustness against multiple adversarial attacks compared to previous methods. Our regularization method can improve adversarial robustness in both natural and adversarial training frameworks. Besides, CausalIB can be applied to both Convolutional Neural Networks and Vision Transformers as a plug-and-play module. Our code is available at https://github.com/HydrogenWasser/CausalIB.
Jun Yan 0013, Huan Hua, Weiquan Huang, Wancheng Ge, Jiancheng Yang
IEEE Trans. Inf. Forensics Secur.6
2025 Frenet-Serret Frame-Based Decomposition for Part Segmentation of 3-D Curvilinear Structures
abstract
Accurate segmentation of anatomical substructures within 3D curvilinear structures in medical imaging remains challenging due to their complex geometry and the scarcity of diverse, large-scale datasets for algorithm development and evaluation. In this paper, we use dendritic spine segmentation as a case study and address these challenges by introducing a novel Frenet-Serret Frame-based Decomposition, which decomposes 3D curvilinear structures into a globally smooth continuous curve that captures the overall shape, and a cylindrical primitive that encodes local geometric properties. This approach leverages Frenet-Serret Frames and arc length parameterization to preserve essential geometric features while reducing representational complexity, facilitating data-efficient learning, improved segmentation accuracy, and generalization on 3D curvilinear structures. To rigorously evaluate our method, we introduce two datasets: CurviSeg, a synthetic dataset for 3D curvilinear structure segmentation that validates our method's key properties, and DenSpineEM, a benchmark for dendritic spine segmentation, which comprises 4,476 manually annotated spines from 70 dendrites across three public electron microscopy datasets, covering multiple brain regions and species. Our experiments on DenSpineEM demonstrate exceptional cross-region and cross-species generalization: models trained on the mouse somatosensory cortex subset achieve 94.43% Dice, maintaining strong performance in zero-shot segmentation on both mouse visual cortex (95.61% Dice) and human frontal lobe (86.63% Dice) subsets. Moreover, we test the generalizability of our method on the IntrA dataset, where it achieves 77.08% Dice (5.29% higher than prior arts) on intracranial aneurysm segmentation from entire artery models. These findings demonstrate the potential of our approach for accurately analyzing complex curvilinear structures across diverse medical imaging fields. Our dataset, code, and models are available at https://github.com/VCG/FFD4DenSpineEM to support future research.
Shixuan Gu, Jason Ken Adhinarta, Mikhail Bessmeltsev, Jiancheng Yang, Yongjie Jessica Zhang, Daniel Berger, Jeff Lichtman, Hanspeter Pfister, Donglai Wei 0001
IEEE Trans. Medical Imaging4
2025 Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac Challenge
abstract
Rib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis.
Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni
IEEE Trans. Medical Imaging1
2024 Generating Anatomically Accurate Heart Structures via Neural Implicit Fields
Jiancheng Yang, Ekaterina Sedykh, Jason Ken Adhinarta, Hieu Le 0001, Pascal Fua
MICCAI (1)1
2024 Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance
Weiyi Zhang 0004, Siyu Huang, Jiancheng Yang, ZongYuan Ge, Yingfeng Zheng, Danli Shi, Mingguang He
MICCAI (1)3
2024 SGDA: Towards 3-D Universal Pulmonary Nodule Detection via Slice Grouped Domain Attention
abstract
Lung cancer is the leading cause of cancer death worldwide. The best solution for lung cancer is to diagnose the pulmonary nodules in the early stage, which is usually accomplished with the aid of thoracic computed tomography (CT). As deep learning thrives, convolutional neural networks (CNNs) have been introduced into pulmonary nodule detection to help doctors in this labor-intensive task and demonstrated to be very effective. However, the current pulmonary nodule detection methods are usually domain-specific, and cannot satisfy the requirement of working in diverse real-world scenarios. To address this issue, we propose a slice grouped domain attention (SGDA) module to enhance the generalization capability of the pulmonary nodule detection networks. This attention module works in the axial, coronal, and sagittal directions. In each direction, we divide the input feature into groups, and for each group, we utilize a universal adapter bank to capture the feature subspaces of the domains spanned by all pulmonary nodule datasets. Then the bank outputs are combined from the perspective of domain to modulate the input group. Extensive experiments demonstrate that SGDA enables substantially better multi-domain pulmonary nodule detection performance compared with the state-of-the-art multi-domain learning methods.
Rui Xu 0031, Zhi Liu 0002, Yong Luo 0002, Han Hu 0003, Li Shen 0008, Bo Du 0001, Kaiming Kuang, Jiancheng Yang
IEEE Trans. Comput. Biol. Bioinform.8
2024 : A Large-Scale Benchmark for Rib Labeling and Anatomical Centerline Extraction
abstract
Automatic rib labeling and anatomical centerline extraction are common prerequisites for various clinical applications. Prior studies either use in-house datasets that are inaccessible to communities, or focus on rib segmentation that neglects the clinical significance of rib labeling. To address these issues, we extend our prior dataset (RibSeg) on the binary rib segmentation task to a comprehensive benchmark, named RibSeg v2, with 660 CT scans (15,466 individual ribs in total) and annotations manually inspected by experts for rib labeling and anatomical centerline extraction. Based on the RibSeg v2, we develop a pipeline including deep learning-based methods for rib labeling, and a skeletonization-based method for centerline extraction. To improve computational efficiency, we propose a sparse point cloud representation of CT scans and compare it with standard dense voxel grids. Moreover, we design and analyze evaluation metrics to address the key challenges of each task. Our dataset, code, and model are available online to facilitate open research at https://github.com/M3DV/RibSeg.
Shixuan Gu, Donglai Wei 0001, Jason Ken Adhinarta, Kaiming Kuang, Yongjie Jessica Zhang, Hanspeter Pfister, Bingbing Ni, Jiancheng Yang, Ming Li 0005
IEEE Trans. Medical Imaging9
2024 GMILT: A Novel Transformer Network That Can Noninvasively Predict EGFR Mutation Status
abstract
Noninvasively and accurately predicting the epidermal growth factor receptor (EGFR) mutation status is a clinically vital problem. Moreover, further identifying the most suspicious area related to the EGFR mutation status can guide the biopsy to avoid false negatives. Deep learning methods based on computed tomography (CT) images may improve the noninvasive prediction of EGFR mutation status and potentially help clinicians guide biopsies by visual methods. Inspired by the potential inherent links between EGFR mutation status and invasiveness information, we hypothesized that the predictive performance of a deep learning network can be improved through extra utilization of the invasiveness information. Here, we created a novel explainable transformer network for EGFR classification named gated multiple instance learning transformer (GMILT) by integrating multi-instance learning and discriminative weakly supervised feature learning. Pathological invasiveness information was first introduced into the multitask model as embeddings. GMILT was trained and validated on a total of 512 patients with adenocarcinoma and tested on three datasets (the internal test dataset, the external test dataset, and The Cancer Imaging Archive (TCIA) public dataset). The performance (area under the curve (AUC) =0.772 on the internal test dataset) of GMILT exceeded that of previously published methods and radiomics-based methods (i.e., random forest and support vector machine) and attained a preferable generalization ability (AUC =0.856 in the TCIA test dataset and AUC =0.756 in the external dataset). A diameter-based subgroup analysis further verified the efficiency of our model (most of the AUCs exceeded 0.772) to noninvasively predict EGFR mutation status from computed tomography (CT) images. In addition, because our method also identified the "core area" of the most suspicious area related to the EGFR mutation status, it has the potential ability to guide biopsies.
Wei Zhao 0040, Weidao Chen, Du Lei, Jiancheng Yang, Yanjing Chen, Yingjia Jiang, Jiangfen Wu, Bingbing Ni, Yeqi Sun, Yingli Sun, Ming Li 0005, Jun Liu 0075
IEEE Trans. Neural Networks Learn. Syst.5
2023 Efficient Lung Nodule Detection via 3D Deep Learning with Shifted Convolutions
abstract
The high computational costs of deep convolutional neural networks hinder their deployment in real-world applications, including pulmonary nodule detection from CT scans where large 3D image sizes amplify the issue. This paper presents a novel 3D method to detect pulmonary nodules, based on anchor-free U-shaped networks, AFNet. A shifted convolution is further introduced to replace standard 3D convolutions, which reduces both the model sizes and FLOPs (floating-point operations). The shift operator is parameter-free, enabling 3D context fusion between CT slices using 2D convolutions. Extensive experiments on a large-scale lung nodule detection dataset validate the effectiveness of the proposed methods. The AFNet backbone is first proven to be comparable to the previous state of the art (e.g., NoduleNet). We then show that the proposed method with shifted convolutions balances model complexity and performance better than several lightweight methods, and generalizes well with different backbones. As an example, compared to the vanilla model, AFNet with shifted convolutions increases average FROC by 3.08% and reduces FLOPs (floating-point operations) and parameters by 62.40% and 66.62%, respectively.
Xiaohuan Kuang, Kang Yuan, Bo Du 0001, Jiancheng Yang
IJCNN4
2023 Scale-Aware Test-Time Click Adaptation for Pulmonary Nodule and Mass Segmentation
Jiancheng Yang, Yongchao Xu, Li Zhang 0085, Bo Du 0001
MICCAI (3)2
2023 Topology Repairing of Disconnected Pulmonary Airways and Vessels: Baselines and a Dataset
Ziqiao Weng, Jiancheng Yang, Dongnan Liu, Tom Weidong Cai
MICCAI (7)2
2023 Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu
Medical Image Anal.29
2022 ImplicitAtlas: Learning Deformable Shape Templates in Medical Imaging
abstract
Deep implicit shape models have become popular in the computer vision community at large but less so for biomed-ical applications. This is in part because large training databases do not exist and in part because biomedical an-notations are often noisy. In this paper, we show that by introducing templates within the deep learning pipeline we can overcome these problems. The proposed framework, named ImplicitAtlas, represents a shape as a deformation field from a learned template field, where multiple templates could be integrated to improve the shape representation ca-pacity at negligible computational cost. Extensive experi-ments on three medical shape datasets prove the superiority over current implicit representation methods.
Jiancheng Yang, Udaranga Wickramasinghe, Bingbing Ni, Pascal Fua
CVPR1
2022 Representation-Agnostic Shape Fields
Jiancheng Yang, Linguo Li, Teng Li 0001, Bingbing Ni, Wenjun Zhang 0001
ICLR2
2022 What Makes for Automatic Reconstruction of Pulmonary Segments
Kaiming Kuang, Li Zhang 0085, Hongwei Li 0004, Bo Du 0001, Jiancheng Yang
MICCAI (1)7
2022 Weakly Supervised Volumetric Image Segmentation with Deformed Templates
Udaranga Wickramasinghe, Patrick M. Jensen, Mian Shah, Jiancheng Yang, Pascal Fua
MICCAI (5)4
2022 LSSANet: A Long Short Slice-Aware Network for Pulmonary Nodule Detection
Rui Xu 0031, Yong Luo 0002, Bo Du 0001, Kaiming Kuang, Jiancheng Yang
MICCAI (1)5
2022 Neural Annotation Refinement: Development of a New 3D Dataset for Adrenal Gland Analysis
Jiancheng Yang, Udaranga Wickramasinghe, Qikui Zhu, Bingbing Ni, Pascal Fua
MICCAI (4)1
2022 SelfMix: A Self-adaptive Data Augmentation Method for Lesion Segmentation
Qikui Zhu, Jiancheng Yang, Shuo Li 0001
MICCAI (4)4
2021 3D Human Action Representation Learning via Cross-View Consistency Pursuit
abstract
In this work, we propose a Cross-view Contrastive Learning framework for unsupervised 3D skeleton-based action Representation (CrosSCLR), by leveraging multi-view complementary supervision signal. CrosSCLR consists of both single-view contrastive learning (Skeleton-CLR) and cross-view consistent knowledge mining (CVC-KM) modules, integrated in a collaborative learning manner. It is noted that CVC-KM works in such a way that high-confidence positive/negative samples and their distributions are exchanged among views according to their embedding similarity, ensuring cross-view consistency in terms of contrastive context, i.e., similar distributions. Extensive experiments show that CrosSCLR achieves remarkable action recognition results on NTU-60 and NTU-120 datasets under unsupervised settings, with observed higher-quality action representations. Our code is available at https://github.com/LinguoLi/CrosSCLR.
Linguo Li, Minsi Wang, Bingbing Ni, Jiancheng Yang, Wenjun Zhang 0001
CVPR5
2021 Shape Self-Correction for Unsupervised Point Cloud Understanding
abstract
We develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervised manner, we first design a shape-disorganizing module to destroy certain local shape parts of an object. Then the destroyed shape and the normal shape are sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct them to restore the shape to normal. To perform better in these two associated pretext tasks, the network is constrained to capture useful shape features from the object, which indicates that the point cloud network encodes rich geometric and contextual information. The learned feature extractor transfers well to downstream classification and segmentation tasks. Experimental results on ModelNet, ScanNet and ShapeNetPart demonstrate that our method achieves state-of-the-art performance among unsupervised methods. Our framework can be applied to a wide range of deep learning networks for point cloud analysis and we show experimentally that pre-training with our framework significantly boosts the performance of supervised models.
Ye Chen 0006, Jinxian Liu, Bingbing Ni, Jiancheng Yang, Teng Li 0001, Qi Tian 0001
ICCV5
2021 RibSeg Dataset and Strong Point Cloud Baselines for Rib Segmentation from CT Scans
Jiancheng Yang, Shixuan Gu, Donglai Wei 0001, Hanspeter Pfister, Bingbing Ni
MICCAI (1)1
2021 Asymmetric 3D Context Fusion for Universal Lesion Detection
Jiancheng Yang, Kaiming Kuang, Zudi Lin, Hanspeter Pfister, Bingbing Ni
MICCAI (5)1
2021 Decoupled gradient harmonized detector for partial annotation: Application to signet ring cell detection
Tiancheng Lin 0001, Yuanfan Guo, Canqian Yang, Jiancheng Yang, Yi Xu 0001
Neurocomputing4
2021 Reinventing 2D Convolutions for 3D Images
abstract
There have been considerable debates over 2D and 3D representation learning on 3D medical images. 2D approaches could benefit from large-scale 2D pretraining, whereas they are generally weak in capturing large 3D contexts. 3D approaches are natively strong in 3D contexts, however few publicly available 3D medical dataset is large and diverse enough for universal 3D pretraining. Even for hybrid (2D + 3D) approaches, the intrinsic disadvantages within the 2D/3D parts still exist. In this study, we bridge the gap between 2D and 3D convolutions by reinventing the 2D convolutions. We propose ACS (axial-coronal-sagittal) convolutions to perform natively 3D representation learning, while utilizing the pretrained weights on 2D datasets. In ACS convolutions, 2D convolution kernels are split by channel into three parts, and convoluted separately on the three views (axial, coronal and sagittal) of 3D representations. Theoretically, ANY 2D CNN (ResNet, DenseNet, or DeepLab) is able to be converted into a 3D ACS CNN, with pretrained weight of a same parameter size. Extensive experiments validate the consistent superiority of the pretrained ACS CNNs, over the 2D/3D CNN counterparts with/without pretraining. Even without pretraining, the ACS convolution can be used as a plug-and-play replacement of standard 3D convolution, with smaller model size and less computation.
Jiancheng Yang, Jingwei Xu 0005, Canqian Yang, Guozheng Xu, Bingbing Ni
IEEE J. Biomed. Health Informatics1
2020 Deep Kinematics Analysis for Monocular 3D Human Pose Estimation
abstract
For monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxiliary losses in previous researches. Thus it still remains challenging about how to effectively utilize the power of human prior knowledge for this task. In this paper, we propose to address above issue in a systematic view. Firstly, we show that optimizing the kinematics structure of noisy 2D inputs is critical to obtain accurate 3D estimations. Secondly, based on corrected 2D joints, we further explicitly decompose articulated motion with human topology, which leads to more compact 3D static structure easier for estimation. Finally, temporal refinement emphasizing the validity of 3D dynamic structure is naturally developed to pursue more accurate result. Above three steps are seamlessly integrated into deep neural models, which form a deep kinematics analysis pipeline concurrently considering the static/dynamic structure of 2D inputs and 3D outputs. Extensive experiments show that proposed framework achieves state-of-the-art performance on two widely used 3D human action datasets. Meanwhile, targeted ablation study shows that each former step is critical for the latter one to obtain promising results.
Jingwei Xu 0005, Zhenbo Yu, Bingbing Ni, Jiancheng Yang, Xiaokang Yang 0001, Wenjun Zhang 0001
CVPR4
2020 Stochastic Label Refinery: Toward Better Target Label Distribution
abstract
This paper proposes a simple yet effective strategy for improving deep supervised learning, named Stochastic Label Refinery (SLR), by refining training labels to more informative labels. When training a neural network, target distributions (or ground-truth) are typically “hard”, which means the target label of each category consists of only 0 and 1. However, the fixed “hard” target distributions do not capture association between categories or that between objects. In this study, instead of using the hard target distributions, we iteratively generate “soft” target label distributions for training the neural networks, which leads to better performances. The soft target distributions are obtained via an Expectation-Maximization (EM) iteration, where the “true” target distributions and the learned models are regarded as hidden variables. In E step, the models are optimized to approximate the target distributions on stochastic splits of training data; In M step, the target distributions are updated with predicted pseudo-label on leave-out splits. Extensive experiments on classification and ordinal regression tasks, empirically prove that the refined target distribution consistently leads to considerable performance improvements even applied on competitive baselines. Notably, in DeepDR 2020 Diabetic Retinopathy Grading (DeepDRiD) challenge, our method improves the quadratic weighted kappa on official validation set from 0.8247 to 0.8348 and achieves a state-of-the-art score on online test set. The proposed SLR technique is easy to implement and practically applicable.
Jiancheng Yang, Bingbing Ni
ICPR2
2020 Learning Tumor Growth via Follow-Up Volume Prediction for Lung Nodules
Jiancheng Yang, Jingwei Xu 0005, Xiaodan Ye, Guangyu Tao, Xueqian Xie, Guixue Liu
MICCAI (6)2
2020 MIA-Prognosis: A Deep Learning Framework to Predict Therapy Response
Jiancheng Yang, Kaiming Kuang, Tiancheng Lin 0001, Junjun He, Bingbing Ni
MICCAI (2)1
2020 Hierarchical Classification of Pulmonary Lesions: A Large-Scale Radio-Pathomics Study
Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yunlang She, Chang Chen 0010
MICCAI (6)1
2020 AlignShift: Bridging the Gap of Imaging Thickness in 3D Anisotropic Volumes
Jiancheng Yang, Jingwei Xu 0005, Xiaodan Ye, Guangyu Tao, Bingbing Ni
MICCAI (4)1
2020 Learning Black-Box Attackers with Transferable Priors and Query Feedback
abstract
This paper addresses the challenging black-box adversarial attack problem, where only classification confidence of a victim model is available. Inspired by consistency of visual saliency between different vision models, a surrogate model is expected to improve the attack performance via transferability. By combining transferability-based and query-based black-box attack, we propose a surprisingly simple baseline approach (named SimBA++) using the surrogate model, which significantly outperforms several state-of-the-art methods. Moreover, to efficiently utilize the query feedback, we update the surrogate model in a novel learning scheme, named High-Order Gradient Approximation (HOGA). By constructing a high-order gradient computation graph, we update the surrogate model to approximate the victim model in both forward and backward pass. The SimBA++ and HOGA result in Learnable Black-Box Attack (LeBA), which surpasses previous state of the art by considerable margins: the proposed LeBA significantly reduces queries, while keeping higher attack success rates close to 100% in extensive ImageNet experiments, including attacking vision benchmarks and defensive models. Code is open source at https://github.com/TrustworthyDL/LeBA.
Jiancheng Yang, Yangzhou Jiang, Bingbing Ni, Chenglong Zhao
NeurIPS1
2019 Modeling Point Clouds With Self-Attention and Gumbel Subset Sampling
abstract
Geometric deep learning is increasingly important thanks to the popularity of 3D sensors. Inspired by the recent advances in NLP domain, the self-attention transformer is introduced to consume the point clouds. We develop Point Attention Transformers (PATs), using a parameter-efficient Group Shuffle Attention (GSA) to replace the costly Multi-Head Attention. We demonstrate its ability to process size-varying inputs, and prove its permutation equivariance. Besides, prior work uses heuristics dependence on the input data (e.g., Furthest Point Sampling) to hierarchically select subsets of input points. Thereby, we for the first time propose an end-to-end learnable and task-agnostic sampling operation, named Gumbel Subset Sampling (GSS), to select a representative subset of input points. Equipped with Gumbel-Softmax, it produces a "soft" continuous subset in training phase, and a "hard" discrete subset in test phase. By selecting representative subsets in a hierarchical fashion, the networks learn a stronger representation of the input sets with lower computation cost. Experiments on classification and segmentation benchmarks show the effectiveness and efficiency of our methods. Furthermore, we propose a novel application, to process event camera stream as point clouds, and achieve a state-of-the-art performance on DVS128 Gesture Dataset.
Jiancheng Yang, Bingbing Ni, Linguo Li, Jinxian Liu, Mengdie Zhou, Qi Tian 0001
CVPR1
2019 Dynamic Points Agglomeration for Hierarchical Point Sets Learning
abstract
Many previous works on point sets learning achieve excellent performance with hierarchical architecture. Their strategies towards points agglomeration, however, only perform points sampling and grouping in original Euclidean space in a fixed way. These heuristic and task-irrelevant strategies severely limit their ability to adapt to more varied scenarios. To this end, we develop a novel hierarchical point sets learning architecture, with dynamic points agglomeration. By exploiting the relation of points in semantic space, a module based on graph convolution network is designed to learn a soft points cluster agglomeration. We construct a hierarchical architecture that gradually agglomerates points by stacking this learnable and lightweight module. In contrast to fixed points agglomeration strategy, our method can handle more diverse situations robustly and efficiently. Moreover, we propose a parameter sharing scheme for reducing memory usage and computational burden induced by the agglomeration module. Extensive experimental results on several point cloud analytic tasks, including classification and segmentation, well demonstrate the superior performance of our dynamic hierarchical learning framework over current state-of-the-art methods.
Jinxian Liu, Bingbing Ni, Caiyuan Li, Jiancheng Yang, Qi Tian 0001
ICCV4
2019 Probabilistic Radiomics: Ambiguous Diagnosis with Controllable Shape Analysis
Jiancheng Yang, Rongyao Fang, Bingbing Ni, Yi Xu 0001, Linguo Li
MICCAI (6)1