Jieneng Chen

dblp:238/1489 · DBLP profile ↗
← Back
34ranked-venue papers
6as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021
YearPublicationVenuePosition
2026 Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
Zhao-Yang Wang, Jieneng Chen, Yuxiang Guo 0001, Jiang Liu 0014, Rama Chellappa
FG2
2026 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
abstract
Reconstructing animatable 3D animals from videos traditionally depends on sparse semantic keypoints to fit parametric models. Acquiring these keypoints is labor-intensive, and detectors trained on limited animal datasets are often unreliable. We propose 4D-Animal, a keypoint-free framework that reconstructs animatable 3D animals directly from videos. Our method employs a dense feature network to map 2D image representations to SMAL parameters, improving both efficiency and stability. Additionally, we introduce a hierarchical alignment strategy that leverages silhouette, part-level, pixel-level, and temporal cues from pretrained 2D models, ensuring accurate and temporally coherent reconstructions. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines on dog dataset. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal.
Shanshan Zhong, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang 0025, Qihao Liu, Alan L. Yuille, Jieneng Chen
WACV9
2025 SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
abstract
Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial reasoning. This limitation stems from the scarcity of 3D training data and the bias in current model designs toward 2D data. In this paper, we systematically study the impact of 3D-informed data, architecture, and training setups, introducing SpatialLLM, a large multi-modal model with advanced 3D spatial reasoning abilities. To address data limitations, we develop two types of 3D-informed training datasets: (1) 3D-informed probing data focused on object’s 3D location and orientation, and (2) 3D-informed conversation data for complex spatial relationships. Notably, we are the first to curate VQA data that incorporate 3D orientation relationships on real images. Furthermore, we systematically integrate these two types of training data with the architectural and training designs of LMMs, providing a roadmap for optimal design aimed at achieving superior 3D reasoning capabilities. Our SpatialLLM advances machines toward highly capable 3D-informed reasoning, surpassing GPT-4o performance by 8.7%. Our systematic empirical design and the resulting findings offer valuable insights for future research in this direction. Our project page is available at this link.
Wufei Ma, Luoxin Ye, Celso de Melo, Alan L. Yuille, Jieneng Chen
CVPR5
2025 Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Models
abstract
Although large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a framework to comprehensively evaluate 6D spatial reasoning across varying complexities. To address this limitation, we present Spatial457, a scalable and unbiased synthetic dataset designed with 4 key capability for spatial reasoning: multi-object recognition, 2D location, 3D location, and 3D orientation. We develop a cascading evaluation structure, constructing 7 question types across 5 difficulty levels that range from basic single object recognition to our new proposed complex 6D spatial reasoning tasks. We evaluated various large multimodal models (LMMs) on Spatial457, observing a general decline in performance as task complexity increases, particularly in 3D reasoning and 6D spatial tasks. To quantify these challenges, we introduce the Relative Performance Dropping Rate (RPDR), highlighting key weaknesses in 3D reasoning capabilities. Leveraging the unbiased attribute design of our dataset, we also uncover prediction biases across different attributes, with similar patterns observed in real-world image settings.1The code is released in https://github.com/XingruiWang/Spatial457.
Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso de Melo, Jieneng Chen, Alan L. Yuille
CVPR5
2025 UniGait: A Unified Transformer-based Multitask Framework for Gait Analysis in the Wild
abstract
Gait recognition is a rapidly emerging and significant area of biometrics, leveraging the unique walking patterns of individuals to perform personal identification and facilitate healthcare monitoring, such as elderly care, fall detection, etc. While existing gait recognition methods perform well in indoor, or short-range environments, their effectiveness diminishes significantly when applied to unconstrained outdoor scenarios. Challenges such as environmental turbulence, occlusion, varying viewing angles contribute to this performance drop. To address these challenges and enhance gait recognition accuracy in real-world settings, while also expanding the functionality of gait features for healthcare applications, we propose a unified multitask framework called UniGait. UniGait is designed to perform a comprehensive range of gait analysis tasks, including gait recognition and estimation of gait-related human attributes. UniGait is built upon a transformer-based architecture, which leverages the power of a cross-attention mechanism to simultaneously process multiple sub-tasks. This multitask learning approach allows the model to extract more robust gait features by jointly learning gait recognition and human attribute estimation, leading to improved overall performance. We report the results of extensive experiments and analysis on large-scale, real-world datasets collected under challenging conditions, including long-range (up to 1000 meters) and high-pitch angles (including UAV-based data). The results demonstrated state-of-the-art performance, highlighting the potential of UniGait for deployment in real-world applications, making it a valuable tool for a range of biometric and healthcare monitoring scenarios.
Zhao-Yang Wang, Jiang Liu 0014, Yuxiang Guo 0001, Jieneng Chen, Rama Chellappa
FG4
2025 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmark
abstract
3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/.
Wufei Ma, Guofeng Zhang 0025, Yu-Cheng Chou, Jieneng Chen, Celso de Melo, Alan L. Yuille
ICCV5
2025 Medical World Model
Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang 0016, Rama Chellappa, Zongwei Zhou, Alan L. Yuille, Lei Zhu 0003, Jieneng Chen
ICCV11
2025 GenEx: Generating an Explorable World
abstract
Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step toward this goal by introducing *GenEx*, a system capable of planning complex embodied world exploration, guided by its generative imagination that forms expectations about the surrounding environments. *GenEx* generates high-quality, continuous 360-degree virtual environments, achieving robust loop consistency and active 3D mapping over extended trajectories. Leveraging generative imagination, GPT-assisted agents can undertake complex embodied tasks, including goal-agnostic exploration and goal-driven navigation. Agents utilize imagined observations to update their beliefs, simulate potential outcomes, and enhance their decision-making. Training on the synthetic urban dataset *GenEx-DB* and evaluation on *GenEx-EQA* demonstrate that our approach significantly improves agents' planning capabilities, providing a transformative platform toward intelligent, imaginative embodied exploration.
Taiming Lu, Tianmin Shu, Alan L. Yuille, Daniel Khashabi, Jieneng Chen
ICLR5
2025 Learning Segmentation from Radiology Reports
Pedro R. A. S. Bassi, Jieneng Chen, Zheren Zhu, Sergio Decherchi, Andrea Cavalli, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou
MICCAI (5)3
2025 Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models
abstract
Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-quality image-text pairs, requiring millions of GPU hours. This paper introduces the Vision-Language-Vision **(VLV)** auto-encoder framework, which strategically leverages key pretrained components: a vision encoder, the decoder of a Text-to-Image (T2I) diffusion model, and subsequently, a Large Language Model (LLM). Specifically, we establish an information bottleneck by regularizing the language representation space, achieved through freezing the pretrained T2I diffusion decoder. Our VLV pipeline effectively distills knowledge from the text-conditioned diffusion model using continuous embeddings, demonstrating comprehensive semantic understanding via high-quality reconstructions. Furthermore, by fine-tuning a pretrained LLM to decode the intermediate language representations into detailed descriptions, we construct a state-of-the-art (SoTA) captioner comparable to leading models like GPT-4o and Gemini 2.0 Flash. Our method demonstrates exceptional cost-efficiency and significantly reduces data requirements; by primarily utilizing single-modal images for training and maximizing the utility of existing pretrained models (image encoder, T2I diffusion model, and LLM), it circumvents the need for massive paired image-text datasets, keeping the total training expenditure under $1,000 USD.
Tiezheng Zhang, Yu-Cheng Chou, Jieneng Chen, Alan L. Yuille, Junfei Xiao
NeurIPS4
2025 VM-Gait: Multi-Modal 3D Representation Based on Virtual Marker for Gait Recognition
abstract
Gait recognition plays a vital role in biometric applications by analyzing the unique characteristics of an individ-ual's walking pattern. Methods based on 2D representations, such as silhouettes and skeletons, are increasingly being developed to learn the shape features and joint dy-namic movements. Nevertheless, the effectiveness of 2D representation-based methods is impeded by factors such as changes in viewpoint, partial occlusion, and noisy en-vironments. 3D representation-based methods can complement 2D representation-based approaches by providing more precise dynamic body shapes and motion information, along with increased robustness against changes in viewpoint and partial occlusion. However, the complex-ity of acquiring accurate 3D representations and the chal-lenges associated with extracting dynamic topological features from sequences of 3D representations hinder the de-velopment of 3D representations-based methods. In this pa-per, we present VM-Gait, a novel multi-modal gait recognition framework that harnesses the advantages of integrating both 2D and 3D representations. Furthermore, we in-troduce a new 3D representation, Virtual Marker, into gait recognition to efficiently learn topological features from 3D representations, avoiding the computational complexi-ties inherent in directly learning from 3D representations like 3D meshes or 3D point clouds. Extensive experiments demonstrate that the proposed framework effectively learns and fuses discriminative information from different gait modalities, enhancing gait recognition performance.
Zhao-Yang Wang, Jiang Liu 0014, Jieneng Chen, Rama Chellappa
WACV3
2024 ViTamin: Designing Scalable Vision Models in the Vision-Language Era
abstract
Recent breakthroughs in vision-language models (VLMs) start a new page in the vision community. The VLMs provide stronger and more generalizable feature embeddings compared to those from ImageNet-pretrained models, thanks to the training on the large-scale Internet image-text pairs. However, despite the amazing achievement from the VLMs, vanilla Vision Transformers (ViTs) remain the default choice for the image encoder. Although pure transformer proves its effectiveness in the text encoding area, it remains questionable whether it is also the case for image encoding, especially considering that various types of networks are proposed on the ImageNet benchmark, which, unfortunately, are rarely studied in VLMs. Due to small data/model scale, the original conclusions of model design on ImageNet can be limited and biased. In this paper, we aim at building an evaluation protocol of vision models in the vision-language era under the contrastive language-image pretraining (CLIP) framework. We provide a comprehensive way to benchmark different vision models, covering their zero-shot performance and scalability in both model and training data sizes. To this end, we introduce ViTamin, a new vision models tailored for VLMs. ViTamin-L significantly outperforms ViT-L by 2.0% ImageNet zero-shot accuracy, when using the same publicly available DataComp-1B dataset and the same OpenCLIP training scheme. ViTamin-L presents promising results on 60 diverse benchmarks, including classification, retrieval, open-vocabulary detection and segmentation, and large multi-modal models. When further scaling up the model size, our ViTamin-XL with only 436M parameters attains 82.9% ImageNet zero-shot accuracy, surpassing 82.0% achieved by EVA-E that has ten times more parameters (4.4B).
Jieneng Chen, Qihang Yu, Xiaohui Shen, Alan L. Yuille, Liang-Chieh Chen
CVPR1
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS6
2024 Efficient Large Multi-modal Models via Visual Context Compression
abstract
While significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLMs (MLLMs) has remained a largely overlooked area. In this work, we present the study on the analysis of redundancy concerning visual tokens and efficient training within these models. Our initial experiments show that eliminating up to 70% of visual tokens at the testing stage by simply average pooling only leads to a minimal 3% reduction in visual question answering accuracy on the GQA benchmark, indicating significant redundancy in visual context. Addressing this, we introduce Visual Context Compressor, which reduces the number of visual tokens to enhance training and inference efficiency without sacrificing performance. To minimize information loss caused by the compression on visual tokens while maintaining training efficiency, we develop LLaVolta as a light and staged training scheme that incorporates stage-wise visual context compression to progressively compress the visual tokens from heavily to lightly compression during training, yielding no loss of information when testing. Extensive experiments demonstrate that our approach enhances the performance of MLLMs in both image-language and video-language understanding, while also significantly cutting training costs and improving inference efficiency.
Jieneng Chen, Luoxin Ye, Ju He, Daniel Khashabi, Alan L. Yuille
NeurIPS1
2024 TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers
abstract
Medical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively.
Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou
Medical Image Anal.1
2024 SpecTr: Spectral Transformer for Microscopic Hyperspectral Pathology Image Segmentation
abstract
Hyperspectral imaging (HSI) unlocks the huge potential to a wide variety of applications relying on high-precision pathology image segmentation, such as computational pathology. It can acquire biochemical properties even invisible to naked eyes from histological specimens. Since 1) spectra contain discriminative and continuous patterns for differentiating tissues/cells, and 2) the discriminability of spectra relies on both fine-grained relations in the high-resolution spectrum and coarse relations in the low-resolution spectrum, the key to achieving high-precision hyperspectral pathology image segmentation is to felicitously model the intra- and inter-scale context especially for spectra. In this paper, we propose a spectral transformer (SpecTr) for hyperspectral pathology image segmentation, which first captures global context for intra-scale spectral features, and subsequently extract coarse and fine-grained discriminative spectral information from inter-scale features, respectively. To learn intra-scale spectral context, we propose a Spectral Attentive Module (SAM). Unlike the existing Transformer model that is designed for modalities such as natural images, our proposed SAM is efficient in capturing sparse and pivotal spectral context while avoiding the heterogeneous underlying distributions and noises of different bands. Besides, to reduce the computational complexity of the HSI segmentation model, we further propose a global-local attention module to effectively learn a condensed spectral feature. Experiments show that HSIs can become a more powerful image modality for understanding microscopic pathology images than RGB images, and the proposed SpecTr outperforms other competing methods for hyperspectral pathology image segmentation, with an improvement of 3% compared with the popular 3D-nnUNet and other transformer-based methods. Our code is available at https://github.com/DeepMed-Lab-ECNU/SpecTr.
Boxiang Yun, Bai Ying Lei, Jieneng Chen, Song Qiu, Wei Shen 0002, Qingli Li, Yan Wang 0033
IEEE Trans. Circuits Syst. Video Technol.3
2023 Compositor: Bottom-Up Clustering and Compositing for Robust Part and Object Segmentation
abstract
In this work, we present a robust approach for joint part and object segmentation. Specifically, we reformulate object and part segmentation as an optimization problem and build a hierarchical feature representation including pixel, part, and object-level embeddings to solve it in a bottom-up clustering manner. Pixels are grouped into several clusters where the part-level embeddings serve as cluster centers. Afterwards, object masks are obtained by compositing the part proposals. This bottom-up interaction is shown to be effective in integrating information from lower semantic levels to higher semantic levels. Based on that, our novel approach Compositor produces part and object segmentation masks simultaneously while improving the mask quality. Compositor achieves state-of-the-art performance on PartImageNet and Pascal-Part by outperforming previous methods by around 0.9% and 1.3% on PartImageNet, 0.4% and 1.7% on Pascal-Part in terms of part and object mIoU and demonstrates better robustness against occlusion by around 4.4% and 7.1% on part and object respectively.
Ju He, Jieneng Chen, Ming-Xian Lin, Qihang Yu, Alan L. Yuille
CVPR2
2023 Label-Free Liver Tumor Segmentation
abstract
We demonstrate that AI models can accurately segment liver tumors without the need for manual annotation by using synthetic tumors in CT scans. Our synthetic tumors have two intriguing advantages: (I) realistic in shape and texture, which even medical professionals can confuse with real tumors; (II) effective for training AI models, which can perform liver tumor segmentation similarly to the model trained on real tumors—this result is exciting because no existing work, using synthetic tumors only, has thus far reached a similar or even close performance to real tumors. This result also implies that manual efforts for annotating tumors voxel by voxel (which took years to create) can be significantly reduced in the future. Moreover, our synthetic tumors can automatically generate many examples of small (or even tiny) synthetic tumors and have the potential to im-prove the success rate of detecting small liver tumors, which is critical for detecting the early stages of cancer. In addition to enriching the training data, our synthesizing strategy also enables us to rigorously assess the AI robustness.
Qixin Hu, Yixiong Chen, Junfei Xiao, Shuwen Sun, Jieneng Chen, Alan L. Yuille, Zongwei Zhou
CVPR5
2023 CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT Scans
abstract
Human readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI’s clinical adoption. A certain number of AI models need to be assembled nontrivially to match the diagnostic process of a human reading a CT scan. In this paper, we construct a Unified Tumor Transformer (CancerUniT) model to jointly detect tumor existence & location and diagnose tumor characteristics for eight major cancers in CT scans. CancerUniT is a query-based Mask Transformer model with the output of multi-tumor prediction. We decouple the object queries into organ queries, tumor detection queries and tumor diagnosis queries, and further establish hierarchical relationships among the three groups. This clinically-inspired architecture effectively assists inter- and intra-organ representation learning of tumors and facilitates the resolution of these complex, anatomically related multi-organ cancer image reading tasks. CancerUniT is trained end-to-end using a curated large-scale CT images of 10,042 patients including eight major types of cancers and occurring non-cancer tumors (all are pathology-confirmed with 3D tumor masks annotated by radiologists). On the test set of 631 patients, CancerUniT has demonstrated strong performance under a set of clinically relevant evaluation metrics, substantially outperforming both multi-disease methods and an assembly of eight single-organ expert models in tumor detection, segmentation, and diagnosis. This moves one step closer towards a universal high performance cancer screening tool.
Jieneng Chen, Yingda Xia, Jiawen Yao, Ke Yan 0006, Le Lu 0001, Fakai Wang, Bo Zhou 0009, Mingyan Qiu, Qihang Yu, Mingze Yuan, Wei Fang 0005, Yuxing Tang, Minfeng Xu, Xianghua Ye, Xiaoli Yin, Xin Chen 0058, Jingren Zhou 0001, Alan L. Yuille, Zaiyi Liu, Ling Zhang 0002
ICCV1
2023 CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection
abstract
An increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited to segmenting specific organs/tumors and ignore the semantics of anatomical structures, nor can they be extended to novel domains. To address these issues, we propose the CLIP-Driven Universal Model, which incorporates text embedding learned from Contrastive Language-Image Pre-training (CLIP) to segmentation models. This CLIP-based label encoding captures anatomical relationships, enabling the model to learn a structured feature embedding and segment 25 organs and 6 types of tumors. The proposed model is developed from an assembly of 14 datasets, using a total of 3,410 CT scans for training and then evaluated on 6,162 external CT scans from 3 additional datasets. We rank first on the Medical Segmentation Decathlon (MSD) public leaderboard and achieve state-of-the-art results on Beyond The Cranial Vault (BTCV). Additionally, the Universal Model is computationally more efficient (6× faster) compared with dataset-specific models, generalized better to CT scans from varying sites, and shows stronger transfer learning performance on novel tasks.
Jie Liu 0044, Yixiao Zhang 0001, Jieneng Chen, Junfei Xiao, Yongyi Lu, Bennett A. Landman, Yixuan Yuan, Alan L. Yuille, Yucheng Tang, Zongwei Zhou
ICCV3
2022 TransFG: A Transformer Architecture for Fine-Grained Recognition
abstract
Fine-grained visual classification (FGVC) which aims at recognizing objects from subcategories is a very challenging task due to the inherently subtle inter-class differences. Most existing works mainly tackle this problem by reusing the backbone network to extract features of detected discriminative regions. However, this strategy inevitably complicates the pipeline and pushes the proposed regions to contain most parts of the objects thus fails to locate the really important parts. Recently, vision transformer (ViT) shows its strong performance in the traditional classification task. The self-attention mechanism of the transformer links every patch token to the classification token. In this work, we first evaluate the effectiveness of the ViT framework in the fine-grained recognition setting. Then motivated by the strength of the attention link can be intuitively considered as an indicator of the importance of tokens, we further propose a novel Part Selection Module that can be applied to most of the transformer architectures where we integrate all raw attention weights of the transformer into an attention map for guiding the network to effectively and accurately select discriminative image patches and compute their relations. A contrastive loss is applied to enlarge the distance between feature representations of confusing classes. We name the augmented transformer-based model TransFG and demonstrate the value of it by conducting experiments on five popular fine-grained benchmarks where we achieve state-of-the-art performance. Qualitative results are presented for better understanding of our model.
Ju He, Jieneng Chen, Adam Kortylewski, Yutong Bai, Changhu Wang
AAAI2
2022 TransMix: Attend to Mix for Vision Transformers
abstract
Mixup-based augmentation has been found to be effective for generalizing models during training, especially for Vision Transformers (ViTs) since they can easily overfit. However, previous mixup-based methods have an underlying prior knowledge that the linearly interpolated ratio of targets should be kept the same as the ratio proposed in input interpolation. This may lead to a strange phenomenon that sometimes there is no valid object in the mixed image due to the random process in augmentation but there is still response in the label space. To bridge such gap between the input and label spaces, we propose TransMix, which mixes labels based on the attention maps of Vision Transformers. The confidence of the label will be larger if the corresponding input image is weighted higher by the attention map. TransMix is embarrassingly simple and can be implemented in just a few lines of code without introducing any extra parameters and FLOPs to ViT-based models. Experimental results show that our method can consistently improve various ViT-based models at scales on ImageNet classification. After pre-trained with TransMix on ImageNet, the ViT-based models also demonstrate better transferability to semantic segmentation, object detection and instance segmentation. TransMix also exhibits to be more robust when evaluating on 4 different benchmarks. Code is publicly available at https://github.com/Beckschen/TransMix.
Jieneng Chen, Shuyang Sun, Ju He, Philip Torr 0001, Alan L. Yuille, Song Bai 0001
CVPR1
2022 PartImageNet: A Large, High-Quality Dataset of Parts
Ju He, Shaokang Yang, Adam Kortylewski, Xiaoding Yuan, Jieneng Chen, Qihang Yu, Alan L. Yuille
ECCV (8)6
2022 WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
Xiangde Luo, Wenjun Liao, Jianghong Xiao, Jieneng Chen, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Dimitris N. Metaxas, Guotai Wang, Shaoting Zhang 0001
Medical Image Anal.4
2022 SCPM-Net: An anchor-free 3D lung nodule detection network using sphere representation and center points matching
Xiangde Luo, Tao Song 0002, Guotai Wang, Jieneng Chen, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.4
2022 Semi-supervised medical image segmentation via uncertainty rectified pyramid consistency
Xiangde Luo, Guotai Wang, Wenjun Liao, Jieneng Chen, Tao Song 0002, Shichuan Zhang, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.4
2022 NeuroIV: Neuromorphic Vision Meets Intelligent Vehicle Towards Safe Driving With a New Database and Baseline Evaluations
abstract
Neuromorphic vision sensors such as the Dynamic and Active-pixel Vision Sensor (DAVIS) using silicon retina are inspired by biological vision, they generate streams of asynchronous events to indicate local log-intensity brightness changes. Their properties of high temporal resolution, low-bandwidth, lightweight computation, and low-latency make them a good fit for many applications of motion perception in the intelligent vehicle. However, as a younger and smaller research field compared to classical computer vision, neuromorphic vision is rarely connected with the intelligent vehicle. For this purpose, we present three novel datasets recorded with DAVIS sensors and depth sensor for the distracted driving research and focus on driver drowsiness detection, driver gaze-zone recognition, and driver hand-gesture recognition. To facilitate the comparison with classical computer vision, we record the RGB, depth and infrared data with a depth sensor simultaneously. The total volume of this dataset has 27360 samples. To unlock the potential of neuromorphic vision on the intelligent vehicle, we utilize three popular event-encoding methods to convert asynchronous event slices to event-frames and adapt state-of-the-art convolutional architectures to extensively evaluate their performances on this dataset. Together with qualitative and quantitative results, this work provides a new database and baseline evaluations named NeuroIV in cross-cutting areas of neuromorphic vision and intelligent vehicle.
Guang Chen 0001, Fa Wang, Lin Hong, Jörg Conradt, Jieneng Chen, Zhenyan Zhang, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.6
2021 Semi-supervised Medical Image Segmentation through Dual-task Consistency
abstract
Deep learning-based semi-supervised learning (SSL) algorithms have led to promising results in medical images segmentation and can alleviate doctors' expensive annotations by leveraging unlabeled data. However, most of the existing SSL algorithms in literature tend to regularize the model training by perturbing networks and/or data. Observing that multi/dual-task learning attends to various levels of information which have inherent prediction perturbation, we ask the question in this work: can we explicitly build task-level regularization rather than implicitly constructing networks- and/or data-level perturbation and then regularization for SSL? To answer this question, we propose a novel dual-task-consistency semi-supervised framework for the first time. Concretely, we use a dual-task deep network that jointly predicts a pixel-wise segmentation map and a geometry-aware level set representation of the target. The level set representation is converted to an approximated segmentation map through a differentiable task transform layer. Simultaneously, we introduce a dual-task consistency regularization between the level set-derived segmentation maps and directly predicted segmentation maps for both labeled and unlabeled data. Extensive experiments on two public datasets show that our method can largely improve the performance by incorporating the unlabeled data. Meanwhile, our framework outperforms the state-of-the-art semi-supervised learning methods.
Xiangde Luo, Jieneng Chen, Tao Song 0002, Guotai Wang
AAAI2
2021 Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
Jieneng Chen, Ke Yan 0006, Youbao Tang, Shuwen Sun, Qiuping Liu, Lingyun Huang, Jing Xiao 0006, Alan L. Yuille, Ya Zhang 0002, Le Lu 0001
MICCAI (7)1
2021 Fully Test-Time Adaptation for Image Segmentation
Minhao Hu, Tao Song 0002, Yujun Gu, Xiangde Luo, Jieneng Chen, Ya Zhang 0002, Shaoting Zhang 0001
MICCAI (3)5
2021 Efficient Semi-supervised Gross Target Volume of Nasopharyngeal Carcinoma Segmentation via Uncertainty Rectified Pyramid Consistency
Xiangde Luo, Wenjun Liao, Jieneng Chen, Tao Song 0002, Shichuan Zhang, Nianyong Chen, Guotai Wang, Shaoting Zhang 0001
MICCAI (2)3
2020 Deep Distance Transform for Tubular Structure Segmentation in CT Scans
abstract
Tubular structure segmentation in medical images, e.g., segmenting vessels in CT scans, serves as a vital step in the use of computers to aid in screening early stages of related diseases. But automatic tubular structure segmentation in CT scans is a challenging problem, due to issues such as poor contrast, noise and complicated background. A tubular structure usually has a cylinder-like shape which can be well represented by its skeleton and cross-sectional radii (scales). Inspired by this, we propose a geometry-aware tubular structure segmentation method, Deep Distance Transform (DDT), which combines intuitions from the classical distance transform for skeletonization and modern deep segmentation networks. DDT first learns a multi-task network to predict a segmentation mask for a tubular structure and a distance map. Each value in the map represents the distance from each tubular structure voxel to the tubular structure surface. Then the segmentation mask is refined by leveraging the shape prior reconstructed from the distance map. We apply our DDT on six medical image datasets. Results show that (1) DDT can boost tubular structure segmentation performance significantly (e.g., over 13% DSC improvement for pancreatic duct segmentation), and (2) DDT additionally provides a geometrical measurement for a tubular structure, which is important for clinical diagnosis (e.g., the cross-sectional scale of a pancreatic duct can be an indicator for pancreatic cancer).
Yan Wang 0033, Fengze Liu, Jieneng Chen, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille
CVPR4
2020 CPM-Net: A 3D Center-Points Matching Network for Pulmonary Nodule Detection in CT Scans
Tao Song 0002, Jieneng Chen, Xiangde Luo, Yechong Huang, Xinglong Liu, Zhaoxiang Ye, Huaqiang Sheng, Shaoting Zhang 0001, Guotai Wang
MICCAI (6)2
2020 Coarse-to-Fine Adversarial Networks and Zone-Based Uncertainty Analysis for NK/T-Cell Lymphoma Segmentation in CT/PET Images
abstract
Extranodal natural killer/T cell lymphoma (ENKL), nasal type is a kind of rare disease with a low survival rate that primarily affects Asian and South American populations. Segmentation of ENKL lesions is crucial for clinical decision support and treatment planning. This paper is the first study on computer-aided diagnosis systems for the ENKL segmentation problem. We propose an automatic, coarse-to-fine approach for ENKL segmentation using adversarial networks. In the coarse stage, we extract the region of interest bounding the lesions utilizing a segmentation neural network. In the fine stage, we use an adversarial segmentation network and further introduce a multi-scale L1loss function to drive the network to learn both global and local features. The generator and discriminator are alternately trained by backpropagation in an adversarial fashion in a min-max game. Furthermore, we present the first exploration of zone-based uncertainty estimates based on Monte Carlo dropout technique in the context of deep networks for medical image segmentation. Specifically, we propose the uncertainty criteria based on the lesion and the background, and then linearly normalize them to a specific interval. This is not only the crucial criterion for evaluating the superiority of the algorithm, but also permits subsequent optimization by engineers and revision by clinicians after quantitatively understanding the main source of uncertainty from the background or the lesion zone. Experimental results demonstrate that the proposed method is more effective and lesion-zone stable than state-of-the-art deep-learning based segmentation model.
Xiaobin Hu, Jieneng Chen, Hongwei Li 0004, Diana Waldmannstetter, Yu Zhao 0009, Kuangyu Shi, Bjoern Menze
IEEE J. Biomed. Health Informatics3