Linshan Wu

dblp:318/8264 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 MG-3D: Multi-grained knowledge-enhanced vision-language pre-training for 3D medical image analysis
Xuefeng Ni, Linshan Wu, Jiaxin Zhuang, Qiong Wang 0001, Mingxiang Wu, Varut Vardhanabhuti, Lihai Zhang, Hanyu Gao, Hao Chen 0011
Medical Image Anal.2
2026 Large-Scale 3D Medical Image Pre-Training With Geometric Context Priors
abstract
The scarcity of annotations poses a significant challenge in medical image analysis, which demands extensive efforts from radiologists, especially for high-dimension 3D medical images. Large-scale pre-training has emerged as a promising label-efficient solution, owing to the utilization of large-scale data, large models, and advanced pre-training techniques. However, its development in medical images remains underexplored. The primary challenge lies in harnessing large-scale unlabeled data and learning high-level semantics without annotations. We observe that 3D medical images exhibit consistent geometric context, i.e., consistent geometric relations between different organs, which leads to a promising way for learning consistent representations. Motivated by this, we introduce a simple-yet-effective Volume Contrast (VoCo) framework to leverage geometric context priors for self-supervision. Given an input volume, we extract base crops from different regions to construct positive and negative pairs for contrastive learning. Then we predict the contextual position of a random crop by contrasting its similarity to the base crops. In this way, VoCo implicitly encodes the inherent geometric context into model representations, facilitating high-level semantic learning without annotations. To assess effectiveness, we (1) introduce PreCT-160 K, the largest medical image pre-training dataset to date, which comprises 160 K Computed Tomography (CT) volumes covering diverse anatomic structures; (2) investigate scaling laws and propose guidelines for tailoring different model sizes to various medical tasks; (3) build a comprehensive benchmark encompassing 51 medical tasks, including segmentation, classification, registration, and vision-language. Extensive experiments highlight the superiority of VoCo, showcasing promising transferability to unseen modalities and datasets. VoCo notably enhances performance on datasets with limited labeled cases and significantly expedites fine-tuning convergence.
Linshan Wu, Jiaxin Zhuang, Hao Chen 0011
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Bio2Vol: Adapting 2D Biomedical Foundation Models for Volumetric Medical Image Segmentation
Jiaxin Zhuang, Linshan Wu, Xuefeng Ni, Xi Wang 0013, Liansheng Wang 0002, Hao Chen 0011
MICCAI (6)2
2025 Modeling the Label Distributions for Weakly-Supervised Semantic Segmentation
abstract
Weakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models by weak labels, which is receiving significant attention due to its low annotation cost. Existing approaches focus on generating pseudo labels for supervision while largely ignoring to leverage the inherent semantic correlation among different pseudo labels. We observe that pseudo-labeled pixels that are close to each other in the feature space are more likely to share the same class, and those closer to the distribution centers tend to have higher confidence. Motivated by this, we propose to model the underlying label distributions and employ cross-label constraints to generate more accurate pseudo labels. In this paper, we develop a unified WSSS framework named Adaptive Gaussian Mixtures Model, which leverages a GMM to model the label distributions. Specifically, we calculate the feature distribution centers of pseudo-labeled pixels and build the GMM by measuring the distance between the centers and each pseudo-labeled pixel. Then, we introduce an Online Expectation-Maximization (OEM) algorithm and a novel maximization loss to optimize the GMM adaptively, aiming to learn more discriminative decision boundaries between different class-wise Gaussian mixtures. Based on the label distributions, we leverage the GMM to generate high-quality pseudo labels for more reliable supervision. Our framework is capable of solving different forms of weak labels: image-level labels, points, scribbles, blocks, and bounding-boxes. Extensive experiments on PASCAL, COCO, Cityscapes, and ADE20 K datasets demonstrate that our framework can effectively provide more reliable supervision and outperform the state-of-the-art methods under all settings.
Linshan Wu, Zhun Zhong, Jiayi Ma 0001, Yunchao Wei, Hao Chen 0011, Leyuan Fang, Shutao Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Exploring Spatio-Temporal Carbon Emission Across Passenger Car Trajectory Data
abstract
Carbon emissions caused by passenger cars in cities are essentially responsible for severe climate change and serious environmental problems. Exploring carbon emissions from passenger cars helps to control urban pollution and achieve urban sustainability. However, it is a challenging task to foresee the spatio-temporal distribution of carbon emission from passenger cars, as the following technical issues remain. i) Vehicle carbon emissions contain complex spatial interactions and temporal dynamics. How to collaboratively integrate such spatial-temporal correlations for carbon emission prediction is not yet resolved. ii) Given the mobility of passenger cars, the hidden dependencies inherent in traffic density are not properly addressed in predicting carbon emissions from passenger cars. To tackle these issues, we propose a Collaborative Spatial-temporal Network (CSTNet) for implementing carbon emissions prediction by using passenger car trajectory data. Within the proposed method, we devote to extract collaborative properties that stem from a multi-view graph structure together with parallel input of carbon emission and traffic density. Then, we design a spatial-temporal convolutional block for both carbon emission and traffic density, which constitutes of temporal gate convolution, spatial convolution and temporal attention mechanism. Following that, an interaction layer between carbon emission and traffic density is proposed to handle their internal dependencies, and further model spatial relationships between the features. Besides, we identify several global factors and embed them for final prediction with a collaborative fusion. Experimental results on the real-world passenger car trajectory dataset demonstrate that the proposed method outperforms the baselines with a roughly 7%-11% improvement.
Zhu Xiao, Bo Liu 0104, Linshan Wu, Hongbo Jiang 0001, Beihao Xia, Tao Li 0056, Cassandra C. Wang
IEEE Trans. Intell. Transp. Syst.3
2025 MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image Analysis
abstract
The Vision Transformer (ViT) has demonstrated remarkable performance in Self-Supervised Learning (SSL) for 3D medical image analysis. Masked AutoEncoder (MAE) for feature pre-training can further unleash the potential of ViT on various medical vision tasks. However, due to large spatial sizes with much higher dimensions of 3D medical images, the lack of hierarchical design for MAE may hinder the performance of downstream tasks. In this paper, we propose a novel Mask in Mask (MiM) pre-training framework for 3D medical images, which aims to advance MAE by learning discriminative representation from hierarchical visual tokens across varying scales. We introduce multiple levels of granularity for masked inputs from the volume, which are then reconstructed simultaneously ranging at both fine and coarse levels. Additionally, a cross-level alignment mechanism is applied to adjacent level volumes to enforce anatomical similarity hierarchically. Furthermore, we adopt a hybrid backbone to enhance the hierarchical representation learning efficiently during the pre-training. MiM was pre-trained on a large scale of available 3D volumetric images, i.e., Computed Tomography (CT) images containing various body parts. Extensive experiments on twelve public datasets demonstrate the superiority of MiM over other SSL methods in organ/tumor segmentation and disease classification. We further scale up the MiM to large pre-training datasets with more than 10k volumes, showing that large-scale pre-training can further enhance the performance of downstream tasks. Code is available at https://github.com/JiaxinZhuang/MiM.
Jiaxin Zhuang, Linshan Wu, Qiong Wang 0001, Peng Fei, Varut Vardhanabhuti, Hao Chen 0011
IEEE Trans. Medical Imaging2
2024 VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis
abstract
Self-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We ob-serve that 3D medical images contain relatively consistent contextual position information, i.e., consistent geometric relations between different organs, which leads to a potential way for us to learn consistent semantic representations in pre-training. In this paper, we propose a simple-yet-effective Volume Contrast (VoCo) framework to leverage the contextual position priors for pre-training. Specif-ically, we first generate a group of base crops from different regions while enforcing feature discrepancy among them, where we employ them as class assignments of dif-ferent regions. Then, we randomly crop sub-volumes and predict them belonging to which class (located at which re-gion) by contrasting their similarity to different base crops, which can be seen as predicting contextual positions of different sub-volumes. Through this pretext task, VoCo implic-itly encodes the contextual position priors into model rep-resentations without the guidance of annotations, enabling us to effectively improve the performance of downstream tasks that require high-level semantics. Extensive exper-imental results on six downstream tasks demonstrate the superior effectiveness of VoCo. Code will be available at httpsu/github.com/luffytls/vo'Co.
Linshan Wu, Jiaxin Zhuang, Hao Chen 0011
CVPR1
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS44
2024 Exploring Intercity Mobility in Urban Agglomeration: Evidence from Private Car Trajectory Data
abstract
In this article, we explore intercity mobility in urban agglomerations by surveying people traveling across cities based on private car trajectory data. Specifically, we first adopt the statistical analysis method to mine the intercity mobility in terms of various metrics of travel trips, so as to gain a preliminary understanding of intercity mobility in urban agglomeration. Then, we utilize the tensor decomposition method to conduct in-depth study on the intercity mobility pattern from the perspectives of complexity and multidimensionality. We construct a 4-D tensor based on private car trajectory and point-of-interest (POI) datasets and define the functional similarity and geographic adjacency between regions. Finally, we design an alternating proximal gradient (APG)-based method to resolve the core tensor and factor matrix, leading to the fine-grained discovery of intercity mobility patterns on administrative divisions in the urban agglomeration. Extensive experiments are conducted to evaluate the analysis of intercity mobility, using a real-world dataset containing one-year private car trajectories from five cities in the selected urban agglomeration. The experiments show that the proposed method successfully captures 20 intercity mobility patterns, in which the factor matrices retrieve the patterns from different dimensions with core tensors characterizing correlations between patterns in factor matrices. Besides, the extracted intercity mobility patterns not only cover administrative areas with frequent intercity interactions, but also contain areas with less intercity interactions. It validates that the intercity mobility is consistent with the regional functions in urban agglomeration.
Zhu Xiao, Linshan Wu, Hongbo Jiang 0001, Zheng Qin 0001, Chengxi Gao, Hongyang Chen 0001, Jiangchuan Liu
IEEE Trans. Comput. Soc. Syst.2
2023 Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures
abstract
Sparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent relation between labeled and unlabeled pixels. In this paper, we observe that pixels that are close to each other in the feature space are more likely to share the same class. Inspired by this, we propose a novel SASS framework, which is equipped with an Adaptive Gaussian Mixture Model (AGMM). Our AGMM can effectively endow reliable supervision for unlabeled pixels based on the distributions of labeled and unlabeled pixels. Specifically, we first build Gaussian mixtures using labeled pixels and their relatively similar unlabeled pixels, where the labeled pixels act as centroids, for modeling the feature distribution of each class. Then, we leverage the reliable information from labeled pixels and adaptively generated GMM predictions to supervise the training of unlabeled pixels, achieving online, dynamic, and robust selfsupervision. In addition, by capturing category-wise Gaussian mixtures, AGMM encourages the model to learn discriminative class decision boundaries in an end-to-end contrastive learning manner. Experimental results conducted on the PASCAL VOC 2012 and Cityscapes datasets demonstrate that our AGMM can establish new state-of-the-art SASS performance. Code is available at https://github.com/Luffy03/AGMM-SASS
Linshan Wu, Zhun Zhong, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Hao Chen 0011
CVPR1
2023 Querying Labeled for Unlabeled: Cross-Image Semantic Consistency Guided Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation aims to learn a semantic segmentation model via limited labeled images and adequate unlabeled images. The key to this task is generating reliable pseudo labels for unlabeled images. Existing methods mainly focus on producing reliable pseudo labels based on the confidence scores of unlabeled images while largely ignoring the use of labeled images with accurate annotations. In this paper, we propose a Cross-Image Semantic Consistency guided Rectifying (CISC-R) approach for semi-supervised semantic segmentation, which explicitly leverages the labeled images to rectify the generated pseudo labels. Our CISC-R is inspired by the fact that images belonging to the same class have a high pixel-level correspondence. Specifically, given an unlabeled image and its initial pseudo labels, we first query a guiding labeled image that shares the same semantic information with the unlabeled image. Then, we estimate the pixel-level similarity between the unlabeled image and the queried labeled image to form a CISC map, which guides us to achieve a reliable pixel-level rectification for the pseudo labels. Extensive experiments on the PASCAL VOC 2012, Cityscapes, and COCO datasets demonstrate that the proposed CISC-R can significantly improve the quality of the pseudo labels and outperform the state-of-the-art methods. Code is available at https://github.com/Luffy03/CISC-R.
Linshan Wu, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Zhun Zhong
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 A Multi-Level Label-Aware Semi-Supervised Framework for Remote Sensing Scene Classification
abstract
Semi-supervised learning (SSL) is a promising approach to reduce the labeling burden in remote sensing scene classification tasks. However, most semi-supervised methods typically exploit the single-level semantic information of unlabeled data, ignoring the multi-level semantic structure prevalent in remote sensing data. The multi-level semantic structure, which contains the correlation of different categories and the multi-granularity semantic information, can help the scene classification model to more accurately measure the feature distance between different categories and more effectively utilize unlabeled data. Therefore, this paper proposes a multi-level label-aware semi-supervised scene classification framework, MLLA, which extends the semantic information captured in unlabeled data from single-level to multi-level to improve the scene classification performance. Specifically, we first propose a multi-level prototype awareness module to capture the multi-level semantic structure underlying remote sensing data. Then, based on this structure, a multi-level pseudo-label generation module is designed to assign multi-level pseudo-labels to the unlabeled data. Finally, by combining the labeled samples and the multi-level pseudo-labeled samples, the scene classification model is progressively trained. The experimental results on three benchmark datasets show that the proposed MLLA achieves excellent performance compared to other semi-supervised classification methods.
Linshan Wu, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.4
2022 Deep Covariance Alignment for Domain Adaptive Remote Sensing Image Segmentation
abstract
Unsupervised domain adaptive (UDA) image segmentation has recently gained increasing attention, aiming to improve the generalization capability for transferring knowledge from the source domain to the target domain. However, in high spatial resolution remote sensing image (RSI), the same category from different domains (e.g., urban and rural) can appear to be totally different with extremely inconsistent distributions, which heavily limits the UDA accuracy. To address this problem, in this article, we propose a novel deep covariance alignment (DCA) model for UDA RSI segmentation. The DCA can explicitly align category features to learn shared domain-invariant discriminative feature representations, which enhance the ability of model generalization. Specifically, a category feature pooling (CFP) module is first used to extract category features by combining coarse outputs and deep features. Then, we leverage a novel covariance regularization (CR) to enforce the intracategory features to be closer and the intercategory features to be further separate. Compared with the existing category alignment methods, our CR aims to regularize the correlation between different dimensions of the features, and thus performs more robustly when dealing with divergent category features of imbalanced and inconsistent distributions. Finally, we propose a stagewise procedure to train the DCA to alleviate error accumulation. Experiments on both rural-to-urban and urban-to-rural scenarios of the LoveDA dataset demonstrate the superiority of our proposed DCA over other state-of-the-art UDA segmentation methods. Code is available athttps://github.com/Luffy03/DCA.
Linshan Wu, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.1
2022 Deep Bilateral Filtering Network for Point-Supervised Semantic Segmentation in Remote Sensing Images
abstract
Semantic segmentation methods based on deep neural networks have achieved great success in recent years. However, training such deep neural networks relies heavily on a large number of images with accurate pixel-level labels, which requires a huge amount of human effort, especially for large-scale remote sensing images. In this paper, we propose a point-based weakly supervised learning framework called the deep bilateral filtering network (DBFNet) for the semantic segmentation of remote sensing images. Compared with pixel-level labels, point annotations are usually sparse and cannot reveal the complete structure of the objects; they also lack boundary information, thus resulting in incomplete prediction within the object and the loss of object boundaries. To address these problems, we incorporate the bilateral filtering technique into deeply learned representations in two respects. First, since a target object contains smooth regions that always belong to the same category, we perform deep bilateral filtering (DBF) to filter the deep features by a nonlinear combination of nearby feature values, which encourages the nearby and similar features to become closer, thus achieving a consistent prediction in the smooth region. In addition, the DBF can distinguish the boundary by enlarging the distance between the features on different sides of the edge, thus preserving the boundary information well. Experimental results on two widely used datasets, the ISPRS 2-D semantic labeling Potsdam and Vaihingen datasets, demonstrate that our proposed DBFNet can achieve a highly competitive performance compared with state-of-the-art fully-supervised methods. Code is available at https://github.com/Luffy03/DBFNet.
Linshan Wu, Leyuan Fang, Jun Yue 0004, Bob Zhang 0001, Pedram Ghamisi
IEEE Trans. Image Process.1