Shandong Wu

dblp:80/2513 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-0770-2203ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author
YearPublicationVenuePosition
2026 Δt-Mamba3D: A Time‑Aware Spatio‑Temporal State‑Space Model for Breast Cancer Risk Prediction
abstract
Longitudinal analysis of sequential radiological images is hampered by a fundamental data challenge: how to effectively model a sequence of high-resolution images captured at irregular time intervals. This data structure contains indispensable spatial and temporal cues that current methods fail to fully exploit. Models often compromise by either collapsing spatial information into vectors or applying spatio-temporal models that are computationally inefficient and incompatible with non-uniform time steps. We address this challenge with Time-Aware $\Delta t$-Mamba3D, a novel state-space architecture adapted for longitudinal medical imaging. Our model simultaneously encodes irregular inter-visit intervals and rich spatio-temporal context while remaining computationally efficient. Its core innovation is a continuous-time selective scanning mechanism that explicitly integrates the true time difference between exams into its state transitions. This is complemented by a multi-scale 3D neighborhood fusion module that robustly captures spatio-temporal relationships. In a comprehensive breast cancer risk prediction benchmark using sequential screening mammogram exams, our model shows superior performance, improving the validation C-index by 2–5 percentage points and achieving higher 1–5 year AUC scores compared to established variants of recurrent, transformer, and state-space models. Thanks to its linear complexity, the model can efficiently process long and complex patient screening histories of mammograms, forming a new framework for longitudinal image analysis.
Zhengbo Zhou, Dooman Arefan, Margarita L. Zuley, Shandong Wu
AAAI4
2025 Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
abstract
Federated prompt learning benefits federated learning with CLIP-like Vision-Language Model's (VLM's) robust representation learning ability through prompt learning. However, current federated prompt learning methods are habitually restricted to the traditional FL paradigm, where the participating clients are generally only allowed to download a single globally aggregated model from the server. While justifiable for training full-sized models under federated settings, in this work, we argue that this paradigm is ill-suited for lightweight prompts. By facilitating the clients to download multiple pre-aggregated prompts as fixed non-local experts, we propose Personalized Federated Mixture of Adaptive Prompts (pFedMoAP), a novel FL framework that personalizes the prompt learning process through the lens of Mixture of Experts (MoE). pFedMoAP implements a local attention-based gating network that learns to generate enhanced text features for better alignment with local image data, benefiting from both local and downloaded non-local adaptive prompt experts. Extensive experiments on 9 datasets under various federated settings demonstrate the efficacy of the proposed pFedMoAP algorithm. The code is available at https://github.com/ljaiverson/pFedMoAP.
Jun Luo 0010, Chen Chen 0001, Shandong Wu
ICLR3
2025 Guest Editorial: Multi-Modal Joint Learning in Healthcare Imaging
Tao Tan 0002, Yue Sun 0001, Shandong Wu
IEEE J. Biomed. Health Informatics4
2024 SAH-NET: Structure-Aware Hierarchical Network for Clustered Microcalcification Classification in Digital Breast Tomosynthesis
abstract
Benign and malignant classification of clustered microcalcifications (MCs) in digital breast tomosynthesis (DBT) is an essential task in computer-aided diagnosis. However, due to the anisotropic resolution of DBT, three-dimensional (3-D) convolutional neural network (CNN)-based methods cannot extract hierarchical features efficiently. Moreover, the sparse distribution of MC points in the cluster makes it difficult for the CNN to extract discriminative structural information for classification. To comprehensively address these challenges, we propose a novel structure-aware hierarchical network (SAH-Net) for benign and malignant classification of clustered MC in a DBT volume. Specifically, the two-dimensional (2-D) group convolution is used to extract intraslice features. The one-to-one correspondence between group convolutions and slices ensures the independence of hierarchical feature extraction. Then, a partial deformable Transformer-based 3-D structural feature learning module is proposed to capture the long-range dependency between MC points in the cluster. We evaluate the proposed method on an in-house dataset with 495 clustered MCs collected from 462 DBT images. Experimental results confirm the validity of our proposed modules. The results also show that the proposed SAH-Net outperforms several other representative methods on this topic, and achieves the best classification result, with an area under the receiver operation curve (AUC) of 86.87%. The implementation of the proposed model is available at https://github.com/sunhaotian130911/SAHNet.
Shandong Wu, Xinjian Chen 0001, Lingji Kong, Xiaodong Yang 0005, You Meng, Shuangqing Chen, Jian Zheng 0001
IEEE Trans. Cybern.2
2023 PGFed: Personalize Each Client's Global Objective for Federated Learning
abstract
Personalized federated learning has received an upsurge of attention due to the mediocre performance of conventional federated learning (FL) over heterogeneous data. Unlike conventional FL which trains a single global consensus model, personalized FL allows different models for different clients. However, existing personalized FL algorithms only implicitly transfer the collaborative knowledge across the federation by embedding the knowledge into the aggregated model or regularization. We observed that this implicit knowledge transfer fails to maximize the potential of each client’s empirical risk toward other clients. Based on our observation, in this work, we propose Personalized Global Federated Learning (PGFed), a novel personalized FL framework that enables each client to personalize its own global objective by explicitly and adaptively aggregating the empirical risks of itself and other clients. To avoid massive (O(N2)) communication overhead and potential privacy leakage while achieving this, each client’s risk is estimated through a first-order approximation for other clients’ adaptive risk aggregation. On top of PGFed, we develop a momentum upgrade, dubbed PGFedMo, to more efficiently utilize clients’ empirical risks. Our extensive experiments on four datasets under different federated settings show consistent improvements of PGFed over previous state-of-the-art methods. The code is publicly available at https://github.com/ljaiverson/pgfed.
Jun Luo 0010, Matías Mendieta, Chen Chen 0001, Shandong Wu
ICCV4
2023 FedPerfix: Towards Partial Model Personalization of Vision Transformers in Federated Learning
abstract
Personalized Federated Learning (PFL) represents a promising solution for decentralized learning in heterogeneous data environments. Partial model personalization has been proposed to improve the efficiency of PFL by selectively updating local model parameters instead of aggregating all of them. However, previous work on partial model personalization has mainly focused on Convolutional Neural Networks (CNNs), leaving a gap in understanding how it can be applied to other popular models such as Vision Transformers (ViTs). In this work, we investigate where and how to partially personalize a ViT model. Specifically, we empirically evaluate the sensitivity to data distribution of each type of layer. Based on the insights that the self-attention layer and the classification head are the most sensitive parts of a ViT, we propose a novel approach called FedPerfix, which leverages plugins to transfer information from the aggregated model to the local client as a personalization. Finally, we evaluate the proposed approach on CIFAR-100, OrganAMNIST, and Office-Home datasets and demonstrate its effectiveness in improving the model’s performance compared to several advanced PFL methods. Code is available at https://github.com/imguangyu/FedPerfix
Guangyu Sun 0004, Matías Mendieta, Jun Luo 0010, Shandong Wu, Chen Chen 0001
ICCV4
2022 Adapt to Adaptation: Learning Personalization for Cross-Silo Federated Learning
abstract
Conventional federated learning (FL) trains one global model for a federation of clients with decentralized data, reducing the privacy risk of centralized training. However, the distribution shift across non-IID datasets, often poses a challenge to this one-model-fits-all solution. Personalized FL aims to mitigate this issue systematically. In this work, we propose APPLE, a personalized cross-silo FL framework that adaptively learns how much each client can benefit from other clients' models. We also introduce a method to flexibly control the focus of training APPLE between global and local objectives. We empirically evaluate our method's convergence and generalization behaviors, and perform extensive experiments on two benchmark datasets and two medical imaging datasets under two non-IID settings. The results show that the proposed personalized FL framework, APPLE, achieves state-of-the-art performance compared to several other personalized FL approaches in the literature. The code is publicly available at https://github.com/ljaiverson/pFL-APPLE.
Jun Luo 0010, Shandong Wu
IJCAI2
2022 A self-training teacher-student model with an automatic label grader for abdominal skeletal muscle segmentation
Degan Hao, Maaz Ahsan, Tariq Salim, Andres Duarte-Rojo, Dadashzadeh Esmaeel, Yudong Zhang 0001, Dooman Arefan, Shandong Wu
Artif. Intell. Medicine8
2022 SurvivalCNN: A deep learning-based method for gastric cancer survival prediction using radiological imaging data and clinicopathological variables
Degan Hao, Qiuxia Feng, Xisheng Liu, Dooman Arefan, Yudong Zhang 0001, Shandong Wu
Artif. Intell. Medicine8
2022 Deep learning of longitudinal mammogram examinations for breast cancer risk prediction
Saba Dadsetan, Dooman Arefan, Wendie A. Berg, Margarita L. Zuley, Jules H. Sumkin, Shandong Wu
Pattern Recognit.6
2021 Radiomics-Informed Deep Curriculum Learning for Breast Cancer Diagnosis
Giacomo Nebbia, Saba Dadsetan, Dooman Arefan, Margarita L. Zuley, Jules H. Sumkin, Heng Huang 0001, Shandong Wu
MICCAI (5)7
2021 Disentangled and Proportional Representation Learning for Multi-view Brain Connectomes
Yanfu Zhang, Liang Zhan, Shandong Wu, Paul M. Thompson, Heng Huang 0001
MICCAI (7)3
2021 3D Context-Aware Convolutional Neural Network for False Positive Reduction in Clustered Microcalcifications Detection
abstract
False positives (FPs) reduction is indispensable for clustered microcalcifications (MCs) detection in digital breast tomosynthesis (DBT), since there might be excessive false candidates in the detection stage. Considering that DBT volume has an anisotropic resolution, we proposed a novel 3D context-aware convolutional neural network (CNN) to reduce FPs, which consists of a 2D intra-slices feature extraction branch and a 3D inter-slice features fusion branch. In particular, 3D anisotropic convolutions were designed to learn representations from DBT volumes and inter-slice information fusion is only performed on the feature map level, which could avoid the influence of anisotropic resolution of DBT volume. The proposed method was evaluated on a large-scale Chinese women population of 877 cases with 1754 DBT volumes and compared with 8 related methods. Experimental results show that the proposed network achieved the best performance with an accuracy of 92.68% for FPs reduction with an AUC of 97.65%, and the FPs are 0.0512 per DBT volume at a sensitivity of 90%. This also proved that making full use of 3D contextual information of DBT volume can improve the performance of the classification algorithm.
Jian Zheng 0001, Shandong Wu, Yunsong Peng, Xiaodong Yang 0005
IEEE J. Biomed. Health Informatics3
2020 Handling imbalanced medical image data: A deep-learning-based one-class classification approach
Shandong Wu
Artif. Intell. Medicine4
2020 Response score of deep learning for out-of-distribution sample detection of medical images
Shandong Wu
J. Biomed. Informatics2
2020 Inaccurate Labels in Weakly-Supervised Deep Learning: Automatic Identification and Correction and Their Impact on Classification Performance
abstract
In data-driven deep learning-based modeling, data quality may substantially influence classification performance. Correct data labeling for deep learning modeling is critical. In weakly-supervised learning, a challenge lies in dealing with potentially inaccurate or mislabeled training data. In this paper, we proposed an automated methodological framework to identify mislabeled data using two metric functions, namely, Cross-entropy Loss that indicates divergence between a prediction and ground truth, and Influence function that reflects the dependence of a model on data. After correcting the identified mislabels, we measured their impact on the classification performance. We also compared the mislabeling effects in three experiments on two different real-world clinical questions. A total of 10,500 images were studied in the contexts of clinical breast density category classification and breast cancer malignancy diagnosis. We used intentionally flipped labels as mislabels to evaluate the proposed method at a varying proportion of mislabeled data included in model training. We also compared the effects of our method to two published schemes for breast density category classification. Experiment results show that when the dataset contains 10% of mislabeled data, our method can automatically identify up to 98% of these mislabeled data by examining/checking the top 30% of the full dataset. Furthermore, we show that correcting the identified mislabels leads to an improvement in the classification performance. Our method provides a feasible solution for weakly-supervised deep learning modeling in dealing with inaccurate labels.
Degan Hao, Lei Zhang 0036, Jules H. Sumkin, Aly A. Mohamed, Shandong Wu
IEEE J. Biomed. Health Informatics5
2020 Guest Editorial: Deep Learning in Ultrasound Imaging
abstract
Among the different imaging modalities, ultrasound is the most widespread modality for visualizing human tissue due to it being low-cost, non-ionizing, real-time with immediate feedback to the sonographer, convenient to operate, widely available and well established, with a very large number of images generated in a single setting. On the other hand, ultrasound imaging suffers from the disadvantage of being user dependent and of variable quality,which makes the automated interpretation of ultrasound images often very difficult. In recent years, algorithms in medical imaging have been significantly improved thanks to the advent of deep learning methods (including convolutional neural networks, recurrent neural networks, autoencoders, or generative adversarial networks). To address the various challenges of automatically processing and interpreting ultrasound images, deep learning techniques have been gradually applied to various types of ultrasound data (such as B-mode ultrasound, Doppler ultrasound, or contrast-enhanced ultrasound), acquired with a range of different probes, with the aim of improving image quality, for organ segmentation, device localization and tracking, for tissue characterization, and ultimately to improve disease diagnosis and therapeutic outcome. The papers in this special section seek to present and highlight the latest development on applying advanced deep learning techniques in ultrasound imaging.
Caifeng Shan, Tao Tan 0002, Shandong Wu, Julia A. Schnabel
IEEE J. Biomed. Health Informatics3
2019 Robust UAV-based tracking using hybrid classifiers
Shandong Wu
Mach. Vis. Appl.3
2019 Object tracking based on Huber loss function
Shiqiang Hu, Shandong Wu
Vis. Comput.3
2015 Visual tracking based on group sparsity learning
Shiqiang Hu, Shandong Wu
Mach. Vis. Appl.3
2012 Atlas-Based Probabilistic Fibroglandular Tissue Segmentation in Breast MRI
Shandong Wu, Susan Weinstein, Despina Kontos
MICCAI (2)1
2011 Action recognition in videos acquired by a moving camera using motion decomposition of Lagrangian particle trajectories
abstract
Recognition of human actions in a video acquired by a moving camera typically requires standard preprocessing steps such as motion compensation, moving object detection and object tracking. The errors from the motion compensation step propagate to the object detection stage, resulting in miss-detections, which further complicates the tracking stage, resulting in cluttered and incorrect tracks. Therefore, action recognition from a moving camera is considered very challenging. In this paper, we propose a novel approach which does not follow the standard steps, and accordingly avoids the aforementioned difficulties. Our approach is based on Lagrangian particle trajectories which are a set of dense trajectories obtained by advecting optical flow over time, thus capturing the ensemble motions of a scene. This is done in frames of unaligned video, and no object detection is required. In order to handle the moving camera, we propose a novel approach based on low rank optimization, where we decompose the trajectories into their camera-induced and object-induced components. Having obtained the relevant object motion trajectories, we compute a compact set of chaotic invariant features which captures the characteristics of the trajectories. Consequently, a SVM is employed to learn and recognize the human actions using the computed motion features. We performed intensive experiments on multiple benchmark datasets and two new aerial datasets called ARG and APHill, and obtained promising results.
Shandong Wu, Omar Oreifej, Mubarak Shah
ICCV1
2010 Chaotic invariants of Lagrangian particle trajectories for anomaly detection in crowded scenes
abstract
A novel method for crowd flow modeling and anomaly detection is proposed for both coherent and incoherent scenes. The novelty is revealed in three aspects. First, it is a unique utilization of particle trajectories for modeling crowded scenes, in which we propose new and efficient representative trajectories for modeling arbitrarily complicated crowd flows. Second, chaotic dynamics are introduced into the crowd context to characterize complicated crowd motions by regulating a set of chaotic invariant features, which are reliably computed and used for detecting anomalies. Third, a probabilistic framework for anomaly detection and localization is formulated. The overall work-flow begins with particle advection based on optical flow. Then particle trajectories are clustered to obtain representative trajectories for a crowd flow. Next, the chaotic dynamics of all representative trajectories are extracted and quantified using chaotic invariants known as maximal Lyapunov exponent and correlation dimension. Probabilistic model is learned from these chaotic feature set, and finally, a maximum likelihood estimation criterion is adopted to identify a query video of a scene as normal or abnormal. Furthermore, an effective anomaly localization algorithm is designed to locate the position and size of an anomaly. Experiments are conducted on known crowd data set, and results show that our method achieves higher accuracy in anomaly detection and can effectively localize anomalies.
Shandong Wu, Brian E. Moore, Mubarak Shah
CVPR1
2010 Motion trajectory reproduction from generalized signature description
Shandong Wu, Youfu Li 0001
Pattern Recognit.1
2009 Probabilistic Cluster Signature for Modeling Motion Classes
abstract
In this paper, a novel 3-D motion trajectory signature is introduced to serve as an effective description to the raw trajectory. More importantly, based on the trajectory signature, a probabilistic model-based cluster signature is further developed for modeling a motion class. The cluster signature is a mixture model-based motion description that is useful for motion class perception, recognition and to benefit a generalized robot task representation. The signature modeling process is supported by integrating the EM and IPRA algorithms. The conducted experiments verified the cluster signature's effectiveness.
Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001
IROS1
2009 Flexible signature descriptions for adaptive motion trajectory representation, perception and recognition
Shandong Wu, Youfu Li 0001
Pattern Recognit.1
2008 A hierarchical motion trajectory signature descriptor
abstract
Motion trajectory is a compact clue for motion characterization. However, it is normally used directly in its raw data form in most work and effective trajectory description is lacking. In this paper, we propose a novel hierarchical motion trajectory signature descriptor, which can not only fully capture motion features for detailed perception, but also can be used for probabilistic fast recognition. The hierarchy enables the signature to exhibit high functional adaptability meeting different application requirements. At the first-level, differential invariants are employed to describe trajectory features and a nonlinear signature warping method is developed to perceive and recognize trajectories. The second-level signature is the condensation of the first-level signature by applying PCA based dimension optimization. It behaves more efficiently in recognition based on the Gaussian Mixture modeling and Bayesian classifier. The conducted experiments verified the signature’s effectiveness.
Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001
ICRA1
2008 Invariant signature description and trajectory reproduction for robot Learning by Demonstration
abstract
In most reported works about robot learning by demonstration (LbD), the demonstration is normally limited to simple gestures or grasp actions. In this paper, motion trajectory oriented LbD is studied in which free form 3-D motion trajectory is extracted to characterize certain human demonstrations. We propose to build effective description to motion trajectories to be learned by a robot instead of learning the raw trajectory data. A novel signature descriptor is formulated which serves as a generic and invariant description for motion trajectories. More importantly, a trajectory reproduction algorithm based on the learned signature is investigated to enable a robot to repeat/follow the reproduced trajectory instance. Experiments are reported to show the signature description and the reproduction algorithm for further application to the LbD.
Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001
IROS1