Shuchao Pang

dblp:146/9250 · also Shu-Chao Pang · DBLP profile ↗
← Back
42ranked-venue papers
14as first author
35since 2021 · last 2026
0000-0002-5668-833XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Security and privacy · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal Robust Prompt Distillation for 3D Point Cloud Models
abstract
Adversarial attacks pose a significant threat to learning-based 3D point cloud models, critically undermining their reliability in security-sensitive applications. Existing defense methods often suffer from (1) high computational overhead and (2) poor generalization ability across diverse attack types. To bridge these gaps, we propose a novel yet efficient teacher-student framework, namely Multimodal Robust Prompt Distillation (MRPD) for distilling robust 3D point cloud model. It learns lightweight prompts by aligning student point cloud model's features with robust embeddings from three distinct teachers: a vision model processing depth projections, a high-performance 3D model, and a text encoder. To ensure a reliable knowledge transfer, this distillation is guided by a confidence-gated mechanism which dynamically balances the contribution of all input modalities. Notably, since the distillation is all during the training stage, there is no additional computational cost at inference. Extensive experiments demonstrate that MRPD substantially outperforms state-of-the-art defense methods against a wide range of white-box and black-box attacks, while even achieving better performance on clean data. Our work presents a new, practical paradigm for building robust 3D vision systems by efficiently harnessing multimodal knowledge.
Anan Du, Yongbin Zhou, Shuchao Pang
AAAI6
2026 TIDE: Making Task-Agnostic Backdoors Harder to Erase in Pre-trained Language Models
Zhigang Lu 0001, Bing Li 0002, Anan Du, Shuchao Pang
ACISP (2)5
2026 QuEST: Quantization-Conditioned Efficient Stealthy Trojan
abstract
Quantization-conditioned backdoor attacks, which exploit model quantization states to trigger malicious behavior, pose a hidden threat to deep learning security. However, current studies ignore the feasibility of attacks,i.e., detection visibility, and computational overhead, leading to significant constraints on their practical deployment in real-world adversarial scenarios. To address these limitations, we propose Quantization-conditioned Efficient Stealthy Trojan (QuEST), a novel framework that enhances both the stealth and efficiency of backdoor attacks under quantization constraints. For enhancing stealthiness, we design a stealth-optimized training scheme that benefits from the parametric backdoor injection and trigger scaling augmentation to maintain the consistency of model behavior during attack. In this way, the defender will fail to capture the suspicious behavior differences for detection due to the made efforts in both model-side and data-side. To improve efficiency, we introduce information-guided parameter sharing, which utilizes parameter redundancy analysis and Fisher divergence metrics to identify a minimal amount of quantization-preserved parameters for back-door injection. These parameters are strategically shared between the malicious and benign models, enabling concurrent training and substantially reducing overall training time. Extensive experiments demonstrate that QuEST maintains competitive attack success rates while improving stealth performance by 18.75% and reducing computational costs by 26.66% on average compared to state-of-the-art methods, highlighting QuEST’s potential for more practical adversarial deployments in real-world scenarios. Our code is available here.
Shuchao Pang, Jiakai Wang, Yunhuai Liu, Xianglong Liu 0001, Yongbin Zhou
IEEE Trans. Inf. Forensics Secur.2
2026 SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision Language Models
Shuchao Pang, Xiyu Zeng, Siyuan Liang 0004, Chuanting Zhang, Enguang Liu, Basem Shihada, Yongbin Zhou, Minhui Xue 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Ciard: Cyclic Iterative Adversarial Robustness Distillation
abstract
Adversarial robustness distillation (ARD) aims to transfer both performance and robustness from teacher model to lightweight student model, enabling resilient performance on resource-constrained scenarios. Though existing ARD approaches enhance student model's robustness, the inevitable by-product leads to the degraded performance on clean examples. We summarize the causes of this problem inherent in existing methods with dual-teacher framework as: 1. The divergent optimization objectives of dual-teacher models, i.e., the clean and robust teachers, impede effective knowledge transfer to the student model, and 2. The iteratively generated adversarial examples during training lead to performance deterioration of the robust teacher model. To address these challenges, we propose a novel Cyclic Iterative ARD (CIARD) method with two key innovations: a. A multi-teacher framework with contrastive push-loss alignment to resolve conflicts in dual-teacher optimization objectives, and b. Continuous adversarial retraining to maintain dynamic teacher robustness against performance degradation from the varying adversarial examples. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that CIARD achieves remarkable performance with an average 3.53 improvement in adversarial defense rates across various attack scenarios and a 5.87 increase in clean sample accuracy, establishing a new benchmark for balancing model robustness and generalization. Our code is available at https://github.com/eminentgu/CIARD
Shuchao Pang, Anan Du, Yunhuai Liu, Yongbin Zhou
ICCV2
2025 Towards a 3D Transfer-Based Black-Box Attack via Critical Feature Guidance
abstract
Deep neural networks for 3D point clouds have been demonstrated to be vulnerable to adversarial examples. Previous 3D adversarial attack methods often exploit certain information about the target models, such as model parameters or outputs, to generate adversarial point clouds. However, in realistic scenarios, it is challenging to obtain any information about the target models under conditions of absolute security. Therefore, we focus on transfer-based attacks, where generating adversarial point clouds does not require any information about the target models. Based on our observation that the critical features used for point cloud classification are consistent across different DNN architectures, we propose CFG, a novel transfer-based black-box attack method that improves the transferability of adversarial point clouds via the proposed Critical Feature Guidance. Specifically, our method regularizes the search of adversarial point clouds by computing the importance of the extracted features, prioritizing the corruption of critical features that are likely to be adopted by diverse architectures. Further, we explicitly constrain the maximum deviation extent of the generated adversarial point clouds in the loss function to ensure their imperceptibility. Extensive experiments conducted on the ModelNet40 and ScanObjectNN benchmark datasets demonstrate that the proposed CFG outperforms the state-of-the-art attack methods by a large margin.
Shuchao Pang, Zhenghan Chen, Siyuan Liang 0004, Anan Du, Yongbin Zhou
ICCV1
2025 GAP-Diff: Protecting JPEG-Compressed Images from Diffusion-based Facial Customization
Shuchao Pang, Zhigang Lu 0001, Yongbin Zhou, Minhui Xue 0001
NDSS2
2025 One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention Head
abstract
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in tasks requiring multimodal understanding. However, recent studies indicate that LVLMs are more vulnerable than LLMs to unsafe inputs and prone to generating harmful content. Existing defense strategies primarily include fine-tuning, input sanitization, and output intervention. Although these approaches provide a certain level of protection, they tend to be resource-intensive and struggle to effectively counter sophisticated attack techniques. To tackle such issues, we propose One-head Defense (Oh Defense), a novel yet simple approach utilizing LVLMs' internal safety capabilities. Through systematic analysis of the attention mechanisms, we discover that LVLMs' safety capabilities are concentrated within specific attention heads that respond differently to safe or unsafe inputs. Further exploration reveals that a single critical attention head can effectively serve as a safety guard, providing a strong discriminative signal that amplifies the model's inherent safety capabilities. Hence, the Oh Defense requires no additional training or external modules, making it computationally efficient while effectively reactivating suppressed safety mechanisms. Extensive experiments across diverse LVLM architectures and unsafe datasets validate our approach, i.e., the Oh Defense achieves near-perfect defense success rates (> 98\%) for unsafe inputs while maintaining low false positive rates (< 5\%) for safe content. The source code is available at https://github.com/AIASLab/Oh-Defense.
Junhao Xia, Shuchao Pang, Zhigang Lu 0001, Bing Li 0002, Yongbin Zhou, Minhui Xue 0001
NeurIPS3
2025 Reconstruction of Differentially Private Text Sanitization via Large Language Models
abstract
Differential privacy (DP) is the de facto privacy standard against privacy leakage attacks, including many recently discovered ones against large language models (LLMs). However, we discovered that LLMs could reconstruct the altered/removed privacy from given DP-sanitized prompts. We propose two attacks (black-box and white-box) based on the accessibility to LLMs and show that LLMs could connect the pair of DPsanitized text and the corresponding private training data of LLMs by giving sample text pairs as instructions (in the blackbox attacks) or fine-tuning data (in the white-box attacks). To illustrate our findings, we conduct comprehensive experiments on modern LLMs (e.g., LLaMA-2, LLaMA-3, ChatGPT-3.5, ChatGPT-4, ChatGPT-4o, Claude-3, Claude-3.5, OPT, GPT-Neo, GPT-J, Gemma-2, and Pythia) using commonly used datasets (such as WikiMIA, Pile-CC, and Pile-Wiki) against both wordlevel and sentence-level DP. The experimental results show promising recovery rates, e.g., the black-box attacks against the word-level DP over WikiMIA dataset gave 72.18% on LLaMA2 (70B), 82.39% on LLaMA-3 (70B), 75.35% on Gemma-2, 91.2% on ChatGPT-4o, and 94.01% on Claude-3.5 (Sonnet). More urgently, this study indicates that these well-known LLMs have emerged as a new security risk for existing DP text sanitization approaches in the current environment.
Shuchao Pang, Zhigang Lu 0001, Haichen Wang, Peng Fu 0008, Yongbin Zhou, Minhui Xue 0001
RAID1
2025 CDSRNP: Cross-Domain Sequential Recommendation via Neural Process
abstract
Cross-Domain Sequential Recommendation (CDSR) is a hot topic in sequence-based user interest modeling, which aims at utilizing a single model to predict the next items for different domains. To tackle the CDSR, many methods are focused on domain overlapped users’ behaviors fitting, which heavily relies on the same user’s different-domain item sequences collaborating signals to capture the synergy of cross-domain item-item correlation. Indeed, these overlapped users occupy a small fraction of the entire user set only, which introduces a strong assumption that the small group of domain overlapped users is enough to represent all domain user behavior characteristics. However, intuitively, such a suggestion is biased, and the insufficient learning paradigm in non-overlapped users will inevitably limit model performance. Further, it is not trivial to model non-overlapped user behaviors in CDSR because there are no other domain behaviors to collaborate with, which causes the observed single-domain users’ behavior sequences to be hard to contribute to cross-domain knowledge mining. Considering such a phenomenon, we raise a challenging and unexplored question: How to unleash the potential of non-overlapped users’ behaviors to empower CDSR? To this end, we propose a novel CDSR framework with Neural Processes (NP), briefly termed CDSRNP, where NP combines the advantages of meta-learning and stochastic processes. As a meta-learning based method, we first sample some observed overlapped users’ behaviors as the support set to empower query users’ prediction. Next, we employ the NP principle to align the cross-domain correlation prior/posterior distributions generated by support/query user sets, thus the query user (e.g., non-overlapped user) behaviors sequence could also establish a straight bridge to connect other domain items. Additionally, we design a fine-grained interest adaptive layer to identify the users’ interests to enhance prediction. Experimental results illustrate that CDSRNP1 outperforms state-of-the-art methods in two real-world datasets.
Jiangxia Cao, Yiwen Gao 0001, Yunhuai Liu, Shuchao Pang
SDM5
2025 CLEAR: A Clean-Label Backdoor Attack via Representation-Guided Trigger Embedding
abstract
Recent studies have shown that although DNNs perform well on visual tasks, they are still vulnerable to clean-label backdoor attacks. As poisoned samples come from the target class, the model learns both original and trigger features, weakening the association between the trigger and the target label and thereby reducing the attack success rate (ASR). To address this issue, we propose a novel clean-label backdoor attack framework, named CLEAR, i.e., Clean-Label Embedding Attack with Representation-Guidance, which strengthens the correlation between triggers and target labels while maintaining high stealth. Specifically, the "benign-label push" mechanism perturbs clean samples selected from the target class (the samples designated for poisoning), pushing them away from their original class in the feature space, while the "target-label pull" mechanism pulls the poisoned sample representations closer to the target class through saliency-guided embedding. CLEAR consists of three key components: integrating a diffusion model with PGD to generate natural but semantically perturbed adversarial samples; selecting backdoor-susceptible samples near the decision boundary based on classification loss and embedding triggers into high-saliency regions identified using a Feature Pyramid Network (FPN) combined with a local self-attention mechanism. Experimental results on CIFAR10 and GTSRB with ResNet18 and VGG16 demonstrate that CLEAR generally improves ASR while maintaining strong stealth.
Zhan Wu, Shuchao Pang
SMC4
2025 Controllable text-to-3D multi-object generation via integrating layout and multiview patterns
Shaorong Sun, Shuchao Pang, Yazhou Yao, Xiaoshui Huang
Comput. Graph.2
2025 A position-aware sets based weakly supervised framework for whole-slide subtype classification
Jiuman Song, Bo Yu 0013, Lele Cong, Xianling Cong, Hongyan Sun, Shuchao Pang, Hechang Chen
Eng. Appl. Artif. Intell.8
2025 PriDM: Effective and Universal Private Data Recovery via Diffusion Models
abstract
Deep models excel in analyzing image data. However, recent studies on Black-Box Model Inversion (MI) Attacks against image models have revealed the potential to recover concealed (via specific masks) private training images using publicly available images from the same domain as the training data. This study introduces PriDM, a novel diffusion model-based MI attack, illustrating the increased vulnerability of image models. PriDM leverages range-null space decomposition to extract essential range-space information and incorporates it into the diffusion model's sampling process. This enables the recovery of private information from arbitrarily masked images relying solely on images only aligned with the same machine-learning tasks as the target model. To demonstrate PriDM's effectiveness, we conducted experiments with various adversary background knowledge, including different public dataset domains and image masks. Results show PriDM produces recovered images of significantly higher quality, approximately twice as good as existing methods. Moreover, in scenarios involving complex backgrounds, PriDM outperforms the state-of-the-art by approximately 70%. In specific background knowledge scenarios, such as compressed and blurred images, our method achieves an almost 100% success rate. Additionally, PriDM performs well with real-world background knowledge including individuals wearing masks and randomly masked face images, which are not considered by existing works.
Shuchao Pang, Yihang Rao, Zhigang Lu 0001, Haichen Wang, Yongbin Zhou, Minhui Xue 0001
IEEE Trans. Dependable Secur. Comput.1
2024 UniADS: Universal Architecture-Distiller Search for Distillation Gap
abstract
In this paper, we present UniADS, the first Universal Architecture-Distiller Search framework for co-optimizing student architecture and distillation policies. Teacher-student distillation gap limits the distillation gains. Previous approaches seek to discover the ideal student architecture while ignoring distillation settings. In UniADS, we construct a comprehensive search space encompassing an architectural search for student models, knowledge transformations in distillation strategies, distance functions, loss weights, and other vital settings. To efficiently explore the search space, we utilize the NSGA-II genetic algorithm for better crossover and mutation configurations and employ the Successive Halving algorithm for search space pruning, resulting in improved search efficiency and promising results. Extensive experiments are performed on different teacher-student pairs using CIFAR-100 and ImageNet datasets. The experimental results consistently demonstrate the superiority of our method over existing approaches. Furthermore, we provide a detailed analysis of the search results, examining the impact of each variable and extracting valuable insights and practical guidance for distillation design and implementation.
Zhenghan Chen, Yihang Rao, Lujun Li 0001, Shuchao Pang
AAAI6
2024 A Unified Deep Learning-Based EEG Biometric Authentication System for Cross-Session Scenarios
Yijing Gong, Min Wang 0009, Yu Zhang 0217, Wenjie Zhang 0001, Shuchao Pang
ADMA (4)5
2024 ADDM: Adversarial Defenses with Diffusion Model for Medical Imaging Data Mining
Yimin He, Shuchao Pang, Anan Du, Hechang Chen, Lele Cong, Mehmet A. Orgun
ADMA (4)2
2024 Enhancing Content-based Recommendation via Large Language Model
abstract
In real-world applications, users express different behaviors when they interact with different items, including implicit click/like interactions, and explicit comments/reviews interactions. Nevertheless, almost all recommender works are focused on how to describe user preferences by the implicit click/like interactions, to find the synergy of people. For the content-based explicit comments/reviews interactions, some works attempt to utilize them to mine the semantic knowledge to enhance recommender models. However, they still neglect the following two points: (1) The content semantic is a universal world knowledge; how do we extract the multi-aspect semantic information to empower different domains? (2) The user/item ID feature is a fundamental element for recommender models; how do we align the ID and content semantic feature space? In this paper, we propose a 'plugin' semantic knowledge transferring method LoID, which includes two major components: (1) LoRA-based large language model pretraining to extract multi-aspect semantic information; (2) ID-based contrastive objective to align their feature spaces. We conduct extensive experiments with SOTA baselines to demonstrate superiority of our method LoID.
Qianqian Xie, Jiangxia Cao, Shuchao Pang
CIKM5
2024 Dynamic Multimodal Prompt Tuning: Boost Few-Shot Learning with VLM-Guided Point Cloud Models
abstract
Few-shot learning is crucial for downstream tasks involving point clouds, given the challenge of obtaining sufficient datasets due to extensive collecting and labeling efforts. Pre-trained VLM-Guided point cloud models, containing abundant knowledge, can compensate for the scarcity of training data, potentially leading to very good performance. However, adapting these pre-trained point cloud models to specific few-shot learning tasks is challenging due to their huge number of parameters and high computational cost. To this end, we propose a novel Dynamic Multimodal Prompt Tuning method, named DMMPT, for boosting few-shot learning with pre-trained VLM-Guided point cloud models. Specifically, we build a dynamic knowledge collector capable of gathering task- and data-related information from various modalities. Then, a multimodal prompt generator is constructed to integrate collected dynamic knowledge and generate multimodal prompts, which efficiently direct pre-trained VLM-guided point cloud models toward few-shot learning tasks and address the issue of limited training data. Our method is evaluated on benchmark datasets not only in a standard N-way K-shot few-shot learning setting, but also in a more challenging setting with all classes and K-shot few-shot learning. Notably, our method outperforms other prompt-tuning techniques, achieving highly competitive results comparable to full fine-tuning methods while significantly enhancing computational efficiency.
Shuchao Pang, Anan Du, Jixiang Miao, Jorge Díez 0001
ECAI2
2024 Veil Privacy on Visual Data: Concealing Privacy for Humans, Unveiling for DNNs
Shuchao Pang, Ruhao Ma, Bing Li 0002, Yongbin Zhou, Yazhou Yao
ECCV (83)1
2024 Foster Adaptivity and Balance in Learning with Noisy Labels
Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Shuchao Pang, Yucheng Wang 0013, Yazhou Yao
ECCV (27)4
2024 Anti-AsynDGAN: Black-box Membership Inference Attacks Against Medical Distributed Generation Models
abstract
Recently, Distributed Asynchronized Discriminator GAN (AsynDGAN), one of the representatives of distributed generative adversarial networks, has received much attention in the medical field. It provides a collaborative model training approach to protect raw medical data from leaving the local medical entity. However, does raw data not leave the local client to ensure absolute data security? And potential privacy risks associated with distributed generative adversarial networks have not been systematically analyzed. In this paper, we design and implement two targeted membership inference attacks against AsynDGAN, i.e., global and local black-box attacks, by investigating the possible vulnerabilities during the training and deployment of AsynDGAN, respectively. Experimental results on three widely-used and multi-modality medical datasets show that AsynDGAN leads to serious leakage of sensitive information and does not effectively protect private-sensitive medical data.
Shuchao Pang
IJCNN2
2024 TriEn-Net: Non-parametric Representation Learning for Large-Scale Point Cloud Semantic Segmentation
Jixiang Miao, Anan Du, Shuchao Pang
PRCV (6)5
2024 dp-promise: Differentially Private Diffusion Probabilistic Models for Image Synthesis
Haichen Wang, Shuchao Pang, Zhigang Lu 0001, Yihang Rao, Yongbin Zhou, Minhui Xue 0001
USENIX Security Symposium2
2024 An end-to-end weakly supervised learning framework for cancer subtype classification using histopathological slides
Hongren Zhou, Hechang Chen, Bo Yu 0013, Shuchao Pang, Xianling Cong, Lele Cong
Expert Syst. Appl.4
2024 MAFFN-SAT: 3-D Point Cloud Defense via Multiview Adaptive Feature Fusion and Smooth Adversarial Training
abstract
Adversarial attacks pose a significant threat to deep neural networks (DNNs) used for 3-D point cloud classification, especially in safety-critical applications. While previous works have proposed several defense model architectures and adversarial training strategies, they often either fall short in capturing the intricate geometric and topological aspects of point cloud data or grapple with challenges pertaining to model convergence. To solve these problems, in this article, we propose an innovative point cloud defense framework, called MAFFN-SAT, which contains a multiview adaptive feature fusion network (MAFFN) along with a smooth adversarial training (SAT) strategy. Specifically, we construct a multiview defense module to obtain multiview features in MAFFN, which uses geometric proximity and spatial queries to comprehensively explore the inherent characteristics of point cloud data. Subsequently, an adaptive feature fusion module is designed to integrate the multiview features. Furthermore, we introduce SAT, which uses an optimized regularization to measure the information divergence between two probability distributions, guiding the model to develop a smoother decision boundary, thereby more robust to adversarial attacks. Extensive experiments conducted on three benchmark datasets demonstrate the robustness of our approach against various attacks. Remarkably, our defense framework achieves 15.34% performance improvement under point dropping attacks on the ModelNet40 dataset. Our implementation:https://github.com/shenyu234/MAFFN-SAT.
Anan Du, Jue Zhang 0001, Yiwen Gao 0001, Shuchao Pang
IEEE Trans. Geosci. Remote. Sens.5
2024 PCL: Point Contrast and Labeling for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation is a fundamental task in 3D scene understanding and has recently achieved remarkable progress. The success of existing approaches is attributed to recent advanced deep networks for point clouds and the availability of a large amount of labeled training data. However, creating such fully annotated training datasets for supervised point cloud semantic segmentation methods is a time-consuming and labor-intensive process, which increases the difficulty of extending supervised approaches to new application scenarios. To alleviate the data-hungry nature of deep learning, we propose PCL, the point contrast and labeling framework for weakly supervised point cloud semantic segmentation with small percentages of point-level annotations. The core idea of this method is to exploit contrastive learning to help learn a larger number of discriminative feature representations with limited annotations. By introducing two types of contrastive relationships, cross-sample point contrast and low-level similarity-based point contrast, our proposed framework can directly regularize the learned feature space, considering not only the low-level similarity within each point cloud but also the discriminative semantics within and across point clouds on both labeled and unlabeled points via pseudo labels. In addition, we propose a pseudo label refinery module to generate robust and reliable pseudo labels online, reducing the negative impact of incorrect pseudo labels. Our method achieves state-of-the-art performance on a diverse set of label-efficient semantic segmentation tasks.
Anan Du, Tianfei Zhou, Shuchao Pang, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.3
2023 Point-Level Label-Free Segmentation Framework for 3D Point Cloud Semantic Mining
Anan Du, Shuchao Pang, Mehmet A. Orgun
ADMA (1)2
2023 UniMOS: A Universal Framework For Multi-Organ Segmentation Over Label-Constrained Datasets
abstract
Machine learning models for medical images can help physicians diagnose and manage diseases. However, due to the fact that medical image annotation requires a great deal of manpower and expertise, as well as the fact that clinical departments perform image annotation based on task orientation, there is the problem of having fewer medical image annotation data with more unlabeled data and having many datasets that annotate only a single organ. In this paper, we present UniMOS, the first universal framework for achieving the utilization of fully and partially labeled images as well as unlabeled images. Specifically, we construct a Multi-Organ Segmentation (MOS) module over fully/partially labeled data as the basenet and designed a new target adaptive loss. Furthermore, we incorporate a semi-supervised training module that combines consistent regularization and pseudo-labeling techniques on unlabeled data, which significantly improves the segmentation of unlabeled data. Experiments show that the framework exhibits excellent performance in several medical image segmentation tasks compared to other advanced methods, and also significantly improves data utilization and reduces annotation cost. Code and models are available at: https://github.com/lw8807001/UniMOS.
Sheng Shao, Junyi Qu, Shuchao Pang, Mehmet A. Orgun
BIBM4
2023 A Multimodal Adversarial Database: Towards A Comprehensive Assessment of Adversarial Attacks and Defenses on Medical Images
abstract
Deep learning models have been widely applied in many fields, including medical image analysis and computer-aided disease diagnosis. However, these models are easily fooled by adversarial attacks from some created adversarial examples which are hardly distinguished by humans. In this paper, we implement a comprehensive assessment of six popular adversarial attacks on four multimodal medical image datasets using two main deep learning-based target models. Moreover, in order to evaluate the capability of defense, two new defense methods are leveraged to cope with medical adversarial attacks. More importantly, we also build and release a big multimodal medical adversarial database (including four medical adversarial datasets) with 712,596 examples to facilitate future research of adversarial attacks and defenses in the multimodal medical image field. Extensive experiments indicate that all-sided adversarial attacks like BIM are still scarce under different evaluation metrics and defenses are not universally successful.
Junyao Hu, Yimin He, Shuchao Pang, Ruhao Ma, Anan Du
DSAA4
2023 Data and knowledge co-driving for cancer subtype classification on multi-scale histopathological slides
Bo Yu 0013, Hechang Chen, Yunke Zhang, Lele Cong, Shuchao Pang, Hongren Zhou, Xianling Cong
Knowl. Based Syst.5
2023 Beyond CNNs: Exploiting Further Inherent Symmetries in Medical Image Segmentation
abstract
Automatic tumor or lesion segmentation is a crucial step in medical image analysis for computer-aided diagnosis. Although the existing methods based on convolutional neural networks (CNNs) have achieved the state-of-the-art performance, many challenges still remain in medical tumor segmentation. This is because, although the human visual system can detect symmetries in 2-D images effectively, regular CNNs can only exploit translation invariance, overlooking further inherent symmetries existing in medical images, such as rotations and reflections. To solve this problem, we propose a novel group equivariant segmentation framework by encoding those inherent symmetries for learning more precise representations. First, kernel-based equivariant operations are devised on each orientation, which allows it to effectively address the gaps of learning symmetries in existing approaches. Then, to keep segmentation networks globally equivariant, we design distinctive group layers with layer-wise symmetry constraints. Finally, based on our novel framework, extensive experiments conducted on real-world clinical data demonstrate that a group equivariant Res-UNet (called GER-UNet) outperforms its regular CNN-based counterpart and the state-of-the-art segmentation methods in the tasks of hepatic tumor segmentation, COVID-19 lung infection segmentation, and retinal vessel detection. More importantly, the newly built GER-UNet also shows potential in reducing the sample complexity and the redundancy of filters, upgrading current segmentation CNNs, and delineating organs on other medical imaging modalities.
Shuchao Pang, Anan Du, Mehmet A. Orgun, Yan Wang 0002, Quan Z. Sheng, Shoujin Wang, Xiaoshui Huang, Zhenmei Yu
IEEE Trans. Cybern.1
2022 ConTenNet: Quantum Tensor-augmented Convolutional Representations for Breast Cancer Histopathological Image Classification
abstract
In recent years, deep convolutional neural networks (CNNs) have been spectacularly successful in the classification and diagnosis of breast cancer and its histopathological images. However, for CNNs, the whole learning process requires high computational complexity, a large number of parameters, and loss of certain global feature information. Meanwhile, the flexibility of tensor networks (TNs) algorithms to machine learning leads to creativity in devising new approaches. In this paper, we propose a novel framework named ConTenNet based on the pre-trained CNNs and quantum TNs (QTNs) to address the weaknesses in CNNs. We propose ConTenNet on the BreakHis dataset, and the experiments show that our model competes with the state-of-the-art methods on both original and normalized images with lower computational complexity, a less number of parameters, and global feature information. Moreover, we adopt the color normalization method to avoid the interference of color in model learning, using the gradient-weighted class activation mapping (Grad-CAM) to prove the necessity of color normalization and the reliability of model learning.
Hong Lai, Jinshu Ma, Shuchao Pang
BIBM4
2022 Training radiomics-based CNNs for clinical outcome prediction: Challenges, strategies and findings
Shuchao Pang, Matthew Field, Jason Dowling, Shalini K. Vinod, Lois Holloway, Arcot Sowmya
Artif. Intell. Medicine1
2021 Tumor attention networks: Better feature selection, better tumor segmentation
Shuchao Pang, Anan Du, Mehmet A. Orgun, Zhenmei Yu
Neural Networks1
2020 Exploring Long-Short-Term Context For Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation attracts numerous attention following the success of the point-based convolution neural network. Due to the ambiguity of the point-based feature, many methods study on integrating contextual information to solve the ambiguous problem. However, the extracted context is severely limited to the small input blocks. Few prior works exploit contextual information beyond the blocks to capture long-range dependencies. To address this limitation, we propose a novel long-short-term context framework, which adopts a long-short-term feature bank to exploit both the local context within each block and the long-range context beyond the current task block. The proposed framework is flexible and easy to be combined with existing models, thereby enables existing models to capture the larger range context. Extensive experiments demonstrate that the proposed model achieves improved segmentation performance, and augmenting existing models with a long-short-term feature bank consistently increases the performance.
Anan Du, Shuchao Pang, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICIP2
2020 Correlation Matters: Multi-scale Fine-Grained Contextual Information Extraction for Hepatic Tumor Segmentation
Shuchao Pang, Anan Du, Zhenmei Yu, Mehmet A. Orgun
PAKDD (1)1
2020 Weakly supervised learning for image keypoint matching using graph convolutional networks
Shuchao Pang, Anan Du, Mehmet A. Orgun, Hechang Chen
Knowl. Based Syst.1
2019 Fast and Accurate Lung Tumor Spotting and Segmentation for Boundary Delineation on CT Slices in a Coarse-to-Fine Framework
Shuchao Pang, Anan Du, Xiaoli He, Jorge Díez 0001, Mehmet A. Orgun
ICONIP (4)1
2018 Deep Learning and Preference Learning for Object Tracking: A Combined Approach
Shuchao Pang, Juan José del Coz, Zhezhou Yu, Oscar Luaces, Jorge Díez 0001
Neural Process. Lett.1
2017 Deep learning to frame objects for visual target tracking
Shuchao Pang, Juan José del Coz, Zhezhou Yu, Oscar Luaces, Jorge Díez 0001
Eng. Appl. Artif. Intell.1
2016 Combining Deep Learning and Preference Learning for Object Tracking
Shuchao Pang, Juan José del Coz, Zhezhou Yu, Oscar Luaces, Jorge Díez 0001
ICONIP (3)1