Xinshan Zhu

dblp:17/3633 · DBLP profile ↗
← Back
34ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0003-2060-9932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Security and privacy · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Selection of the best electrical impedance tomography image using a new attribute weighting c-means clustering algorithm
Shenglu Yue, Xinshan Zhu, Ming Zeng 0001
Eng. Appl. Artif. Intell.2
2026 Pedestrian trajectory prediction using multi-cue transformer
Yanlong Tian, Xiaoting Fan, Zhong Zhang 0001, Xinshan Zhu
J. Vis. Commun. Image Represent.6
2026 DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-Spoofing
abstract
The overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks.
Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Continual Deepfake Detection Based on Multi-Perspective Sample Selection Mechanism
abstract
The rapid development and malicious use of deepfakes pose a significant crisis of trust. To cope with the evolving deepfake technologies, an increasing number of detection methods adopt the continual learning paradigm, but they often suffer from catastrophic forgetting. Although replay-based methods mitigate this issue by storing a portion of samples from historical tasks, their sample selection strategies usually rely on a single metric, which may lead to the omission of critical samples and consequently hinder the construction of a robust instance memory bank. In this paper, we propose a novel Multi-perspective Sample Selection Mechanism (MSSM) for continual deepfake detection, which jointly evaluates prediction error, temporal instability, and sample diversity to preserve informative and challenging samples in the instance memory bank. Furthermore, we design a Hierarchical Prototype Generation Mechanism (HPGM) that constructs prototypes at both the category and task levels, which are stored in the prototype memory bank. Extensive experiments under two evaluation protocols demonstrate that the proposed method achieves state-of-the-art performance.
Yu Lian, Xinshan Zhu, Di He 0008, Biao Sun 0003
IEEE Signal Process. Lett.2
2026 Asymptotic Critical Transmission Radii in Wireless Ad-Hoc Networks Over a Convex Region
abstract
Critical transmission ranges (or radii) in wireless ad-hoc and sensor networks have been extensively investigated for various performance metrics such as connectivity, coverage, power assignment and energy consumption. However, the regions on which the networks are distributed are typically either squares or disks in existing works, which seriously limits the usage in applications. In this article, we consider a convex region (i.e., a generalisation of squares and disks) on which wireless nodes are uniformly distributed. We have investigated two types of critical transmission radii, defined in terms of$k-$connectivity and the minimum vertex degree, respectively, and have also established their precise asymptotic distributions. These make the previous results obtained under the circumstance of squares or disks special cases of this work. More importantly, our results reveal how the region shape impacts the critical transmission ranges: it is the length of the boundary of the (fixed-area) region that completely determines the transmission ranges. Furthermore, by isodiametric inequality, the smallest critical transmission ranges are achieved when regions are disks only.
Jie Ding 0008, Shuai Ma 0001, Xinshan Zhu
IEEE Trans. Netw.3
2025 Rethinking U-Net: Task-Adaptive Mixture of Skip Connections for Enhanced Medical Image Segmentation
abstract
U-Net is a widely used model for medical image segmentation, renowned for its strong feature extraction capabilities and U-shaped design, which incorporates skip connections to preserve critical information. However, its decoders exhibit information-specific preferences for the supplementary content provided by skip connections, instead of adhering to a strict one-to-one correspondence, which limits its flexibility across diverse tasks. To address this limitation, we propose the Task-Adaptive Mixture of Skip Connections (TA-MoSC) module, inspired by the Mixture of Experts (MoE) framework. TA-MoSC innovatively reinterprets skip connections as a task allocation problem, employing a routing mechanism to adaptively select expert combinations at different decoding stages. By introducing MoE, our approach enhances the sparsity of the model, and lightweight convolutional experts are shared across all skip connection stages, with a Balanced Expert Utilization (BEU) strategy ensuring that all experts are effectively trained, maintaining training balance and preserving computational efficiency. Our approach introduces minimal additional parameters to the original U-Net but significantly enhances its performance and stability. Experiments on GlaS, MoNuSeg, Synapse, and ISIC16 datasets demonstrate state-of-the-art accuracy and better generalization across diverse tasks. Moreover, while this work focuses on medical image segmentation, the proposed method can be seamlessly extended to other segmentation tasks, offering a flexible and efficient solution for diverse applications.
Zichen Luo, Xinshan Zhu
AAAI2
2025 Cross-modal Shared Concept Learning for Text-to-Image Person Retrieval
abstract
Text-to-image person retrieval aims to match target pedestrian images based on a text query. Existing methods mainly learn feature alignment between texts and pedestrian images from global and local perspectives. However, they ignore the matching ambiguity problem caused by different modal characteristics during the alignment learning, thereby resulting in sub-optimal performance. In this paper, we propose a cross-modal Shared Concept Learning (SCL) method, which reformulates alignment as learning shared cross-modal concepts in order to improve the semantic consistency between modalities. Specifically, we design a Shared Concept Perception (SCP) module to capture shared cross-modal concepts from global-local perspectives using decoupled visual semantic concepts guided by text concepts within the identity. Furthermore, we propose a Visual-driven Concept Matching (VCM) module to learn the identity invariance of concepts in each text under the guidance of the visual information. Extensive experiments on three text-to-image person retrieval benchmarks demonstrate the effectiveness and superiority of SCL.
Di He 0008, Xinshan Zhu, Zhong Zhang 0001
ICME2
2025 Unsupervised Person Reidentification Using Stripe-Driven Fusion Transformer Network
abstract
In recent years, some methods utilize a transformer as the backbone to model the long‐range context dependencies, reflecting a prevailing trend in unsupervised person reidentification (Re‐ID) tasks. However, they only explore the global information through interactive learning in the framework of the transformer, which ignores the learning of the part information in the interaction process for pedestrian images. In this study, we present a novel transformer network for unsupervised person Re‐ID, a stripe‐driven fusion transformer (SDFT), designed to simultaneously capture the global interaction and the part interaction when modeling the long‐range context dependencies. Meanwhile, we present a stripe‐driven regularization (SDR) to constrain the part aggregation features and the global features by considering the consistency principle from the aspects of the features and the clusters, aiming to improve the representational capacity of the features. Furthermore, to investigate the relationships between local regions of pedestrian images, we present a stripe‐driven contrastive loss (SDCL) to learn discriminative part features from the perspectives of pedestrian identity and stripes. The proposed method has undergone extensive validations on publicly available unsupervised person Re‐ID benchmarks, and the experimental results confirm its superiority and effectiveness.
Zeyu Zang, Shuang Liu 0001, Zhong Zhang 0001, Xinshan Zhu
IET Softw.5
2025 Cross-Modal Alignment Enhancement Network for Text-to-Image Person Re-Identification
Di He 0008, Xinshan Zhu, Bin Li 0028, Shenglu Yue, Zhong Zhang 0001
IEEE Internet Things J.2
2024 SAMIF: Adapting Segment Anything Model for Image Inpainting Forensics
Xinshan Zhu, Di He 0008, Xin Liao 0001, Biao Sun 0003
ACCV (7)2
2023 A transformer-CNN for deep image inpainting forensics
Xinshan Zhu, Junyan Lu, Honghao Ren, Hongquan Wang, Biao Sun 0003
Vis. Comput.1
2022 Coupling Convolution, Transformer and Graph Embedding for Motor Imagery Brain-Computer Interfaces
abstract
Over the past ten years, convolution neural network (CNN) and self-attention based models (e.g., transformer) have shown extremely competitive performance in the classification of motor imagery (MI) tasks based on electroencephalogram (EEG) signals. CNN exploits local features effectively, while self-attention based models are good at capturing long-distance feature dependencies. In this paper, we propose a hybrid network structure, termed TransEEG, that takes advantage of convolutional operations and self-attention mechanisms to model both local and global dependencies for EEG signal processing. Specifically, EEG channel relationships are exploited to build a graph embedding that further improves signal classification accuracy. We evaluated the performance of TransEEG on two datasets performed MI movements. Experiments have shown that the TransEEG significantly outperformed the previous MI classification methods and achieved state-of-the-art accuracy in subject-specifical scenario.
Zexu Wu, Xinshan Zhu
ISCAS3
2022 Training-Free Deep Generative Networks for Compressed Sensing of Neural Action Potentials
abstract
Energy consumption is an important issue for resource-constrained wireless neural recording applications with limited data bandwidth. Compressed sensing (CS) is a promising framework for addressing this challenge because it can compress data in an energy-efficient way. Recent work has shown that deep neural networks (DNNs) can serve as valuable models for CS of neural action potentials (APs). However, these models typically require impractically large datasets and computational resources for training, and they do not easily generalize to novel circumstances. Here, we propose a new CS framework, termed APGen, for the reconstruction of APs in a training-free manner. It consists of a deep generative network and an analysis sparse regularizer. We validate our method on two in vivo datasets. Even without any training, APGen outperformed model-based and data-driven methods in terms of reconstruction accuracy, computational efficiency, and robustness to AP overlap and misalignment. The computational efficiency of APGen and its ability to perform without training make it an ideal candidate for long-term, resource-constrained, and large-scale wireless neural recording. It may also promote the development of real-time, naturalistic brain-computer interfaces.
Biao Sun 0003, Chaoxu Mu, Zexu Wu, Xinshan Zhu
IEEE Trans. Neural Networks Learn. Syst.4
2021 An efficient ensemble method for missing value imputation in microarray gene expression data
abstract
BACKGROUND: The genomics data analysis has been widely used to study disease genes and drug targets. However, the existence of missing values in genomics datasets poses a significant problem, which severely hinders the use of genomics data. Current imputation methods based on a single learner often explores less known genomic data information for imputation and thus causes the imputation performance loss. RESULTS: In this study, multiple single imputation methods are combined into an imputation method by ensemble learning. In the ensemble method, the bootstrap sampling is applied for predictions of missing values by each component method, and these predictions are weighted and summed to produce the final prediction. The optimal weights are learned from known gene data in the sense of minimizing a cost function about the imputation error. And the expression of the optimal weights is derived in closed form. Additionally, the performance of the ensemble method is analytically investigated, in terms of the sum of squared regression errors. The proposed method is simulated on several typical genomic datasets and compared with the state-of-the-art imputation methods at different noise levels, sample sizes and data missing rates. Experimental results show that the proposed method achieves the improved imputation performance in terms of the imputation accuracy, robustness and generalization. CONCLUSION: The ensemble method possesses the superior imputation performance since it can make use of known data information more efficiently for missing data imputation by integrating diverse imputation methods and learning the integration weights in a data-driven way.
Xinshan Zhu, Chao Ren 0003
BMC Bioinform.1
2021 Multi-Stream Fusion Network With Generalized Smooth L1 Loss for Single Image Dehazing
abstract
Single image dehazing is an important but challenging computer vision problem. For the problem, an end-to-end convolutional neural network, named multi-stream fusion network (MSFNet), is proposed in this paper. MSFNet is built following the encoder-decoder network structure. The encoder is a three-stream network to produce features at three resolution levels. Residual dense blocks (RDBs) are used for feature extraction. The resizing blocks serve as bridges to connect different streams. The features from different streams are fused in a full connection manner by a feature fusion block, with stream-wise and channel-wise attention mechanisms. The decoder directly regresses the dehazed image from coarse to fine by the use of RDBs and the skip connections. To train the network, we design a generalized smooth L1 loss function, which is a parametric loss family and permits to adjust the insensitivity to the outliers by varying the parameter settings. Moreover, to guide MSFNet to capture the valid features in each stream, we propose the multi-scale supervision learning strategy, where the loss at each resolution level is computed and summed as the final loss. Extensive experimental results demonstrate that the proposed MSFNet achieves superior performance on both synthetic and real-world images, as compared with the state-of-the-art single image dehazing methods.
Xinshan Zhu, Shuoshi Li, Yongdong Gan, Yun Zhang 0003, Biao Sun 0003
IEEE Trans. Image Process.1
2020 A Deep Learning Approach in the Discrete Cosine Transform Domain to Median Filtering Forensics
abstract
This letter presents a novel median filtering forensics approach, based on a convolutional neural network (CNN) with an adaptive filtering layer (AFL), which is built in the discrete cosine transform (DCT) domain. Using the proposed AFL, the CNN can determine the main frequency range closely related with the operational traces. Then, to automatically learn the multi-scale manipulation features, a multi-scale convolutional block is developed, exploring a new multi-scale feature fusion strategy based on the maxout function. The resultant features are further processed by a convolutional stream with pooling and batch normalization operations, and finally fed into the classification layer with the Softmax function. Experimental results show that our proposed approach is able to accurately detect the median filtering manipulation and outperforms the state-of-the-art schemes, especially in the scenarios of low image resolution and serious compression loss.
Yixin Liao, Xinshan Zhu, Hongquan Wang
IEEE Signal Process. Lett.3
2018 A deep learning approach to patch-based image inpainting forensics
Xinshan Zhu, Yongjun Qian, Xianfeng Zhao
Signal Process. Image Commun.1
2016 A Note on the k-NN Density Estimate
Xinshan Zhu
IDEAL2
2016 Calculating the response time based on action flow in Stochastic Process Algebra models
abstract
Response time plays an important factor in determining the Service Level Agreement (SLA). For the reason that actual measurement costs a large amount of resource, theoretical/numerical analysis based on Stochastic Process Algebra (SPA) is a good choice to obtain the response time of concurrent systems. Among all SPAs, Performance Evaluation Process Algebra (PEPA) is the most popular one due to its precise semantics. As a result, this paper gives two methods, theoretical and numerical, for analyzing response time between two specified actions. These two methods are restricted in the scenarios that there are no actions can be performed parallelly in a response. In addition, theoretical analysis just applies to small scale models.
Leijie Sha, Xinshan Zhu
SMC3
2016 An EL-LDA based general color harmony model for photo aesthetics assessment
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process.3
2015 Finding more relevance: Propagating similarity on Markov random field for object retrieval
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process. Image Commun.3
2014 Normalized Correlation-Based Quantization Modulation for Robust Watermarking
abstract
A novel quantization watermarking method is presented in this paper, which is developed following the established feature modulation watermarking model. In this method, a feature signal is obtained by computing the normalized correlation (NC) between the host signal and a random signal. Information modulation is carried out on the generated NC by selecting a codeword from the codebook associated with the embedded information. In a simple case, the structured codebooks are designed using uniform quantizers for modulation. The watermarked signal is produced to provide the modulated NC in the sense of minimizing the embedding distortion. The performance of the NC-based quantization modulation (NCQM) is analytically investigated, in terms of the embedding distortion and the decoding error probability in the presence of valumetric scaling and additive noise attacks. Numerical simulations on artificial signals confirm the validity of our analyses and exhibit the performance advantage of NCQM over other modulation techniques. The proposed method is also simulated on real images by using the wavelet-based implementations, where the host signal is constructed by the detail coefficients of wavelet decomposition at the third level and transformed into the NC feature signal for the information modulation. Experimental results show that the proposed NCQM not only achieves the improved watermark imperceptibility and a higher embedding capacity in high-noise regimes, but also is more robust to a wide range of attacks, e.g., valumetric scaling, Gaussian filtering, additive noise, Gamma correction, and Gray-level transformations, as compared with the state-of-the-art watermarking methods.
Xinshan Zhu, Jie Ding 0008, Honghui Dong, Kongfa Hu
IEEE Trans. Multim.1
2012 A novel quantization watermarking scheme by modulating the normalized correlation
abstract
This paper presents a novel quantization based watermarking scheme. Watermark embedding is performed through modulating the normalized correlation between the host vector and a random vector with dither modulation. The watermarked signal is derived to provide the modulated normalized correlation in the sense of minimizing the embedding distortion. The proposed scheme is theoretically invariant to valumetric scaling and can resist stronger noise than the well-known spread transform dither modulation. Numerical simulations on real images show that it achieves the good imperceptibility and strong robustness against a wide range of attacks.
Xinshan Zhu, Shuoling Peng
ICASSP1
2011 Detection for Multiplicative Watermarking in DCT Domain by Cauchy Model
Shuoling Peng, Xinshan Zhu
ICICS3
2011 Analyzing the Performance of Dither Modulation in Presence of Composite Attacks
Xinshan Zhu
ICICS1
2008 Improved quantization index modulation watermarking robust against amplitude scaling and constant change distortions
abstract
The original quantization index modulation watermarking is largely vulnerable to valumetric scaling and constant change attacks. To overcome these two drawbacks, the normalized dither modulation (NDM) is presented in this paper. The main idea of it is to construct a gain-invariant vector with zero mean for quantization. Performance analysis shows that NDM is theoretically invariant to valumetric scaling and constant change and achieves similar performance to dither modulation (DM) in other aspects. Some useful strategies are further provided to improve the performance of NDM. Experiments on real data (images) demonstrate that not only the proposed method is extremely robust to amplitude scaling and constant change, and outperform the original DM subject to several other typical attacks.
Xinshan Zhu, Zhi Tang 0001
ICIP1
2008 Improved quantization index modulation watermarking robust against amplitude scaling distortions
abstract
The original quantization index modulation (QIM) watermarking is largely vulnerable to valumetric scaling. To solve this problem, a gain-invariant vector for quantization is constructed through dividing the host signal by a statistical feature extracted from the target content. With this idea, the improved dither modulation (IM-DM) and spread transform dither modulation (IM-STDM) are developed in this paper. We present the strategies on the choice of the introduced feature and analyze the performance of two improved QIM schemes theoretically. It is shown that the improved QIM is theoretically invariant to valumetric scaling, but becomes sensitive to constant change. Experiments on real data (images) demonstrate that the proposed methods are extremely robust to amplitude scaling and possess the similar performance to the original QIM subject to several other typical attacks.
Xinshan Zhu, Zhi Tang 0001
ICME1
2008 A new spatial perceptual mask for image watermarking
abstract
This paper develops a new mask in the spatial domain for image watermarking. The mask exploits the properties of human visual system with respect to background luminance, the edge and texture masking, which are expressed by certain image features. On the use of the mask, a new double domain watermarking framework is presented. The watermark is embedded into one transform domain, but perceptually shaped in the spatial domain. It allows us to use the features of double domains for embedding. Experimental results demonstrate that the proposed mask has superior performance compared to existing spatial masking schemes, and further achieves the improved performance by incorporating the proposed watermarking framework.
Xinshan Zhu
ICPR1
2007 Steganography using Sensor Noise and Linear Prediction Synthesis Filter
abstract
This paper presents a new approach utilizing the sensor's pattern noise and linear prediction synthesis filter for steganography. The pattern noise is extracted from the images using denoising filter (for example wavelet-based filter). Then the approach introduces the linear prediction synthesis filter, whose parameters are derived from the extracted noise. After being filtered by such a filter, the secret message can be embedded by adapting the characteristics of the sensor's pattern noise. As a result, the embedding process violate little of the natural image statistics, and hence the detectability of steganalytic method is noticeably decreased. The experimental results prove the effectiveness of the new approach.
Xiaoyi Yu, Xinshan Zhu, Noboru Babaguchi
ICIP (2)2
2006 Image-adaptive watermarking based on perceptually shaping watermark blockwise
abstract
In a general additive watermarking model, a watermark signal is perceptually shaped and scaled with a global gain factor before embedding. This paper presents a new image-adaptive watermarking scheme based on perceptually shaping watermark blockwise. Instead of the global gain factor, a localized one is used for each block. And Watson's DCT-based visual model is adopted to measure the distortion of each block introduced by watermark, rather than the whole image. With the given distortion constraint, the maximum output value of linear correlation detector is derived in one block, which demonstrates the reachable maximum robustness in a sense. Meanwhile, an extended perceptually shaped watermarking (EX-PSW) is acquired through making detection value approach that upper limit. It is proved mathematically that EX-PSW outputs higher detection value than perceptually shaped watermarking (PSW) with the same distortion constraint. We also discuss the adjustment strategies of parameters in EX-PSW, which are helpful for improving the local image quality. Experimental results show our scheme provides very good results both in terms of image transparency and robustness.
Xinshan Zhu, Yong Gao 0002, Yan Zhu 0010
AsiaCCS1
2006 Collusion secure convolutional fingerprinting information codes
abstract
Digital Fingerprinting is a technique for the merchant who can embed unique buyer identity marks into digital media copy, and also makes it possible to identify traitors who redistribute their illegal copies. At present, the fingerprinting scheme generally have many difficulties and disadvantages for large-size uses problems involve in the code construction with shorter length and effective traitor tracing. To resolve these problems, this paper presents the definition of Fingerprinting Information Code and a practical construction method by composing of convolutional codes and generally fingerprinting codes based on Boneh-Shaw model. Its decoding algorithm is presented by introducing the ideal of 'Optional Code Subset' and improving Viterbi algorithm. The security properties and performance are proved and analyzed by theory and example. As the results, the proposed scheme has shorter information encoding length and achieves optimal traitor searching in larger number of buyers.
Yan Zhu 0010, Xinshan Zhu
AsiaCCS3
2006 A Novel Multibit Watermarking Scheme Combining Spread Spectrum and Quantization
Xinshan Zhu, Zhi Tang 0001, Liesen Yang
IWDW1
2005 A Voting Method and Its Application in Precise Object Location
Yong Gao 0002, Xinshan Zhu, Xiangsheng Huang, Yangsheng Wang
ACII2
2004 Better Use of Human Visual Model in Watermarking Based on Linear Prediction Synthesis Filter
Xinshan Zhu, Yangsheng Wang
IWDW1