Zhenshan Tan

dblp:253/0308 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0003-3466-5417ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 High-Resolution Image Steganalysis
abstract
Image steganalysis detects hidden information within images. However, existing methods are primarily designed for low-resolution images and they struggle to address the challenges posed by the widespread use of high-resolution images in real-world scenarios such as communications and social media. Under the secure embedding constrained by the square root law, the steganographic noise of high-resolution images is significantly diluted, thereby making “strong decision regions” critical to detection increasingly scarce, whereas the interference effect of “weak decision regions” is relatively prominent. Existing methods treat all areas equally, making it difficult to fully utilize strong decision regions and suppress the negative impact of weak decision regions. To address these issues, we propose the HRIS, a two-phase collaborative optimization framework for high-resolution steganalysis from “discovery” to “utilization”. In the “discovery” phase, we propose a dynamically contribution-guided decision region recognition mechanism. This mechanism employs a difference amplification module to amplify the steganographic noise differences between regions and then leverages a cooperative game-driven dynamic optimization strategy to compute each sub-region's contribution to the prediction. It accurately identifies and reinforces strong decision regions while suppressing interference from weak decision regions, resulting in significantly improved local detection accuracy. In the “utilization” phase, we propose a decision regions-global steganographic features bidirectional collaborative optimization framework that leverages the identified strong and weak decision regions to direct the extraction of global steganographic noise features. These global features are then fed back to refine the local feature representations, enabling collaborative enhancement between local and global analyses. Extensive experimental results demonstrate that our method achieves state-of-the-art performance on high-resolution images.
Xinjue Hu, Zhenshan Tan, Xiang Zhang 0023, Zhangjie Fu 0001
IEEE Trans. Dependable Secur. Comput.2
2025 AdvMap: Crafting Adversarial Maps to Counter AI Aimbot in First-Person Shooter Games
abstract
AI-based automatic aiming cheats (a.k.a., AI aimbots) have proliferated inFirst-Person Shootergames, which grant malicious users an unfair gameplay advantage. Since AI aimbots operate independently of game data and are developed using object detection algorithms, they are difficult to detect with traditional anti-cheating methods. To actively counter AI aimbots, we proposeAdvMap, which introduces invisible adversarial perturbations into game scene elements. In optimizing these adversarial perturbations, we design a mixture-of-misleading loss function that increases the total target confidence score within each misleading bounding box. It mitigates the risk of segment missing even whenAdvMapis obscured, thereby enhancing the robustness. Besides, an L1-norm constraint with a small scale is employed during each update of the adversarial perturbations, which preserves the fidelity of the game scene. In addition, to enable effective adaptation to interact with various elements within game environments, we introduce an image-subspace-based multidirectional optimization strategy. It enables the adversarial perturbations to adaptively fit into each element by leveraging the mapping relationship between the game's 3-D scenes and its corresponding 2-D images. Furthermore, we construct a comprehensive benchmark, which includes various FPS games with different graphics styles and perspectives. Extensive experimental results demonstrate the efficacy of our method in countering various AI aimbot tools on different state-of-the-art object detection methods.
Xianyi Chen, Haoqin Yuan, Zhenshan Tan
IEEE Trans. Games5
2025 Color Decoupling for Multi-Illumination Color Constancy
abstract
Current multi-illumination color constancy methods typically estimate illumination for each pixel directly. However, according to the multi-illumination imaging equation, the color of each pixel is determined by various components, including the innate color of the scene content, the colors of multiple illuminations, and the weightings of these illuminations. Failing to distinguish between these components results in color coupling. On the one hand, there is color coupling between illumination and scene content, where estimations are easily misled by the colors of the content, and the distribution of the estimated illuminations is relatively scattered. On the other hand, there is color coupling between illuminations, where estimations are susceptible to interference from high-frequency and heterogeneous illumination colors, and the local contrast is low. To address color coupling, we propose a Color Decoupling Network (CDNet) that includes a Content Color Awareness Module (CCAM) and a Contrast HArmonization Module (CHAM). CCAM learns scene content color priors, decoupling the colors of content and illuminations by providing the model with the color features of the content, thereby reducing out-of-gamut estimations and enhancing consistency. CHAM constrains feature representation, decoupling illuminants by mutual calibration between adjacent features. CHAM utilizes spatial correlation to make the model more sensitive to the relationships between neighboring features and utilizes illumination disparity degree to guide feature classification. By enhancing the uniqueness of homogeneous illumination features and the distinctiveness of heterogeneous illumination features, CHAM improves local edge contrast. Additionally, by allocating fine-grained margin coefficients to emphasize the soft distinctiveness of similar illumination features, further enhancing local contrast. Extensive experiments on single- and multi-illumination benchmark datasets demonstrate that the proposed method achieves superior performance.
Zhenshan Tan, Zhijiang Li
IEEE Trans. Circuits Syst. Video Technol.2
2025 More Unlabeled Data Does Matter: A Full-Cycle Framework for Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
Semi-supervised semantic segmentation has gained considerable attention due to its ability to leverage large amounts of unlabeled data to enhance model generalization. Although increasing the amount of unlabeled data ideally enhances performance, reality contradicts this notion. Research indicates that the reasons for performance degradation span the entire semi-supervised learning process. These include the inconsistent distribution of unlabeled and labeled data at the data level, the susceptibility to interference from unknown categories at the model level, and the semantic drift problem caused by the training strategy. This article refers to these challenges as the “full-cycle” challenge. To address these issues, we propose a full-cycle framework. At the data level, we develop a consistency-based method, beginning with unlabeled data screening, to prioritize unlabeled data that closely aligns with the labeled data distribution in training. At the model level, we introduce a consistency-based semi-supervised semantic segmentation model capable of identifying unknown category regions and distancing their features from known category features, reducing the interference of unknown classes. At the training strategy level, we adopt a progressive approach that gradually incorporates data from simpler to more complex cases. This strategy improves the model’s adaptability to potential data changes, ensures stability, and improves overall generalization performance. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches, validating its effectiveness in addressing the full-cycle challenge.
Zhenshan Tan, Guo Zhang 0001, Zhijiang Li
IEEE Trans. Geosci. Remote. Sens.2
2025 A Bias Correction Semi-Supervised Semantic Segmentation Framework for Remote Sensing Images
abstract
The clustering assumption is widely adopted in semi-supervised semantic segmentation methods. However, this assumption heavily relies on high-quality feature representation, leading to learning and cognitive biases if it does not hold. Learning bias entails the potential overfitting of labeled data, leading to capturing local features and consequently misclassifying semantic categories. Cognitive bias indicates the network’s susceptibility to interference features. Recently, a consistency-based mechanism has been proposed to address these biases. By subjecting unlabeled data to diverse weak perturbations, it breaks original clustering features, compelling the model to learn more robust and generalized representations. However, when processing complex remote sensing images, these weak perturbations often prove ineffective in the later stages of training. To address this, we propose a bias correction framework (BCF). The BCF begins with a feature consistency enhancement module (CEM) that guides the student model to learn feature representations with greater generalization capability. In the feature decoding stage, we introduce a multidecoder structure with weakly orthogonal weights to maximize feature differences, thereby further reducing learning and cognitive biases. In addition, to improve the confidence of pseudolabels and enhance consistency learning, we design a multidecoder teacher model based on symmetrical knowledge transfer, allowing the diverse and multiangle information learned by the student model to be transferred to the teacher model. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on the ISPRS Potsdam and Vaihingen datasets.
Zhenshan Tan, Yuzhi Zheng, Guo Zhang 0001, Zhijiang Li
IEEE Trans. Geosci. Remote. Sens.2
2024 Co-saliency detection with two-stage co-attention mining and individual calibration
Zhenshan Tan, Xiaodong Gu 0001, Qingrong Cheng
Eng. Appl. Artif. Intell.1
2024 SPNet: Semantic preserving network with semantic constraint and non-semantic calibration for color constancy
Zhijiang Li, Zhenshan Tan
Neurocomputing4
2024 Bridging spatiotemporal feature gap for video salient object detection
Zhenshan Tan, Keyu Wen, Qingrong Cheng, Zhangjie Fu 0001
Knowl. Based Syst.1
2024 Semantic Pre-Alignment and Ranking Learning With Unified Framework for Cross-Modal Retrieval
abstract
Cross-modal retrieval aims at retrieving highly semantic relevant information among multi-modalities. Existing cross-modal retrieval methods mainly explore the semantic consistency between image and text while rarely consider the rankings of positive instances in the retrieval results. Moreover, these methods seldom take into account the cross-interaction between image and text, which leads to the deficiency of learning their semantic relations. In this paper, we propose a Unified framework with Ranking Learning (URL) for cross-modal retrieval. The unified framework consists of three sub-networks, visual network, textual network, and interaction network. Visual network and textual network project the image feature and text feature into their corresponding hidden spaces respectively. Then, the interaction network forces the target image-text representation to align in the common space. For unifying both semantics and rankings, we propose a new optimization paradigm including pre-alignment for semantic knowledge transfer and ranking learning for final retrieval, which can decouple semantic alignment and ranking learning. The former focuses on the semantic pre-alignment optimized by semantic classification and the latter revolves around the retrieval rankings. For the ranking learning, we introduce a cross-AP loss which can directly optimize the retrieval metric average precision for cross-modal retrieval. We conduct experiments on four widely-used benchmarks, including Wikipedia dataset, Pascal Sentence dataset, NUS-WIDE-10k dataset, and PKU XMediaNet dataset respectively. Extensive experimental results show that the proposed method can obtain higher retrieval precision.
Qingrong Cheng, Zhenshan Tan, Keyu Wen, Xiaodong Gu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learn More and Learn Usefully: Truncation Compensation Network for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Semantic segmentation of high-resolution remote-sensing images (HR-RSIs) focuses on classifying each pixel of input images. Recent methods have incorporated a downscaled global image as supplementary input to alleviate global context loss from cropping. Nonetheless, these methods encounter two key challenges: diminished detail in features due to down-sampling of the global auxiliary image, and noise from the same image that reduces the network’s discriminability of useful and useless information. To overcome these challenges, we propose a truncation compensation network (TCNet) for HR-RSI semantic segmentation. TCNet features three pivotal modules: the guidance feature extraction module (GFM), the related-category semantic enhancement module (RSEM), and the global-local contextual cross-fusion module (CFM). GFM focuses on compensating for truncated features in the local image and minimizing noise to emphasize learning of useful information. RSEM enhances discernment of global semantic information by predicting spatial positions of related categories and establishing spatial mappings for each. CFM facilitates local image semantic segmentation with extensive contextual information by transferring information from global to local feature maps. Extensive testing on the ISPRS, BLU, and GID datasets confirms the superior efficiency of TCNet over other approaches.
Zhenshan Tan, Guo Zhang 0001, Zhijiang Li
IEEE Trans. Geosci. Remote. Sens.2
2023 A Unified Video Semantics Extraction and Noise Object Suppression Network for Video Saliency Detection
Zhenshan Tan, Xiaodong Gu 0001
ICANN (7)1
2023 Triplet Spatiotemporal Aggregation Network for Video Saliency Detection
abstract
The effective aggregation of spatiotemporal information to accommodate real-world complex scenes is a fundamental issue in video saliency detection. In this paper, we propose a Triplet Spatiotemporal Aggregation Network (TSAN) to address it from the aggregation of spatiotemporal interaction, spatiotemporal information distribution, and multi-level spatiotemporal features. Firstly, we propose an interactive aggregation gate (IAG) module to model spatial and temporal global context information and perform inter-modal information transfer. Secondly, we employ an information distribution consistency (IDC) module to enhance the consistency of spatiotemporal representation by maximizing the correlation of spatiotemporal high-level features. Finally, we design a multi-level spatiotemporal feature aggregation (MSF) framework to merge cross-level and cross-modal features. These three modules are combined into a unified framework to jointly optimize spatiotemporal information for more precise results. Experimental results on five prevailing datasets show that TSAN outperforms previous competitors.
Zhenshan Tan, Xiaodong Gu 0001
ICME1
2022 UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog
abstract
Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1.
Zhenshan Tan, Qingrong Cheng, Xin Jiang 0002, Qun Liu 0001, Yudong Zhu, Xiaodong Gu 0001
CVPR2
2022 A Unified Multiple Inducible Co-attentions and Edge Guidance Network for Co-saliency Detection
Zhenshan Tan, Xiaodong Gu 0001
ICANN (1)1
2022 Feature Recalibration Network for Salient Object Detection
Zhenshan Tan, Xiaodong Gu 0001
ICANN (4)1
2022 A Unified Two-Stage Group Semantics Propagation and Contrastive Learning Network for Co-Saliency Detection
abstract
Co-saliency detection (CoSOD) aims at discovering the repetitive salient objects from multiple images. Two primary challenges are group semantics extraction and noise object suppression. In this paper, we present a unified Two-stage grOup semantics PropagatIon and Contrastive learning NETwork (TopicNet) for CoSOD. TopicNet can be decomposed into two substructures, including a two-stage group semantics propagation module (TGSP) to address the first challenge and a contrastive learning module (CLM) to address the second challenge. Concretely, for TGSP, we design an image-to-group propagation module (IGP) to capture the consensus representation of intra-group similar features and a group-to-pixel propagation module (GPP) to build the relevancy of consensus representation. For CLM, with the design of positive samples, the semantic consistency is enhanced. With the design of negative samples, the noise objects are suppressed. Experimental results on three prevailing benchmarks reveal that TopicNet outperforms other competitors in terms of various evaluation metrics.
Zhenshan Tan, Keyu Wen, Yuzhuo Qin, Xiaodong Gu 0001
ICME1
2022 Co-saliency detection with intra-group two-stage group semantics propagation and inter-group contrastive learning
Zhenshan Tan, Xiaodong Gu 0001
Knowl. Based Syst.1
2022 Visual context learning based on textual knowledge for image-text retrieval
Yuzhuo Qin, Xiaodong Gu 0001, Zhenshan Tan
Neural Networks3
2021 Depth scale balance saliency detection with connective feature pyramid and edge guidance
Zhenshan Tan, Xiaodong Gu 0001
Appl. Intell.1
2020 Salient Object Detection with Edge Recalibration
Zhenshan Tan, Yikai Hua, Xiaodong Gu 0001
ICANN (1)1
2020 SBN: Scale Balance Network for Accurate Salient Object Detection
abstract
Recent great progress has been made on Salient Object Detection (SOD) by deep Convolutional Neural Networks (CNNs). However, most SOD methods still suffer from scale imbalance issue, which pays more attention on large salient areas but ignores small salient areas though they belong to the same object. To address this issue, this paper proposes a Scale Balance Network (SBN) to accurately locate large salient areas and recognize small salient areas. Firstly, a backbone network specifically designed for object detection is adopted in this paper, which captures larger resolution with more spatial features in deeper layers. Secondly, to focus on the balance between the large salient areas and the small salient areas, this paper proposes a novel Connective Feature Pyramid Module (CFPM) for sufficiently leveraging the multi-scale features and the multi-level features, which includes Feature Coherence Enhancement (FCE) and Feature Progressive Extraction (FPE). FCE is designed to enhance the coherence between high-level and low-level features, and FPE is designed to extract the progressive features in different convolutional layers. Finally, an Edge Enhancement Architecture with Various Kernels (EEAVK) is proposed to refine the edge features. Experimental results on five benchmark datasets show that the proposed method outperforms or achieves consistently superior performance in comparison with other methods under different evaluation metrics.
Zhenshan Tan, Xiaodong Gu 0001
IJCNN1
2019 Directive local color transfer based on dynamic look-up table
abstract
Color transfer in image processing usually suffers from misleading color mapping and loss of details. This paper presents a novel directive local color transfer method based on dynamic look-up table (D-DLT) to solve these problems in two steps. First, a directive mapping between the source and the reference image is established based on the salient detection and the color clusters to obtain directive color transfer intention. Then, dynamic look-up tables are created according to the color clusters to preserve the details, which can suppress pseudo contours and avoid detail loss. Subjective and objective assessments are presented to verify the feasibility and the availability of the proposed approach. Experimental results demonstrate that our proposed method has better performance on natural color images than classical color transfer algorithms. Furthermore, the reference image can be extended to color blocks instead of images.
Zhijiang Li, Zhenshan Tan, Liqin Cao, Lei Jiao 0001, Yanfei Zhong
Signal Process. Image Commun.2