Chengxin Zhao

dblp:258/3442 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 GlyphShield: Document Watermarking for the Physical World via Vector Typeface Synthesis
Yuxing Lu, Han Fang 0004, Sijing Xie, Luyu Yuan, Chengxin Zhao
AAAI7
2025 END^2: Robust Dual-Decoder Watermarking Framework Against Non-Differentiable Distortions
abstract
DNN-based watermarking methods have rapidly advanced, with the ``Encoder-Noise Layer-Decoder'' (END) framework being the most widely used. To ensure end-to-end training, the noise layer in the framework must be differentiable. However, real-world distortions are often non-differentiable, leading to challenges in end-to-end training. Existing solutions only treat the distortion perturbation as additive noise, which does not fully integrate the effect of distortion in training. To better incorporate non-differentiable distortions into training, we propose a novel dual-decoder architecture (END^2). Unlike conventional END architecture, our method employs two structurally identical decoders: the Teacher Decoder, processing pure watermarked images, and the Student Decoder, handling distortion-perturbed images. The gradient is backpropagated only through the Teacher Decoder branch to optimize the encoder thus bypassing the problem of non-differentiability. To ensure resistance to arbitrary distortions, we enforce alignment of the two decoders' feature representations by maximizing the cosine similarity between their intermediate vectors on a hypersphere. Extensive experiments demonstrate that our scheme outperforms state-of-the-art algorithms under various non-differentiable distortions. Moreover, even without the differentiability constraint, our method surpasses baselines with a differentiable noise layer. Our approach is effective and easily implementable across all END architectures, enhancing practicality and generalizability.
Han Fang 0004, Yuxing Lu, Chengxin Zhao
AAAI4
2025 AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions
abstract
Deep watermarking models optimize robustness by incorporating distortions between the encoder and decoder. To tackle non-differentiable distortions, current methods only train the decoder with distorted images, which breaks the joint optimization of the encoder-decoder, resulting in suboptimal performance. To address this problem, we propose an Adversarial Distortion Domain Translation (AD2T) method by treating the distortion as an image-to-image translation task. AD2T adopts conditional GANs to learn the non-differentiable distortion mappings. It employs generators to transform the encoded image into the distorted one to bridge the encoder-decoder for joint optimization. We also supervise the GANs to generate challenging distorted samples to augment the watermarking model via adversarial training. This further improves the model robustness by minimizing the maximum decoding loss. Extensive experiments demonstrate the superiority of our method when tested on non-differentiable distortions, including lossy compression and style transfers. Codes are released here: https://github.com/zcx-language/AdversarialDistortionDomainTranslation.
Chengxin Zhao, Jiazhong Chen, Han Fang 0004, Zongyi Li, Sijing Xie
ICASSP1
2025 Ultra-high Resolution Watermarking Framework Resistant to Extreme Cropping and Scaling
abstract
Recent developments in DNN-based image watermarking techniques have achieved impressive results in protecting digital content. However, most existing methods are constrained to low-resolution images as they need to encode the entire image, leading to prohibitive memory and computational costs when applied to high-resolution images. Moreover, they lack robustness to distortions prevalent in large-image transmission, such as extreme scaling and random cropping. To address these issues, we propose a novel watermarking method based on implicit neural representations (INRs). Leveraging the properties of INRs, our method employs resolution-independent coordinate sampling mechanism to generate watermarks pixel-wise, achieving ultra-high resolution watermark generation with fixed and limited memory and computational resources. This design ensures strong robustness in watermark extraction, even under extreme cropping and scaling distortions. Additionally, we introduce a hierarchical multi-scale coordinate embedding and a low-rank watermark injection strategy to ensure high-quality watermark generation and robust decoding. Experimental results demonstrate that our method significantly outperforms existing schemes in terms of both robustness and computational efficiency while preserving high image quality. Our approach achieves an accuracy greater than 98\% in watermark extraction with only 0.4\% of the image area in 2K images. These results highlight the effectiveness of our method, making it a promising solution for large-scale and high-resolution image watermarking applications.
Luyu Yuan, Han Fang 0004, Yuxing Lu, Sijing Xie, Chengxin Zhao
NeurIPS7
2025 GSyncCode: Geometry Synchronous Hidden Code for One-step Photography Decoding
abstract
Invisible hyperlinks and hidden barcodes have recently emerged as a hot topic in offline-to-online messaging, where an invisible message or barcode is embedded in an image and can be decoded via camera shooting. Current schemes involve a two-step decoding process: starting with vertex localization of the embedded region to correct the perspective distortion introduced by shooting, followed by decoding the message from the corrected region. However, vertex localization can be complex and time-consuming, which affects the efficiency and accuracy of message decoding. To address this issue, this article proposes a geometry synchronous decoding scheme called GSyncCode, allowing for one-step extraction of a Data Matrix code from the photograph. Instead of correction before decoding, GSyncCode directly decodes a geometry-transformed Data Matrix that is synchronized with the embedded region. A barcode scanner is then used to efficiently retrieve messages. We design a Haar transform-based encoder HaarUNet and a HaarLoss visual function to select the key component of the Data Matrix for embedding. They improve the visual quality of the embedded image by reducing redundant embedding signals. Extensive simulated and real-world experiments demonstrate the superiority of GSyncCode in both decoding efficiency and accuracy. Our codes are published at: https://github.com/zcx-language/GSyncCode .
Chengxin Zhao, Jialie Shen 0001, Han Fang 0004, Sijing Xie, Yaokun Fang, Zongyi Li, Ping Li 0021
ACM Trans. Multim. Comput. Commun. Appl.1
2024 DITW: A High-Performance Deep-Independent Template-Based Watermarking
abstract
Watermarking algorithms based on deep Convolutional Neural Networks (CNN) have been extensively studied and shown to effectively improve performance. Most deep watermarking algorithms are dependent on the participation of host images, which results in more time and computing resources for embedding watermarks. In order to reduce the complexity of watermark embedding while retaining the performance optimization brought by CNN, we propose a high-performance Deep-Independent Template-based Watermarking (DITW). The proposed method generates deep templated-watermarks based on secret messages independently, without the involvement of host images. Then the embedding process is implemented through additive operation, which greatly improves the efficiency of embedding watermarks. To improve the performance of the network, We design the pixel-change loss and a learnable Embedding Strength (ES) Matrix Adaptor which replaces the universal ES. Through extensive experiments, we demonstrate that our scheme outperforms the existing state-of-the-art deep template-based watermarking algorithms in terms of imperceptibility and capacity without robustness degradation.
Yaokun Fang, Changxi Huang, Chengxin Zhao, Xunjie Lin, Jinlong Guo
ICASSP3
2024 Picking watermarks from noise (PWFN): an improved robust watermarking model against intensive distortions
abstract
Digital watermarking is the process of embedding secret information by altering images in an undetectable way to the human eye. To increase the robustness of the model, many deep learning-based watermarking methods use the encoder-noise-decoder architecture by adding different noises to the noise layer. The decoder then extracts the watermarked information from the distorted image. However, this method can only resist weak noise attacks. To improve the robustness of the decoder against stronger noise, this paper proposes to introduce a denoise module between the noise layer and the decoder. The module aims to reduce noise and recover some of the information lost caused by distortion. Additionally, the paper introduces the SE module to fuse the watermarking information pixel-wise and channel dimensions-wise, improving the encoder’s efficiency. Experimental results show that our proposed method is comparable to existing models and outperforms state-of-the-art under different noise intensities. In addition, ablation experiments show the superiority of our proposed module.
Sijing Xie, Chengxin Zhao, Wei Li 0151
ICME2
2024 SSyncOA: Self-synchronizing Object-aligned Watermarking to Resist Crop-paste Attacks
abstract
Modern image processing tools can easily crop local objects from images and paste them elsewhere. The challenge posed by this crop-paste attack is that it breaks the synchronization of the image watermark by inducing multiple superimposed desynchronization distortions. Existing image watermarking methods can only resist a single type of desynchronization attack and are inapplicable to this scenario. Finding that the key to resisting the crop-paste attack lies in the geometrically robust features of the object itself, this paper proposes a Self-Synchronizing Object-Aligned watermarking scheme, called SSyncOA. Specifically, we design a self-synchronization process that normalizes the watermark region, the centroid, the principal direction, and the minimum bounding square of the object during encoding and decoding to achieve synchronization of cropping, translation, rotation, and scaling, respectively. In cooperation with SSync, we propose an object-aligned watermarking method that embeds and extracts watermark messages only from the object region. This is achieved by training the watermarking model end-to-end with crop-paste attacks introduced between the encoder and decoder. Extensive experiments illustrate the impact of different desynchronization distortions on the trained watermark model, as well as the superior performance of our method compared to other SOTAs.
Chengxin Zhao, Sijing Xie, Han Fang 0004, Yaokun Fang
ICME1
2024 DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
abstract
Embedding invisible hyperlinks or hidden codes in images to replace QR codes has become a hot topic recently. This technology requires first localizing the embedded region in the captured photos before decoding. Existing methods that train models to find the invisible embedded region struggle to obtain accurate localization results, leading to degraded decoding accuracy. This limitation is primarily because the CNN network is sensitive to low-frequency signals, while the embedded signal is typically in the high-frequency form. Based on this, this paper proposes a Dual-Branch Dual-Head (DBDH) neural network tailored for the precise localization of invisible embedded regions. Specifically, DBDH uses a low-level texture branch containing 62 high-pass filters to capture the high-frequency signals induced by embedding. A high-level context branch is used to extract discriminative features between the embedded and normal regions. DBDH employs a detection head to directly detect the four vertices of the embedding region. In addition, we introduce an extra segmentation head to segment the mask of the embedding region during training. The segmentation head provides pixel-level supervision for model learning, facilitating better learning of the embedded signals. Based on two state-of-the-art invisible offline-to-online messaging methods, we construct two datasets and augmentation strategies for training and testing localization models. Extensive experiments demonstrate the superior performance of the proposed DBDH over existing methods.
Chengxin Zhao, Sijing Xie, Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen
IJCNN1
2024 Background adaptive PosMarker based on online generation and detection for locating watermarked regions in photographs
Chengxin Zhao, Jinlong Guo, Zhenghai He
J. Vis. Commun. Image Represent.1
2024 Knowledge Consistency Distillation for Weakly Supervised One Step Person Search
abstract
Weakly supervised person search targets to detect and identify a person with only bounding box annotations. Recent approaches have focused on learning person relations in a single model, ignoring the conflicts between the detection and Re-ID heads, along with the influence of background elements, which may lead to noisy pseudo labels and inaccurate Re-ID features. To address this challenge, we introduce a novel framework named Knowledge Consistency Distillation (KCD) for weakly supervised person search, which explores the capabilities of an advanced unsupervised person re-identification (Re-ID) model to mitigate the conflicts and background influences. We propose hierarchical consistency alignments, including feature-level, cluster-level, and instance-level consistency alignment, to synchronize the knowledge from the state-of-the-art unsupervised Re-ID model. Specifically, the feature-level consistency aligns the feature through both context and relation alignment. The cluster-level consistency aligns the teacher cluster information by reusing its OIM module. To tackle the inconsistency problem between student instances and teacher cluster centroids, we incorporate pseudo-label refinement to assist the student model in comprehending the teacher’s knowledge at cluster-level while mitigating the negative effects of noisy labels. Finally, an instance-level consistency loss weighted by the similarity between the instance and its corresponding cluster is proposed to align the positive instance correlations. Our approach aims to train a one-step weakly supervised model for person search by exploiting the characteristics of unsupervised person Re-ID. Extensive experiments illustrate that our method achieves state-of-the-art performance on two widely-used person search datasets, CUHK-SYSU and PRW. Our code will be available on GitHub athttps://github.com/zongyi1999/KCD.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Chengxin Zhao, Qian Wang 0001, Shijuan Huang
IEEE Trans. Circuits Syst. Video Technol.6
2024 Gait Recognition With Multi-Level Skeleton-Guided Refinement
abstract
Existing methods combining skeleton and silhouette representations demonstrate explicit effectiveness for gait recognition. However, current related methods simply combine the video-level representations of model-based skeleton data and gait silhouettes for retrieval. Therefore, diverse skeleton information is not fully exploited in existing related works: Firstly, the position and movement of bones are not clear from individual silhouettes. This indicates that the frame-level interaction between features of skeletons and silhouettes is critical, which is ignored by previous methods. Secondly, diverse part-level skeleton-guided gait features are not fully captured in existing related approaches. To solve the above issues, we present a novel framework with multi-level skeleton-guided refinement, including frame-level, part-level, and video-level skeleton-guided refinement, for comprehensive skeleton-aided gait representation learning. First, two modules are proposed for frame-level skeleton-guided refinement. Specifically, Visual Skeleton Enhanced Backbone (VSEB) is proposed to visually highlight the global and part-level skeleton regions for the feature of each silhouette frame. Moreover, Cross-Visual-Model Frame-level Interaction (CVMFI) is proposed to further transfer the model-based skeleton information to features of the visual modalities. Secondly, part-level visual and model-based skeleton features are utilized to refine the final gait representation. Concretely, in VSEB, Part Skeleton Enhance Network (PSEN) is proposed to visually enhance the position and movement of part-level skeletons. In addition, Semantic Part Pooling (SPP) is proposed for capturing the model-based skeleton features of different semantic parts. Finally, as the video-level skeleton-guided refinement, multimodal video-level features are combined to boost the final recognition performance. Extensive experimental results on prevailing datasets demonstrate that our approach outperforms most existing methods, including the skeleton-aided multi-modal methods. With the multi-level refinement guided by the skeleton modalities, the framework is expected to provide a deeper understanding of skeleton-aided gait recognition.
Runsheng Wang, Yuxuan Shi 0002, Zongyi Li, Chengxin Zhao, Bohao Wei, He Li 0052, Ping Li 0021
IEEE Trans. Multim.5
2024 Viewpoint Disentangling and Generation for Unsupervised Object Re-ID
abstract
Unsupervised object Re-ID aims to learn discriminative identity features from a fully unlabeled dataset to solve the open-class re-identification problem. Satisfying results have been achieved in existing unsupervised Re-ID methods, primarily trained with pseudo-labels created by feature clustering. However, the viewpoint variation of objects is the key challenge, introducing noisy labels in the clustering process. To address this problem, a novel viewpoint disentangling and generation framework (VDG) is proposed to learn viewpoint-invariant ID features, including a disentangling and generation module, as well as a contrastive learning module. First, we design an ID encoder to map the viewpoint and identity features into the latent space. Second, a generator is used to disentangle view features and synthesize images with different orientations. Especially, the well-trained encoder serves as a pre-trained feature extractor in the contrastive learning module. Third, a viewpoint-aware loss and a class-level loss are integrated to facilitate contrastive learning between original and novel views. The generation of novel view images and the application of viewpoint-aware contrastive loss mutually assist model learning viewpoint-invariant ID features. Extensive experiments on Market-1501, DukeMTMC, MSMT17, and VeRi-776 demonstrate the effectiveness of the proposed VDG framework, as well as its superiority over the existing state-of-the-art approaches. The VDG model also demonstrates high quality in the image generation tasks.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Boyuan Liu, Runsheng Wang, Chengxin Zhao
ACM Trans. Multim. Comput. Commun. Appl.7
2023 MEGL: Multi-Experts Guided Learning Network for Single Camera Training Person Re-Identification
abstract
The time-saving single-camera training(SCT) person re-identification aims to learn camera-invariant information without cross-camera pedestrian annotations. To address this challenging task, we propose a novel approach called Multi-Experts Guided Learning Network (MEGL-Net) for SCT-ReID that can obtain features not influenced by camera views at the global and local levels under the guidance of multi-camera experts. Firstly, to obtain camera-invariant features, an adaptive feature integration module (AFI) is introduced to adaptively integrate expert-guided features from different camera branches. Then, the proposed camera-local interactive module (CLI) facilitates interaction between the local branch and the camera experts branch for automatically extracting discriminative, domain-invariant features at a fine-grained level. Finally, our framework aggregates expert-guided features with global features and enhanced local features in the testing stage for pedestrian retrieval. Under the Market-SCT and Duke-SCT datasets, experimental results demonstrate that our approach significantly improves ReID performance and outperforms existing state-of-the-art (SOTA) methods.
He Li 0052, Yuxuan Shi 0002, Zongyi Li, Runsheng Wang, Chengxin Zhao, Ping Li 0021
ICIP6
2023 Deep Unsupervised Hashing with Selective Semantic Mining
abstract
Most of the existing unsupervised hashing methods usually construct semantic similarity structure to guide hashing learning. However, due to the lack of filtering of useless information, some wrong guiding information in the similarity structure may damage the retrieval performance. Besides, some works adopt the framework of contrastive learning to preserve the discriminative semantic information that is more important for the hashing task. But such a training strategy may incorrectly embed some semantically similar samples far away due to the absence of manual label supervision, thus producing sub-optimal hash codes. To solve the aforementioned problems, we propose a novel method named Deep Selective Semantic Mining Hashing (DSSMH). Specifically, with the prior knowledge obtained by clustering, we select semantically correct image pairs with high confidence to alleviate the guidance of wrong information and correct sampling bias in contrastive learning. Extensive experiments demonstrate that DSSMH outperforms existing state-of-the-art methods.
Chuang Zhao 0001, Yuxuan Shi 0002, Chengxin Zhao, Jiazhong Chen
ICME4
2021 Learning Diverse Local Patterns for Deepfake Detection with Image-level Supervision
abstract
To prevent the Deepfake-like videos from spreading, researchers have proposed many anti-forgery methods. However, most approaches require pixel-wise annotation, which conflicts with the real Deepfake detection scenario. To make full use of the image-level label, we propose a Local-Prediction framework that indirectly allows the image-level label to supervise local regions. To further enrich the local feature, we introduced the Local-Diversity concept in the Deepfake detection field for the first time. We proposed the Local-Diversity Loss based on the motivation that regional pattern differences can provide semi-supervised information during training. Compared to the previous method, our approach limits each classification unit's receptive field and enriches the feature diversity. In the experiment, our method is evaluated on three benchmark datasets of four widely-used manipulation types. The result shows that the Local-Prediction framework is beneficial to different CNN backbones and achieved significant performance. The proposed LD loss enriches the learned patterns of binary classifiers. Furthermore, we provide visualization and ablation studies to understand the mechanism.
Junrui Huang, Chengxin Zhao, Yutong Yao, Jiazhong Chen, Ping Li 0021
IJCNN3
2020 PRO: A periodical reset optimized page migration scheme for hybrid memory system
Na Niu, Fangfa Fu, Jiacai Yuan, Fengchang Lai, Chengxin Zhao, Jinxiang Wang 0001
J. Syst. Archit.6