VLDB 2026 Research / reviewers in the wild / expert
Sijing Xie
dblp:329/9825
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GlyphShield: Document Watermarking for the Physical World via Vector Typeface Synthesis
Yuxing Lu, Han Fang 0004, Sijing Xie, Luyu Yuan, Chengxin Zhao |
AAAI | 5 |
| 2026 | FedLoDrop: Federated LoRA With Dropout for Generalized LLM Fine-TuningabstractFine-tuning (FT) large language models (LLMs) is crucial for adapting general-purpose models to specific tasks, enhancing accuracy and relevance with minimal resources. To further enhance generalization ability while reducing training costs, this paper proposes Federated LoRA with Dropout (FedLoDrop), a new framework that applies dropout to the rows and columns of the trainable matrix in Federated LoRA. A generalization error bound and convergence analysis under sparsity regularization are obtained, which elucidate the fundamental trade-off between underfitting and overfitting. The error bound reveals that a higher dropout rate increases model sparsity, thereby lowering the upper bound of pointwise hypothesis stability (PHS). While this reduces the gap between empirical and generalization errors, it also incurs a higher empirical error, which, together with the gap, determines the overall generalization error. On the other hand, though dropout reduces communication costs, deploying FedLoDrop at the network edge still faces challenges due to limited network resources. To address this issue, an optimization problem is formulated to minimize the upper bound of the generalization error, by jointly optimizing the dropout rate and resource allocation subject to the latency and per-device energy consumption constraints. To solve this problem, a branch-and-bound (B&B)-based method is proposed to obtain its globally optimal solution. Moreover, to reduce the high computational complexity of the B&B-based method, a penalized successive convex approximation (P-SCA)-based algorithm is proposed to efficiently obtain its high-quality suboptimal solution. Finally, numerical results demonstrate the effectiveness of the proposed approach in mitigating overfitting and improving the generalization capability. Sijing Xie, Dingzhu Wen, Changsheng You, Qimei Chen, Mehdi Bennis, Kaibin Huang |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | Integrated Sensing, Communication, and Computation for Over-the-Air Federated Edge LearningabstractThis paper studies an over-the-air federated edge learning (Air-FEEL) system with integrated sensing, communication, and computation (ISCC), in which one edge server coordinates multiple edge devices to wirelessly sense the objects and use the sensing data to collaboratively train a machine learning model for recognition tasks. In this system, over-the-air computation (AirComp) is employed to enable one-shot model aggregation from edge devices. Under this setup, we analyze the convergence behavior of the ISCC-enabled Air-FEEL in terms of the loss function degradation, by particularly taking into account the wireless sensing noise during the training data acquisition and the AirComp distortions during the over-the-air model aggregation. The result theoretically shows that sensing, communication, and computation compete for network resources to jointly decide the convergence rate. Based on the analysis, we design the ISCC parameters under the target of maximizing the loss function degradation while ensuring the latency and energy budgets in each round. The challenge lies on the tightly coupled processes of sensing, communication, and computation among different devices. To tackle the challenge, we derive a low-complexity ISCC algorithm by alternately optimizing the batch size control and the network resource allocation. It is found that for each device, less sensing power should be consumed if a larger batch of data samples is obtained and vice versa. Besides, with a given batch size, the optimal computation speed of one device is the minimum one that satisfies the latency constraint. Numerical results based on a human motion recognition task verify the theoretical convergence analysis and show that the proposed ISCC algorithm well coordinates the batch size control and resource allocation among sensing, communication, and computation to enhance the learning performance. Dingzhu Wen, Sijing Xie, Xiaowen Cao 0001, Yuanhao Cui, Jie Xu 0002, Yuanming Shi, Shuguang Cui |
IEEE Trans. Wirel. Commun. | 2 |
| 2026 | Federated Dropout: Convergence Analysis and Resource AllocationabstractFederated Dropout is an efficient technique to overcome both communication and computation bottlenecks for deploying federated learning at the network edge. In each training round, an edge device only needs to update and transmit a sub-model, which is generated by the typical method of dropout in deep learning, and thus effectively reduces the per-round latency. \textcolor{blue}{However, the theoretical convergence analysis for Federated Dropout is still lacking in the literature, particularly regarding the quantitative influence of dropout rate on convergence}. To address this issue, by using the Taylor expansion method, we mathematically show that the gradient variance increases with a scaling factor of $γ/(1-γ)$, with $γ\in [0, θ)$ denoting the dropout rate and $θ$ being the maximum dropout rate ensuring the loss function reduction. Based on the above approximation, we provide the convergence analysis for Federated Dropout. Specifically, it is shown that a larger dropout rate of each device leads to a slower convergence rate. This provides a theoretical foundation for reducing the convergence latency by making a tradeoff between the per-round latency and the overall rounds till convergence. Moreover, a low-complexity algorithm is proposed to jointly optimize the dropout rate and the bandwidth allocation for minimizing the loss function in all rounds under a given per-round latency and limited network resources. Finally, numerical results are provided to verify the effectiveness of the proposed algorithm. Sijing Xie, Dingzhu Wen, Changsheng You, Tharmalingam Ratnarajah, Kaibin Huang |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable DistortionsabstractDeep watermarking models optimize robustness by incorporating distortions between the encoder and decoder. To tackle non-differentiable distortions, current methods only train the decoder with distorted images, which breaks the joint optimization of the encoder-decoder, resulting in suboptimal performance. To address this problem, we propose an Adversarial Distortion Domain Translation (AD2T) method by treating the distortion as an image-to-image translation task. AD2T adopts conditional GANs to learn the non-differentiable distortion mappings. It employs generators to transform the encoded image into the distorted one to bridge the encoder-decoder for joint optimization. We also supervise the GANs to generate challenging distorted samples to augment the watermarking model via adversarial training. This further improves the model robustness by minimizing the maximum decoding loss. Extensive experiments demonstrate the superiority of our method when tested on non-differentiable distortions, including lossy compression and style transfers. Codes are released here: https://github.com/zcx-language/AdversarialDistortionDomainTranslation. Chengxin Zhao, Jiazhong Chen, Han Fang 0004, Zongyi Li, Sijing Xie |
ICASSP | 6 |
| 2025 | Ultra-high Resolution Watermarking Framework Resistant to Extreme Cropping and ScalingabstractRecent developments in DNN-based image watermarking techniques have achieved impressive results in protecting digital content. However, most existing methods are constrained to low-resolution images as they need to encode the entire image, leading to prohibitive memory and computational costs when applied to high-resolution images. Moreover, they lack robustness to distortions prevalent in large-image transmission, such as extreme scaling and random cropping. To address these issues, we propose a novel watermarking method based on implicit neural representations (INRs). Leveraging the properties of INRs, our method employs resolution-independent coordinate sampling mechanism to generate watermarks pixel-wise, achieving ultra-high resolution watermark generation with fixed and limited memory and computational resources. This design ensures strong robustness in watermark extraction, even under extreme cropping and scaling distortions. Additionally, we introduce a hierarchical multi-scale coordinate embedding and a low-rank watermark injection strategy to ensure high-quality watermark generation and robust decoding. Experimental results demonstrate that our method significantly outperforms existing schemes in terms of both robustness and computational efficiency while preserving high image quality. Our approach achieves an accuracy greater than 98\% in watermark extraction with only 0.4\% of the image area in 2K images. These results highlight the effectiveness of our method, making it a promising solution for large-scale and high-resolution image watermarking applications. Luyu Yuan, Han Fang 0004, Yuxing Lu, Sijing Xie, Chengxin Zhao |
NeurIPS | 6 |
| 2025 | Federated LoRA with Dropout: An Efficient and Overfitting Control Approach for LLM Fine-TuningabstractThis paper introduces the Federated LoRA with Dropout (FedLoDrop) framework, designed to enhance generalization performance for downstream tasks at the network edge while simultaneously reducing overhead. Within this framework, we derive a generalization error bound under sparsity regularization, elucidating the theoretical principles that balance underfitting and overfitting. Our analysis shows that a higher dropout rate increases sparsity, lowering the Pointwise Hypothesis Stability (PHS) upper bound and narrowing the gap between empirical and generalization errors. However, this also leads to a higher empirical error, which, together with the gap, contributes to the total generalization error. Consequently, we formulate an optimization problem that jointly considers dropout rate and resource allocation, aiming to minimize the upper bound of the generalization error. Finally, numerical results demonstrate the effectiveness of the proposed approach in mitigating overfitting and enhancing generalization capabilities. Sijing Xie, Changsheng You, Qimei Chen, Dingzhu Wen |
PIMRC | 1 |
| 2025 | GSyncCode: Geometry Synchronous Hidden Code for One-step Photography DecodingabstractInvisible hyperlinks and hidden barcodes have recently emerged as a hot topic in offline-to-online messaging, where an invisible message or barcode is embedded in an image and can be decoded via camera shooting. Current schemes involve a two-step decoding process: starting with vertex localization of the embedded region to correct the perspective distortion introduced by shooting, followed by decoding the message from the corrected region. However, vertex localization can be complex and time-consuming, which affects the efficiency and accuracy of message decoding. To address this issue, this article proposes a geometry synchronous decoding scheme called GSyncCode, allowing for one-step extraction of a Data Matrix code from the photograph. Instead of correction before decoding, GSyncCode directly decodes a geometry-transformed Data Matrix that is synchronized with the embedded region. A barcode scanner is then used to efficiently retrieve messages. We design a Haar transform-based encoder HaarUNet and a HaarLoss visual function to select the key component of the Data Matrix for embedding. They improve the visual quality of the embedded image by reducing redundant embedding signals. Extensive simulated and real-world experiments demonstrate the superiority of GSyncCode in both decoding efficiency and accuracy. Our codes are published at: https://github.com/zcx-language/GSyncCode . Chengxin Zhao, Jialie Shen 0001, Han Fang 0004, Sijing Xie, Yaokun Fang, Zongyi Li, Ping Li 0021 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Convergence Analysis for Federated DropoutabstractFederated dropout on the weight is an efficient technique to overcome both communication and computation bottlenecks for deploying federated learning at the network edge. However, the theoretical analysis for Federated Dropout is still lacking in the literature, due to the challenge arising from the gradient bias. To address this issue, by using the Taylor expansion method, we mathematically show that the gradient vector with dropout can be approximated as an unbiased estimation of that without dropout; while its gradient variance increases with a scaling factor of γ/(1 − γ), with γ ∈ [0,θ) denoting the dropout rate and θ being the maximum dropout rate ensuring the loss function reduction. Based on the above approximation, we provide the loss function analysis for Federated Dropout. Specifically, it is shown that a larger dropout rate of each device leads to a slower convergence rate. Finally, numerical results are provided to verify the effects of dropout rate on convergence in both underfitting and overfitting scenarios. Sijing Xie, Dingzhu Wen, Changsheng You, Tharmalingam Ratnarajah, Kaibin Huang |
GLOBECOM | 1 |
| 2024 | InvertedFontNet: Font Watermarking based on Perturbing Style ManifoldabstractWe introduce InvertedFontNet, a pioneering unsupervised font watermarking framework leveraging the font style manifold. Previous methodologies typically necessitate manual intervention or are confined by specialized font data and media. In contrast, our approach introduces an algorithm that exclusively relies on unlabeled font image data, enabling the embedding of extensive watermark information across diverse fonts. The algorithm strategically modifies font spatial structures by manipulating style manifolds, facilitating the embedding of watermarks via subtle glyph perturbations. Experimental results reveal the robustness of our algorithm against prevalent digital noise attacks, demonstrating superior detection accuracy compared to existing schemes. Chenxin Zhao, Sijing Xie |
ICASSP | 3 |
| 2024 | Picking watermarks from noise (PWFN): an improved robust watermarking model against intensive distortionsabstractDigital watermarking is the process of embedding secret information by altering images in an undetectable way to the human eye. To increase the robustness of the model, many deep learning-based watermarking methods use the encoder-noise-decoder architecture by adding different noises to the noise layer. The decoder then extracts the watermarked information from the distorted image. However, this method can only resist weak noise attacks. To improve the robustness of the decoder against stronger noise, this paper proposes to introduce a denoise module between the noise layer and the decoder. The module aims to reduce noise and recover some of the information lost caused by distortion. Additionally, the paper introduces the SE module to fuse the watermarking information pixel-wise and channel dimensions-wise, improving the encoder’s efficiency. Experimental results show that our proposed method is comparable to existing models and outperforms state-of-the-art under different noise intensities. In addition, ablation experiments show the superiority of our proposed module. Sijing Xie, Chengxin Zhao, Wei Li 0151 |
ICME | 1 |
| 2024 | SSyncOA: Self-synchronizing Object-aligned Watermarking to Resist Crop-paste AttacksabstractModern image processing tools can easily crop local objects from images and paste them elsewhere. The challenge posed by this crop-paste attack is that it breaks the synchronization of the image watermark by inducing multiple superimposed desynchronization distortions. Existing image watermarking methods can only resist a single type of desynchronization attack and are inapplicable to this scenario. Finding that the key to resisting the crop-paste attack lies in the geometrically robust features of the object itself, this paper proposes a Self-Synchronizing Object-Aligned watermarking scheme, called SSyncOA. Specifically, we design a self-synchronization process that normalizes the watermark region, the centroid, the principal direction, and the minimum bounding square of the object during encoding and decoding to achieve synchronization of cropping, translation, rotation, and scaling, respectively. In cooperation with SSync, we propose an object-aligned watermarking method that embeds and extracts watermark messages only from the object region. This is achieved by training the watermarking model end-to-end with crop-paste attacks introduced between the encoder and decoder. Extensive experiments illustrate the impact of different desynchronization distortions on the trained watermark model, as well as the superior performance of our method compared to other SOTAs. Chengxin Zhao, Sijing Xie, Han Fang 0004, Yaokun Fang |
ICME | 3 |
| 2024 | DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions LocalizationabstractEmbedding invisible hyperlinks or hidden codes in images to replace QR codes has become a hot topic recently. This technology requires first localizing the embedded region in the captured photos before decoding. Existing methods that train models to find the invisible embedded region struggle to obtain accurate localization results, leading to degraded decoding accuracy. This limitation is primarily because the CNN network is sensitive to low-frequency signals, while the embedded signal is typically in the high-frequency form. Based on this, this paper proposes a Dual-Branch Dual-Head (DBDH) neural network tailored for the precise localization of invisible embedded regions. Specifically, DBDH uses a low-level texture branch containing 62 high-pass filters to capture the high-frequency signals induced by embedding. A high-level context branch is used to extract discriminative features between the embedded and normal regions. DBDH employs a detection head to directly detect the four vertices of the embedding region. In addition, we introduce an extra segmentation head to segment the mask of the embedding region during training. The segmentation head provides pixel-level supervision for model learning, facilitating better learning of the embedded signals. Based on two state-of-the-art invisible offline-to-online messaging methods, we construct two datasets and augmentation strategies for training and testing localization models. Extensive experiments demonstrate the superior performance of the proposed DBDH over existing methods. Chengxin Zhao, Sijing Xie, Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen |
IJCNN | 3 |