EDBT 2026 Demo / reviewers in the wild / expert
Shuguang Yuan 0003
dblp:274/6479
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-1312-8762ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARGH-Mark: Anchor-Synchronized Watermarking with Hamming Correction for Robust and Quality-Preserving LLM AttributionabstractThe proliferation of large language models has intensified demands for reliable content attribution, yet existing watermarking techniques face a fundamental trilemma: they cannot simultaneously optimize for robustness against attacks, minimal text quality degradation, and detection efficiency. To resolve this challenge, we propose ARGH-Mark, a novel watermarking framework that integrates three synergistic innovations: (1) Anchor-synchronized phase recovery for maintaining detection integrity under insertion/deletion attacks, (2) RG-balanced vocabulary modulation that dynamically partitions lexicons via contextual hashing to preserve generation quality, and (3) Hamming-based error correction enabling single-bit error rectification through algebraic coding. Comprehensive evaluations across question answering (ELI5), summarization (CNN/DailyMail), and text generation (C4) demonstrate state-of-the-art performance: the proposed ARGH-Mark framework achieves near-perfect match rate and bit accuracy across diverse configurations, while preserving the quality of the generated text. It significantly reduces detection latency, enabling real-time extraction, and maintains high robustness against token tampering attacks through integrated Hamming error correction, ensuring reliable attribution in adversarial settings. ARGH-Mark achieves a new Pareto frontier in the watermarking design space and advances trustworthy deployment of generative AI in alignment-critical applications. He Li 0010, Xiaojun Chen 0004, Jingcheng He, Zhendong Zhao, Shuguang Yuan 0003, Yunfei Yang 0001 |
AAAI | 5 |
| 2026 | You Can Have a Second Chance: Unbiased and Multi-bit Watermarking for Diffusion Language Models with Regret-based RemaskingabstractThe rapid development of Diffusion Language Models (DLMs) raises concerns about watermarking for DLM-generated detection.However, existing sequential LLM watermarking cannot be directly applied to DLMs, as DLMs' generation order is arbitrary.While emerging studies adapt biased LLM watermarking to DLMs by temporarily predicting the watermark prefix, they suffer from degraded quality and unstable watermarking due to bias accumulation and prediction errors.Besides, they cannot carry multi-bit watermarks.In this paper, we propose unbiased multi-bit watermarking for DLMs.We introduce a stability-aware constraint that allows watermarking only in stable contexts and a bit-controlled, unbiased modulation to preserve the original DLM output distribution, achieving stable watermarking with minimal quality impact.To enhance detection robustness, we design a Regret-based Remasking, which grants a "second chance" for unwatermarked tokens to be regenerated.It can seamlessly integrate into DLM inference with no added diffusion steps and latency.Experiments across DLMs and various tasks show that our scheme is effective, achieving superior generation quality compared to baselines while maintaining high detection accuracy and multi-bit capacity.Our code is available here. Dongyang Liang, Jing Yu 0007, Shuguang Yuan 0003, Chi Chen 0001 |
ACL (1) | 4 |
| 2026 | Privacy-preserving for user-uploaded images and text in Vision-Language Models
Zixiang Liu, Chi Chen 0001, Shuguang Yuan 0003, Weilong Huang, Xiaojie Zhu, Peizhuo Lv |
Comput. Secur. | 3 |
| 2026 | Detecting photo-taking actions in surveillance videos based on CPU-only devicesabstractAbstract Taking photos of sensitive facilities and sensitive information in no photography area may cause sensitive information leakage if not discovered in time. Employing action recognition models to detect instances of photography can effectively prevent information leakage. Current action recognition models have shown unsatisfactory performance in detecting photo-taking actions in surveillance videos, and their reliance on GPU devices hinder their practicality. This paper presents a novel approach to address the detection of photo-taking actions. The method utilizes object detection to filter out background data and incorporates human pose estimation to extract human skeleton data. By combining these AI techniques, the method enables accurate recognition of photo-taking actions. We introduce a novel technique called self-annotation that enables the model to focus on the crucial elements associated with photo-taking actions. Additionally, we introduce a new alarm mechanism that leads to a 69 $$\%$$ % reduction in false positives while maintaining the same level of recall by integrating the labels over a period to recognize actions. Compared with traditional action recognition approaches, our method is more flexible and lightweight in actual engineering applications. Moreover, our model is capable of running on CPU-only devices. Experimental results show that our model achieves a precision of 91 $$\%$$ % on our dataset. Zixiang Liu, Peisong Shen, Chi Chen 0001, Shuguang Yuan 0003, Xiaojie Zhu, Houzhe Wang |
Cybersecur. | 4 |
| 2026 | An Unbiased and Robust Privacy-Preserving Fingerprinting Scheme for Relational DatabasesabstractSharing relational databases is essential in today’s data-driven world for fostering collaboration, enhancing efficiency, and enabling real-time data access. However, privacy and copyright issues arise when sharing privacy-sensitive or valuable data. Additionally, high utility is required in shared data to enable accurate data mining and analysis. Entry-level differentially private fingerprinting schemes (DPFS) could address these concerns. In a DPFS, data can be securely shared without leaking original values while still supporting accurate analysis. Moreover, detectable fingerprints can deter unauthorized redistribution. However, existing DPFSs often lack utility—due to format changes and entry-wise bias—or robustness, as fingerprints can be removed undetected. In this paper, we propose an unbiased and robust differential privacy-based fingerprinting scheme (DPFS), which ensures that the fingerprinted copy remains an unbiased estimate of the original data. By incorporating differential privacy noise, our scheme effectively mitigates alteration, collusion, and hybrid attacks. Our DPFS satisfies ϵ-entry-level differential privacy, enabling clients to conduct unbiased analysis. To improve robustness, we design group-based fingerprint detection, which estimates the mean of injected noise per group with error tolerance. We provide a theoretical robustness analysis and propose a method for achieving optimal robustness. Experiments on four real-world databases show that our scheme consistently detects fingerprints and improves accuracy by up to 20% on machine learning tasks compared to existing DPFSs. Shujie Cui, Hui Cui 0001, Jiabao Qiu, Shuguang Yuan 0003, Xiaojie Zhu, Jing Yu 0007, Chi Chen 0001, Xun Yi |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | TSALockMark: An Asymmetric and Robust Watermarking Scheme for Relational Databases with Distortion Constraints
Shuguang Yuan 0003, Jing Yu 0007, Zhaochen Li, Chi Chen 0001 |
DASFAA (5) | 2 |
| 2025 | MSAnony: A dynamic anonymization algorithm supporting on-demand update through Merging and SplitabstractPrivacy-preserving data publishing has garnered widespread concern. Anonymization as a kind of privacy protection technique can safeguard individual privacy. In the context of publishing data, benign data updates occur frequently in the associated legitimate administrations and applications. In order to address updates, dynamic anonymization algorithms have been proposed. However, existing algorithms are unable to adequately generalize datasets, resulting in higher information loss. Furthermore, none of these algorithms involves attribute updates so far, and they only focus on record updates. This paper designs an anonymization algorithm, MSAnony, which is the first to address attribute updates problem. The algorithm introduces a novel data structure called Generalization Tag for efficiently re-generalizing. Moreover, it supports dynamic insertion, deletion and modification of records and attributes by basic operations Merging and Split. In a real-world dataset, MSAnony has good data utility compared with Mondrian, Incognito, Flash and Clustering under record update scenarios. Moreover, it can address attribute insertion and deletion under different parameters. Yulin Yuan, Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
ISCC | 2 |
| 2025 | An Efficient White-box LLM Watermarking for IP Protection on Online Market PlatformsabstractOnline market platforms serve as a central hub for sharing and deploying AI models among researchers, developers, and companies. In this context, watermarking techniques are essential to protect intellectual property (IP), preventing unauthorized use and duplication of large language models (LLMs). Two key challenges arise: (i) These platforms host diverse LLMs, yet current watermarking techniques are only tailored to specific models, such as fine-tuned or quantized LLMs. (ii) Efficient watermarking is critical. However, traditional methods require substantial data and costly hardware, which limits their feasibility. In this paper, we propose an efficient white-box LLM watermarking technique called ELLMark. This method treats LLMs as multi-layered matrices while embedding watermarks only relies on modifying the model's weights. To preserve LLMs' performance, it filters weights by correlations with the activation magnitudes and downstream tasks, then modifies weights as minimal as possible via histogram modulation. Notably, all phases are training-free with low hardware resources, making it efficient for online platforms. We conduct extensive experiments to evaluate the effectiveness of ELLMark on LLaMA-3, OPT, and Phi-3 LLMs. The results demonstrate that it achieves 100% success in watermark detection while preserving model performance. Moreover, the preprocessing, encoding, and decoding processes remain efficient, taking less than 7 minutes, 12 minutes, and 18 seconds, respectively, for models with 80B parameters. Lastly, it exhibits robustness against parameter overwriting, re-watermarking, forging, fine-tuning, and pruning attacks. Shuguang Yuan 0003, Xingyu Su, Peizhuo Lv, Weiji Xue, Jing Yu 0007, Xiaojie Zhu, Chi Chen 0001 |
KDD (2) | 1 |
| 2025 | ColorFP: Improving AI-Generated Text Detection via Fixed Vocabulary Partitioning and Half-Bit Fingerprinting
He Li 0010, Xiaojun Chen 0004, Yunfei Yang 0001, Zhendong Zhao, Shuguang Yuan 0003 |
PRICAI (4) | 5 |
| 2025 | Unitanony: a fine-grained and practical anonymization framework for better data utilityabstractAbstract In order to share data without revealing private information, privacy-preserving data publishing techniques are proposed. K-anonymity and l-diversity secure against identity and attribute disclosure. Anonymization algorithms enforce the above models and are willing to reach two primary goals: achieving the privacy objective while maximizing data utility. Even though anonymization has been studied for decades, finding efficient techniques to improve data utility is an open question. It is a crucial challenge that impacts many anonymized data on the web, cloud, and IoT environments. However, some factors incur huge information loss for existing works. The original dataset may be transformed into generalized data to an excessive extent. To address this problem, we give a new framework and propose a heuristic algorithm called UnitAnony. It builds a full-coverage hierarchy for more generalization candidates and proposes an interval-mapping technique for fine-grained generalization extent. However, these improvements raise another challenge. It is the vast cost because more generalization will derive many operations for forming new values, grouping records, and verifying anonymization models. Therefore, we designed a data structure unit to generalize records at low costs and implemented a skipping strategy to execute the algorithm within an acceptable time. Besides, our algorithm can support many models. By evaluating well-known K-anonymity and l-diversity on real-world datasets, i.e., Adult and Census datasets, the experimental results demonstrate that our algorithm outperforms the existing algorithms (e.g., Mondrian, Top-down, Improved-Clustering, Flash, and Incognito) regarding data utility and effectiveness. Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
Cybersecur. | 1 |
| 2025 | A study of watermarking techniques for data publishing
Shuguang Yuan 0003, Jing Yu 0007, Zhaochen Li, Jiabao Qiu, Chi Chen 0001 |
Multim. Tools Appl. | 1 |
| 2024 | Don't Abandon the Primary Key: A High-Synchronization and Robust Virtual Primary Key Scheme for Watermarking Relational Databases
Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
ICICS (2) | 2 |
| 2024 | GPSSE: A GPU-Accelerated Dynamic SSE Scheme with Efficient Batch UpdatingabstractDynamic Searchable Symmetric Encryption (DSSE) allows cloud users to securely retrieve and update their data while outsourcing it to untrusted cloud service providers. Although extensive research efforts in recent years have notably improved the retrieval efficiency of DSSE, there remains potential for enhancing update efficiency, particularly in large-scale dataset updating scenarios. Therefore, we proposed a pioneering scheme called GPSSE (GPU-Accelerated Dynamic Searchable Symmetric Encryption Scheme) to explore accelerating batch updates of DSSE through GPU. In our design, we break the traditional chain-based data structure and build independent label-based data blocks to store entries. It facilitates the parallel updating of entries. Moreover, GPSSE achieves forward privacy and Type II backward privacy. In addition, we formally prove the security of the proposed GPSSE and show its practicality by conducting experiments using the publicly well-known Enron Email dataset. Experimental results demonstrate that GPSSE outperforms 40× to 194156× than other state-of-the-art schemes in updating efficiency. Jiancong Zhou, Xiaojie Zhu, Shuguang Yuan 0003, Chi Chen 0001 |
ISPA | 4 |
| 2022 | An Attribute-attack-proof Watermarking Technique for Relational DatabaseabstractProving ownership rights on relational databases is an important issue. The robust watermarking technique could claim ownership by insertion information about the data owner. Hence, it is vital to improving the robustness of watermarking technique in that intruders could launch types of attacks to corrupt the inserted watermark. Furthermore, attributes are explicit and operable objectives to destroy the watermark. To my knowledge, there does not exist a comprehensive solution to resist attribute attack. In this paper, we propose a robust watermarking technique that is robust against subset and attribute attacks. The novelties lie in several points: applying the classifier to reorder watermarked attributes, designing a secret sharing mechanism to duplicate watermark independently on each attribute, and proposing twice majority voting to correct errors caused by attacks for improving the accuracy of watermark detection. In addition, our technique has features of blind, key-based, incrementally updatable, and low false hit rate. Experiments show that our algorithm is robust against subset and attribute attacks compared with AHK, DEW, and KSF algorithms. Moreover, it is efficient with running time in both insertion and detection phases. Shuguang Yuan 0003, Chi Chen 0001, Jing Yu 0007 |
TrustCom | 1 |
| 2020 | Verify a Valid Message in Single Tuple: A Watermarking Technique for Relational Database
Shuguang Yuan 0003, Jing Yu 0007, Peisong Shen, Chi Chen 0001 |
DASFAA (1) | 1 |