EDBT 2026 Demo / reviewers in the wild / expert
Chi Chen 0001
dblp:21/1794-1
· DBLP profile ↗
28ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-5491-0542ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 15 · 13 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | You Can Have a Second Chance: Unbiased and Multi-bit Watermarking for Diffusion Language Models with Regret-based RemaskingabstractThe rapid development of Diffusion Language Models (DLMs) raises concerns about watermarking for DLM-generated detection.However, existing sequential LLM watermarking cannot be directly applied to DLMs, as DLMs' generation order is arbitrary.While emerging studies adapt biased LLM watermarking to DLMs by temporarily predicting the watermark prefix, they suffer from degraded quality and unstable watermarking due to bias accumulation and prediction errors.Besides, they cannot carry multi-bit watermarks.In this paper, we propose unbiased multi-bit watermarking for DLMs.We introduce a stability-aware constraint that allows watermarking only in stable contexts and a bit-controlled, unbiased modulation to preserve the original DLM output distribution, achieving stable watermarking with minimal quality impact.To enhance detection robustness, we design a Regret-based Remasking, which grants a "second chance" for unwatermarked tokens to be regenerated.It can seamlessly integrate into DLM inference with no added diffusion steps and latency.Experiments across DLMs and various tasks show that our scheme is effective, achieving superior generation quality compared to baselines while maintaining high detection accuracy and multi-bit capacity.Our code is available here. Dongyang Liang, Jing Yu 0007, Shuguang Yuan 0003, Chi Chen 0001 |
ACL (1) | 5 |
| 2026 | Privacy-preserving for user-uploaded images and text in Vision-Language Models
Zixiang Liu, Chi Chen 0001, Shuguang Yuan 0003, Weilong Huang, Xiaojie Zhu, Peizhuo Lv |
Comput. Secur. | 2 |
| 2026 | Detecting photo-taking actions in surveillance videos based on CPU-only devicesabstractAbstract Taking photos of sensitive facilities and sensitive information in no photography area may cause sensitive information leakage if not discovered in time. Employing action recognition models to detect instances of photography can effectively prevent information leakage. Current action recognition models have shown unsatisfactory performance in detecting photo-taking actions in surveillance videos, and their reliance on GPU devices hinder their practicality. This paper presents a novel approach to address the detection of photo-taking actions. The method utilizes object detection to filter out background data and incorporates human pose estimation to extract human skeleton data. By combining these AI techniques, the method enables accurate recognition of photo-taking actions. We introduce a novel technique called self-annotation that enables the model to focus on the crucial elements associated with photo-taking actions. Additionally, we introduce a new alarm mechanism that leads to a 69 $$\%$$ % reduction in false positives while maintaining the same level of recall by integrating the labels over a period to recognize actions. Compared with traditional action recognition approaches, our method is more flexible and lightweight in actual engineering applications. Moreover, our model is capable of running on CPU-only devices. Experimental results show that our model achieves a precision of 91 $$\%$$ % on our dataset. Zixiang Liu, Peisong Shen, Chi Chen 0001, Shuguang Yuan 0003, Xiaojie Zhu, Houzhe Wang |
Cybersecur. | 3 |
| 2026 | An Unbiased and Robust Privacy-Preserving Fingerprinting Scheme for Relational DatabasesabstractSharing relational databases is essential in today’s data-driven world for fostering collaboration, enhancing efficiency, and enabling real-time data access. However, privacy and copyright issues arise when sharing privacy-sensitive or valuable data. Additionally, high utility is required in shared data to enable accurate data mining and analysis. Entry-level differentially private fingerprinting schemes (DPFS) could address these concerns. In a DPFS, data can be securely shared without leaking original values while still supporting accurate analysis. Moreover, detectable fingerprints can deter unauthorized redistribution. However, existing DPFSs often lack utility—due to format changes and entry-wise bias—or robustness, as fingerprints can be removed undetected. In this paper, we propose an unbiased and robust differential privacy-based fingerprinting scheme (DPFS), which ensures that the fingerprinted copy remains an unbiased estimate of the original data. By incorporating differential privacy noise, our scheme effectively mitigates alteration, collusion, and hybrid attacks. Our DPFS satisfies ϵ-entry-level differential privacy, enabling clients to conduct unbiased analysis. To improve robustness, we design group-based fingerprint detection, which estimates the mean of injected noise per group with error tolerance. We provide a theoretical robustness analysis and propose a method for achieving optimal robustness. Experiments on four real-world databases show that our scheme consistently detects fingerprints and improves accuracy by up to 20% on machine learning tasks compared to existing DPFSs. Shujie Cui, Hui Cui 0001, Jiabao Qiu, Shuguang Yuan 0003, Xiaojie Zhu, Jing Yu 0007, Chi Chen 0001, Xun Yi |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | TSALockMark: An Asymmetric and Robust Watermarking Scheme for Relational Databases with Distortion Constraints
Shuguang Yuan 0003, Jing Yu 0007, Zhaochen Li, Chi Chen 0001 |
DASFAA (5) | 6 |
| 2025 | MSAnony: A dynamic anonymization algorithm supporting on-demand update through Merging and SplitabstractPrivacy-preserving data publishing has garnered widespread concern. Anonymization as a kind of privacy protection technique can safeguard individual privacy. In the context of publishing data, benign data updates occur frequently in the associated legitimate administrations and applications. In order to address updates, dynamic anonymization algorithms have been proposed. However, existing algorithms are unable to adequately generalize datasets, resulting in higher information loss. Furthermore, none of these algorithms involves attribute updates so far, and they only focus on record updates. This paper designs an anonymization algorithm, MSAnony, which is the first to address attribute updates problem. The algorithm introduces a novel data structure called Generalization Tag for efficiently re-generalizing. Moreover, it supports dynamic insertion, deletion and modification of records and attributes by basic operations Merging and Split. In a real-world dataset, MSAnony has good data utility compared with Mondrian, Incognito, Flash and Clustering under record update scenarios. Moreover, it can address attribute insertion and deletion under different parameters. Yulin Yuan, Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
ISCC | 4 |
| 2025 | An Efficient White-box LLM Watermarking for IP Protection on Online Market PlatformsabstractOnline market platforms serve as a central hub for sharing and deploying AI models among researchers, developers, and companies. In this context, watermarking techniques are essential to protect intellectual property (IP), preventing unauthorized use and duplication of large language models (LLMs). Two key challenges arise: (i) These platforms host diverse LLMs, yet current watermarking techniques are only tailored to specific models, such as fine-tuned or quantized LLMs. (ii) Efficient watermarking is critical. However, traditional methods require substantial data and costly hardware, which limits their feasibility. In this paper, we propose an efficient white-box LLM watermarking technique called ELLMark. This method treats LLMs as multi-layered matrices while embedding watermarks only relies on modifying the model's weights. To preserve LLMs' performance, it filters weights by correlations with the activation magnitudes and downstream tasks, then modifies weights as minimal as possible via histogram modulation. Notably, all phases are training-free with low hardware resources, making it efficient for online platforms. We conduct extensive experiments to evaluate the effectiveness of ELLMark on LLaMA-3, OPT, and Phi-3 LLMs. The results demonstrate that it achieves 100% success in watermark detection while preserving model performance. Moreover, the preprocessing, encoding, and decoding processes remain efficient, taking less than 7 minutes, 12 minutes, and 18 seconds, respectively, for models with 80B parameters. Lastly, it exhibits robustness against parameter overwriting, re-watermarking, forging, fine-tuning, and pruning attacks. Shuguang Yuan 0003, Xingyu Su, Peizhuo Lv, Weiji Xue, Jing Yu 0007, Xiaojie Zhu, Chi Chen 0001 |
KDD (2) | 7 |
| 2025 | SafeMLLM: Extending Safety Alignment from Single-Modal LLMs to Multimodal LLMsabstractMultimodal large language models (MLLMs) are capable of processing diverse multimodal inputs to generate informative and contextually relevant responses. In this paper, alignment involves safeguarding the model against producing harmful or inappropriate content and ensuring its outputs are aligned with ethical guidelines. Existing alignment works have primarily focused on making the outputs of large language models(LLMs) harmless. However, the integration of multimodal information into MLLM inputs can lead to unintended or undesirable responses, thereby undermining the original alignment mechanisms of LLMs. When image, audio, or video information is contained in queries, MLLMs cannot guarantee the safe alignment of responses. In this study, we first collect VAdvBench to demonstrate that the additional multimodal inputs can easily break the safeguards of LLMs in MLLMs, highlighting the security vulnerabilities in these MLLMs. To address this issue, we propose SafeMLLM, a safety alignment framework designed for MLLMs. SafeMLLM aligns multimodal content by transforming multimodal inputs into a unified space, enabling MLLM to effectively block multimodal harmful information. In addition, since the current multimodal safety alignment works only consider images as additional modality inputs and lack benchmarks for other modalities like audio and video, we propose a pipeline to collect and synthesize a multimodal safety alignment benchmark, MLGuard, to evaluate the safety alignment performance on multimodal queries. The experiment results on MLGuard demonstrate that our SafeMLLM effectively rejects harmful instructions containing multimodal queries of image, audio, and video, ensuring robust safety alignment across diverse multimodal inputs. The code is available at https://github.com/jiangdi841/SafeMLLM/tree/main. Xiaojie Zhu, Chi Chen 0001 |
TrustCom | 3 |
| 2025 | A review of privacy-preserving biometric identification and authentication protocols
Peisong Shen, Xiaojie Zhu, Xue Tian, Chi Chen 0001 |
Comput. Secur. | 5 |
| 2025 | Reusable and robust fuzzy extractor for CRS-dependent sourcesabstractAbstract Fuzzy extractors allow for the extraction and reproduction of a nearly uniform string from a noisy and non-uniform source. Reusable and robust fuzzy extractors further require that the output string should remain pseudorandom under multiple extractions and any modification of public value should be detectable. Existing constructions of reusable and robust fuzzy extractors are all designed in the Common Reference String (CRS) model and work only for CRS-independent sources. In this work, we introduce a construction of reusable and robust fuzzy extractor for CRS-dependent sources. Our construction is built upon some well-studied cryptography, including a collision resistant hash function, a symmetric key encryption that keeps secure with respect to auxiliary input, a public key encryption and simulation-sound non-interactive zero-knowledge argument. We also present some instantiations of these primitives from Learning with Parity Noise (LPN) assumption and Short Integer Solution (SIS) assumption. These instantiations result in the first reusable and robust fuzzy extractor for CRS-dependent sources that can tolerate linear errors. Yucheng Ma, Peisong Shen, Xue Tian, Kewei Lv, Chi Chen 0001 |
Cybersecur. | 5 |
| 2025 | Unitanony: a fine-grained and practical anonymization framework for better data utilityabstractAbstract In order to share data without revealing private information, privacy-preserving data publishing techniques are proposed. K-anonymity and l-diversity secure against identity and attribute disclosure. Anonymization algorithms enforce the above models and are willing to reach two primary goals: achieving the privacy objective while maximizing data utility. Even though anonymization has been studied for decades, finding efficient techniques to improve data utility is an open question. It is a crucial challenge that impacts many anonymized data on the web, cloud, and IoT environments. However, some factors incur huge information loss for existing works. The original dataset may be transformed into generalized data to an excessive extent. To address this problem, we give a new framework and propose a heuristic algorithm called UnitAnony. It builds a full-coverage hierarchy for more generalization candidates and proposes an interval-mapping technique for fine-grained generalization extent. However, these improvements raise another challenge. It is the vast cost because more generalization will derive many operations for forming new values, grouping records, and verifying anonymization models. Therefore, we designed a data structure unit to generalize records at low costs and implemented a skipping strategy to execute the algorithm within an acceptable time. Besides, our algorithm can support many models. By evaluating well-known K-anonymity and l-diversity on real-world datasets, i.e., Adult and Census datasets, the experimental results demonstrate that our algorithm outperforms the existing algorithms (e.g., Mondrian, Top-down, Improved-Clustering, Flash, and Incognito) regarding data utility and effectiveness. Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
Cybersecur. | 4 |
| 2025 | A study of watermarking techniques for data publishing
Shuguang Yuan 0003, Jing Yu 0007, Zhaochen Li, Jiabao Qiu, Chi Chen 0001 |
Multim. Tools Appl. | 8 |
| 2024 | Don't Abandon the Primary Key: A High-Synchronization and Robust Virtual Primary Key Scheme for Watermarking Relational Databases
Shuguang Yuan 0003, Jing Yu 0007, Chi Chen 0001 |
ICICS (2) | 6 |
| 2024 | One-Factor Cancelable Biometric Template Protection Scheme for Real-Valued Features
Ruoqi Zhang, Peisong Shen, Kewei Lv, Chi Chen 0001 |
ICPR (29) | 4 |
| 2024 | GPSSE: A GPU-Accelerated Dynamic SSE Scheme with Efficient Batch UpdatingabstractDynamic Searchable Symmetric Encryption (DSSE) allows cloud users to securely retrieve and update their data while outsourcing it to untrusted cloud service providers. Although extensive research efforts in recent years have notably improved the retrieval efficiency of DSSE, there remains potential for enhancing update efficiency, particularly in large-scale dataset updating scenarios. Therefore, we proposed a pioneering scheme called GPSSE (GPU-Accelerated Dynamic Searchable Symmetric Encryption Scheme) to explore accelerating batch updates of DSSE through GPU. In our design, we break the traditional chain-based data structure and build independent label-based data blocks to store entries. It facilitates the parallel updating of entries. Moreover, GPSSE achieves forward privacy and Type II backward privacy. In addition, we formally prove the security of the proposed GPSSE and show its practicality by conducting experiments using the publicly well-known Enron Email dataset. Experimental results demonstrate that GPSSE outperforms 40× to 194156× than other state-of-the-art schemes in updating efficiency. Jiancong Zhou, Xiaojie Zhu, Shuguang Yuan 0003, Chi Chen 0001 |
ISPA | 5 |
| 2024 | Dynamic group fuzzy extractorabstractAbstract The group fuzzy extractor allows group users to extract and reproduce group cryptographic keys from their individual non-uniform random sources. It can be easily used in group-oriented cryptographic applications. However, current group fuzzy extractors are not dynamic, i.e. they spend a large cost when dealing with user revocation. In this work, we propose the formal definition and construction of dynamic group fuzzy extractor (DGFE) to address this issue. For the revocation, DGFE allows unrevoked group users to reproduce updated group keys from the existing group help data. Meanwhile, it prevents any revoked group user from generating new group keys using the previously authorized individual help data. We propose a DGFE construction based on the revocable group signature. Furthermore, we give formal proofs of reusability, anonymity and traceability of our construction. Kaini Chen, Peisong Shen, Kewei Lv, Xue Tian, Chi Chen 0001 |
Cybersecur. | 5 |
| 2023 | General Constructions of Fuzzy Extractors for Continuous Sources
Yucheng Ma, Peisong Shen, Kewei Lv, Xue Tian, Chi Chen 0001 |
Inscrypt (1) | 5 |
| 2023 | Cancelable biometric schemes for Euclidean metric and Cosine metricabstractAbstract The handy biometric data is a double-edged sword, paving the way of the prosperity of biometric authentication systems but bringing the personal privacy concern. To alleviate the concern, various biometric template protection schemes are proposed to protect the biometric template from information leakage. The preponderance of existing proposals is based on Hamming metric, which ignores the fact that predominantly deployed biometric recognition systems (e.g. face, voice, gait) generate real-valued templates, more applicable to Euclidean metric and Cosine metric. Moreover, since the emergence of similarity-based attacks, those schemes are not secure under a stolen-token setting. In this paper, we propose a succinct biometric template protection scheme to address such a challenge. The proposed scheme is designed for Euclidean metric and Cosine metric instead of Hamming distance. Mainly, the succinct biometric template protection scheme consists of distance-preserving, one-way, and obfuscation modules. To be specific, we adopt location sensitive hash function to realize the distance-preserving and one-way properties simultaneously and use the modulo operation to implement many-to-one mapping. We also thoroughly analyze the proposed scheme in three aspects: irreversibility, unlinkability and revocability. Moreover, comprehensive experiments are conducted on publicly known face databases. All the results show the effectiveness of the proposed scheme. Yubing Jiang, Peisong Shen, Xiaojie Zhu, Chi Chen 0001 |
Cybersecur. | 6 |
| 2022 | Secure Sketch and Fuzzy Extractor with Imperfect Randomness: An Information-Theoretic Study
Kaini Chen, Peisong Shen, Kewei Lv, Chi Chen 0001 |
ICICS | 4 |
| 2022 | A Secure and Practical Sample-then-lock Scheme for Iris RecognitionabstractSample-then-lock construction is a reusable fuzzy extractor for low-entropy sources. When applied on iris recognition scenarios, many subsets of an iris-code are used to lock the cryptographic key. The security of this construction relies on the entropy of subsets of iris codes. Simhadri et al. reported a security level of 32 bits on iris sources. In this paper, we propose two kinds of attacks to crack existing sample-then-lock schemes. Exploiting the low-entropy subsets, our attacks can break the locked key and the enrollment iris-code respectively in less than 220brute force attempts. To protect from these proposed attacks, we design an improved sample-then-lock scheme. More precisely, our scheme employs stability and discriminability to select high-entropy subsets to lock the genuine secret, and conceals genuine locker by a large amount of chaff lockers. Our experiment verifies that existing schemes are vulnerable to the proposed attacks with a security level of less than 20 bits, while our scheme can resist these attacks with a security level of more than 100 bits when number of genuine subsets is 106. Peisong Shen, Kaini Chen, Yucheng Ma, Chi Chen 0001 |
ICPR | 5 |
| 2022 | An Attribute-attack-proof Watermarking Technique for Relational DatabaseabstractProving ownership rights on relational databases is an important issue. The robust watermarking technique could claim ownership by insertion information about the data owner. Hence, it is vital to improving the robustness of watermarking technique in that intruders could launch types of attacks to corrupt the inserted watermark. Furthermore, attributes are explicit and operable objectives to destroy the watermark. To my knowledge, there does not exist a comprehensive solution to resist attribute attack. In this paper, we propose a robust watermarking technique that is robust against subset and attribute attacks. The novelties lie in several points: applying the classifier to reorder watermarked attributes, designing a secret sharing mechanism to duplicate watermark independently on each attribute, and proposing twice majority voting to correct errors caused by attacks for improving the accuracy of watermark detection. In addition, our technique has features of blind, key-based, incrementally updatable, and low false hit rate. Experiments show that our algorithm is robust against subset and attribute attacks compared with AHK, DEW, and KSF algorithms. Moreover, it is efficient with running time in both insertion and detection phases. Shuguang Yuan 0003, Chi Chen 0001, Jing Yu 0007 |
TrustCom | 2 |
| 2020 | Verify a Valid Message in Single Tuple: A Watermarking Technique for Relational Database
Shuguang Yuan 0003, Jing Yu 0007, Peisong Shen, Chi Chen 0001 |
DASFAA (1) | 4 |
| 2019 | A Performance-Optimization Method for Reusable Fuzzy Extractor Based on Block Error Distribution of Iris Trait
Peisong Shen, Chi Chen 0001 |
SecureComm (2) | 3 |
| 2018 | A Robust Iris Segmentation Using Fully Convolutional Network with Dilated ConvolutionsabstractIris segmentation is a critical part in iris recognition systems. It segments the acquired image into iris and non-iris parts. It is the foundation of subsequent processing. The errors in this stage are propagated to subsequent processing stages, which will affecting the recognition rate of the whole system. A majority of iris segmentation algorithms require a significant amount of user cooperation during image acquisition process to provide good segmentation performance. However, the quality of iris images can not be guaranteed. When an iris image is acquired under non-ideal conditions (e.g., bad illumination, uncooperative subject, occluded iris, etc.), segmentation becomes a challenging task. In this paper, we present a more robust iris segmentation method using fully convolutional network (FCN) with dilated convolutions. We reduce the downsampling factor of the FCN model, and use the dilated convolutions to extract the more global features, which makes our method better at dealing with details. Moreover, our model supports end to end prediction, it does not need any pre-processing, such as adjusting the image to a fixed size. We used three datasets for training and testing, including CASIA-iris-interval-v4, UBIRIS v2 and IITD Delhi datasets. Experiments show that our model greatly reduced the error rate of the current state-of-the-arts by 79%, 84% and 79% on the CASIA-iris-interval-v4, IITD Delhi and UBIRIS v2 datasets respectively. Peisong Shen, Chi Chen 0001 |
ISM | 3 |
| 2017 | Privacy-Preserving Relevance Ranking Scheme and Its Application in Multi-keyword Searchable Encryption
Peisong Shen, Chi Chen 0001, Xiaojie Zhu |
SecureComm | 2 |
| 2016 | An Efficient Privacy-Preserving Ranked Keyword Search MethodabstractCloud data owners prefer to outsource documents in an encrypted form for the purpose of privacy preserving. Therefore it is essential to develop efficient and reliable ciphertext search techniques. One challenge is that the relationship between documents will be normally concealed in the process of encryption, which will lead to significant search accuracy performance degradation. Also the volume of data in data centers has experienced a dramatic growth. This will make it even more challenging to design ciphertext search schemes that can provide efficient and reliable online information retrieval on large volume of encrypted data. In this paper, a hierarchical clustering method is proposed to support more search semantics and also to meet the demand for fast ciphertext search within a big data environment. The proposed hierarchical approach clusters the documents based on the minimum relevance threshold, and then partitions the resulting clusters into sub-clusters until the constraint on the maximum size of cluster is reached. In the search phase, this approach can reach a linear computational complexity against an exponential size increase of document collection. In order to verify the authenticity of search results, a structure called minimum hash sub-tree is designed in this paper. Experiments have been conducted using the collection set built from the IEEE Xplore. The results show that with a sharp increase of documents in the dataset the search time of the proposed method increases linearly whereas the search time of the traditional method increases exponentially. Furthermore, the proposed method has an advantage over the traditional method in the rank privacy and relevance of retrieved documents. Chi Chen 0001, Xiaojie Zhu, Peisong Shen, Jiankun Hu, Song Guo 0001, Zahir Tari, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | SOLS: A scheme for outsourced location based service
Chi Chen 0001, Xiaojie Zhu, Peisong Shen, Jing Yu 0007, Hong Zou, Jiankun Hu |
J. Netw. Comput. Appl. | 1 |
| 2008 | Research on Malicious Transaction Processing Method of Database SystemabstractRecovery from information attacks is difficult because DBMS is not designed to deal with malicious committed transactions. A few existing methods developed for this purpose rely on operation logs, which can't express the dependency between different transactions directly. These methods usually use rollback mechanism and abandon results of innocent transactions to maintain correctness, which may indeed be used as an approach to realize DOS attack. Hence, it's necessary to find out the malicious transaction and subsequent transactions depending on it precisely. In this paper, the definition of transaction recovery log is presented and each log item records the actions taken in one transaction, by which, we can calculate transactions' dependency directly. Based on the log model and the algorithm for log's creation, the dependency calculation and data recovery algorithm are studied, which are proofed to be complete and correct. Using transaction recovery log and the algorithm, database system can significantly enhance the performance of recovery for defensive information warfare. Chi Chen 0001, Dengguo Feng, Min Zhang 0043, He-qun Xian |
WAIM | 1 |