VLDB 2026 Research / reviewers in the wild / expert
Kaiyi Pang
dblp:361/4092
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0248-171XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Framework of Semantic-Based Text SteganographyabstractText steganography helps protect citizens' freedom of speech and privacy in cyberspace by constructing innocent-looking texts to evade censorship and surveillance. Existing methods, especially generative ones, rely on delicate character-level manipulations, making them vulnerable to failure even under minor alterations. This paper introduces a novel framework that shifts from character-level to semantic space-based information hiding, which greatly improves improving robustness and reliability. The framework comprises three phases: preparation, where a stable semantic space is designed for hiding information; embedding, where a reversible codebook maps binary messages to semantemes via semantic encoding and source coding; and synthesis, where large language models act as multi-agent systems to produce stegotexts that preserve semantic consistency. Extensive experiments show that our method matches state-of-the-art techniques in hiding capacity and text quality, while significantly surpassing them in robustness-achieving at least an 81.2% higher correct message rate under three types of attacks. These findings highlight the strong potential of semantic-based text steganography. Jinshuai Yang, Minhao Bai, Kaiyi Pang, Yue Gao 0003, Yongfeng Huang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | WinStega: An Adaptive Robust Enhancement Framework for Generative Linguistic SteganographyabstractWith the increasing prevalence of surveillance, safeguarding personal privacy has become a critical concern. To protect privacy, various linguistic steganography methods have been developed to conceal private information within seemingly innocuous text for covert communication. However, these methods are highly sensitive to alterations in the stego text, rendering the extraction of secret information impossible if any changes occur, thus limiting their practical application. In this paper, we introduce WinStega, an adaptive and robust linguistic steganography method that employs substring decoding to withstand edit attacks. WinStega embeds secret messages discontinuously using a sliding window approach, incorporating entropy-based constraints to enhance imperceptibility while preserving linguistic quality. This plug-and-play method does not require additional model training and can be implemented during the inference stage, enhancing the robustness of stego texts. Extensive evaluations using three language models demonstrate that WinStega produces high linguistic quality, imperceptible stegotexts and can partially recover secret messages even under adversarial attacks. Kaiyi Pang, Minhao Bai, Jinshuai Yang, Minghu Jiang, Yongfeng Huang 0001 |
ICASSP | 1 |
| 2025 | Reinforcement Learning-based Copyright Protection Watermarking for Large Language ModelabstractWith the widespread application of large language models (LLMs) in the field of natural language processing (NLP), copyright protection issues are becoming increasingly important.As an effective means of copyright protection, watermarking technology can help developers and users prove the copyright ownership of the model.However, existing watermarking methods struggle to optimize both watermark effectiveness and model performance simultaneously.To overcome this challenge, in this paper, we propose a backdoor watermarking method named CRMark based on Chain-of-Thought (CoT) and reinforcement learning.This method embeds backdoor symbols into the prompt of the datasets and adds copyright information as the watermark into the model response.Reinforcement learning is further employed to alleviate the performance degradation of the watermark model in normal tasks.Experimental results show that the CRMark does not reduce the model's original task performance while effectively maintaining the effectiveness of backdoor watermarks, with the watermark success rate of up to 96.5%. Shengnan Guo 0008, Kaiyi Pang, Zhongliang Yang, Yu Qing, Yongfeng Huang 0001 |
IH&MMSec | 2 |
| 2025 | Provably Robust and Secure Steganography in Asymmetric Resource ScenarioabstractTo circumvent the unbridled and ever-encroaching surveillance and censorship in cyberspace, steganography has garnered attention for its ability to hide private information in innocent-looking carriers. Current provably secure steganography approaches require a pair of encoder and decoder to hide and extract private messages, both of which must run the same model with the same input to obtain identical distributions. These requirements pose significant challenges to the practical implementation of steganography, including limited access to powerful hardware and the intolerance of any changes to the shared input. To relax the limitation of hardware and solve the challenge of vulnerable shared input, a novel and practically significant scenario with asymmetric resource should be considered, where only the encoder is high-resource and accessible to powerful models while the decoder can only read the stegano-graphic carriers without any other model's input. This paper proposes a novel provably robust and secure steganography framework for the asymmetric resource setting. Specifically, the encoder uses various permutations of distribution to hide secret bits, while the decoder relies on a sampling function to extract the hidden bits by guessing the permutation used. Further, the sampling function only takes the steganographic carrier as input, which makes the decoder independent of model's input and model itself. A comprehensive assessment of applying our framework to generative models substantiates its effectiveness. Our implementation demonstrates robustness when transmitting over binary symmetric channels with errors. Minhao Bai, Jinshuai Yang, Kaiyi Pang, Zhen Yang 0015, Yongfeng Huang 0001 |
SP | 3 |
| 2025 | Shimmer: a Provably Secure Steganography Based on Entropy Collecting Mechanism
Minhao Bai, Kaiyi Pang, Guorui Liao, Jinshuai Yang, Yongfeng Huang 0001 |
USENIX Security Symposium | 2 |
| 2025 | A plug-and-play method for linguistic alignment in language models
Kaiyi Pang, Minhao Bai, Jinshuai Yang, Yue Gao 0003, Minghu Jiang, Yongfeng Huang 0001 |
Knowl. Based Syst. | 1 |
| 2025 | ModelShield: Adaptive and Robust Watermark Against Model Extraction AttackabstractLarge language models (LLMs) demonstrate general intelligence across a variety of machine learning tasks, thereby enhancing the commercial value of their intellectual property (IP). To protect this IP, model owners typically allow user access only in a black-box manner, however, adversaries can still utilize model extraction attacks to steal the model intelligence encoded in model generation. Watermarking technology offers a promising solution for defending against such attacks by embedding unique identifiers into the model-generated content. However, existing watermarking methods often compromise the quality of generated content due to heuristic alterations and lack robust mechanisms to counteract adversarial strategies, thus limiting their practicality in real-world scenarios. In this paper, we introduce an adaptive and robust watermarking method (named ModelShield) to protect the IP of LLMs. Our method incorporates a self-watermarking mechanism that allows LLMs to autonomously insert watermarks into their generated content to avoid the degradation of model content. We also propose a robust watermark detection mechanism capable of effectively identifying watermark signals under the interference of varying adversarial strategies. Besides, ModelShield is a plug-and-play method that does not require additional model training, enhancing its applicability in LLM deployments. Extensive evaluations on two real-world datasets and three LLMs demonstrate that our method surpasses existing methods in terms of defense effectiveness and robustness while significantly reducing the degradation of watermarking on the model-generated content. Kaiyi Pang, Tao Qi 0001, Chuhan Wu, Minhao Bai, Minghu Jiang, Yongfeng Huang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Enhancing Steganography of Generative Image Based on Image RetouchingabstractSteganography, which hides messages within innocent-looking carriers, is an essential technique to protect data privacy. The rapid advancement of generative models makes AI-generated images a potential steganographic carrier. However, the distortion resulting from the embedding of messages makes it difficult for steganographic images to escape deep-learning based detection, especially since such distortion is more pronounced in generative images. To improve the imperceptibility of steganographic generative images, in this paper we propose a cover-source switching based steganographic scheme employing image retouching. Specifically, we utilize an image-adaptive LUTs (LookUp Tables) model to generate the LUT required to retouch the cover image. The generated LUT is then applied to the stego image, effectively obfuscating the steganographic behavior. To ensure accurate extraction, we introduce wet cost to mark ambiguous elements that are strictly prohibited from modification. Experimental results show that our scheme can significantly improve the imperceptibility of the steganographic generative images. Leveraging the reproducibility of generative images, we are able to embed secrets at the pixel level, resulting in higher payload and extraction accuracy compared to existing cover-source switching based steganographic methods. Yue Gao 0003, Jinshuai Yang, Cheng Chen 0049, Kaiyi Pang, Yongfeng Huang 0001 |
ICASSP | 4 |
| 2024 | FREmax: A Simple Method Towards Truly Secure Generative Linguistic SteganographyabstractGenerative Linguistic Steganography (GLS) is applied to protect privacy against excessive censorship by employing Language Models (LMs) to hide privacy messages in texts. To effectively circumvent censorship, GLS generates steganographic texts (stegos) that closely resemble normal human texts (covers) as possible. However, due to the inherent distribution difference between LM-generated text and human covers, existing methods that simply use LMs to generate stegos face challenges in achieving sufficient imperceptibility. To narrow the gap between stegos and covers, this paper proposes a distribution reformation method named ${\mathbf{Frequency}}$ ${\mathbf{REformed}}$ ${\mathbf{Softmax}}$ $\left( {{\mathbf{FREmax}}} \right)$. ${\mathbf{FREmax}}$ generates highly imperceptible stegos aligned with human text by reforming the softmax function in the generation stage of LMs. This reformation is based on the frequency distribution of tokens in the human corpus, ensuring that the distribution of LM-generated stegos closely resembles that of humans. Extensive experimental results show that ${\mathbf{FREmax}}$ improves the linguistic quality and imperceptibility of the generated stegos, providing a valuable remedy to existing GLS methods .1 Kaiyi Pang, Minhao Bai, Jinshuai Yang, Huili Wang 0001, Minghu Jiang, Yongfeng Huang 0001 |
ICASSP | 1 |
| 2024 | Co-Stega: Collaborative Linguistic Steganography for the Low Capacity Challenge in Social MediaabstractSocial media platforms, with their extensive and real-time text data, are important application environments for linguistic generative steganography. However, the fragmented and context-constrained nature of social media text leads to a notably low capacity for hiding messages, making linguistic steganography impractical in real social media platforms. More frustratingly, even for high-capacity linguistic generative steganography, the upper bound of capacity is limited to a low level when required to meet a slightly strict security level under Cachin's model. To overcome the low capacity challenge, we identified an indicator of capacity that is independent of any specific steganography method to analyze the origins of this challenge, then we proposed a novel linguistic steganography framework named Collaborative Steganography (Co-Stega). Co-Stega utilizes existing texts and contextual relevance between texts in social media, collaboratively embedding secret messages in an existing text and its contextually related text via efficient retrieval and generation respectively. Additionally, we proposed an innovative and simple technique called "Entropy Enhancement Strategy", which effectively increases the entropy of generated text, thereby enhancing capacity further. Our evaluation shows that Co-Stega significantly improves capacity and maintains text quality, making it a valuable extension for linguistic generative steganography in social media platforms. Guorui Liao, Jinshuai Yang, Kaiyi Pang, Yongfeng Huang 0001 |
IH&MMSec | 3 |
| 2023 | CATS: Connection-Aware and Interaction-Based Text Steganalysis in Social Networks
Kaiyi Pang, Jinshuai Yang, Yue Gao 0003, Minhao Bai, Zhongliang Yang, Minghu Jiang, Yongfeng Huang 0001 |
ICONIP (5) | 1 |