Jinshuai Yang

dblp:267/5357 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0002-6293-1981ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Security and privacy · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust Streaming Tensor Train Completion for Dynamic Space-based Spectrum Situation Map Construction
Jinshuai Yang, Ruifeng Xiao
ICC1
2026 Context-fused emotional flow modeling for dynamic affective recognition in dialogue interactions
Zhixiao Qi, Jinshuai Yang, Dandan Tao, Chengqi Li, Minghu Jiang, Yongfeng Huang 0001
Expert Syst. Appl.3
2026 A Novel Framework of Semantic-Based Text Steganography
abstract
Text steganography helps protect citizens' freedom of speech and privacy in cyberspace by constructing innocent-looking texts to evade censorship and surveillance. Existing methods, especially generative ones, rely on delicate character-level manipulations, making them vulnerable to failure even under minor alterations. This paper introduces a novel framework that shifts from character-level to semantic space-based information hiding, which greatly improves improving robustness and reliability. The framework comprises three phases: preparation, where a stable semantic space is designed for hiding information; embedding, where a reversible codebook maps binary messages to semantemes via semantic encoding and source coding; and synthesis, where large language models act as multi-agent systems to produce stegotexts that preserve semantic consistency. Extensive experiments show that our method matches state-of-the-art techniques in hiding capacity and text quality, while significantly surpassing them in robustness-achieving at least an 81.2% higher correct message rate under three types of attacks. These findings highlight the strong potential of semantic-based text steganography.
Jinshuai Yang, Minhao Bai, Kaiyi Pang, Yue Gao 0003, Yongfeng Huang 0001
IEEE Trans. Dependable Secur. Comput.1
2025 Integrating Textual and Emotional Dynamics for Accurate Detection of Mental Health Disorders in Social Media
Zhixiao Qi, Jinshuai Yang, Zhechen Wei, Congqi Wang, Shizhong Yang, Minghu Jiang, Yongfeng Huang 0001
CogSci3
2025 WinStega: An Adaptive Robust Enhancement Framework for Generative Linguistic Steganography
abstract
With the increasing prevalence of surveillance, safeguarding personal privacy has become a critical concern. To protect privacy, various linguistic steganography methods have been developed to conceal private information within seemingly innocuous text for covert communication. However, these methods are highly sensitive to alterations in the stego text, rendering the extraction of secret information impossible if any changes occur, thus limiting their practical application. In this paper, we introduce WinStega, an adaptive and robust linguistic steganography method that employs substring decoding to withstand edit attacks. WinStega embeds secret messages discontinuously using a sliding window approach, incorporating entropy-based constraints to enhance imperceptibility while preserving linguistic quality. This plug-and-play method does not require additional model training and can be implemented during the inference stage, enhancing the robustness of stego texts. Extensive evaluations using three language models demonstrate that WinStega produces high linguistic quality, imperceptible stegotexts and can partially recover secret messages even under adversarial attacks.
Kaiyi Pang, Minhao Bai, Jinshuai Yang, Minghu Jiang, Yongfeng Huang 0001
ICASSP3
2025 Provably Robust and Secure Steganography in Asymmetric Resource Scenario
abstract
To circumvent the unbridled and ever-encroaching surveillance and censorship in cyberspace, steganography has garnered attention for its ability to hide private information in innocent-looking carriers. Current provably secure steganography approaches require a pair of encoder and decoder to hide and extract private messages, both of which must run the same model with the same input to obtain identical distributions. These requirements pose significant challenges to the practical implementation of steganography, including limited access to powerful hardware and the intolerance of any changes to the shared input. To relax the limitation of hardware and solve the challenge of vulnerable shared input, a novel and practically significant scenario with asymmetric resource should be considered, where only the encoder is high-resource and accessible to powerful models while the decoder can only read the stegano-graphic carriers without any other model's input. This paper proposes a novel provably robust and secure steganography framework for the asymmetric resource setting. Specifically, the encoder uses various permutations of distribution to hide secret bits, while the decoder relies on a sampling function to extract the hidden bits by guessing the permutation used. Further, the sampling function only takes the steganographic carrier as input, which makes the decoder independent of model's input and model itself. A comprehensive assessment of applying our framework to generative models substantiates its effectiveness. Our implementation demonstrates robustness when transmitting over binary symmetric channels with errors.
Minhao Bai, Jinshuai Yang, Kaiyi Pang, Zhen Yang 0015, Yongfeng Huang 0001
SP2
2025 Shimmer: a Provably Secure Steganography Based on Entropy Collecting Mechanism
Minhao Bai, Kaiyi Pang, Guorui Liao, Jinshuai Yang, Yongfeng Huang 0001
USENIX Security Symposium4
2025 A Framework for Designing Provably Secure Steganography
Guorui Liao, Jinshuai Yang, Weizhi Shao, Yongfeng Huang 0001
USENIX Security Symposium2
2025 A plug-and-play method for linguistic alignment in language models
Kaiyi Pang, Minhao Bai, Jinshuai Yang, Yue Gao 0003, Minghu Jiang, Yongfeng Huang 0001
Knowl. Based Syst.3
2025 Class-Aware Adversarial Unsupervised Domain Adaptation for Linguistic Steganalysis
abstract
Recent advancements in deep learning have significantly improved linguistic steganalysis, but challenges persist when labeled samples are scarce in the target domain. Existing cross-domain linguistic steganalysis methods seek to improve model generalization by minimizing the domain discrepancy between the source and target domains. However, these steganalysis methods often struggle with incorrect alignment between stego and cover texts in both domains, which hampers the generalization of steganalysis models. Additionally, they struggle to capture domain-specific features of the target domain, reducing the effectiveness of steganalysis models in discriminating stego texts. To address these issues, we propose a novel Class-aware Adversarial unsupervised Domain Adaptation (CADA) method, which operates in two stages. In the first stage, Class-aware Adversarial Pre-Training (CAPT), we design the Weighted Class-Aware Domain Distance (WCADD) to leverage class information of stego and cover texts. This ensures accurate class-aware alignment across domains. In the CAPT stage, the steganalysis model is pre-trained with WCADD, Class-Aware Adversarial Training (CAAT), and Class-Aware Label Smoothing (CALS) to enhance its ability to extract domain-invariant features, thereby improving its generalization. In the second stage, Class-aware Fine-Tuning (CFT), we employ the pre-trained steganalysis model alongside the Class-Aware Progressive Strategy (CAPS) to generate pseudo-labels for the target domain. Fine-tuning the model with these pseudo-labels enhances its ability to recognize domain-specific features, thereby improving its performance in discriminating stego texts within the target domain. Extensive experiments demonstrate that our proposed method outperforms the existing baseline methods.
Zhen Yang 0015, Yufei Luo, Jinshuai Yang, Ru Zhang 0002, Yongfeng Huang 0001
IEEE Trans. Inf. Forensics Secur.3
2024 RedditEM: Unveiling Diachronic Semantic Shifts in Social Network Discourse
Sixing Wu, Jinshuai Yang, Minghu Jiang, Yongfeng Huang 0001
ACML3
2024 Enhancing Steganography of Generative Image Based on Image Retouching
abstract
Steganography, which hides messages within innocent-looking carriers, is an essential technique to protect data privacy. The rapid advancement of generative models makes AI-generated images a potential steganographic carrier. However, the distortion resulting from the embedding of messages makes it difficult for steganographic images to escape deep-learning based detection, especially since such distortion is more pronounced in generative images. To improve the imperceptibility of steganographic generative images, in this paper we propose a cover-source switching based steganographic scheme employing image retouching. Specifically, we utilize an image-adaptive LUTs (LookUp Tables) model to generate the LUT required to retouch the cover image. The generated LUT is then applied to the stego image, effectively obfuscating the steganographic behavior. To ensure accurate extraction, we introduce wet cost to mark ambiguous elements that are strictly prohibited from modification. Experimental results show that our scheme can significantly improve the imperceptibility of the steganographic generative images. Leveraging the reproducibility of generative images, we are able to embed secrets at the pixel level, resulting in higher payload and extraction accuracy compared to existing cover-source switching based steganographic methods.
Yue Gao 0003, Jinshuai Yang, Cheng Chen 0049, Kaiyi Pang, Yongfeng Huang 0001
ICASSP2
2024 FREmax: A Simple Method Towards Truly Secure Generative Linguistic Steganography
abstract
Generative Linguistic Steganography (GLS) is applied to protect privacy against excessive censorship by employing Language Models (LMs) to hide privacy messages in texts. To effectively circumvent censorship, GLS generates steganographic texts (stegos) that closely resemble normal human texts (covers) as possible. However, due to the inherent distribution difference between LM-generated text and human covers, existing methods that simply use LMs to generate stegos face challenges in achieving sufficient imperceptibility. To narrow the gap between stegos and covers, this paper proposes a distribution reformation method named ${\mathbf{Frequency}}$ ${\mathbf{REformed}}$ ${\mathbf{Softmax}}$ $\left( {{\mathbf{FREmax}}} \right)$. ${\mathbf{FREmax}}$ generates highly imperceptible stegos aligned with human text by reforming the softmax function in the generation stage of LMs. This reformation is based on the frequency distribution of tokens in the human corpus, ensuring that the distribution of LM-generated stegos closely resembles that of humans. Extensive experimental results show that ${\mathbf{FREmax}}$ improves the linguistic quality and imperceptibility of the generated stegos, providing a valuable remedy to existing GLS methods .1
Kaiyi Pang, Minhao Bai, Jinshuai Yang, Huili Wang 0001, Minghu Jiang, Yongfeng Huang 0001
ICASSP3
2024 Co-Stega: Collaborative Linguistic Steganography for the Low Capacity Challenge in Social Media
abstract
Social media platforms, with their extensive and real-time text data, are important application environments for linguistic generative steganography. However, the fragmented and context-constrained nature of social media text leads to a notably low capacity for hiding messages, making linguistic steganography impractical in real social media platforms. More frustratingly, even for high-capacity linguistic generative steganography, the upper bound of capacity is limited to a low level when required to meet a slightly strict security level under Cachin's model. To overcome the low capacity challenge, we identified an indicator of capacity that is independent of any specific steganography method to analyze the origins of this challenge, then we proposed a novel linguistic steganography framework named Collaborative Steganography (Co-Stega). Co-Stega utilizes existing texts and contextual relevance between texts in social media, collaboratively embedding secret messages in an existing text and its contextually related text via efficient retrieval and generation respectively. Additionally, we proposed an innovative and simple technique called "Entropy Enhancement Strategy", which effectively increases the entropy of generated text, thereby enhancing capacity further. Our evaluation shows that Co-Stega significantly improves capacity and maintains text quality, making it a valuable extension for linguistic generative steganography in social media platforms.
Guorui Liao, Jinshuai Yang, Kaiyi Pang, Yongfeng Huang 0001
IH&MMSec2
2024 A machine reading comprehension framework for recognizing emotion cause in conversations
Yexuan Zhang, Sixing Wu, Jinshuai Yang, Xuanmei Qin, Lizhi Ying, Minghu Jiang, Yongfeng Huang 0001
Knowl. Based Syst.4
2024 FET-LM: Flow-Enhanced Variational Autoencoder for Topic-Guided Language Modeling
abstract
Variational autoencoder (VAE) is widely used in tasks of unsupervised text generation due to its potential of deriving meaningful latent spaces, which, however, often assumes that the distribution of texts follows a common yet poor-expressed isotropic Gaussian. In real-life scenarios, sentences with different semantics may not follow simple isotropic Gaussian. Instead, they are very likely to follow a more intricate and diverse distribution due to the inconformity of different topics in texts. Considering this, we propose a flow-enhanced VAE for topic-guided language modeling (FET-LM). The proposed FET-LM models topic and sequence latent separately, and it adopts a normalized flow composed of householder transformations for sequence posterior modeling, which can better approximate complex text distributions. FET-LM further leverages a neural latent topic component by considering learned sequence knowledge, which not only eases the burden of learning topic without supervision but also guides the sequence component to coalesce topic information during training. To make the generated texts more correlative to topics, we additionally assign the topic encoder to play the role of a discriminator. Encouraging results on abundant automatic metrics and three generation tasks demonstrate that the FET-LM not only learns interpretable sequence and topic representations but also is fully capable of generating high-quality paragraphs that are semantically consistent.
Haoqin Tu, Zhongliang Yang, Jinshuai Yang, Linna Zhou, Yongfeng Huang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 LINK: Linguistic Steganalysis Framework with External Knowledge
abstract
Linguistic steganalysis is the technology to distinguish whether looking-innocent texts hide covert (possibly hazardous) messages. Traditional methods, dominantly focusing on internal linguistic difference in texts, are seriously challenged by the recent linguistic steganography technology that can reduce the difference to near zero. However, even via the most advanced linguistic steganography methods, due to the random and uncontrollable message bits, steganographic texts may express content against common sense knowledge. To fully employ this defect of linguistic steganography, we propose LINK, a novel Linguistic steganalysis framework with the help of external Knowledge. We link texts to the external knowledge database, and employ Graph Neural Networks (GNNs) to translate linked knowledge into knowledge features, while linguistic features will be captured by the same modules from existing methods. Knowledge features and linguistic features will be combined to make final decisions. Extensive experimental results show that owing to additional external knowledge, the proposed framework can effectively compensate for the shortcomings of existing methods.1
Jinshuai Yang, Zhongliang Yang, Xinrui Ge, Yue Gao 0003, Yongfeng Huang 0001
ICASSP1
2023 Minimizing Distortion in Steganography via Adaptive Language Model Tuning
Cheng Chen 0049, Jinshuai Yang, Yue Gao 0003, Huili Wang 0001, Yongfeng Huang 0001
ICONIP (12)2
2023 CATS: Connection-Aware and Interaction-Based Text Steganalysis in Social Networks
Kaiyi Pang, Jinshuai Yang, Yue Gao 0003, Minhao Bai, Zhongliang Yang, Minghu Jiang, Yongfeng Huang 0001
ICONIP (5)2
2023 Hi-Stega: A Hierarchical Linguistic Steganography Framework Combining Retrieval and Generation
Huili Wang 0001, Zhongliang Yang, Jinshuai Yang, Yue Gao 0003, Yongfeng Huang 0001
ICONIP (5)3
2023 DNA Synthetic Steganography Based on Conditional Probability Adaptive Coding
abstract
Steganography is an important technology for ensuring the security of cyberspace and the privacy of communications. In the last decade, emerging biotechnology has made it possible for DNA to be used as a promising steganographic carrier with high hidden capacity, high imperceptibility and high feasibility. However, severe statistical distortion might appear in steganographic carriers generated by existing DNA steganographies when they are compared with the natural ones. Therefore, efforts are being made to seek an advanced strategy to generate quasi-natural steganographic carriers with a strong anti-steganalysis capability. In this work, we first thoroughly analyze and model the numerous complicated statistical properties that exist in natural DNA chains, and then utilize the LSTM model to learn the serialized statistical properties. After obtaining an optimal sequence model that highly satisfies the statistical properties of natural DNA chains, we utilize the Adaptive Dynamic Grouping (ADG) algorithm to perform information hiding. In addition, we have carried out experimental analysis and verification from the perspectives of perceptual-imperceptibility, statistical-imperceptibility, and anti-steganalysis capability, all of which show that our proposed steganography method vastly outperforms previous DNA steganographic methods, taking a successful step towards achieving higher security DNA steganography.
Chenwei Huang, Zhongliang Yang, Zhiwen Hu, Jinshuai Yang, Haochen Qi, Lei Zheng 0008
IEEE Trans. Inf. Forensics Secur.4
2023 Linguistic Steganalysis in Few-Shot Scenario
abstract
Due to the widespread use of text in cyberspace, linguistic steganography, which hides secret information into normal texts, develops quickly in these years. While linguistic steganography protects users’ privacy, it also has the risk of being abused to endanger network security. Therefore, its corresponding detection technology, namely linguistic steganalysis, has attracted more and more researchers’ attention in the past several years. However, most of the current linguistic steganalysis methods rely heavily on a large number of labeled samples, which presents a significant gap from real-world scenarios where labeled steganographic samples are difficult to obtain. In this paper, we proposed the Pre-trained Language model with Self-training for Few-shot Linguistic Steganalysis (LSFLS) method which effectively copes with few-shot linguistic steganalysis through a small number of labeled samples and some auxiliary unlabeled samples. Numerous experiments have proved that the proposed method can achieve high detection accuracy of linguistic steganalysis when only a few labeled samples are provided (even less than 10), significantly improving the detection ability of existing methods in few-shot scenario. Furthermore, the experimental results demonstrate that the proposed method can maintain good detection capability in the case of data source mismatch and label unbalance. We believe that our work will greatly advance the practical application of linguistic steganalysis techniques.
Huili Wang 0001, Zhongliang Yang, Jinshuai Yang, Cheng Chen 0049, Yongfeng Huang 0001
IEEE Trans. Inf. Forensics Secur.3
2023 Linguistic Steganalysis Toward Social Network
abstract
With the rapid development of the internet and social media, linguistic steganography can be easily abused in social networks to make considerable damage to varied aspects like personal privacy, network virus and national defense. Currently, considerable linguistic steganalysis methods are proposed to detect harmful steganographic carriers. However, almost all the existing methods fail in real social networks, since they are only devoted to the linguistic features that are extreme insufficient owing to the extreme sparsity and extreme fragmentation challenges of real social networks. In this paper, we attempt to fill the long-standing gap that the datasets and effective methods are absent for hunting steganographic texts in social network scenarios. Concretely, we construct a dataset called Stego-Sandbox to simulate the real social network scenarios, which contains texts and their relation. And we propose an effective linguistic steganalysis framework integrating linguistic features contained in texts and context features represented by these connections. Extensive experimental results demonstrate owing to the captured context features, our proposed framework can effectively compensate for shortcomings of these existing methods and tremendously improve their detection ability in real social network scenarios.
Jinshuai Yang, Zhongliang Yang, Haoqin Tu, Yongfeng Huang 0001
IEEE Trans. Inf. Forensics Secur.1
2022 PCAE: A framework of plug-in conditional auto-encoder for controllable text generation
Haoqin Tu, Zhongliang Yang, Jinshuai Yang, Si-yu Zhang 0001, Yongfeng Huang 0001
Knowl. Based Syst.3
2022 SeSy: Linguistic Steganalysis Framework Integrating Semantic and Syntactic Features
abstract
With the rapid development of natural language processing technology and linguistic steganography, linguistic steganalysis gains considerable interest in recent years. Current advanced methods dominantly focus on statistical features in semantic view yet ignore syntax structure of text, which leads to limited performance to some newly statistically indistinguishable steganography algorithms. To fill this gap, in this paper, we propose a novel linguistic steganalysis framework named SeSy to integrate bothsemantic andsyntactic features. Specifically, we propose to employ transformer-architecture language model as semantics extractor and leverage a graph attention network to retain syntactic features. Extensive experimental results show that owing to additional syntactic information, the SeSy framework effectively brings about remarkable improvement to current advanced linguistic steganalysis methods.
Jinshuai Yang, Zhongliang Yang, Si-yu Zhang 0001, Haoqin Tu, Yongfeng Huang 0001
IEEE Signal Process. Lett.1
2021 Linguistic Steganography: From Symbolic Space to Semantic Space
abstract
Previous works about linguistic steganography such as synonym substitution and sampling-based methods usually manipulate observed symbols explicitly to conceal secret information, which may give rise to security risks. In this letter, in order to preclude straightforward operation on observed symbols, we explored generation-based linguistic steganography in latent space by means of encoding secret messages in the selection of implicit attributes (semanteme) of natural language. We proposed a novel framework of linguistic semantic steganography based on rejection sampling strategy. Concretely, we utilized controllable text generation model for embedding and semantic classifier for extraction. In experiments, a model based on CTRL and BERT is implemented for further quantitative assessment. Results reveal that our approach is able to achieve satisfactory efficiency as well as nearly perfect imperceptibility. Our code is available at https://github.com/YangzlTHU/Linguistic-Steganography-and-Steganalysis/tree/master/Steganography/Linguistic-Semantic-Steganography.
Si-yu Zhang 0001, Zhongliang Yang, Jinshuai Yang, Yongfeng Huang 0001
IEEE Signal Process. Lett.3