Ru Zhang 0002

dblp:80/446-2 · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0001-6641-3236ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Security and privacy · 8 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
abstract
Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks show that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model on a single consumer-grade GPU, greatly lowering the barrier to deploying VLA model.
Yihao Wang 0006, Pengxiang Ding, Can Cui 0008, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao 0008, Pengxu Hou, Siteng Huang, Ru Zhang 0002
AAAI14
2026 User profile constructed by multiple attributes for optimizing linguistic steganalysis in social networks
Yihao Wang 0006, Ru Zhang 0002
Expert Syst. Appl.5
2026 CLIProv: A contrastive log-to-intelligence multimodal approach for threat detection and provenance analysis
Ru Zhang 0002, Wanguo Zhao
Future Gener. Comput. Syst.2
2026 Aggregated text steganalysis toward social network based on efficient multi-perspective feature fusion
Ru Zhang 0002, Yongfeng Huang 0001
Knowl. Based Syst.2
2026 TASDF-Stega: High Capacity Secure Text-Audio Joint Steganography Using Diffusion Latent Space
abstract
Provably secure steganography ensures indistinguishability between stego and cover carrier through mathematical proofs. However, existing methods face limited embedding capacity and distribution synchronization challenges, especially at high embedding rates. To address these issues, we propose TASDF-Stega, a text-audio joint steganography method based on the latent space of diffusion models, which achieves high capacity and provable security. First, we design an encrypted steganographic mapping module with adaptive arithmetic decoding, which efficiently embeds secret information into the latent space while preserving the distribution. Second, a reversible secret diffusion mechanism enables high-capacity embedding and precise extraction. Moreover, to resolve the problem of distribution parameter synchronization in practical communication, we introduce an audio-assisted joint encode module. This design ensures accurate reconstruction of the diffusion inverse process and avoids cumulative extraction errors. Experimental results on multiple datasets demonstrate that TASDF-Stega achieves provable security, the outperforms state-of-the-art methods in embedding capacity and imperceptibility.
Zhen Yang 0015, Yelei Wang, Yufei Luo, Ru Zhang 0002
IEEE Signal Process. Lett.5
2025 GLoCIM: Global-view Long Chain Interest Modeling for news recommendation
abstract
Accurately recommending candidate news articles to users has always been the core challenge of news recommendation system. News recommendations often require modeling of user interest to match candidate news. Recent efforts have primarily focused on extracting local subgraph information in a global click graph constructed by the clicked news sequence of all users. However, the computational complexity of extracting global click graph information has hindered the ability to utilize far-reaching linkage which is hidden between two distant nodes in global click graph collaboratively among similar users. To overcome the problem above, we propose a Global-view Long Chain Interests Modeling for news recommendation (GLoCIM), which combines neighbor interest with long chain interest distilled from a global click graph, leveraging the collaboration among similar users to enhance news recommendation. We therefore design a long chain selection algorithm and long chain interest encoder to obtain global-view long chain interest from the global click graph. We design a gated network to integrate long chain interest with neighbor interest to achieve the collaborative interest among similar users. Subsequently we aggregate it with local news category-enhanced representation to generate final user representation. Then candidate news representation can be formed to match user representation to achieve news recommendation. Experimental results on real-world datasets validate the effectiveness of our method to improve the performance of news recommendation.
Zhen Yang 0015, Tao Qi 0001, Tianyun Zhang, Ru Zhang 0002, Yongfeng Huang 0001
COLING6
2025 MM-LogVec: System Log Anomaly Detection Method Based on Multimodal Representation Learning
abstract
Advanced persistent threats (APTs) pose significant risks to national infrastructure and corporate security. System logs record interactions between system entities, which are widely used for APT detection. However, the complex syntax and intricate relationships in system logs pose significant challenges to the performance of deep learning models. This paper proposes MM-LogVec, a novel anomaly detection method based on multimodal feature representation. It introduces three pre-training tasks to learn joint embeddings of logs and security events, establishing semantic associations between log operations and security events through cross-modal feature alignment, thereby enhancing the understanding of complex behaviors. Additionally, a lightweight graph autoencoder is used to model interaction relationships among entities, enabling entity-level anomaly detection via feature reconstruction. Evaluations across 10 APT scenarios show MM-LogVec’s exceptional performance, with a true positive rate of 98.19% and a false positive rate of 3.93%, while improving processing speed by a factor of 9 compared to the state-of-the-art AIRTAG.
Ru Zhang 0002
ICASSP2
2025 An Adversarial Perturbation Generation Method for Image Anti-Forensics Based on Dual-Path Spatial Attention GAN
abstract
Adversarial attacks are essential for evaluating the robustness of deep learning-based forensics, revealing potential vulnerabilities. However, most existing adversarial sample generation methods face significant trade-offs between anti-forensic ability, transferability, and visual quality, as they typically apply perturbations either uniformly across entire images or modify only a limited number of arbitrary pixels. This paper proposes a novel method for generating anti-forensic images through a salient region-focused adversarial GAN based on meta-learning. By developing a dual-path perturbation generation model, we enable the generation of inconspicuous perturbations based on the spatial attention module. During the model’s training process, the perturbation generator uses a multi-task training strategy based on meta-learning to enhance anti-forensics transferability. Experimental results demonstrate that the proposed method outperforms state-of-the-art anti-forensic methods in maintaining rich image details while achieving higher anti-forensic ability.
Yihong Lu, Ru Zhang 0002
ICASSP3
2025 Clustering-Driven Pseudo-Labeling in Source-Free Domain Adaptation for Linguistic Steganalysis
abstract
Linguistic steganalysis often encounters domain shift in practice, mainly due to differences in text sources and steganographic methods.These variations create distribution discrepancies between the training (source domain) and test (target domain) sets, reducing detection accuracy.Most existing domain adaptation methods for linguistic steganalysis rely on labeled source domain data to alleviate these issues, but due to data privacy or transmission costs, source domain data is often unavailable in many real-world scenarios.Without access to this data, models cannot directly compare feature distributions between the source and target domains, hindering the model's ability to learn the target domain's features and ultimately affecting detection performance for stego texts.In this paper, we propose a Clustering-driven Pseudo-labeling method for Source-free domain adaptation in Linguistic Steganalysis (CPSLS).During the adaptation phase, we leverage the clustering structure of the target domain data to generate pseudo-labels, helping the model identify stego features in the target domain.Additionally, we use a weighted classification loss function to reduce the impact of incorrect pseudo-labels.To prevent the model from overlooking the diversity between stego and cover texts during optimization, we introduce a prediction diversity loss, improving the model's ability to differentiate between the two.Experimental results show that CPSLS not only has stronger practical applicability but also outperforms existing domain adaptation linguistic steganalysis in terms of detection accuracy.
Yufei Luo, Zhen Yang 0015, Yelei Wang, Ru Zhang 0002, Yongfeng Huang 0001
IH&MMSec5
2025 Bilingual generated text detection through semantic and statistical analysis
abstract
The release of Large Language Models (LLMs) has achieved human-level text generation, leading to malicious uses such as disinformation propagation and academic dishonesty. Existing research has faced substantial challenges in low detection rates and poor generalization on multilingual generated text and short text. To fill these gaps, in this paper, we propose a generic bilingual generated text detection model to integrate semantic and statistical features, which exhibits proficiency in English and Chinese. To obtain fine-grained features, we employ the multilingual pre-trained language model xlm-RoBERTa to extract the CLS vector as overall semantic features, integrating with statistical features log rank, probability, and cumulative probability for detection. Moreover, Shapley additive explanations (SHAP) serves to interpret the decision-making process. The experimental results demonstrate significant advancements over baselines, notably with the F1 score improvements exceeding 10% and 5% on the English and Chinese HC3 sentence-level datasets, respectively. Our proposed method exhibits higher generalization for advanced LLMs and out-of-domain datasets with a 91.13% F1 score, thereby providing a more robust solution for detecting generated text.
Chenxi Min, Ru Zhang 0002
Intell. Data Anal.2
2025 Linguistic Steganalysis via LLMs: Two Modes for Efficient Detection of Strongly Concealed Stego
abstract
To detect stego (steganographic text) in complex scenarios, linguistic steganalysis (LS) with various motivations has been proposed and achieved excellent performance. However, with the development of generative steganography, some stegos have strong concealment, especially after the emergence of LLMs-based steganography, the existing LS has low detection or cannot detect them. We designed a novel LS with two modes called LS-G/C. In the generation mode, we created an LS-task “description” and used the generation ability of LLM to explain whether texts to be detected are stegos. On this basis, we rethought the principle of LS and LLMs, and proposed the classification mode. In this mode, LS-G/C deleted the LS-task “description” and used the “causalLM” LLMs to extract steganographic features. The LS features can be extracted by only one pass of the model, and a linear layer with initialization weights is added to obtain the classification probability. Experiments on strongly concealed stegos show that LS-G/C significantly improves detection and reaches SOTA performance. Additionally, LS-G/C in classification mode greatly reduces training time while maintaining high performance.
Yihao Wang 0006, Ru Zhang 0002
IEEE Signal Process. Lett.3
2025 Multi-Classification of Linguistic Steganography Driven by Large Language Models
abstract
In linguistic steganalysis (LS), the fundamental requirement is to detect steganographic text (stego) effectively, and existing LS methods have shown excellent detection performance even in complex scenarios. However, different steganography employs various grammatical rules and syntactic structures, which result in distinct text representations. This diversity complicates the identification of stego in mixed datasets. A major challenge in LS is pinpointing the specific steganography used, which is crucial for developing effective extraction algorithms. Thus, this paper proposes a multi-classification method based on the Large Language Models (LLMs) called LSMC. This approach utilizes the strengths of LLMs in semantic understanding and contextual analysis, allowing for accurate classification. Experimental results show that the LSMC method can efficiently perform multi-classification, precisely discerning the steganography used in mixed datasets with up to 20 categories. This provides a feasible solution for fully deciphering steganography in subsequent stages.
Jie Wang 0098, Yihao Wang 0006, Ru Zhang 0002
IEEE Signal Process. Lett.4
2025 Linguistic Steganalysis via Text Dual Attention Fusing Statistical and Multi-Layer Semantic Features
abstract
Linguistic steganalysis faces the challenge of increasingly high-quality stego text, making it difficult to distinguish these text from cover text. The two main issues with current methods are: 1) deep learning models tend to overfit, which hurts their ability to apply to new situations, and 2) feature fusion models don't mix different types of features effectively, leading to poor results. In this letter, we propose aTextDualAttention linguistic steganalysis methodFusingStatistical andMulti-layerSemantic features (TDA-FSMS). TDA-FSMS firstly extracts multi-layer semantic features using different encoder layers from Enhanced Representation through knowledge integration (ERNIE), combining shallow and deep features to relieve potential overfitting. TDA-FSMS also designs a text dual attention network that simultaneously maps both multi-layer semantic and statistical features into a shared high-dimensional space to bring in a smooth feature fusion. Experimental results show that the text dual attention and the multi-layer semantic fusion enable TDA-FSMS to improve steganalysis performance than existing methods.
Zhen Yang 0015, Zhongliang Yang, Ru Zhang 0002
IEEE Signal Process. Lett.5
2025 Class-Aware Adversarial Unsupervised Domain Adaptation for Linguistic Steganalysis
abstract
Recent advancements in deep learning have significantly improved linguistic steganalysis, but challenges persist when labeled samples are scarce in the target domain. Existing cross-domain linguistic steganalysis methods seek to improve model generalization by minimizing the domain discrepancy between the source and target domains. However, these steganalysis methods often struggle with incorrect alignment between stego and cover texts in both domains, which hampers the generalization of steganalysis models. Additionally, they struggle to capture domain-specific features of the target domain, reducing the effectiveness of steganalysis models in discriminating stego texts. To address these issues, we propose a novel Class-aware Adversarial unsupervised Domain Adaptation (CADA) method, which operates in two stages. In the first stage, Class-aware Adversarial Pre-Training (CAPT), we design the Weighted Class-Aware Domain Distance (WCADD) to leverage class information of stego and cover texts. This ensures accurate class-aware alignment across domains. In the CAPT stage, the steganalysis model is pre-trained with WCADD, Class-Aware Adversarial Training (CAAT), and Class-Aware Label Smoothing (CALS) to enhance its ability to extract domain-invariant features, thereby improving its generalization. In the second stage, Class-aware Fine-Tuning (CFT), we employ the pre-trained steganalysis model alongside the Class-Aware Progressive Strategy (CAPS) to generate pseudo-labels for the target domain. Fine-tuning the model with these pseudo-labels enhances its ability to recognize domain-specific features, thereby improving its performance in discriminating stego texts within the target domain. Extensive experiments demonstrate that our proposed method outperforms the existing baseline methods.
Zhen Yang 0015, Yufei Luo, Jinshuai Yang, Ru Zhang 0002, Yongfeng Huang 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Multi-Grained Multimodal Interaction Network for Sentiment Analysis
abstract
Multimodal sentiment analysis aims to utilize different modalities including language, visual, and audio to identify human emotions in videos. Multimodal interaciton mechanism is the key challenge. Previous works lack modeling of multimodal interaction at different grain levels, and does not suppress redundant information in multimodal interaction. This leads to incomplete multimodal representation with noisy information. To address these issues, we propose Multi-grained Multimodal Interaction Network (MMIN) to provide a more complete view of multimodal representation. Coarse-grained Interaction Network (CIN) exploits the unique characteristics of different modalities at a coarse-grained level and adversarial learning is used to reduce redundancy. Fine-grained Interaction Network (FIN) employ sparse-attention mechanism to capture fine-grained interactions between multimodal sequences across distinct time steps and reduce irrelevant fine-grained multimodal interaction. Experimental results on two public datasets demonstrate the effectiveness of our model in multimodal sentiment analysis.
Lingyong Fang, Gongshen Liu, Ru Zhang 0002
ICASSP3
2024 An Images Regeneration Method for CG Anti-Forensics Based on Sensor Device Trace
abstract
Sensor device trace are the specific characteristics of natural images (NI), which are also the primary focus of many advanced forensic algorithms. However, most existing anti-forensic methods only consider the pixel-level differences between NI and computer-generated graphics (CG), while ignoring sensor device trace in NI. In this paper, we propose a novel anti-forensic method called SDTNet, which captures the trace of sensor device and integrates them with CG. Initially, the image content is suppressed by applying a high-pass filter, followed by the use of multiple convolution layers in an inverse manner. Subsequently, the sensor device trace is extracted using the Siamese network. Finally, this extracted trace is fused with CG at the feature level. The proposed method enables the generation of anti-forensic images that more accurately simulate the trace of NI device. Compared with the existing methods, the proposed method demonstrates superior performance in both visual quality and anti-forensic capabilities.
Yihong Lu, Ru Zhang 0002
ICME3
2024 Attack Behavior Extraction Based on Heterogeneous Threat Intelligence Graphs and Data Augmentation
abstract
Recently, cyber attacks have become increasingly complex and diverse. Threat intelligence is being systematically integrated to enhance threat detection capabilities. Tactics, Techniques, and Procedures (TTPs), as an advanced form of threat intelligence, provide detailed information about attackers’ methods of operation. This information holds significant value for understanding attacker behavior patterns and implementing proactive defense measures. In recent times, machine learning methods have been applied to identify TTPs. However, annotating TTPs requires experts with profound knowledge of offensive and defensive tactics. This results in a limited amount of available labeled data, limiting the effectiveness of machine learning models. To address this issue, we propose a method for extracting attack behavior based on heterogeneous threat intelligence graph and data augmentation. Firstly, a large amount of unlabeled threat intelligence text is utilized to train a masked language model based on Bidirectional Encoder Representation from Transformers (BERT). By predicting masked tokens within the annotated text, new text is generated that aligns with the textual characteristics of threat intelligence. Secondly, to effectively leverage contextual information in threat intelligence, we construct a Heterogeneous Threat Intelligence Graph (HTIG), modeling threat entities and their associated relationships. Simultaneously, to alleviate the issue of feature sparsity, a graph attention network is employed to learn embedded representations of nodes in the HTIG, facilitating the extraction of TTPs. Experimental results on the ATT&CK dataset indicate that our approach significantly enhances the efficiency of TTP identification.
Ru Zhang 0002
IJCNN2
2024 A Semantic Controllable Long Text Steganography Framework Based on LLM Prompt Engineering and Knowledge Graph
abstract
With ongoing advancements in natural language technology, text steganography has achieved notable progress. However, existing methods primarily concentrate on the probability distribution between words, often overlooking comprehensive control over text semantics. Particularly in the case of longer texts, these methods struggle to preserve coherence and contextual consistency, thereby increasing the risk of detection in practical applications. To effectively improve steganography security, we propose a semantic controllable long-text steganography framework based on prompt engineering and knowledge graph (KG) integration, obviating supplementary training. This framework leverages triplets from the KG and task descriptions to construct prompts, directing the large language model (LLM) to generate text that aligns with the triplet content. Subsequently, the model effectively embeds secret information by encoding the candidate pools established around the sampled target words. The experimental results demonstrate that our framework ensures the concealment of steganographic text while maintaining the relevance and consistency of the content as expected. Moreover, it can be flexibly adapted to various application scenarios, showcasing its potential and advantages in practical implementations.
Ru Zhang 0002
IEEE Signal Process. Lett.2
2024 Generative Image Steganography Based on Guidance Feature Distribution
abstract
Without modification, generative steganography is more secure than modification-based steganography. However, existing generative steganography methods still have limitations, such as low embedding capacity and poor quality. To solve these issues, a synthesis-based generative steganographic model is proposed in this article. In the image synthesis task, guidance features are utilized to synthesize images with specific styles and attributes. Due to the consistency of the guidance features before and after image synthesis, the features can be used as cover for steganography. The proposed model adopts the mean and standard deviation to quantify the distribution of guidance features, enabling the secret hiding within different trends of the feature distribution. By controlling the statistical dispersion of the embedded guidance features through the mean and standard deviation, the original feature distribution is preserved, and the synthesized image maintains good generation quality. The space of guidance features contains styles and attribute descriptions of various images, offering a large space for information hiding. According to the experimental results, compared with existing steganography, the proposed steganographic model achieves better quality and hidden capacity with strong robustness.
Youqiang Sun, Ru Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2023 A Robust Generative Image Steganography Method based on Guidance Features in Image Synthesis
abstract
The existing generative steganography methods have the limitations of the low capacity and poor stego quality. The target image which are served as guidance features and used to translate the image from the original to special one during the synthetic process. These guidance features are abundant, stable and do not contain the identity information which can be used as cover in steganography. This paper proposed a generative image steganography by using the guidance features in image synthesis, and a secret fusion algorithm is proposed to solve the problems of guidance features embedding and extraction errors. Due to the robustness of styles and attributes, the embedded guidance features can be extracted directly in receiver side from the synthesized image without code book or database. Compared with the existing generative steganography methods, the proposed method can achieve a higher security and quality while maintaining a larger embedding capacity.
Youqiang Sun, Ru Zhang 0002
ICME3
2023 Neural Linguistic Steganography with Controllable Security
abstract
Information hiding is an art and science with a long history and is widely used in covert communication. There are many ways to hide secret data in image, audio, and video. However, relatively few systems can hide information in text. Generative text steganography is a promising topic in natural language text infor-mation hiding. Previous generative text steganography methods use a fixed candidate pool generation rule, and they cannot effec-tively control the security of the generated text. The perceptual-imperceptibility and statistical-imperceptibility conflict effect also causes the poor quality of the steganographic text generated by previous generative text steganography methods. Moreover, pre-vious generative text steganography approaches barely discuss the robustness of steganographic text. This paper proposes a security controllable text steganography method that can generate natural-looking steganographic text with a statistical distribution that matches the natural language distribution. The proposed method combines the metrics of per-ceptual-imperceptibility and statistical-imperceptibility to calcu-late the combined distortion. It selects the tokens with the smallest combined distortion to construct a candidate pool at each time step. Moreover, the maximum combined distortion threshold is set when embedding secret messages to ensure controllable security. We conducted several experiments to evaluate the proposed model from the perspectives of embedding rate, perceptual-impercepti-bility, statistical-imperceptibility, and anti-attack ability. The ex-perimental results show that the proposed method can generate smooth and readable steganographic sentences with good re-sistance to steganalysis and high robustness.
Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju
IJCNN3
2023 Robust Secret Data Hiding for Transformer-based Neural Machine Translation
abstract
Hiding secret information in text is a research area of significant importance and a great challenge. In recent years, there have been huge developments and exciting advances in generation-based text information hiding techniques. Current generative text information hiding methods mainly establish correspondence between token and secret bits based on probability distributions given by language models. However, the semantic control of such methods is weak, and their robustness is not discussed. In this paper, we investigate an end-to-end generation-based text information hiding scheme. The proposed method uses a sequence-to-sequence model with adversarial training as a machine translation model. It converts the secret information into an embedding vector to be added to each position of the hidden state representation of the source language text, which in turn allows the model to automatically learn to produce translation results with the embedded secret information without using fixed rules. The semantics of the text with embedded secret messages obtained by translation can be controlled by the meaning of the source language text. Our experiments show that the proposed method can embed the secret message into the translation results with little loss of the translation quality and is robust to active attacks such as word deletion or synonym substitution.
Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju
IJCNN3
2023 Large capacity generative image steganography via image style transfer and feature-wise deep fusion
Youqiang Sun, Ru Zhang 0002
Appl. Intell.3
2023 V-A3tS: A rapid text steganalysis method based on position information and variable parameter multi-head self-attention controlled by length
Yihao Wang 0006, Ru Zhang 0002
J. Inf. Secur. Appl.2
2023 RLS-DTS: Reinforcement-Learning Linguistic Steganalysis in Distribution-Transformed Scenario
abstract
When the data undergo a distribution change, existing linguistic steganalysis often struggles to effectively capture the statistical characteristics of the transformed cover or stego, resulting in a drop in performances. To address this issue and fully use the information from the original data before the change, this letter proposes a reinforcement learning-based method for linguistic steganalysis. This method employs an agent (steganalyzer) to interact within an observation space, enabling adaptation to the characteristics of the transformed data and capturing steganalysis features. Specifically, we map the texts to the GloVe observation space and construct an agent comprising Actor module and Critic module to provide action, state, and other information. In the pre-training phase, agent trains and reinforces the Actor and Critic modules using the original data before the change. In the fine-tuning phase, agent optimizes these two modules to extract steganalysis feature through reinforcement training with instant reward in the transformed data. Experiments show that the proposed method exhibits better performances than the baseline in the transformed scenarios. Furthermore, this method offers a more autonomous training solution for linguistic steganalysis.
Yihao Wang 0006, Ru Zhang 0002
IEEE Signal Process. Lett.2
2023 Linguistic Steganalysis by Enhancing and Integrating Local and Global Features
abstract
With the improvement of steganography, the difference of statistical distribution caused by information hiding is getting smaller and smaller, increasing the difficulty of steganalysis. The existing steganalysis models value different features equally, while the differences in the importance of high-dimensional features are ignored. That is the influences of feature quality on model performance are not considered. To fill the gap, a novel linguistic steganalysis by enhancing and integrating local and global features is proposed in this letter. It extracts and integrates the features of two dimensions, namely local semantic features and global long-term dependencies, to construct a joint feature map. Then, to improve the quality of features, a group-wise enhancement mechanism is employed, which divides features into multiple groups, enhancing important features in each group while weakening the less important ones by generating an important coefficient for each sub-feature in each semantic group. Finally, the enhanced features are used to further extract high-quality text representation to help the classification module distinguish cover and stego texts more accurately. The experimental results show that the proposed model has better detection performance in multiple hidden scenarios. The ablation experiment manifests that using integrated and enhanced features to extract high-quality text representations can effectively improve the discrimination ability of the model.
Ru Zhang 0002
IEEE Signal Process. Lett.2
2022 Sense-aware BERT and Multi-task Fine-tuning for Multimodal Sentiment Analysis
abstract
Humans convey emotions through verbal and non-verbal signals when communicating face-to-face. Pre-trained language model such as BERT can be fine-tuned to improve the performance of various downstream tasks including sentiment analysis. However, most prior works about BERT fine-tuning contains only textual unimodal data and lacks information from sense organs, such as audio and visual signals, which are crucial for sentiment analysis. In this paper, we propose Sense-aware BERT (SenBERT) which allows sense information integrated with BERT during fine-tuning. In particular, we exploit multimodal multi-head attention to capture the interaction between unaligned multimodal data. Additionally, due to the variable information richness of different modalities, multimodal network may be dominated by some modalities during training process, so we propose unimodal sentiment analysis auxiliary tasks for multi-task learning which forces the model to focus on all modalities. We conduct experiments on CMU-MOSI and CMU-MOSEI datasets for multimodal sentiment analysis. The results show the superior performance of SenBERT on all the metrics over previous baselines.
Lingyong Fang, Gongshen Liu, Ru Zhang 0002
IJCNN3
2022 Two statistical traffic features for certain APT group identification
Wenxin Sun, Ru Zhang 0002, Xingjie Huang, Jin Pang
J. Inf. Secur. Appl.6
2022 Linguistic Steganalysis Merging Semantic and Statistical Features
abstract
With the rapid development of Natural Language Processing (NLP), more and more linguistic steganography methods have appeared in recent years, which may bring great challenges to the protection of cyberspace security. Due to the powerful feature extraction capabilities of Deep neural networks (DNN) to learn semantic features of large volumes of text, traditional steganalysis methods using manual features have gradually evolved into DNN-based methods. However, whether these DNN-based steganalysis methods can extract enough carrier features to achieve efficient steganalysis so that they can completely replace traditional methods based on handcrafted features remains an open question. To explore the answer, in this paper, we propose a new steganalysis method to integrate semantic and statistical features. We use BERT to extract semantic features and TF-IDF with AutoEncoder to obtain statistical features of the input text. Finally, we design a fusion mechanism to combine these two features. The experimental results show that due to the addition of statistical features, the proposed model can significantly improve the detection performance over current DNN-based linguistic steganalysis models.
Shengnan Guo 0008, Zhongliang Yang, Weike You, Ru Zhang 0002
IEEE Signal Process. Lett.5
2021 Carrier Robust Reversible Watermark Model Based on Image Block Chain Authentication
Ru Zhang 0002, Xianxu Li, Kaifeng Zhao 0003, Xue Cheng
ICIG (1)2
2020 Depth-Wise Separable Convolutions and Multi-Level Pooling for an Efficient Spatial CNN-Based Steganalysis
abstract
For steganalysis, many studies showed that convolutional neural network (CNN) has better performances than the two-part structure of traditional machine learning methods. Existing CNN architectures use various tricks to improve the performance of steganalysis, such as fixed convolutional kernels, the absolute value layer, data augmentation and the domain knowledge. However, some designing of the network structure were not extensively studied so far, such as different convolutions (inception, xception, etc.) and variety ways of pooling(spatial pyramid pooling, etc.). In this paper, we focus on designing a new CNN network structure to improve detection accuracy of spatial-domain steganography. First, we use$3\times 3$kernels instead of the traditional$5\times 5$kernels and optimize convolution kernels in the preprocessing layer. The smaller convolution kernels are used to reduce the number of parameters and model the features in a small local region. Next, we use separable convolutions to utilize channel correlation of the residuals, compress the image content and increase the signal-to-noise ratio (between the stego signal and the image signal). Then, we use spatial pyramid pooling (SPP) to aggregate the local features and enhance the representation ability of features by multi-level pooling. Finally, data augmentation is adopted to further improve network performance. The experimental results show that the proposed CNN structure is significantly better than other five methods such as SRM, Ye-Net, Xu-Net, Yedroudj-Net and SRNet, when it is used to detect three spatial algorithms such as WOW, S-UNIWARD and HILL with a wide variety of datasets and payloads.
Ru Zhang 0002, Gongshen Liu
IEEE Trans. Inf. Forensics Secur.1
2019 A high capacity reversible data hiding scheme for encrypted covers based on histogram shifting
Ru Zhang 0002, Chunjing Lu
J. Inf. Secur. Appl.1
2019 Invisible steganography via generative adversarial networks
abstract
Nowadays, there are plenty of works introducing convolutional neural networks (CNNs) to the steganalysis and exceeding conventional steganalysis algorithms. These works have shown the improving potential of deep learning in information hiding domain. There are also several works based on deep learning to do image steganography, but these works still have problems in capacity, invisibility and security. In this paper, we propose a novel CNN architecture named as ISGAN to conceal a secret gray image into a color cover image on the sender side and exactly extract the secret image out on the receiver side. There are three contributions in our work: (i) we improve the invisibility by hiding the secret image only in the Y channel of the cover image; (ii) We introduce the generative adversarial networks to strengthen the security by minimizing the divergence between the empirical probability distributions of stego images and natural images. (iii) In order to associate with the human visual system better, we construct a mixed loss function which is more appropriate for steganography to generate more realistic stego images and reveal out more better secret images. Experiment results show that ISGAN can achieve start-of-art performances on LFW, PASCAL-VOC12 and ImageNet datasets.
Ru Zhang 0002, Shiqi Dong
Multim. Tools Appl.1
2017 Dynamic data integrity auditing for secure outsourcing in the cloud
abstract
Summary Cloud servers provide cloud users with storage service and allow cloud users to access their files anytime. To guarantee security of the stored files, auditors need to periodically verify data block correctness. In the existing integrity verification schemes, there are few protocols to support the users' identity anonymity and the data block dynamic operation simultaneously. In this paper, we present an efficient and anonymous identity‐based integrity auditing protocol, which supports data dynamic operation and can be extended to support batch auditing in the multifile or multiuser setting. Our scheme not only resists forgery, replace, and replay attacks but also maintains users' anonymity, which is not discussed in other related techniques. The computation efficiency of auditor is improved a lot. Comparing with Zhang's efficient identity‐based public auditing scheme, our scheme is more suitable for actual application scenario with large‐scale storage system.
Jinxia Wei, Ru Zhang 0002, Xinxin Niu, Yuangang Yao
Concurr. Comput. Pract. Exp.2
2017 Constructing APT Attack Scenarios Based on Intrusion Kill Chain and Fuzzy Clustering
abstract
The APT attack on the Internet is becoming more serious, and most of intrusion detection systems can only generate alarms to some steps of APT attack and cannot identify the pattern of the APT attack. To detect APT attack, many researchers established attack models and then correlated IDS logs with the attack models. However, the accuracy of detection deeply relied on the integrity of models. In this paper, we propose a new method to construct APT attack scenarios by mining IDS security logs. These APT attack scenarios can be further used for the APT detection. First, we classify all the attack events by purpose of phase of the intrusion kill chain. Then we add the attack event dimension to fuzzy clustering, correlate IDS alarm logs with fuzzy clustering, and generate the attack sequence set. Next, we delete the bug attack sequences to clean the set. Finally, we use the nonaftereffect property of probability transfer matrix to construct attack scenarios by mining the attack sequence set. Experiments show that the proposed method can construct the APT attack scenarios by mining IDS alarm logs, and the constructed scenarios match the actual situation so that they can be used for APT attack detection.
Ru Zhang 0002, Yanyu Huo, Fangyu Weng
Secur. Commun. Networks1
2016 Detecting False Information of Social Network in Big Data
Ru Zhang 0002, Yuangang Yao
CollaborateCom4
2011 Research of Spatial Domain Image Digital Watermarking Payload
Jiafa Mao, Ru Zhang 0002, Xinxin Niu, Yixian Yang, Linna Zhou
EURASIP J. Inf. Secur.2