Liming Zhai

dblp:149/3381 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-3229-056XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Security and privacy · 6 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Adversarial rain attack and defensive deraining for DNN perception
Liming Zhai, Qing Guo 0003, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Shengchao Qin, Yang Liu 0003
Neural Networks1
2025 Generalized Local Optimality for Video Steganalysis in Motion Vector Domain
abstract
Video steganography that conceals secret data into motion vectors (MVs) is a popular covert communication technique. The local optimality of MVs is an intrinsic property in video coding, and any modifications to the MVs will inevitably destroy this optimality, making it a sensitive indicator of steganography. Thus the local optimality is commonly used to design features in video steganalysis. However, the local optimality in existing works is often estimated inaccurately or by using an unreasonable assumption, limiting its capability in steganalysis. In this article, we propose to estimate the local optimality in a more reasonable and comprehensive fashion, and generalize the local optimality in two aspects. First, we generalize the local optimality from a static estimation to a dynamic one by considering the variability of predicted motion vectors (PMVs). Second, we generalize the local optimality from MV domain to PMV domain by leveraging the statistical anomaly of PMVs. Based on the two generalizations that ensure a more accurate estimation of local optimality from more views, we construct new types of steganalytic features and also propose feature symmetrization rules to reduce feature dimension. Extensive experiments demonstrate the superiority of the proposed features, which achieve state-of-the-art accuracy and robustness under various conditions.
Liming Zhai, Lina Wang 0001, Yanzhen Ren, Yang Liu 0003
IEEE Trans. Dependable Secur. Comput.1
2025 APFT: Adaptive Phoneme Filter Template to Generate Anti-Compression Speech Adversarial Example in Real-Time
abstract
Automatic Speech Recognition (ASR) systems are widely used for speech censoring. Speech Adversarial Example (AE) offers a novel approach to protect speech privacy by forcing ASR to mistranscribe. However, existing speech AE faces two challenges in real-time voice communication scenarios, such as IP telephone, voice chat, or video conference, it cannot be generated in real-time, and its defensive capability is significantly reduced after the essential audio compression for network transmission. In this paper, we proposeAdaptive Phoneme Filter Template (APFT)to address these issues. The key features of APFT include: 1)Phoneme-level Templatesfor universal AE generation in real-time, 2)Filter, which eliminates redundant signals to improve compression robustness. 3)Adaptive Band Filtering, which limits the attack area from the frequency band without affecting the attack effectiveness and improves speech quality. The comprehensive experimental results show that APFT has four advantages: 1) Real-time Generation, with AE generation time below 1.1ms for 1s speech; 2) Compression Robustness, achieving a WER of 0.64 under AAC and Opus codecs; 3) Transferability, with an average WER of 0.72 across datasets and ASR systems; 4) Stealthiness, achieving a MOS of 4.07 for high-quality speech. In addition, the experiment on Telegram voice calls further proves the practical applicability of APFT. The demo of APFT can be obtained in https://yihuan-qaq.github.io/APFT.github.io/.
Yihuan Huang, Yanzhen Ren, Zongkun Sun, Liming Zhai, Jingmin Wang, Wuyang Liu
IEEE Trans. Inf. Forensics Secur.4
2024 FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage Flooding
abstract
This paper explores the detection of frame-wise instances of violence in both audio and visual modalities, where only clip-level labels are available. Previous works selected fixed value of frames for objective optimization to model frame-level features, and applied straightforward fusion strategy to aggregate audio and visual information. However, these two issues, namely Constant Frames Selection and Vulnerable Fusion, significantly impair the network’s detection performance. To address these issues, we present a novel framework called Frame-wise Cluster Contrast with Modality-stage Flooding (FCC-MF). Our contributions include: 1) We propose Frame-wise Cluster Contrast, which leverages unsupervised clustering for pseudo-labeling frames and triplet loss for contrastive learning to allow for dynamic frame-wise discrimination. 2) We propose Modality-stage Flooding, a two-stage flooding approach with the higher loss flooding level assigned to uni-modal features, which prevents over-memorization of redundant uni-modal data and promotes effective aggregation of multi-modal information. Our FCC-MF framework yields a promising average precision of 84.24% on the XD-Violence dataset, which performs favorably against previous SOTA methods. Extensive ablation studies exhibit that our FCC-MF framework produces finer frame-level violence discrimination ability and generalizable audio-visual fusion features.
Jiaqing He, Yanzhen Ren, Liming Zhai, Wuyang Liu
ICASSP3
2024 VBH-GNN: Variational Bayesian Heterogeneous Graph Neural Networks for Cross-subject Emotion Recognition
abstract
The research on human emotion under electroencephalogram (EEG) is an emerging field in which cross-subject emotion recognition (ER) is a promising but challenging task. Many approaches attempt to find emotionally relevant domain-invariant features using domain adaptation (DA) to improve the accuracy of cross-subject ER. However, two problems still exist with these methods. First, only single-modal data (EEG) is utilized, ignoring the complementarity between multi-modal physiological signals. Second, these methods aim to completely match the signal features between different domains, which is difficult due to the extreme individual differences of EEG. To solve these problems, we introduce the complementarity of multi-modal physiological signals and propose a new method for cross-subject ER that does not align the distribution of signal features but rather the distribution of spatio-temporal relationships between features. We design a Variational Bayesian Heterogeneous Graph Neural Network (VBH-GNN) with Relationship Distribution Adaptation (RDA). The RDA first aligns the domains by expressing the model space as a posterior distribution of a heterogeneous graph for a given source domain. Then, the RDA transforms the heterogeneous graph into an emotion-specific graph to further align the domains for the downstream ER task. Extensive experiments on two public datasets, DEAP and Dreamer, show that our VBH-GNN outperforms state-of-the-art methods in cross-subject scenarios.
Xinliang Zhou, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003
ICLR4
2024 VSGT: Variational Spatial and Gaussian Temporal Graph Models for EEG-based Emotion Recognition
Xinliang Zhou, Jiaping Xiao, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003
IJCAI5
2023 Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice Conversion
abstract
Voice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can prevent the misuse of VC, but unfortunately has not been extensively studied. In this paper, we are the first to investigate the speaker traceability for VC and propose a traceable VC framework named VoxTracer. Our VoxTracer is similar to but beyond the paradigm of audio watermarking. We first use unique speaker embedding to represent speaker identity. Then we design a VAE-Glow structure, in which the hiding process imperceptibly integrates the source speaker identity into the VC, and the tracing process accurately recovers the source speaker identity and even the source speech in spite of severe speech quality degradation. To address the speech mismatch between the hiding and tracing processes affected by different distortions, we also adopt an asynchronous training strategy to optimize the VAE-Glow models. The VoxTracer is versatile enough to be applied to arbitrary VC methods and popular audio coding standards. Extensive experiments demonstrate that the VoxTracer achieves not only high imperceptibility in hiding, but also nearly 100% tracing accuracy against various types of audio lossy compressions (AAC, MP3, Opus and SILK) with a broad range of bitrates (16 kbps - 128 kbps) even in a very short time duration (0.74s). Our source code is available at https://github.com/hongchengzhu/VoxTracer.
Yanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun, Rubing Shen, Lina Wang 0001
ACM Multimedia3
2022 A3GAN: Attribute-Aware Anonymization Networks for Face De-identification
abstract
Face de-identification (De-ID) removes face identity information in face images to avoid personal privacy leakage. Existing face De-ID breaks the raw identity by cutting out the face regions and recovering the corrupted regions via deep generators, which inevitably affect the generation quality and cannot control generation results according to subsequent intelligent tasks (eg., facial expression recognition). In this work, for the first attempt, we think the face De-ID from the perspective of attribute editing and propose an attribute-aware anonymization network (A3GAN) by formulating face De-ID as a joint task of semantic suppression and controllable attribute injection. Intuitively, the semantic suppression removes the identity-sensitive information in embeddings while the controllable attribute injection automatically edits the raw face along the attributes that benefit De-ID. To this end, we first design a multi-scale semantic suppression network with a novel suppressive convolution unit (SCU), which can remove the face identity along multi-level deep features progressively. Then, we propose an attribute-aware injective network (AINet) that can generate De-ID-sensitive attributes in a controllable way (i.e., specifying which attributes can be changed and which cannot) and inject them into the latent code of the raw face. Moreover, to enable effective training, we design a new anonymization loss to let the injected attributes shift far away from the original ones. We perform comprehensive experiments on four datasets covering four different intelligent tasks including face verification, face detection, facial expression recognition, and fatigue detection, all of which demonstrate the superiority of our face De-ID over state-of-the-art methods.
Liming Zhai, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yi Estelle Wang, Yang Liu 0003
ACM Multimedia1
2022 Progressive selection-channel networks for image steganalysis
abstract
Steganalysis is a detection technology against steganography that embeds secret data into digital media carriers. The selection channel, which indicates the embedding details of steganography, is well recognized in boosting the detection performance of image steganalysis. However, nearly all the selection channels are constructed in a hand-crafted manner, even when they are incorporated into end-to-end deep steganalytic networks, for which the embedding rate and steganographic algorithms also need to be predetermined. Such prior knowledge is usually assumed completely known in existing literature, which is obviously unreasonable and impractical. To address this issue, we propose to automatically learn the selection channels for deep learning-based image steganalysis in a progressive way. Specifically, we divide the image steganalysis task into two phases: selection channel estimation and steganalytic detection. For the first phase, we design a multistage progressive network, which enables the learning of selection channels in a coarse-to-fine fashion. For the second phase, we integrate the learned selection channels into the multilayers of the steganalytic network, allowing full exploitation of selection channels for accurate detection. Extensive experiments demonstrate that the proposed method can learn the selection channels rapidly and precisely, and also significantly improve the detection accuracy of the existing state-of-the-art steganographic network without any prior knowledge.
Tian Wu 0004, Lina Wang 0001, Liming Zhai, Canming Fang, Mingcheng Zhang
Int. J. Intell. Syst.3
2022 An Effective Imbalanced JPEG Steganalysis Scheme Based on Adaptive Cost-Sensitive Feature Learning
abstract
Steganalysis in real-world application often exhibit skewed sample distribution which poses a massive challenge for steganography detection. Conventional steganalysis algorithms are not effective when the training data distribution is imbalanced, and may fail in the scenario of imbalanced data distribution. To address imbalanced data distribution issue in steganalysis, a novel framework termed adaptive cost-sensitive feature learning via F-measure maximization is proposed, which is inspired by the fact that F-measure is a more suitable performance metric compared to accuracy for imbalanced data. We investigate the adaptive cost-sensitive strategy by generating and assigning different weight to each instance with misclassification occurrence. This scheme adaptively determines the weights according to the intra-class and inter-class costs from the imbalanced distribution. Features corresponding to the largest F-measure can be obtained by solving a series of adaptive cost-sensitive feature learning problems with optimization theory. In this way, the learned features are the most representative features between the cover and stego images so that imbalanced steganalysis can significantly alleviate. Extensive experiments on various imbalanced steganalysis tasks show the superiority of the proposed method over the state-of-the-art methods, and it can recognize more minority samples and has excellent classification performance.
Ju Jia, Liming Zhai, Weixiang Ren, Lina Wang 0001, Yanzhen Ren
IEEE Trans. Knowl. Data Eng.2
2020 Learning selection channels for image steganalysis in spatial domain
Weixiang Ren, Liming Zhai, Ju Jia, Lina Wang 0001, Lefei Zhang
Neurocomputing2
2020 Transferable heterogeneous feature subspace learning for JPEG mismatched steganalysis
Ju Jia, Liming Zhai, Weixiang Ren, Lina Wang 0001, Yanzhen Ren, Lefei Zhang
Pattern Recognit.2
2020 Universal Detection of Video Steganography in Multiple Domains Based on the Consistency of Motion Vectors
abstract
Digital video provides various types of embedding domains, which lead to a great diversity in video steganography. However, in the detection of video steganography, the existing video steganalytic features all specialize in a particular domain, and are hardly to detect the steganography in other embedding domains. In this paper, we propose a universal feature set which is capable of detecting the video steganography in multiple domains. Two popular embedding domains, i.e., partition mode (PM) domain and motion vector (MV) domain, are considered for steganalysis. The idea is based on the observation that the MVs of the sub-blocks in the same macroblock are usually different from each other, and they will tend to be consistent in values after the MV modifications or PM modifications. Thus the consistency of MVs can be used as an evidence for the steganographic embedding in two domains, and finally a 12-dimensional feature set is designed for universal detection. Extensive experiments are conducted to demonstrate the effectiveness of the proposed feature set. The results show that our feature set achieves superior universality and accuracy in both PM domain and MV domain, and even performs well in mismatched domains, where the detection model trained in one domain can directly be used to attack the steganography in another domain. Besides, the low complexity of the proposed feature set also indicates its advantage in real-time video steganalysis.
Liming Zhai, Lina Wang 0001, Yanzhen Ren
IEEE Trans. Inf. Forensics Secur.1
2019 Multi-domain Embedding Strategies for Video Steganography by Combining Partition Modes and Motion Vectors
abstract
Digital video has various types of entities, which are utilized as embedding domains to hide messages in steganography. However, nearly all video steganography uses only one type of embedding domain, resulting in limited embedding capacity and potential security risks. In this paper, we firstly propose to embed in multi-domains for video steganography by combining partition modes (PMs) and motion vectors (MVs). The multi-domain embedding (MDE) aims to spread the modifications to different embedding domains for achieving higher undetectability. The key issue of MDE is the interactions of entities across domains. To this end, we design two MDE strategies, which hide data in PM domain and MV domain by sequential embedding and simultaneous embedding respectively. These two strategies can be applied to existing steganography within a distortion-minimization framework. Experiments show that the MDE strategies achieve a significant improvement in security performance against targeted steganalysis and fusion based steganalysis.
Liming Zhai, Lina Wang 0001, Yanzhen Ren
ICME1
2019 Designing Non-additive Distortions for JPEG Steganography Based on Blocking Artifacts Reduction
Yubo Lu, Liming Zhai, Lina Wang 0001
IWDW2
2019 A posterior evaluation algorithm of steganalysis accuracy inspired by residual co-occurrence probability
Lina Wang 0001, Liming Zhai, Yanzhen Ren, Bo Du 0001
Pattern Recognit.3
2017 Combined and Calibrated Features for Steganalysis of Motion Vector-Based Steganography in H.264/AVC
abstract
This paper presents a novel feature set for steganalysis of motion vector-based steganography in H.264/AVC. First, the influence of steganographic embedding on the sum of absolute difference (SAD) and the motion vector difference (MVD) is analyzed, and then the statistical characteristics of these two aspects are combined to design features. In terms of SAD, the macroblock partition modes are used to measure the quantization distortion, and by using the optimality of SAD in neighborhood, the partition based neighborhood optimal probability features are extracted. In terms of MVD, it has been proved that MVD is better in feature construction than neighboring motion vector difference (NMVD) which has been widely used by traditional steganalyzers, and thus the inter and intra co-occurrence features are constructed based on the distribution of two components of neighboring MVDs and the distribution of two components of the same MVD. Finally, the combined features are enhanced by window optimal calibration, which utilizes the optimality of both SAD and MVD in a local window area. Experiments on various conditions demonstrate that the proposed scheme generally achieves a more accurate detection than current methods especially for videos encoded in variable block size and high quantization parameter values, and exhibits strong universality in applications.
Liming Zhai, Lina Wang 0001, Yanzhen Ren
IH&MMSec1
2014 Video steganalysis based on subtractive probability of optimal matching feature
abstract
This paper presents a novel motion vector (MV) steganalysis method. MV-based steganographic methods exploite the variability of MV to embed messages by modifying MV slightly. However, we have noticed that the modified MVs after steganography cannot follow the optimal matching rule which is the target of motion estimation. It means that steganographic methods conflict with the basic principle of video compression. Aiming at this difference, we proposed a steganalysis feature based on Subtractive Probability of Optimal Matching(SPOM), which statistics the MV's Probability of the Optimal matching (POM) around its neighbors, and extract the classification feature by subtracting the POM of the test video and its recompressed video. Experiment results show that the proposed feature is sensitive to MV-based steganography methods, and outperforms the other methods, especially for high temporal activity video.
Yanzhen Ren, Liming Zhai, Lina Wang 0001
IH&MMSec2