EDBT 2026 Demo / reviewers in the wild / expert
Yanzhen Ren
dblp:76/10143
· DBLP profile ↗
44ranked-venue papers
15as first author
27since 2021 · last 2026
0000-0003-0799-5082ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 13 since 2021Security and privacy · 17 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Backdoor samples detection based on perturbation discrepancy consistency in pre-trained language models
Zuquan Peng, Jianming Fu, Lixin Zou, Yanzhen Ren, Guojun Peng |
Neural Networks | 5 |
| 2026 | Decoupled Neural Audio Steganography for Adaptive Sender-Side Model UpdatesabstractNeural network–based steganography has garnered considerable attention for its strong security. However, existing approaches often suffer from excessive coupling between the embedding and extraction networks: the sender and receiver must employ paired models and maintain strict synchronization. Such synchronization not only complicates deployment but also introduces more severe potential risks of information leakage. To overcome this limitation, we propose a synchronization-free steganographic framework based on decoupled neural embedding networks, following the destruction–restoration principle. In our design, message embedding is realized through a destruction operation, while recovery is achieved using a neural network from the audio restoration domain. This decoupled architecture allows the sender to upgrade, replace, or randomize the embedding network—thus enabling dynamic model changes—without impairing the receiver’s ability to correctly extract the hidden message. As a result, synchronization-related vulnerabilities are fundamentally eliminated. Experimental results demonstrate that even under dynamic changes in the embedding network, the hidden information can still be reliably extracted, confirming both the effectiveness and enhanced security of the proposed approach. Qiyang Xiao, Yanzhen Ren, Lina Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | UnVC: Protecting Your Voiceprint by Generative Adversarial Speech
Zongkun Sun, Yihuan Huang, Yanzhen Ren, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency DomainabstractSpeaker recognition (SR) systems are particularly vulnerable to adversarial example (AE) attacks. To mitigate these attacks, AE detection systems are typically integrated into SR systems. To overcome the limitations of low detection accuracy, poor generalization, and high latency in existing schemes, this paper proposes FreqSense, an AE detection scheme based on frequency distribution features. FreqSense detects a variety of unknown AE attacks with low latency, and provides interpretability in its detection process. The basic idea of FreqSense is that AE typically introduce carefully designed noise in specific frequency bands that are associated with highly distinctive speaker identities. Therefore, leveraging the distributional variations in these frequency bands can effectively distinguish between AE and benign audio. FreqSense models frequency distribution features by integrating time-frequency transformation technology with a self-attention mechanism and employs a neural network-based classifier to distinguish between AE and benign audio. Experimental results show that FreqSense achieves an overall detection accuracy of 99.2%, surpassing state-of-the-art (SOTA) schemes by 27.2%. When confronting unknown AE attacks, FreqSense achieves a detection accuracy of 98.3% with a latency of just 0.0014 seconds. Yihuan Huang, Yanzhen Ren, Weiping Tu, Yuhong Yang 0001 |
ICASSP | 3 |
| 2025 | Improving Speech Enhancement by Cross- and Sub-band Processing with State Space ModelabstractRecently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additionally, when processing each time frame of the time-frequency representation, the SSM may forget certain high-frequency information of low energy, making the restoration of structure in the high-frequency bands challenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba). To assist the SSM in handling different sub-band features flexibly, we propose a band split block that splits the full-band into four sub-bands with different widths based on their information similarity. We then allocate independent weights to each sub-band, thereby reducing the inference burden on the SSM. Furthermore, to mitigate the forgetting of low-energy information in the high-frequency bands by the SSM, we introduce a spectrum restoration block that enhances the representation of the cross-band features from multiple perspectives. Experimental results on the DNS Challenge 2021 dataset demonstrate that CSMamba outperforms several state-of-the-art (SOTA) speech enhancement methods in three objective evaluation metrics with fewer parameters. Jizhen Li, Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren |
ICASSP | 6 |
| 2025 | Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio CodecsabstractExisting end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality during quantization, leading to a quantization error distribution that fails to reflect the actual importance of latent features. They are also sensitive to unusual data points (outliers) because they use Mean Squared Error (MSE) to measure quantization errors, which can increase quantization noise or spectral artifacts. To address these limitations, we propose AW-CEQCodec which integrates an Attention Weighting (AW) module and a Conditional Entropy-driven Quantization (CEQ) loss. The AW enhances key regions of latent features before quantization, enabling more accurate quantizing critical features and reducing their quantization errors. After quantization, it restores global details from dequantized features, improving overall reconstruction. Moreover, the CEQ minimizes the uncertainty between latent and quantized features, effectively reflecting the distortion introduced by the quantization module. Experimental results on the CodecSuperb-STL dataset demonstrate that our method consistently outperforms baseline approaches, achieving superior audio quality at bitrates as low as 0.5 kbps, confirming its effectiveness in minimizing distortion and preserving perceptual quality. The reconstruction audio samples can be find at https://huazhi1024.github.io/first-page. Weiping Tu, Yuhong Yang 0001, Xinmeng Xu, Yanzhen Ren |
ICASSP | 5 |
| 2025 | Lombard-VLD: Voice Liveness Detection Based on Human Auditory FeedbackabstractVoice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance. Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu |
SP | 3 |
| 2025 | Generalized Local Optimality for Video Steganalysis in Motion Vector DomainabstractVideo steganography that conceals secret data into motion vectors (MVs) is a popular covert communication technique. The local optimality of MVs is an intrinsic property in video coding, and any modifications to the MVs will inevitably destroy this optimality, making it a sensitive indicator of steganography. Thus the local optimality is commonly used to design features in video steganalysis. However, the local optimality in existing works is often estimated inaccurately or by using an unreasonable assumption, limiting its capability in steganalysis. In this article, we propose to estimate the local optimality in a more reasonable and comprehensive fashion, and generalize the local optimality in two aspects. First, we generalize the local optimality from a static estimation to a dynamic one by considering the variability of predicted motion vectors (PMVs). Second, we generalize the local optimality from MV domain to PMV domain by leveraging the statistical anomaly of PMVs. Based on the two generalizations that ensure a more accurate estimation of local optimality from more views, we construct new types of steganalytic features and also propose feature symmetrization rules to reduce feature dimension. Extensive experiments demonstrate the superiority of the proposed features, which achieve state-of-the-art accuracy and robustness under various conditions. Liming Zhai, Lina Wang 0001, Yanzhen Ren, Yang Liu 0003 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | APFT: Adaptive Phoneme Filter Template to Generate Anti-Compression Speech Adversarial Example in Real-TimeabstractAutomatic Speech Recognition (ASR) systems are widely used for speech censoring. Speech Adversarial Example (AE) offers a novel approach to protect speech privacy by forcing ASR to mistranscribe. However, existing speech AE faces two challenges in real-time voice communication scenarios, such as IP telephone, voice chat, or video conference, it cannot be generated in real-time, and its defensive capability is significantly reduced after the essential audio compression for network transmission. In this paper, we proposeAdaptive Phoneme Filter Template (APFT)to address these issues. The key features of APFT include: 1)Phoneme-level Templatesfor universal AE generation in real-time, 2)Filter, which eliminates redundant signals to improve compression robustness. 3)Adaptive Band Filtering, which limits the attack area from the frequency band without affecting the attack effectiveness and improves speech quality. The comprehensive experimental results show that APFT has four advantages: 1) Real-time Generation, with AE generation time below 1.1ms for 1s speech; 2) Compression Robustness, achieving a WER of 0.64 under AAC and Opus codecs; 3) Transferability, with an average WER of 0.72 across datasets and ASR systems; 4) Stealthiness, achieving a MOS of 4.07 for high-quality speech. In addition, the experiment on Telegram voice calls further proves the practical applicability of APFT. The demo of APFT can be obtained in https://yihuan-qaq.github.io/APFT.github.io/. Yihuan Huang, Yanzhen Ren, Zongkun Sun, Liming Zhai, Jingmin Wang, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Provably Secure and Robust Audio Steganography Under Multi-Format Low-Bitrate CompressionabstractWith the rapid advancement of audio generation models, research on audio steganography has entered a new phase of opportunity. Nevertheless, most existing generative steganographic approaches focus primarily on security while neglecting the compression and transcoding processes that are common in real-world communication. This oversight leads to two major issues: the introduction of verification mechanisms would violate its security proof assumptions, and quantization-based compression markedly reduces message extraction accuracy. In this work, we propose a robust audio steganography method that preserves provable security under various compression conditions. The security of our method relies exclusively on the latent space following a fixed distribution, which is independent of the embedded message. The proposed encoding–decoding scheme supports a tunable trade-off between capacity and robustness, allowing the sacrifice of partial capacity to reinforce robustness. Theoretical analysis shows that even with redundant error-checking codes, the latent distribution remains invariant after message embedding, thereby preserving both steganographic security and generation quality while ensuring practical applicability. Experimental results demonstrate that our method maintains message extraction accuracy under both MP3 and AAC compression and re-compression across bitrates of 160 kbps, 128 kbps, 64 kbps, 48 kbps, and even 32 kbps. Qiyang Xiao, Yanzhen Ren, Lina Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | VFD-Net: Vocoder Fingerprints Detection for Fake AudioabstractWith the rapid development of audio deepfake technology, the credibility and authenticity of public opinion is facing a formidable challenge. Since vocoder is the key component of audio deepfake and leaves distinctive fingerprint features, we propose VFD-Net (Vocoder Fingerprints Detection Net), a new vocoder architectures attribution scheme, which is based on patch-wise supervised contrastive learning (PCL) to capture the global consistency of the vocoder fingerprints and to improve the detection performance in cross-set testing and audio compression scenario. PCL brings patches belonging to the same vocoder class closer together in the representation space, while pushing patches from different vocoder classes further apart. Comparative experimental results show that the average accuracy of our proposed outperforms state-of-the-art 30%-45% under cross-set testing and AAC compression circumstances. Furthermore, our proposed approach achieves a 83.67% average accuracy in short-term fake audio detection within one second. It can be used to detect partially fake audio by analyzing the consistency of vocoder fingerprints. Junlong Deng, Yanzhen Ren, Hongcheng Zhu, Zongkun Sun |
ICASSP | 2 |
| 2024 | FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage FloodingabstractThis paper explores the detection of frame-wise instances of violence in both audio and visual modalities, where only clip-level labels are available. Previous works selected fixed value of frames for objective optimization to model frame-level features, and applied straightforward fusion strategy to aggregate audio and visual information. However, these two issues, namely Constant Frames Selection and Vulnerable Fusion, significantly impair the network’s detection performance. To address these issues, we present a novel framework called Frame-wise Cluster Contrast with Modality-stage Flooding (FCC-MF). Our contributions include: 1) We propose Frame-wise Cluster Contrast, which leverages unsupervised clustering for pseudo-labeling frames and triplet loss for contrastive learning to allow for dynamic frame-wise discrimination. 2) We propose Modality-stage Flooding, a two-stage flooding approach with the higher loss flooding level assigned to uni-modal features, which prevents over-memorization of redundant uni-modal data and promotes effective aggregation of multi-modal information. Our FCC-MF framework yields a promising average precision of 84.24% on the XD-Violence dataset, which performs favorably against previous SOTA methods. Extensive ablation studies exhibit that our FCC-MF framework produces finer frame-level violence discrimination ability and generalizable audio-visual fusion features. Jiaqing He, Yanzhen Ren, Liming Zhai, Wuyang Liu |
ICASSP | 2 |
| 2024 | Semantic Proximity Alignment: Towards Human Perception-Consistent Audio Tagging by Aligning with Label Text DescriptionabstractMost audio tagging models are trained with one-hot labels as supervised information. However, one-hot labels treat all sound events equally, ignoring the semantic hierarchy and proximity relationships between sound events. In contrast, the event descriptions contains richer information, describing the distance between different sound events with semantic proximity. In this paper, we explore the impact of training audio tagging models with auxiliary text descriptions of sound events. By aligning the audio features with the text features of corresponding labels, we inject the hierarchy and proximity information of sound events into audio encoders, improving the performance while making the prediction more consistent with human perception. We refer to this approach as Semantic Proximity Alignment (SPA). We use Ontology-aware mean Average Precision (OmAP) as the main evaluation metric for the models. OmAP reweights the false positives based on Audioset ontology distance and is more consistent with human perception compared to mAP. Experimental results show that the audio tagging models trained with SPA achieve higher OmAP compared to models trained with one-hot labels solely (+1.8 OmAP). Human evaluations also demonstrate that the predictions of SPA models are more consistent with human perception. Wuyang Liu, Yanzhen Ren |
ICASSP | 2 |
| 2024 | FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder
Rubing Shen, Yanzhen Ren, Zongkun Sun |
INTERSPEECH | 2 |
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 6 |
| 2024 | AFPM: A Low-Cost and Universal Adversarial Defense for Speaker Recognition SystemsabstractSpeaker recognition systems (SRSs) are commonly used for biometric identification. However, these systems are vulnerable to adversarial attacks. Several defenses have been proposed but they require high costs in terms of additional data and computational resources to ensure robustness. To address these issues, this paper proposes a low-cost input reconstruction defense method called adaptive F-ratio-based partial masking (AFPM), which utilizes a robust feature extraction process to guarantee high defensibility. The underlying distribution of non-robust features is explored and filtered out by partial masking (PM), which helps maintain a low defense construction cost. An F-ratio-based PM (FPM) defense strategy is proposed by integrating the F-ratio, which reflects the weight of each frequency band for distinguishing between speakers, to balance classification accuracy and defensiveness. AFPM, which introduces an adaptive threshold calculation algorithm to FPM, is proposed to achieve further improved defensiveness and flexibility. Comparative experimental results show that AFPM is low-cost, highly defensive and universal. The construction process of AFPM does not involve training and its implementation does not require the protected SRSs to be retrained, only fine-tuned. While maintaining the classification accuracy at 99.42%, the average defense capability of AFPM against five white-box adaptive attacks is 90.89%, which is 9.23% better than that of the low-cost input reconstruction defense method and 3.77% better than that of the high-cost Parallel WaveGAN (PWG) defense approach. Against grey- and black-box adaptive attacks, FAKEBOB and Kenansville, AFPM reaches maximum defense effects of 96.01% and 74.49%, respectively, surpassing PWG by 4.5% and 65.82%. Furthermore, AFPM is universal and capable of protecting various SRSs against different attack strengths. Zongkun Sun, Yanzhen Ren, Yihuan Huang, Wuyang Liu, Hongcheng Zhu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Super-Resolution Time-of-Flight Estimation for Ranging via Multi-Band SplicingabstractIndoor accurate ranging using WiFi signals is a challenge since it is affected by the signal bandwidth, indoor environment, and phase distortions introduced by the underlying hardware. In this paper, we improve channel impulse response (CIR) resolution by splicing channel state information (CSI) measurements from multiple non-contiguous bands. However, most previous efforts have directly used non-contiguous CSI measurements for parameter estimation, which leads to multi-band gain loss and ranging performance degradation. To solve this problem, we propose a novel time of flight (TOF) estimation scheme that includes two stages. In the first stage, we build an equivalent optimization model of the inverse non-uniform discrete Fourier transform (INDFT) to transform the CSI samples to the CIR. In the second stage, we construct a time domain super-resolution TOF estimation model and combine the estimation results of the first stage to obtain the distance information using a subspace projection method. Finally, we conduct the simulation using commercial-grade wireless ray propagation modeling software. Based on the simulation result, the ranging error of our approach is less than 10 cm 85 % with the spliced bandwidth of 240 MHz. Zengshan Tian, Yanzhen Ren, Ze Li 0003, Xingqing Cheng |
GLOBECOM | 2 |
| 2023 | Attention Mixup: An Accurate Mixup Scheme Based On Interpretable Attention Mechanism for Multi-Label Audio ClassificationabstractMixup proves to be an efficient data augmentation method on audio classification tasks. Original mixup scheme directly mixes the waveform of two random samples, which not only ignores the temporal distribution of the sound events but may also interfere with the original sound events in another sample. This paper proposes Attention MixUp (AMU), which only selects those segments that contain sound events for mixup, rather than simply mixing the entire sample. AMU utilizes the attention maps of pretrained audio classification Vision Transformer (ViT) to filter out the patches on the spectrogram that are useful for classification and then selects the regions for mixup according to three different strategies. Experimental results show a remarkable improvement (+1.9 mAP) on state-of-the-art Audioset classification methods with either CNN or ViT backbone. Further experiments show that AMU achieves the performance gain by improving the accuracy on short events (0.1s to 2s) by an average of 6.8% while keeping the accuracy on longer events. Wuyang Liu, Yanzhen Ren |
ICASSP | 2 |
| 2023 | Learning From Single-Expert Annotated Labels for Automatic Sleep StagingabstractExisting automatic sleep staging algorithms rely on accurately labeled data. However, due to the subjectivity of sleep experts, accurate labels must be obtained through joint labeling by multiple experts, which results in high time and labor costs. In this work, we treat labels mislabeled by a single expert as noisy labels and first propose SE-ASS, an automatic sleep staging learning framework based on single-expert annotated data. Since multiple models tend to produce inconsistent predictions for instances with incorrect labels during training, we use two networks with the same structure but different initializations and regularize them with a prediction consistency loss to prevent overfitting to noisy labels. Furthermore, we use a contrastive loss between models to enhance the exploration of feature representations without relying on potentially noisy labels. Our results on two publicly available datasets show that SE-ASS can effectively improve the performance of automatic sleep staging models trained on single-expert annotated datasets. Zhiheng Luan, Yanzhen Ren, Xiuping Yang, Weiping Tu, Yuhong Yang 0001 |
ICASSP | 2 |
| 2023 | ONEI: Unveiling Route and Phase of Breathing from Snoring Sounds
Baoai Han, Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren |
ICONIP (9) | 10 |
| 2023 | A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao 0007, Xiuping Yang, Weiping Tu, Weiyan Yi, Yuhong Yang 0001, Yanzhen Ren |
INTERSPEECH | 9 |
| 2023 | Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice ConversionabstractVoice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can prevent the misuse of VC, but unfortunately has not been extensively studied. In this paper, we are the first to investigate the speaker traceability for VC and propose a traceable VC framework named VoxTracer. Our VoxTracer is similar to but beyond the paradigm of audio watermarking. We first use unique speaker embedding to represent speaker identity. Then we design a VAE-Glow structure, in which the hiding process imperceptibly integrates the source speaker identity into the VC, and the tracing process accurately recovers the source speaker identity and even the source speech in spite of severe speech quality degradation. To address the speech mismatch between the hiding and tracing processes affected by different distortions, we also adopt an asynchronous training strategy to optimize the VAE-Glow models. The VoxTracer is versatile enough to be applied to arbitrary VC methods and popular audio coding standards. Extensive experiments demonstrate that the VoxTracer achieves not only high imperceptibility in hiding, but also nearly 100% tracing accuracy against various types of audio lossy compressions (AAC, MP3, Opus and SILK) with a broad range of bitrates (16 kbps - 128 kbps) even in a very short time duration (0.74s). Our source code is available at https://github.com/hongchengzhu/VoxTracer. Yanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun, Rubing Shen, Lina Wang 0001 |
ACM Multimedia | 1 |
| 2023 | A Universal Audio Steganalysis Scheme Based on Multiscale Spectrograms and DeepResNetabstractGiven the popularity of audio and video applications, compressed audio has become an important carrier of covert communication on the Internet. Many novel compressed audio steganography schemes have emerged that offer good hiding capability and aural concealment. In this paper, a universal steganalysis scheme called MultiSpecNet is proposed to detect steganography based on multiple embedding domains (advanced audio coding (AAC) and MPEG-1 Audio Layer III (MP3)), which are currently the two most popular compressed audio standards. The basic idea is that modification of either domain by a steganography scheme will change the time-frequency relationship of the audio signal after decoding. The proposed approach adopts the spectrogram as the input feature to extract richer information. DeepResNet is used to learn the distinguishing feature representations, and multiscale spectrograms are used to enrich the feature diversity. The experimental results show that the proposed scheme is effective at detecting different steganography schemes based on the AAC and MP3 embedding domains. The detection accuracy of the proposed scheme is higher than that achieved by other state-of-the-art schemes. Using spectrograms as the input, DeepResNet achieves better performance than schemes using quantized modified discrete cosine transform (MDCT) coefficients and mel-spectrogram, although the quantized MDCT coefficient is the parameter modified by the steganography schemes directly and mel-spectrogram is very popular and effective for general audio signal analysis. To the best of our knowledge, this work is the first audio steganalysis scheme that can detect multiple steganography schemes in both the MP3 and AAC embedding domains. The method proposed in this paper can be extended to audio steganalysis for other codecs or for audio forensics purposes. Yanzhen Ren, Dengkai Liu, Qiaochu Xiong, Jianming Fu, Lina Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | An Effective Imbalanced JPEG Steganalysis Scheme Based on Adaptive Cost-Sensitive Feature LearningabstractSteganalysis in real-world application often exhibit skewed sample distribution which poses a massive challenge for steganography detection. Conventional steganalysis algorithms are not effective when the training data distribution is imbalanced, and may fail in the scenario of imbalanced data distribution. To address imbalanced data distribution issue in steganalysis, a novel framework termed adaptive cost-sensitive feature learning via F-measure maximization is proposed, which is inspired by the fact that F-measure is a more suitable performance metric compared to accuracy for imbalanced data. We investigate the adaptive cost-sensitive strategy by generating and assigning different weight to each instance with misclassification occurrence. This scheme adaptively determines the weights according to the intra-class and inter-class costs from the imbalanced distribution. Features corresponding to the largest F-measure can be obtained by solving a series of adaptive cost-sensitive feature learning problems with optimization theory. In this way, the learned features are the most representative features between the cover and stego images so that imbalanced steganalysis can significantly alleviate. Extensive experiments on various imbalanced steganalysis tasks show the superiority of the proposed method over the state-of-the-art methods, and it can recognize more minority samples and has excellent classification performance. Ju Jia, Liming Zhai, Weixiang Ren, Lina Wang 0001, Yanzhen Ren |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | Recalibrated Bandpass Filtering On Temporal Waveform For Audio Spoof DetectionabstractDeepfake techniques mislead people’s cognition with high-quality fake videos and audios, and speech synthesis is an important tool to implement cognitive attacks, which mainly include Text-To-Speech (TTS) and Voice Conversion (VC). In this paper, we propose a method for audio spoof detection based on frequency band recalibration via sinc convolution and squeeze-excitation module, extracting features from the temporal waveform and emphasizing the frequency bands that are more useful on this task. Experimental results show that the proposed method outperforms other similar methods by 18.6% with an average EER of 7.23%, and achieve better generalizability on the detection of unseen spoofing methods, while the size of the model is reduced by 30.8%. Yanzhen Ren, Wuyang Liu, Dengkai Liu, Lina Wang 0001 |
ICIP | 1 |
| 2021 | Using Contrastive Learning to Improve the Performance of Steganalysis Schemes
Yanzhen Ren, Lina Wang 0001 |
IWDW | 1 |
| 2021 | Secure AAC steganography scheme based on multi-view statistical distortion (SofMvD)
Yanzhen Ren, Sen Cai, Lina Wang 0001 |
J. Inf. Secur. Appl. | 1 |
| 2020 | Self-Supervised Spoofing Audio Detection Scheme
Ziyue Jiang 0001, Hongcheng Zhu, Wenbing Ding, Yanzhen Ren |
INTERSPEECH | 5 |
| 2020 | Transferable heterogeneous feature subspace learning for JPEG mismatched steganalysis
Ju Jia, Liming Zhai, Weixiang Ren, Lina Wang 0001, Yanzhen Ren, Lefei Zhang |
Pattern Recognit. | 5 |
| 2020 | Universal Detection of Video Steganography in Multiple Domains Based on the Consistency of Motion VectorsabstractDigital video provides various types of embedding domains, which lead to a great diversity in video steganography. However, in the detection of video steganography, the existing video steganalytic features all specialize in a particular domain, and are hardly to detect the steganography in other embedding domains. In this paper, we propose a universal feature set which is capable of detecting the video steganography in multiple domains. Two popular embedding domains, i.e., partition mode (PM) domain and motion vector (MV) domain, are considered for steganalysis. The idea is based on the observation that the MVs of the sub-blocks in the same macroblock are usually different from each other, and they will tend to be consistent in values after the MV modifications or PM modifications. Thus the consistency of MVs can be used as an evidence for the steganographic embedding in two domains, and finally a 12-dimensional feature set is designed for universal detection. Extensive experiments are conducted to demonstrate the effectiveness of the proposed feature set. The results show that our feature set achieves superior universality and accuracy in both PM domain and MV domain, and even performs well in mismatched domains, where the detection model trained in one domain can directly be used to attack the steganography in another domain. Besides, the low complexity of the proposed feature set also indicates its advantage in real-time video steganalysis. Liming Zhai, Lina Wang 0001, Yanzhen Ren |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Multi-domain Embedding Strategies for Video Steganography by Combining Partition Modes and Motion VectorsabstractDigital video has various types of entities, which are utilized as embedding domains to hide messages in steganography. However, nearly all video steganography uses only one type of embedding domain, resulting in limited embedding capacity and potential security risks. In this paper, we firstly propose to embed in multi-domains for video steganography by combining partition modes (PMs) and motion vectors (MVs). The multi-domain embedding (MDE) aims to spread the modifications to different embedding domains for achieving higher undetectability. The key issue of MDE is the interactions of entities across domains. To this end, we design two MDE strategies, which hide data in PM domain and MV domain by sequential embedding and simultaneous embedding respectively. These two strategies can be applied to existing steganography within a distortion-minimization framework. Experiments show that the MDE strategies achieve a significant improvement in security performance against targeted steganalysis and fusion based steganalysis. Liming Zhai, Lina Wang 0001, Yanzhen Ren |
ICME | 3 |
| 2019 | Replay attack detection based on distortion by loudspeaker for voice authentication
Yanzhen Ren, Zhong Fang, Dengkai Liu, Chang Wen Chen |
Multim. Tools Appl. | 1 |
| 2019 | An AMR adaptive steganographic scheme based on the pitch delay of unvoiced speech
Yanzhen Ren, Dengkai Liu, Lina Wang 0001 |
Multim. Tools Appl. | 1 |
| 2019 | A posterior evaluation algorithm of steganalysis accuracy inspired by residual co-occurrence probability
Lina Wang 0001, Liming Zhai, Yanzhen Ren, Bo Du 0001 |
Pattern Recognit. | 4 |
| 2019 | A Secure AMR Fixed Codebook Steganographic Scheme Based on Pulse Distribution ModelabstractAdaptive multi-rate (AMR), a popular audio compression standard, is widely used in mobile communication and mobile Internet applications and has become a novel carrier for hiding information. To improve the statistical security, this paper presents a steganographic scheme in the AMR fixed codebook (FCB) domain based on the pulse distribution model (PDM-AFS), which is obtained from the distribution characteristics of the FCB value in the cover audio. The pulse positions in stego audio are controlled by message encoding and random masking to make the statistical distribution of the FCB parameters close to that of the cover audio. The experimental results show that the statistical security of the proposed scheme is better than that of the existing schemes. Furthermore, the hiding capacity is maintained compared with the existing schemes. The average hiding capacity can reach 2.06 kbps at an audio compression rate of 12.2 kbps, and the auditory concealment is good. To the best of our knowledge, this is the first secure AMR FCB steganographic scheme that improves the statistical security based on the distribution model of the cover audio. This scheme can be extended to other audio compression codecs under the principle of algebraic code excited linear prediction (ACELP), such as G.723.1 and G.729. Yanzhen Ren, Hanyi Yang, Hongxia Wu, Weiping Tu, Lina Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | An AMR adaptive steganography algorithm based on minimizing distortion
Yanzhen Ren, Hongxia Wu, Lina Wang 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Combined and Calibrated Features for Steganalysis of Motion Vector-Based Steganography in H.264/AVCabstractThis paper presents a novel feature set for steganalysis of motion vector-based steganography in H.264/AVC. First, the influence of steganographic embedding on the sum of absolute difference (SAD) and the motion vector difference (MVD) is analyzed, and then the statistical characteristics of these two aspects are combined to design features. In terms of SAD, the macroblock partition modes are used to measure the quantization distortion, and by using the optimality of SAD in neighborhood, the partition based neighborhood optimal probability features are extracted. In terms of MVD, it has been proved that MVD is better in feature construction than neighboring motion vector difference (NMVD) which has been widely used by traditional steganalyzers, and thus the inter and intra co-occurrence features are constructed based on the distribution of two components of neighboring MVDs and the distribution of two components of the same MVD. Finally, the combined features are enhanced by window optimal calibration, which utilizes the optimality of both SAD and MVD in a local window area. Experiments on various conditions demonstrate that the proposed scheme generally achieves a more accurate detection than current methods especially for videos encoded in variable block size and high quantization parameter values, and exhibits strong universality in applications. Liming Zhai, Lina Wang 0001, Yanzhen Ren |
IH&MMSec | 3 |
| 2017 | A Steganalysis Scheme for AAC Audio Based on MDCT Difference Between Intra and Inter Frame
Yanzhen Ren, Qiaochu Xiong, Lina Wang 0001 |
IWDW | 1 |
| 2017 | PFD - A Flexible Higher-Order Masking SchemeabstractBased on the idea of secret sharing, masking is one of the most popular countermeasure to prevent side channel attacks (SCAs). Despite the redundant time and resource consumption, the existing masking schemes have constant speed and resources, and thus unsuitable for different applications with variable demand for time or space. Motivated by the reconfiguration technology of programmable hardware and disjunctive normal form expression of any logic function, we define a random variable logic circuit to reach the same security for any-order masking schemes. During the encryption, we induce random sequences and utilize them as configuration sequences to generate variable logic circuits, whose results are independent from the original and divided into several shares. We call our new approach polynomial function division (PFD) masking. Furthermore, we analyze the effectiveness and proof the security of PFD in theory. Our experiments using PFD on the advanced encryption standard (AES) algorithm show that the space complexity is almost as small as an implementation of the original AES without any countermeasure. Moreover, due to the flexible structure of PFD, the cost-to-efficiency ratio of PFD is much lower than state-of-the art in software, and its flexibility is coin with the reconfigurable chip. Ming Tang 0002, Zhipeng Guo 0002, Annelie Heuser, Yanzhen Ren, Jean-Luc Danger |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | AMR Steganalysis Based on Second-Order Difference of Pitch DelayabstractThis paper presents a novel steganalysis scheme for the detection of adaptive multi-rate (AMR) audio steganography. AMR audio codec is used widely in mobile communication and mobile Internet system. Due to the modifiability of pitch delay, several AMR steganography schemes based on pitch delay modulation have emerged gradually and have high capacity and good imperceptibility. Based on the difference between cover and stego AMR audios on the continuity of adjacent pitch delay, this paper proposes the matrix of the second-order difference of pitch delay (MSDPD) steganalysis features by calculating the Markov transition probability MSDPD, and uses the calibration method to estimate the cover's features to get the calibrated MSDPD features to improve the accuracy of the scheme. Support vector machine is used as the steganalyzer to test the performance of our proposed scheme. The experimental results show that the correct detection rate of our proposed method is more than 85% when the embedding bit rate is 30% or above, and can reach above 85% for cover audios. The results of contrast experiment show that the performance of our proposed method is better than the existing method, especially on low embedding rate. The method can be extended to other CELP codec, such as G.723.1 and G.729. Yanzhen Ren, Lina Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Steganalysis of AAC using calibrated Markov model of adjacent codebookabstractAAC(Advanced Audio Coding) is the most popular audio compression standard and used widely in recent years. The steganography schemes of AAC emerged gradually. This paper presents a novel steganalysis method to attack the steganography of Huffman codebook, which hide information by modifying the codebook of each scale factor band(SFB), and have good imperceptivity and security. Based on the correlation of neighboring SFBs' codebook, the paper proposes to extract the Markov transition probability of adjacent SFBs' codebook as steganalysis feature, and adopt calibration to improve the accuracy. Extensive experiments demonstrate the effectiveness of the proposed methods. To the best of our knowledge, this piece of work is the first one to detect AAC steganography of Huffman codebook. Yanzhen Ren, Qiaochu Xiong, Lina Wang 0001 |
ICASSP | 1 |
| 2016 | Detection of double MP3 compression Based on Difference of Calibration Histogram
Yanzhen Ren, Mengdi Fan, Dengpan Ye, Lina Wang 0001 |
Multim. Tools Appl. | 1 |
| 2015 | AMR Steganalysis Based on the Probability of Same Pulse PositionabstractThis paper presents a method for detection of adaptive multirate (AMR) audio steganography. AMR audio codec is an audio data compression scheme optimized for speech coding, and widely used in some mobile telecommunications system. The AMR audio steganography schemes are emerging recently and they embed secret messages by modifying the nonzero pulse positions which are determined by fixed codebook search in AMR compression procedure. Those methods have high embedding capacity and good imperceptivity. We have observed that those steganography schemes will cause the probability of same pulse positions in the same track increasing. Based on this phenomenon, this paper presents a set of steganalysis features of the probability of same pulse position. The support vector machine is applied to the proposed features and used as the steganalyzer. The performance of the scheme is tested on a database containing ~140714 audios. Experimental results show that the correct detection rate of our proposed method is 90% when the embedding bit rate is 30% or above, and can reach above 85% for cover audios. Yanzhen Ren, Tingting Cai, Ming Tang 0002, Lina Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Video steganalysis based on subtractive probability of optimal matching featureabstractThis paper presents a novel motion vector (MV) steganalysis method. MV-based steganographic methods exploite the variability of MV to embed messages by modifying MV slightly. However, we have noticed that the modified MVs after steganography cannot follow the optimal matching rule which is the target of motion estimation. It means that steganographic methods conflict with the basic principle of video compression. Aiming at this difference, we proposed a steganalysis feature based on Subtractive Probability of Optimal Matching(SPOM), which statistics the MV's Probability of the Optimal matching (POM) around its neighbors, and extract the classification feature by subtracting the POM of the test video and its recompressed video. Experiment results show that the proposed feature is sensitive to MV-based steganography methods, and outperforms the other methods, especially for high temporal activity video. Yanzhen Ren, Liming Zhai, Lina Wang 0001 |
IH&MMSec | 1 |