Hongcheng Zhu

dblp:277/3794 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Planning-Operation Coordinated Mitigation for Load Redistribution Attacks in Optimal Power Flow With Phase Shifting Transformers
abstract
In this article, we propose a planning-operation coordinated mitigation scheme for load redistribution (LR) attacks to overcome the deficiencies of separately designed phase shifting transformer-based mitigation strategies. Specifically, the interactions amongst the defender, attacker, and system are formulated as a trilevel optimization, where the deployment of defense devices and phase shift angles can be optimized according to possible operation state. Based on the proposed load similarity metric, a clustering-based approximate solution is designed to reduce the computational complexity caused by the integration of planning and operation stages. Simulation results on the IEEE 14-bus and 30-bus test systems verify the performance of the proposed mitigation scheme and the clustering-based approximate solution method.
Hongcheng Zhu, Chensheng Liu, Ming Yang 0023, Xin Wang 0044, Ruilong Deng, Yang Tang 0001, Chengnian Long
IEEE Trans. Ind. Informatics1
2025 Lombard-VLD: Voice Liveness Detection Based on Human Auditory Feedback
abstract
Voice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance.
Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu
SP1
2024 VFD-Net: Vocoder Fingerprints Detection for Fake Audio
abstract
With the rapid development of audio deepfake technology, the credibility and authenticity of public opinion is facing a formidable challenge. Since vocoder is the key component of audio deepfake and leaves distinctive fingerprint features, we propose VFD-Net (Vocoder Fingerprints Detection Net), a new vocoder architectures attribution scheme, which is based on patch-wise supervised contrastive learning (PCL) to capture the global consistency of the vocoder fingerprints and to improve the detection performance in cross-set testing and audio compression scenario. PCL brings patches belonging to the same vocoder class closer together in the representation space, while pushing patches from different vocoder classes further apart. Comparative experimental results show that the average accuracy of our proposed outperforms state-of-the-art 30%-45% under cross-set testing and AAC compression circumstances. Furthermore, our proposed approach achieves a 83.67% average accuracy in short-term fake audio detection within one second. It can be used to detect partially fake audio by analyzing the consistency of vocoder fingerprints.
Junlong Deng, Yanzhen Ren, Hongcheng Zhu, Zongkun Sun
ICASSP4
2024 Parameter-Estimate-First False Data Injection Attacks in AC State Estimation Deployed With Moving Target Defense
abstract
Enabled by the widely deployed distributed flexible alternating current transmission system (D-FACTS) devices in practical systems, moving target defense (MTD) has been considered as an effective way to detect stealthy false data injection (FDI) attacks by actively changing branch parameters. However, existing MTD methods heavily depend on the assumption that opponents can not timely obtain the newly changed branch parameters. In this paper, a parameter-estimate-first FDI (PEF-FDI) attack is proposed to reveal vulnerabilities of MTD methods in AC state estimation, which can bypass bad data detectors in the existence of MTD. Specifically, a PEF-FDI attack model is proposed to timely construct attack vector and stealthily misguide the results of alternating current (AC) state estimation in the presence of MTD. Requirements of constructing PEF-FDI attacks on eavesdropped measurements are deduced to reveal the limitation on capability of attackers. Simulations in the IEEE 118-bus system verify the performance of the proposed PEF-FDI attacks.
Chensheng Liu, Yuanqi Li, Hongcheng Zhu, Yang Tang 0001, Wenli Du
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 AFPM: A Low-Cost and Universal Adversarial Defense for Speaker Recognition Systems
abstract
Speaker recognition systems (SRSs) are commonly used for biometric identification. However, these systems are vulnerable to adversarial attacks. Several defenses have been proposed but they require high costs in terms of additional data and computational resources to ensure robustness. To address these issues, this paper proposes a low-cost input reconstruction defense method called adaptive F-ratio-based partial masking (AFPM), which utilizes a robust feature extraction process to guarantee high defensibility. The underlying distribution of non-robust features is explored and filtered out by partial masking (PM), which helps maintain a low defense construction cost. An F-ratio-based PM (FPM) defense strategy is proposed by integrating the F-ratio, which reflects the weight of each frequency band for distinguishing between speakers, to balance classification accuracy and defensiveness. AFPM, which introduces an adaptive threshold calculation algorithm to FPM, is proposed to achieve further improved defensiveness and flexibility. Comparative experimental results show that AFPM is low-cost, highly defensive and universal. The construction process of AFPM does not involve training and its implementation does not require the protected SRSs to be retrained, only fine-tuned. While maintaining the classification accuracy at 99.42%, the average defense capability of AFPM against five white-box adaptive attacks is 90.89%, which is 9.23% better than that of the low-cost input reconstruction defense method and 3.77% better than that of the high-cost Parallel WaveGAN (PWG) defense approach. Against grey- and black-box adaptive attacks, FAKEBOB and Kenansville, AFPM reaches maximum defense effects of 96.01% and 74.49%, respectively, surpassing PWG by 4.5% and 65.82%. Furthermore, AFPM is universal and capable of protecting various SRSs against different attack strengths.
Zongkun Sun, Yanzhen Ren, Yihuan Huang, Wuyang Liu, Hongcheng Zhu
IEEE Trans. Inf. Forensics Secur.5
2023 Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice Conversion
abstract
Voice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can prevent the misuse of VC, but unfortunately has not been extensively studied. In this paper, we are the first to investigate the speaker traceability for VC and propose a traceable VC framework named VoxTracer. Our VoxTracer is similar to but beyond the paradigm of audio watermarking. We first use unique speaker embedding to represent speaker identity. Then we design a VAE-Glow structure, in which the hiding process imperceptibly integrates the source speaker identity into the VC, and the tracing process accurately recovers the source speaker identity and even the source speech in spite of severe speech quality degradation. To address the speech mismatch between the hiding and tracing processes affected by different distortions, we also adopt an asynchronous training strategy to optimize the VAE-Glow models. The VoxTracer is versatile enough to be applied to arbitrary VC methods and popular audio coding standards. Extensive experiments demonstrate that the VoxTracer achieves not only high imperceptibility in hiding, but also nearly 100% tracing accuracy against various types of audio lossy compressions (AAC, MP3, Opus and SILK) with a broad range of bitrates (16 kbps - 128 kbps) even in a very short time duration (0.74s). Our source code is available at https://github.com/hongchengzhu/VoxTracer.
Yanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun, Rubing Shen, Lina Wang 0001
ACM Multimedia2
2020 Self-Supervised Spoofing Audio Detection Scheme
Ziyue Jiang 0001, Hongcheng Zhu, Wenbing Ding, Yanzhen Ren
INTERSPEECH2