Kesong Wu

dblp:193/1711 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Real-Time Forgery Detection via Dynamic Frequency-Domain Selection and Phoneme Alignment
abstract
The proliferation of deep learning-based video forgery, known as DeepFakes, poses a significant threat to information integrity and digital security, particularly in real-time communication scenarios. Existing detection methods often struggle with three critical challenges: performance degradation due to video compression, a limited capacity to detect subtle audio-visual semantic inconsistencies, and excessive computational latency for real-time deployment. This paper introduces a lightweight, multi-modal framework designed to address these limitations. Our method employs a dual-stream architecture for parallel analysis of visual and auditory signals. Key contributions include: a novel dynamic frequency band selection module that adaptively isolates and enhances faint forgery artifacts from heavily compressed video streams; a phoneme-aligned cross-modal verification system that precisely quantifies lip-audio temporal desynchronization to detect semantic mismatches; and a lightweight hybrid architecture integrating ResNet3D and Mamba for efficient spatiotemporal modeling. Validated on benchmarks like CelebDF and FaceForensics++, our method achieves a state-of-the-art trade-off, delivering accuracy competitive with leading methods at a fraction of the computational cost, thus enabling practical real-time deployment.
Xiaoyu Geng, Kesong Wu
TrustCom4
2025 Hyperchaos and HVS-Adaptive Video Watermarking Embedding
abstract
Watermarking for video playback authorization faces the classic challenge of balancing imperceptibility and robustness, while also maintaining resilience against statistical attacks. This paper introduces a novel scheme that integrates hyperchaos, a human visual system (HVS) model, and asymmetric modulation to address these challenges. First, a four-dimensional hyperchaotic system is constructed to achieve triple dynamic randomization of the watermark information, embedding locations, and embedding strength, thereby enhancing security. Guided by an HVS-based just noticeable distortion (JND) model, a spatio-temporally adaptive embedding strength is then derived, maximizing robustness under strict imperceptibility constraints. Furthermore, a blind extraction mechanism using coefficient-relation modulation is designed, inherently improving resilience against common video processing and malicious attacks. Collectively, these strategies unify imperceptibility, robustness, and security. The experimental results confirm the algorithm’s superior performance against benchmarks. It exhibits stronger resistance to statistical analysis, with a mean Kullback-Leibler (KL) divergence of only 0.003, and enhanced watermark robustness, shown by a 61.8% improvement in normalized correlation (NC). For imperceptibility, it achieves a gain in the peak signal-to-noise ratio (PSNR) over 2.2 dB.
Kesong Wu, Maowei Li, Peng Yang 0009, Jiangtian Nie, Xianbin Cao 0001
TrustCom1