VLDB 2026 Research / reviewers in the wild / expert
Junyan Wu
dblp:223/1777
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 1Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VoiceCloak: A Multi-Dimensional Defense Framework Against Unauthorized Diffusion-Based Voice CloningabstractDiffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery process, but they have been proven incompatible with DMs due to the intricate generative mechanisms of diffusion. To bridge this gap, we introduce VoiceCloak, a multi-dimensional proactive defense framework with the goal of obfuscating speaker identity and degrading perceptual quality in potential unauthorized VC. To achieve these goals, we conduct a focused analysis to identify specific vulnerabilities within DMs, allowing VoiceCloak to disrupt the cloning process by introducing adversarial perturbations into the reference audio. Specifically, to obfuscate speaker identity, VoiceCloak first targets speaker identity by distorting representation learning embeddings to maximize identity variation, which is guided by auditory perception principles. Additionally, VoiceCloak disrupts crucial conditional guidance processes, particularly attention context, thereby preventing the alignment of vocal characteristics that are essential for achieving convincing cloning. Then, to address the second objective, VoiceCloak introduces score magnitude amplification to actively steer the reverse trajectory away from the generation of high-quality speech. Noise-guided semantic corruption is further employed to disrupt structural speech semantics captured by DMs, degrading output quality. Extensive experiments highlight VoiceCloak's outstanding defense success rate against unauthorized diffusion-based voice cloning. Additional audio samples of VoiceCloak are available in demo pages. Qianyue Hu, Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001 |
AAAI | 2 |
| 2026 | Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning FrameworkabstractImage forgery localization aims to precisely identify tampered regions within images, but it commonly depends on costly pixel-level annotations. To alleviate this annotation burden, weakly supervised image forgery localization (WSIFL) has emerged, yet existing methods still achieve limited localization performance as they mainly exploit intra-image consistency clues and lack external semantic guidance to compensate for insufficient supervision information. In this paper, we propose ViLaCo, a vision-language collaborative reasoning framework that introduces auxiliary semantic supervision derived from pre-trained vision-language models (VLMs), enabling accurate pixel-level localization using only image-level labels. Specifically, we first employ a vision-language feature modeling network to jointly extract textual semantics and visual features by leveraging pre-trained VLMs. Next, an adaptive vision-language reasoning network aligns these features through mutual interactions, producing semantically aligned representations. Subsequently, these representations are passed into dual prediction heads, where the coarse head performs image-level classification and the fine head generates pixel-level localization masks, allowing the coarse-grained task to provide guidance for the fine-grained localization. Moreover, a contrastive patch consistency module is introduced to cluster tampered features while separating authentic ones, facilitating more reliable forgery discrimination. Extensive experiments on multiple public datasets demonstrate that ViLaCo substantially outperforms existing WSIFL methods, achieving state-of-the-art performance in both detection and localization accuracy. Ziqi Sheng, Junyan Wu, Wei Lu 0001, Jiantao Zhou 0001 |
AAAI | 2 |
| 2026 | LTFDyG: A learnable temporal function-based dynamic graph neural network with dual-channel encoding
Xinzhi Shi, Chao Li 0022, Xingshuo Han, Shihe Su, Junyan Wu |
Appl. Intell. | 5 |
| 2025 | Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning NetworkabstractAudio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained annotations, which are obtained costly and challenging in real-world scenarios. To meet this challenge, in this paper, we propose a progressive audio-language co-learning network (LOCO) that adopts co-learning and self-supervision manners to prompt localization performance under weak supervision scenarios. Specifically, an audio-language co-learning module is first designed to capture forgery consensus features by aligning semantics from temporal and global perspectives. In this module, forgery-aware prompts are constructed by using utterance-level annotations together with learnable prompts, which can incorporate semantic priors into temporal content features dynamically. In addition, a forgery localization module is applied to produce forgery proposals based on fused forgery-class activation sequences. Finally, a progressive refinement strategy is introduced to generate pseudo frame-level labels and leverage supervised semantic contrastive learning to amplify the semantic distinction between real and fake content, thereby continuously optimizing forgery-aware features. Extensive experiments show that the proposed LOCO achieves SOTA performance on three public benchmarks. Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Shize Guo |
IJCAI | 1 |
| 2025 | A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery LocalizationabstractCurrent researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and challenging to scale for large datasets. To resolve these issues, we present a multimodal deviation perceiving framework for weakly-supervised temporal forgery localization (MDP), which aims to identify temporal partial forged segments using only video-level annotations. The MDP proposes a novel multimodal interaction mechanism (MI) and an extensible deviation perceiving loss to perceive multimodal deviation, which achieves the refined start and end timestamps localization of forged segments. Specifically, MI introduces a temporal property preserving cross-modal attention to measure the relevance between the visual and audio modalities in the probabilistic embedding space. It could identify the inter-modality deviation and construct comprehensive video features for temporal forgery localization. To explore further temporal deviation for weakly-supervised learning, an extensible deviation perceiving loss has been proposed, aiming at enlarging the deviation of adjacent segments of the forged samples and reducing that of genuine samples. Extensive experiments demonstrate the effectiveness of the proposed framework and achieve comparable results to fully-supervised approaches in several evaluation metrics. Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Qian Wang 0002 |
ACM Multimedia | 2 |
| 2024 | Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and LocalizationabstractRecently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a classification purpose and fail to perform meaningful analysis of the start and end timestamps of partial forgery segments. To address this challenge, we introduce a novel coarse-to-fine proposal refinement framework (CFPRF) that incorporates a frame-level detection network (FDN) and a proposal refinement network (PRN) for audio temporal forgery detection and localization. Specifically, the FDN aims to mine informative inconsistency cues between real and fake frames to obtain discriminative features that are beneficial for roughly indicating forgery regions. The PRN is responsible for predicting confidence scores and regression offsets to refine the coarse-grained proposals derived from the FDN. To learn robust discriminative features, we devise a difference-aware feature learning (DAFL) module guided by contrastive representation learning to enlarge the sensitive differences between different frames induced by minor manipulations. We further design a boundary-aware feature enhancement (BAFE) module to capture the contextual information of multiple transition boundaries and guide the interaction between boundary information and temporal features via a cross-attention mechanism. Extensive experiments show that our CFPRF achieves state-of-the-art performance on various datasets, including LAV-DF, ASVS2019PS, and HAD. Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Qian Wang 0002, Xiaochun Cao |
ACM Multimedia | 1 |
| 2024 | Audio Multi-View Spoofing Detection Framework Based on Audio-Text-Emotion CorrelationsabstractIn recent years, audio spoofing detection has received widespread attention for protecting personal privacy and social security. Despite the significant progress achieved in audio single-view spoofing detection, challenges remain with regard to addressing unknown spoofing attacks in realistic scenarios. To solve these challenging problems, in this paper, we introduce a novel audio multi-view spoofing detection framework (AMSDF), whose goal is to capture both intra-view and inter-view cues by measuring correlations within audio multi-view features (i.e., audio-emotion-text) for audio spoofing detection. In general, different view features are inherently interconnected in the real patterns, while they may present unnatural correlations in the spoofing patterns. Therefore, more discriminative cues can be mined by utilizing their complex interactions, which is beneficial to the audio spoofing detection task. To this end, an intra-view graph attention mechanism (IGAM) is first utilized to aggregate each intra-view node within the same view. Subsequently, a heterogeneous graph fusion module (HGFM) is applied to measure correlations within inter-view nodes, which are enhanced with a master node for comprehensive analysis purposes. Finally, a group-based readout scheme (GRS) is designed to capture and preserve the most distinctive cues by leveraging the strengths of different feature sets, thereby effectively distinguishing subtle differences between real and spoofing audio. The experimental results show that our proposed framework can achieve better performance than that of the state-of-the-art methods, especially in realistic scenarios. The code and pre-trained models are available athttps://github.com/ItzJuny/AMSDF. Junyan Wu, Qilin Yin, Ziqi Sheng, Wei Lu 0001, Jiwu Huang, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Emotion recognition in conversations with emotion shift detection based on multi-task learning
Qingqing Gao, Biwei Cao, Tianyun Gu, Xing Bao, Junyan Wu, Bo Liu 0004, Jiuxin Cao |
Knowl. Based Syst. | 6 |
| 2022 | Detecting Multiple Steganography Methods in Speech Streams Using Multi-Encoder NetworkabstractWith the development of speech steganography technology, steganographers are more and more inclined to realize more secure covert communication by combining a series of steganography methods. Thus, this letter presents a novel multi-encoder network (MENet) to achieve more efficient detection of multiple steganography methods. Differing from the previous work, MENet utilizes multiple private encoders to individually model the private features of each coding element, introduces a shared encoder based on an attention mechanism to fuse multiple private features for achieving better feature representation, and finally exploits a shared decoder to reduce feature dimensionality as well as give predictions. Taking the existing state-of-the-art steganography methods as the detection targets, the performance of the proposed steganalysis method is evaluated comprehensively and compared with the state-of-the-art ones. The experimental results show that the detection performance of MENet is overall better than the existing steganalysis methods, especially with low embedding rates and short speech sample lengths. Hui Tian 0002, Junyan Wu, Hanyu Quan, Chin-Chen Chang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Learning Differential Diagnosis of Skin Conditions with Co-occurrence Supervision Using Graph Convolutional Networks
Junyan Wu, Hao Jiang 0010, Anudeep Konda, Yang Zhang 0039 |
MICCAI (2) | 1 |
| 2018 | Deep Learning Based Urban Post-Accidental Congestion PredictionabstractUrban roads tend to cause traffic congestion for a long time after the occurrence of traffic accidents, which greatly affects daily transportations. Therefore, the prediction of the duration of traffic jams caused by traffic accidents can allocate traffic resources more reasonably and effectively, release induced traffic information, avoid secondary congestion, and quickly handle traffic accidents. It is of great significance to the rapid rescue of traffic accidents and to eliminate traffic safety hazards. In response to this hot issue, many scholars have done a lot of researches through numerous models, such as probability distribution and time series, and artificial neural networks. However, these models usually only consider temporal features or are based on shallow networks. Therefore, this work adopts a hybrid deep spatial-temporal residual neural network HD-SP-ResNet to predict the traffic volume and velocity, as well as the road congestion duration after traffic accident, so as to monitor and dispatch real-time traffic, response to the postaccidental congestion in time, in order to reduce the various losses incurred by congestion and improve people's satisfaction with traffic on the road. To verify the effectiveness of the proposed model, we conduct extensive experiments based on the taxi trajectory data and road accident data in Shanghai. The experiment results show that the proposed model can achieve a relatively accurate prediction on traffic volume and velocity, as well as the post-accidental congestion duration. Mingming Lu, Kunfang Zhang, Junyan Wu, Dingwu Tan |
MASS | 3 |