Wenpeng Xing

dblp:295/5626 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0001-5848-9417ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
abstract
Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection.Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for removing these fingerprints remain largely unexplored.Therefore, We present Mismatched Eraser (MEraser), a novel method for effectively removing backdoor-based fingerprints from LLMs while maintaining model performance.Our approach leverages a two-phase fine-tuning strategy utilizing carefully constructed mismatched and clean datasets.Through extensive evaluation across multiple LLM architectures and fingerprinting methods, we demonstrate that MEraser achieves complete fingerprinting removal while maintaining model performance with minimal training data of fewer than 1,000 samples.Furthermore, we introduce a transferable erasure mechanism that enables effective fingerprinting removal across different models without repeated training.In conclusion, our approach provides a practical solution for fingerprinting removal in LLMs, reveals critical vulnerabilities in current fingerprinting techniques, and establishes comprehensive evaluation benchmarks for developing more resilient model protection methods in the future.
Zhenhua Xu 0004, Wenpeng Xing, Xuhong Zhang 0001
ACL (1)4
2025 EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
abstract
The proliferation of large language models (LLMs) has intensified concerns over model theft and license violations, necessitating robust and stealthy ownership verification.Existing fingerprinting methods either require impractical white-box access or introduce detectable statistical anomalies.We propose Ever-Tracer, a novel gray-box fingerprinting framework that ensures stealthy and robust model provenance tracing.EverTracer is the first to repurpose Membership Inference Attacks (MIAs) for defensive use, embedding ownership signals via memorization instead of artificial trigger-output overfitting.It consists of Fingerprint Injection, which fine-tunes the model on any natural language data without detectable artifacts, and Verification, which leverages calibrated probability variation signal to distinguish fingerprinted models.This approach remains robust against adaptive adversaries, including input level modification, and model-level modifications.Extensive experiments across architectures demonstrate Ev-erTracer's state-of-the-art effectiveness, stealthness, and resilience, establishing it as a practical solution for securing LLM intellectual property.Our code and data are publicly available at https://github.com/Xuzhenhua55/EverTracer.
Zhenhua Xu 0004, Wenpeng Xing
EMNLP3
2025 Distill To Detect: Amplifying Anomalies in Backdoor Models through Knowledge Distillation
abstract
Backdoor attacks represent a significant threat to the security of deep learning models. Due to the stealthiness of backdoor attacks, effectively detecting whether a model has been compromised by such attacks remains a major challenge. Previous backdoor detection methods either rely on backdoor datasets to identify anomalies in backdoor samples or depend on prior knowledge of existing backdoor attacks. This results in difficulties in detecting backdoors when backdoor samples are unavailable and leads to poor generalization capabilities when addressing new attack methods. This work proposes a novel approach for detecting backdoor attacks called Distill To Detect (D2D), that does not depend on backdoor samples or any prior knowledge. It utilizes knowledge distillation to amplify more general and universal backdoor anomalies exhibited on clean samples for detection. This approach is not only more efficient but also exhibits strong generalization capabilities, enabling the detection of most backdoor attacks with low time and computational costs. We tested our approach against six backdoor attacks and three different model architectures, demonstrating the effectiveness of our proposed method.
Xuyang Teng, Wenpeng Xing, Chenhao Ye
ICASSP3
2025 IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs
abstract
With the widespread use of large language models (LLMs) in natural language processing, traditional evaluation methods based on static datasets have become inadequate to fully capture their performance and generalization capabilities. To address this challenge, we propose an Iterative Dynamic Evaluation (IDE) framework, which utilizes a multi-agent system to systematically evaluate LLMs. The innovation of IDE lies in its iterative enhancement and comparative selection mechanisms. Through successive rounds of data augmentation, the framework emulates evolutionary processes by progressively increasing sample complexity, while competitive selection ensures only the optimal samples are retained. This iterative optimization and filtering produce increasingly challenging and diverse datasets. Experimental results demonstrate that IDE more effectively exposes the limitations of models in complex tasks compared to traditional methods, providing a more comprehensive foundation for the evaluation, optimization, and application of LLMs.
Wenpeng Xing
ICASSP4
2025 NCDI-Diffusion: Neural Contextual and Directional Inversion for Novel View Synthesis through Diffusion Models
abstract
Novel view synthesis typically requires a comprehensive set of multi-view images for either image-based rendering or scene representation-based optimization. However, achieving high-fidelity novel view rendering often demands a large number of images. To address this limitation, we propose NCDI-Diffusion, a novel diffusion-based view synthesis method that reduces the number of required images by leveraging the prior knowledge embedded in pre-trained diffusion models. Specifically, NCDI-Diffusion encapsulates both the contextual and directional information of a scene by utilizing neural descriptors, which are inversely derived from a limited set of positioned multi-view training images. These descriptors guide the diffusion model's image synthesis process, enabling the generation of high-quality novel views. Empirical results on the Forward-facing Dataset demonstrate the effectiveness of our approach to novel view synthesis.
Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Changting Lin
ICASSP1
2025 Optimizing and Attacking Embodied Intelligence: Instruction Decomposition and Adversarial Robustness
abstract
Embodied intelligence, which enables agents to interact with the physical world, has gained significant attention for its real-world applications. However, these systems face two key challenges: optimizing task instructions and addressing vulnerabilities to adversarial attacks. We propose an AI agent integrating a large language model (LLM), a CLIP model, and two contextual databases. Additionally, we introduce an adversarial attack framework that manipulates prompt attention by replacing high-attention words and appending adversarial suffixes, inducing hazardous behaviors while maintaining visual and semantic plausibility. Our work advances instruction optimization and highlights vulnerabilities in embodied intelligence systems. Experimental results show that our optimization approach generates precise, real-world instructions. At the same time, our adversarial attack achieves higher success rates with reduced computational overhead, advanced optimization, and robustness in embodied intelligence systems.
Minghao Li 0012, Wenpeng Xing, Yong Liu 0029, Wei Zhang 0106
ICME2
2025 IDCNet: Image Decomposition and Cross-View Distillation for Generalizable Deepfake Detection
abstract
Existing deepfake detectors predominantly process entire facial images as input, which limits their sensitivity to local forgery cues due to representation bias and information loss through CNN feature aggregation. To address these limitations, we propose IDCNet, a novel deepfake detection framework based on image decomposition and cross-view distillation. Our key insight is that decomposing images into complementary views enables specialized processing of global and local forgery cues, while cross-view distillation facilitates their mutual enhancement. Specifically, the framework employs a lightweight U-Net generator with a dual-objective mechanism to decompose input images into global content and local detail views, optimized through reconstruction and classification losses. A cross-view distillation strategy is then applied to enhance complementary feature learning between views. Furthermore, to integrate local artifact information into existing detection models without architectural modifications, we propose a feature alignment method. Extensive experiments across 14 forgery methods demonstrate the effectiveness of our approach, achieving up to 4.4% AUC improvement on the CDFV2 dataset compared to state-of-the-art methods. The source code is available at: https://github.com/ wangzhiyuan120/idcnet.
Yuanzhi Yao, Wenpeng Xing, Meng Li 0006
IEEE Trans. Inf. Forensics Secur.5
2025 The Safety Illusion? Testing the Boundaries of Concept Removal in Diffusion Models
abstract
Text-to-image diffusion models are capable of producing high-quality images from textual descriptions; however, they present notable security concerns. These include the potential for generating Not-Safe-For-Work (NSFW) content, replicating artists' styles without authorization, or creating deepfakes. Recent advancements have proposed concept erasure techniques to eliminate sensitive concepts from these models, aiming to mitigate the generation of undesirable content. Nevertheless, the robustness of these techniques against a wide range of adversarial inputs has not been comprehensively investigated. To address this challenge, a novel two-stage optimization attack framework based on adversarial perturbations, referred to as Concept Embedding Adversary (CEA), was proposed in the present study. By leveraging the cross-modal alignment priors of the CLIP model, CEA iteratively adjusts adversarial embedding vectors to approximate the semantic expression of specific target concepts. This process enables the construction of deceptive adversarial prompts that exploit diffusion models, compelling them to regenerate previously erased concepts. The performance of concept erasure methods was evaluated, specifically when dealing with diversified adversarial prompts targeting erased concepts, such as NSFW content, artistic styles, and objects. Extensive experimental results demonstrate that existing concept erasure methods are unable to completely eliminate target concepts. In contrast, the proposed CEA framework exploits residual vulnerabilities within the generative latent space through a two-stage optimization process. By achieving precise cross-modal alignment, CEA attains significantly higher ASR in regenerating erased concepts.
Yixiang Pan, Ting Luo 0001, Wenpeng Xing
IEEE Trans. Image Process.4
2023 CasTensoRF: Cascaded Tensorial Radiance Fields for Novel View Synthesis
abstract
Novel views synthesized from Neural Radiance Fields (NeRF) have reached remarkable rendering quality. However, a 5D radiance field volume is too large to be stored or directly rendered. In order to efficiently reconstruct and manipulate such a high-order tensor, we leverage inspirations from previous tensor decomposition methods, e.g. Tensorial Radiance Fields (TensoRF) and Hierarchical Tucker decomposition. And we propose a Hierarchical Vector-Matrix decomposition (HVMD) framework to learn a sparse approximation of high-order tensors. The proposed HVMD takes advantage of tensor separation and factorization properties and builds a hierarchical scheme that enables a better approximation of the high-order tensor with a very limited number of parameters. Our method achieves better-rendering quality than TensoRF in the NeRF-synthetic dataset given the same model size. The advantage gets more significant when the network parameter number becomes extremely small.
Wenpeng Xing, Jie Chen 0026
ICME1
2023 IRCasTRF: Inverse Rendering by Optimizing Cascaded Tensorial Radiance Fields, Lighting, and Materials From Multi-view Images
abstract
We propose an inverse rendering pipeline that simultaneously reconstructs scene geometry, lighting, and spatially-varying material from a set of multi-view images. Specifically, the proposed pipeline involves volume and physics-based rendering, which are performed separately in two steps: exploration and exploitation. During the exploration step, our method utilizes the compactness of neural radiance fields and a flexible differentiable volume rendering technique to learn an initial volumetric field. Here, we introduce a novel cascaded tensorial radiance field method on top of the Canonical Polyadic (CP) decomposition to boost model compactness beyond conventional methods. In the exploitation step, a shading pass that incorporates a differentiable physics-based shading method is applied to jointly optimize the scene's geometry, spatially-varying materials, and lighting, using image reconstruction loss. Experimental results demonstrate that our proposed inverse rendering pipeline, IRCasTRF, outperforms prior works in inverse rendering quality. The final output is highly compatible with downstream applications like scene editing and advanced simulations. Further details are available on the project page: https://ircasrf.github.io/.
Wenpeng Xing, Jie Chen 0026, Ka Chun Cheung, Simon See
ACM Multimedia1
2022 Temporal-MPI: Enabling Multi-plane Images for Dynamic Scene Modelling via Temporal Basis Learning
Wenpeng Xing, Jie Chen 0026
ECCV (15)1
2022 NEX+: Novel View Synthesis with Neural Regularisation Over Multi-Plane Images
abstract
We propose Nex+, a neural Multi-Plane Image (MPI) representation with alpha denoising for the task of novel view synthesis (NVS). Overfitting to training data is a common challenge for all learning-based models. We propose a novel solution for resolving such issue in the context of NVS with signal denoising-motivated operations over the alpha coefficients of the MPI, without any additional requirements for supervision. Nex+contains a novel 5D Alpha Neural Regulariser (ANR), which favors low-frequency components in the angular domain, i.e., the alpha coefficients’ signal sub-space indicating various viewing directions. ANR’s angular low-frequency property derives from its small number of angular encoding levels and output basis. The regularised alpha in Nex+can model the scene geometry more accurately than Nex, and outperforms other state-of-the-art methods on public datasets for the task of NVS.
Wenpeng Xing, Jie Chen 0026
ICASSP1
2022 MVSPlenOctree: Fast and Generic Reconstruction of Radiance Fields in PlenOctree from Multi-view Stereo
abstract
We present MVSPlenOctree, a novel approach that can efficiently reconstruct radiance fields for view synthesis. Unlike previous scene-specific radiance fields reconstruction methods, we present a generic pipeline that can efficiently reconstruct 360-degree-renderable radiance fields via multi-view stereo (MVS) inference from tens of sparse-spread out images. Our approach leverages variance-based statistic features for MVS inference, and combines this with image based rendering and volume rendering for radiance field reconstruction. We first train a MVS Machine for reasoning scene's density and appearance. Then, based on the spatial hierarchy of the PlenOctree and coarse-to-fine dense sampling mechanism, we design a robust and efficient sampling strategy for PlenOctree reconstruction, which handles occlusion robustly. A 360-degree-renderable radiance fields can be reconstructed in PlenOctree from MVS Machine in an efficient single forward pass. We trained our method on real-world DTU, LLFF datasets, and synthetic datasets. We validate its generalizability by evaluating on the test set of DTU dataset which are unseen in training. In summary, our radiance field reconstruction method is both efficient and generic, a coarse 360-degree-renderable radiance field can be reconstructed in seconds and a dense one within minutes. Please visit the project page for more details: https://derry-xing.github.io/projects/MVSPlenOctree.
Wenpeng Xing, Jie Chen 0026
ACM Multimedia1
2022 Scale-Consistent Fusion: From Heterogeneous Local Sampling to Global Immersive Rendering
abstract
Image-based geometric modeling and novel view synthesis based on sparse large-baseline samplings are challenging but important tasks for emerging multimedia applications such as virtual reality and immersive telepresence. Existing methods fail to produce satisfactory results due to the limitation on inferring reliable depth information over such challenging reference conditions. With the popularization of commercial light field (LF) cameras, capturing LF images (LFIs) is as convenient as taking regular photos, and geometry information can be reliably inferred. This inspires us to use a sparse set of LF captures to render high-quality novel views globally. However, the fusion of LF captures from multiple angles is challenging due to the scale inconsistency caused by various capture settings. To overcome this challenge, we propose a novel scale-consistent volume rescaling algorithm that robustly aligns the disparity probability volumes (DPV) among different captures for scale-consistent global geometry fusion. Based on the fused DPV projected to the target camera frustum, novel learning-based modules (i.e., the attention-guided multi-scale residual fusion module, and the disparity field-guided deep re-regularization module), which comprehensively regularize noisy observations from heterogeneous captures for high-quality rendering of novel LFIs, have been proposed. Both quantitative and qualitative experiments over the Stanford Lytro Multi-view LF dataset show that the proposed method outperforms state-of-the-art methods significantly under different experiment settings for disparity inference and LF synthesis.
Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Qiang Wang 0022, Yike Guo
IEEE Trans. Image Process.1