Xuanyu Zhang 0003

dblp:245/6291-3 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-6713-4500ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021
YearPublicationVenuePosition
2026 VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning
abstract
Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-generated videos remains challenging due to limited generalization, lack of temporal awareness, heavy reliance on large-scale annotated datasets, and the lack of effective interaction with generation models. Most current approaches rely on supervised fine-tuning of vision-language models (VLMs), which often require large-scale annotated datasets and tend to decouple understanding and generation. To address these shortcomings, we propose VQ-Insight, a novel reasoning-style VLM framework for AIGC video quality assessment. Our approach features: (1) a progressive video quality learning scheme that combines image quality warm-up, general task-specific temporal learning, and joint optimization with the video generation model; (2) the design of multi-dimension scoring rewards, preference comparison rewards, and temporal modeling rewards to enhance both generalization and specialization in video quality evaluation. Extensive experiments demonstrate that VQ-Insight consistently outperforms state-of-the-art baselines in preference comparison, multi-dimension scoring, and natural video scoring, bringing significant improvements for video generation tasks.
Xuanyu Zhang 0003, Shijie Zhao 0001, Li Zhang 0006, Jian Zhang 0018
AAAI1
2025 OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking
abstract
With the rapid growth of generative AI and its widespread application in image editing, new risks have emerged regarding the authenticity and integrity of digital content. Existing versatile watermarking approaches suffer from tradeoffs between tamper localization precision and visual quality. Constrained by the limited flexibility of previous framework, their localized watermark must remain fixed across all images. Under AIGC-editing, their copyright extraction accuracy is also unsatisfactory. To address these challenges, we propose OmniGuard, a novel augmented versatile watermarking approach that integrates proactive embedding with passive, blind extraction for robust copyright protection and tamper localization. OmniGuard employs a hybrid forensic framework that enables flexible localization watermark selection and introduces a degradation-aware tamper extraction network for precise localization under challenging conditions. Additionally, a lightweight AIGC-editing simulation layer is designed to enhance robustness across global and local editing. Extensive experiments show that OmniGuard achieves superior fidelity, robustness, and flexibility. Compared to the recent state-of-the-art approach EditGuard, our method outperforms it by 4.25dB in PSNR of the container image, 20.7% in F1-Score under noisy conditions, and 14.8% in average bit accuracy.
Xuanyu Zhang 0003, Zecheng Tang, Zhipei Xu, Runyi Li, Youmin Xu, Bin Chen 0006, Jian Zhang 0018
CVPR1
2025 FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
abstract
The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. Although current image forgery detection and localization (IFDL) methods are generally effective, they tend to face two challenges: \textbf{1)} black-box nature with unknown detection principle, \textbf{2)} limited generalization across diverse tampering methods (e.g., Photoshop, DeepFake, AIGC-Editing). To address these issues, we propose the explainable IFDL task and design FakeShield, a multi-modal framework capable of evaluating image authenticity, generating tampered region masks, and providing a judgment basis based on pixel-level and image-level tampering clues. Additionally, we leverage GPT-4o to enhance existing IFDL datasets, creating the Multi-Modal Tamper Description dataSet (MMTD-Set) for training FakeShield's tampering analysis capabilities. Meanwhile, we incorporate a Domain Tag-guided Explainable Forgery Detection Module (DTE-FDM) and a Multi-modal Forgery Localization Module (MFLM) to address various types of tamper detection interpretation and achieve forgery localization guided by detailed textual descriptions. Extensive experiments demonstrate that FakeShield effectively detects and localizes various tampering techniques, offering an explainable and superior solution compared to previous IFDL methods. The code is available at https://github.com/zhipeixu/FakeShield.
Zhipei Xu, Xuanyu Zhang 0003, Runyi Li, Zecheng Tang, Jian Zhang 0018
ICLR2
2025 SecureGS: Boosting the Security and Fidelity of 3D Gaussian Splatting Steganography
abstract
3D Gaussian Splatting (3DGS) has emerged as a premier method for 3D representation due to its real-time rendering and high-quality outputs, underscoring the critical need to protect the privacy of 3D assets. Traditional NeRF steganography methods fail to address the explicit nature of 3DGS since its point cloud files are publicly accessible. Existing GS steganography solutions mitigate some issues but still struggle with reduced rendering fidelity, increased computational demands, and security flaws, especially in the security of the geometric structure of the visualized point cloud. To address these demands, we propose a \textbf{SecureGS}, a secure and efficient 3DGS steganography framework inspired by Scaffold-GS's anchor point design and neural decoding. SecureGS uses a hybrid decoupled Gaussian encryption mechanism to embed offsets, scales, rotations, and RGB attributes of the hidden 3D Gaussian points in anchor point features, retrievable only by authorized users through privacy-preserving neural networks. To further enhance security, we propose a density region-aware anchor growing and pruning strategy that adaptively locates optimal hiding regions without exposing hidden information. Extensive experiments show that SecureGS significantly surpasses existing GS steganography methods in rendering fidelity, speed, and security.
Xuanyu Zhang 0003, Jiarui Meng, Zhipei Xu, Shuzhou Yang, Yanmin Wu, Ronggang Wang, Jian Zhang 0018
ICLR1
2025 TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
abstract
Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previous methods primarily relied on end-to-end networks to perform single try-on tasks, lacking versatility and flexibility. We propose TalkFashion, an intelligent try-on assistant that leverages the powerful comprehension capabilities of large language models to analyze user instructions and determine which task to execute, thereby activating different processing pipelines accordingly. Additionally, we introduce an instruction-based local repainting model that eliminates the need for users to manually provide masks. With the help of multi-modal models, this approach achieves fully automated local editings, enhancing the flexibility of editing tasks. The experimental results demonstrate better semantic consistency and visual quality compared to the current methods.
Xuanyu Zhang 0003, Jian Zhang 0018
ICME2
2025 Diffusion-Based Hierarchical Image Steganography
abstract
This paper introduces Hierarchical Image Steganography (HIS), a novel method that uses diffusion models to enhance the security and capacity of embedding multiple images into a single container. HIS assigns varying levels of robustness to images based on their importance, ensuring enhanced protection against manipulation. It adeptly exploits the robustness of the Diffusion Model and the reversibility of the Flow Model. Integrating Embed-Flow and Enhance-Flow improves embedding efficiency and image recovery quality, respectively, setting HIS apart from conventional multi-image steganography techniques. This innovative structure can autonomously generate a container image, thereby securely and efficiently concealing multiple images and text. Rigorous subjective and objective evaluations underscore HIS’s advantage in analytical resistance, robustness, and capacity, illustrating its expansive applicability in content safeguarding and privacy fortification.
Youmin Xu, Xuanyu Zhang 0003, Xiandong Meng, Chong Mou, Jian Zhang 0018
ICME2
2025 Self-supervised Scalable Deep Compressed Sensing
Bin Chen 0006, Xuanyu Zhang 0003, Yongbing Zhang 0002, Jian Zhang 0018
Int. J. Comput. Vis.2
2025 DiffLLE: Diffusion-based Domain Calibration for Weak Supervised Low-light Image Enhancement
Shuzhou Yang, Xuanyu Zhang 0003, Yinhuai Wang, Jiwen Yu, Jian Zhang 0018
Int. J. Comput. Vis.2
2024 EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection
abstract
In the era of AI-generated content (AIGC), malicious tampering poses imminent threats to copyright integrity and information security. Current deep image watermarking, while widely accepted for safeguarding visual content, can only protect copyright and ensure traceability. They fall short in localizing increasingly realistic image tampering, potentially leading to trust crises, privacy violations, and legal disputes. To solve this challenge, we propose an innovative proactive forensics framework EditGuard, to unify copyright protection and tamper-agnostic localization, especially for AIGC-based editing methods. It can offer a meticulous embedding of imperceptible watermarks and precise decoding of tampered areas and copyright in-formation. Leveraging our observed fragility and locality of image-into-image steganography, the realization of Edit-Guard can be converted into a united image-bit steganography issue, thus completely decoupling the training process from the tampering types. Extensive experiments verify that our EditGuard balances the tamper localization accuracy, copyright recovery precision, and generalizability to various AIGC-based tampering methods, especially for image forgery that is difficult for the naked eye to detect.
Xuanyu Zhang 0003, Runyi Li, Jiwen Yu, Youmin Xu, Jian Zhang 0018
CVPR1
2024 V2A-Mark: Versatile Deep Visual-Audio Watermarking for Manipulation Localization and Copyright Protection
abstract
AI-generated video has revolutionized short video production, filmmaking, and personalized media, making video local editing an essential tool. However, this progress also blurs the line between reality and fiction, posing challenges in multimedia forensics. To solve this urgent issue, V2A-Mark is proposed to address the limitations of current video tampering forensics, such as poor generalizability, singular function, and single modality focus. Combining the fragility of video-into-video steganography with deep robust watermarking, our method can embed invisible visual-audio localization watermarks and copyright watermarks into the original video frames and audio, enabling precise manipulation localization and copyright protection. We also design a temporal alignment and fusion module and degradation prompt learning to enhance the localization accuracy and decoding robustness. Meanwhile, we introduce a sample-level audio localization method and a cross-modal copyright extraction mechanism to couple the information of audio and video frames. The effectiveness of V2A-Mark has been verified on a visual-audio tampering dataset, emphasizing its superiority in localization precision and copyright accuracy, crucial for the sustainable development of video editing in the AIGC video era.
Xuanyu Zhang 0003, Youmin Xu, Runyi Li, Jiwen Yu, Zhipei Xu, Jian Zhang 0018
ACM Multimedia1
2024 GS-Hider: Hiding Messages into 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has already become the emerging research focus in the fields of 3D scene reconstruction and novel view synthesis. Given that training a 3DGS requires a significant amount of time and computational cost, it is crucial to protect the copyright, integrity, and privacy of such 3D assets. Steganography, as a crucial technique for encrypted transmission and copyright protection, has been extensively studied. However, it still lacks profound exploration targeted at 3DGS. Unlike its predecessor NeRF, 3DGS possesses two distinct features: 1) explicit 3D representation; and 2) real-time rendering speeds. These characteristics result in the 3DGS point cloud files being public and transparent, with each Gaussian point having a clear physical significance. Therefore, ensuring the security and fidelity of the original 3D scene while embedding information into the 3DGS point cloud files is an extremely challenging task. To solve the above-mentioned issue, we first propose a steganography framework for 3DGS, dubbed GS-Hider, which can embed 3D scenes and images into original GS point clouds in an invisible manner and accurately extract the hidden messages. Specifically, we design a coupled secured feature attribute to replace the original 3DGS's spherical harmonics coefficients and then use a scene decoder and a message decoder to disentangle the original RGB scene and the hidden message. Extensive experiments demonstrated that the proposed GS-Hider can effectively conceal multimodal messages without compromising rendering quality and possesses exceptional security, robustness, capacity, and flexibility. Our project is available at: https://xuanyuzhang21.github.io/project/gshider.
Xuanyu Zhang 0003, Jiarui Meng, Runyi Li, Zhipei Xu, Yongbing Zhang 0002, Jian Zhang 0018
NeurIPS1
2024 Progressive Content-Aware Coded Hyperspectral Snapshot Compressive Imaging
abstract
Hyperspectral imaging plays a pivotal role across diverse applications, like remote sensing, medicine, and cytology. The utilization of 2D sensors to acquire 3D hyperspectral images (HSIs) via a coded aperture snapshot spectral imaging (CASSI) system has proven successful, owing to its hardware-friendly implementation and fast sampling speed. Nevertheless, for less spectrally sparse scenes, the use of a single snapshot and unreasonable coded aperture design limits the efficacy of CASSI systems and renders HSI reconstruction more ill-posed, leading to compromised spatial and spectral fidelity. This paper proposes a novel Progressive Content-Aware CASSI (PCA-CASSI) framework, which progressively captures HSIs using multiple optimized content-aware coded apertures and fuses all snapshot measurements for reconstruction. By unlocking the full potential of CASSI systems and elevating their performance ceilings, this framework offers researchers new avenues for improving imaging quality. Furthermore, we develop the RndHRNet, a Range-Null space Decomposition (RND)-inspired deep unfolding network with multiple iterative phases for HSI recovery. Each unfolded recovery phase efficiently exploits the physical information within the coded apertures via explicit RND and adaptively explores the spatial-spectral correlation by dual transformer blocks. Through comprehensive experiments, our approach demonstrates superior performance compared to existing state-of-the-art methods in both the multiple- and single-shot compressive HSI imaging tasks with substantial improvements. Code is available athttps://github.com/xuanyuzhang21/PCA-CASSI.
Xuanyu Zhang 0003, Bin Chen 0006, Wenzhen Zou, Yongbing Zhang 0002, Ruiqin Xiong, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.1
2023 CRoSS: Diffusion Model Makes Controllable, Robust and Secure Image Steganography
abstract
Current image steganography techniques are mainly focused on cover-based methods, which commonly have the risk of leaking secret images and poor robustness against degraded container images. Inspired by recent developments in diffusion models, we discovered that two properties of diffusion models, the ability to achieve translation between two images without training, and robustness to noisy data, can be used to improve security and natural robustness in image steganography tasks. For the choice of diffusion model, we selected Stable Diffusion, a type of conditional diffusion model, and fully utilized the latest tools from open-source communities, such as LoRAs and ControlNets, to improve the controllability and diversity of container images. In summary, we propose a novel image steganography framework, named Controllable, Robust and Secure Image Steganography (CRoSS), which has significant advantages in controllability, robustness, and security compared to cover-based image steganography methods. These benefits are obtained without additional training. To our knowledge, this is the first work to introduce diffusion models to the field of image steganography. In the experimental section, we conducted detailed experiments to demonstrate the advantages of our proposed CRoSS framework in controllability, robustness, and security.
Jiwen Yu, Xuanyu Zhang 0003, Youmin Xu, Jian Zhang 0018
NeurIPS2
2022 HerosNet: Hyperspectral Explicable Reconstruction and Optimal Sampling Deep Network for Snapshot Compressive Imaging
abstract
Hyperspectral imaging is an essential imaging modality for a wide range of applications, especially in remote sensing, agriculture, and medicine. Inspired by existing hyperspectral cameras that are either slow, expensive, or bulky, reconstructing hyperspectral images (HSIs) from a low-budget snapshot measurement has drawn wide attention. By mapping a truncated numerical optimization algorithm into a network with a fixed number of phases, recent deep unfolding networks (DUNs) for spectral snapshot compressive sensing (SCI) have achieved remarkable success. However, DUNs are far from reaching the scope of industrial applications limited by the lack of cross-phase feature interaction and adaptive parameter adjustment. In this paper, we propose a novel Hyperspectral Explicable Reconstruction and Optimal Sampling deep Network for SCI, dubbed HerosNet, which includes several phases under the ISTA-unfolding framework. Each phase can flexibly simulate the sensing matrix and contextually adjust the step size in the gradient descent step, and hierarchically fuse and interact the hidden states of previous phases to effectively recover current HSI frames in the proximal mapping step. Simultaneously, a hardware-friendly optimal binary mask is learned end-to-end to further improve the reconstruction performance. Finally, our HerosNet is validated to outperform the state-of-the-art methods on both simulation and real datasets by large margins. The source code is available at https://github.com/jianzhangcs/HerosNet.
Xuanyu Zhang 0003, Yongbing Zhang 0002, Ruiqin Xiong, Qilin Sun 0001, Jian Zhang 0018
CVPR1