Qi Song 0003

dblp:82/5132-3 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0006-7896-1567ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Creating Blank Canvas Against AI-enabled Image Forgery
abstract
AIGC-based image editing technology has greatly simplified the realistic-level image modification, causing serious potential risks of image forgery. This paper introduces a new approach to tampering detection using the Segment Anything Model (SAM). Instead of training SAM to identify tampered areas, we propose a novel strategy. The entire image is transformed into a blank canvas from the perspective of neural models. Any modifications to this blank canvas would be noticeable to the models. To achieve this idea, we introduce adversarial perturbations to prevent SAM from seeing anything, allowing it to identify forged regions when the image is tampered with. Due to SAM's powerful perceiving capabilities, naive adversarial attacks cannot completely tame SAM. To thoroughly deceive SAM and make it blind to the image, we introduce a frequency-aware optimization strategy, which further enhances the capability of tamper localization. Extensive experimental results demonstrate the effectiveness of our method.
Qi Song 0003, Ziyuan Luo, Renjie Wan
AAAI1
2026 Naturalistic Typographic Attacks on VLM-Based Image Quality Assessment
Ziyuan Luo, Qi Song 0003, Renjie Wan
QoMEX3
2026 Adversarially robust multimedia watermarking via data-centric optimization
Ziyuan Luo, Qi Song 0003, Haoliang Li, Anderson Rocha 0001, Renjie Wan
Pattern Recognit.2
2025 Align 3D Representation and Text Embedding for 3D Content Personalization
abstract
Recent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D personalization approaches predominantly rely on knowledge distillation-based methods, which require computationally expensive retraining procedures. To address this challenge, we propose Invert3D, a novel framework for convenient 3D content personalization. Nowadays, vision-language models such as CLIP enable direct image personalization through aligned vision-text embedding spaces. However, the inherent structural differences between 3D content and 2D images preclude direct application of these techniques to 3D personalization. Our approach bridges this gap by establishing alignment between 3D representations and text embedding spaces. Specifically, we develop a camera-conditioned 3D-to-text inverse mechanism that projects 3D contents into a 3D embedding aligned with text embeddings. This alignment enables efficient manipulation and personalization of 3D content through natural language prompts, eliminating the need for computationally retraining procedures. Extensive experiments demonstrate that Invert3D achieves effective personalization of 3D content.
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
ACM Multimedia1
2025 MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
abstract
The growing popularity of 3D Gaussian Splatting (3DGS) has intensified the need for effective copyright protection. Current 3DGS watermarking methods rely on computationally expensive fine-tuning procedures for each predefined message. We propose the first generalizable watermarking framework that enables efficient protection of Splatter Image-based 3DGS models through a single forward pass. We introduce GaussianBridge that transforms unstructured 3D Gaussians into Splatter Image format, enabling direct neural processing for arbitrary message embedding. To ensure imperceptibility, we design a Gaussian-Uncertainty-Perceptual heatmap prediction strategy for preserving visual quality. For robust message recovery, we develop a dense segmentation-based extraction mechanism that maintains reliable extraction even when watermarked objects occupy minimal regions in rendered views. Project page: https://kevinhuangxf.github.io/marksplatter.
Xiufeng Huang, Ziyuan Luo, Qi Song 0003, Ruofei Wang, Renjie Wan
ACM Multimedia3
2024 Protecting NeRFs' Copyright via Plug-And-Play Watermarking Base Model
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
ECCV (11)1
2024 BFRFormer: Transformer-Based Generator for Real-World Blind Face Restoration
abstract
Blind face restoration is a challenging task due to the unknown and complex degradation. Although face prior-based methods and reference-based methods have recently demonstrated high-quality results, the restored images tend to contain over-smoothed results and lose identity-preserved details when the degradation is severe. It is observed that this is attributed to short-range dependencies, the intrinsic limitation of convolutional neural networks. To model long-range dependencies, we propose a Transformer-based blind face restoration method, named BFRFormer, to reconstruct images with more identity-preserved details in an end-to-end manner. In BFRFormer, to remove blocking artifacts, the wavelet discriminator and aggregated attention module are developed, and spectral normalization and balanced consistency regulation are adaptively applied to address the training instability and over-fitting problem, respectively. Extensive experiments show that our method outperforms state-of-the-art methods on a synthetic dataset and four real-world datasets. The source code, Casia-Test dataset, and pre-trained models is released at https://github.com/s8Znk/BFRFormer.
Guojing Ge, Qi Song 0003, Guibo Zhu, Yuting Zhang 0007, Jinglu Chen, Miao Xin, Ming Tang 0001, Jinqiao Wang
ICASSP2
2024 Geometry Cloak: Preventing TGS-based 3D Reconstruction from Copyrighted Images
abstract
Single-view 3D reconstruction methods like Triplane Gaussian Splatting (TGS) have enabled high-quality 3D model generation from just a single image input within seconds. However, this capability raises concerns about potential misuse, where malicious users could exploit TGS to create unauthorized 3D models from copyrighted images. To prevent such infringement, we propose a novel image protection approach that embeds invisible geometry perturbations, termed ``geometry cloaks'', into images before supplying them to TGS. These carefully crafted perturbations encode a customized message that is revealed when TGS attempts 3D reconstructions of the cloaked image. Unlike conventional adversarial attacks that simply degrade output quality, our method forces TGS to fail the 3D reconstruction in a specific way - by generating an identifiable customized pattern that acts as a watermark. This watermark allows copyright holders to assert ownership over any attempted 3D reconstructions made from their protected images. Extensive experiments have verified the effectiveness of our geometry cloak.
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
NeurIPS1
2023 Degradation Conditioned GAN for Degradation Generalization of Face Restoration Models
abstract
Face restoration models are usually trained on synthetic degraded data to output an image that matches the clean version of itself. Most previous methods use a single model to deal with all the degradation levels, resulting in a domain generalization problem. We explore the value of degradation information and propose a Degradation Conditioned GAN (DeCGAN). The architecture consists of modulated convolution, bias, and fusion modules, inspired by deblurring, denoising and super-resolution. The whole network can be modulated by the degradation levels to achieve delicate and precise restoration effects. Experiments are conducted on conventional and modulated face restoration tasks. DeCGAN can achieve more faithful restoration and better metrics (FID, LPIPS, etc.) than previous methods do. Moreover, our model performs well on real-world low-quality face images.
Qi Song 0003, Wu Shi, Guojing Ge, Liang Chang 0001
ICIP1
2022 TaiSu: A 166M Large-scale High-Quality Dataset for Chinese Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) has been shown to be an efficient method to improve the performance of models on different vision-and-language downstream tasks. Substantial studies have shown that neural networks may be able to learn some general rules about language and visual concepts from a large-scale weakly labeled image-text dataset. However, most of the public cross-modal datasets that contain more than 100M image-text pairs are in English; there is a lack of available large-scale and high-quality Chinese VLP datasets. In this work, we propose a new framework for automatic dataset acquisition and cleaning with which we construct a new large-scale and high-quality cross-modal dataset named as TaiSu, containing 166 million images and 219 million Chinese captions. Compared with the recently released Wukong dataset, our dataset is achieved with much stricter restrictions on the semantic correlation of image-text pairs. We also propose to combine texts collected from the web with texts generated by a pre-trained image-captioning model. To the best of our knowledge, TaiSu is currently the largest publicly accessible Chinese cross-modal dataset. Furthermore, we test our dataset on several vision-language downstream tasks. TaiSu outperforms BriVL by a large margin on the zero-shot image-text retrieval task and zero-shot image classification task. TaiSu also shows better performance than Wukong on the image-retrieval task without using image augmentation for training. Results demonstrate that TaiSu can serve as a promising VLP dataset, both for understanding and generative tasks. More information can be referred to https://github.com/ksOAn6g5/TaiSu.
Guibo Zhu, Qi Song 0003, Guojing Ge, Guanhui Qiao, Ru Peng, Lingxiang Wu, Jinqiao Wang
NeurIPS4