VLDB 2026 Research / reviewers in the wild / expert
Ziyin Zhou
dblp:299/7686
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Security and privacy · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards General Visual-Linguistic Face Forgery DetectionabstractFace manipulation techniques have achieved significant advances, presenting serious challenges to security and social trust. Recent works demonstrate that leveraging multimodal models can enhance the generalization and interpretability of face forgery detection. However, existing annotation approaches, whether through human labeling or direct Multimodal Large Language Model (MLLM) generation, often suffer from hallucination issues, leading to inaccurate text descriptions, especially for high-quality forgeries. To address this, we propose Face Forgery Text Generator (FFTG), a novel annotation pipeline that generates accurate text descriptions by leveraging forgery masks for initial region and type identification, followed by a comprehensive prompting strategy to guide MLLMs in reducing hallucination. We validate our approach through fine-tuning both CLIP with a three-branch training framework combining unimodal and multimodal objectives, and MLLMs with our structured annotations. Experimental results demonstrate that our method not only achieves more accurate annotations with higher region identification accuracy, but also leads to improvements in model performance across various forgery detection benchmarks. Our Codes are available in https://github.com/skJack/VLFFD.git. Ke Sun 0016, Shen Chen 0004, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, Rongrong Ji |
CVPR | 4 |
| 2025 | Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsabstractThe rapid development of AI-generated content (AIGC) technology has led to the misuse of highly realistic AI-generated images (AIGI) in spreading misinformation, posing a threat to public information security. Although existing AIGI detection techniques are generally effective, they face two issues: 1) a lack of human-verifiable explanations, and 2) a lack of generalization in the latest generation technology. To address these issues, we introduce a large-scale and comprehensive dataset, Holmes-Set, which includes the Holmes-SFTSet, an instruction-tuning dataset with explanations on whether images are AI-generated, and the Holmes-DPOSet, a human-aligned preference dataset. Our work introduces an efficient data annotation method called the Multi-Expert Jury, enhancing data generation through structured MLLM explanations and quality control via cross-model evaluation, expert defect filtering, and human preference modification. In addition, we propose Holmes Pipeline, a meticulously designed three-stage training framework comprising visual expert pre-training, supervised fine-tuning, and direct preference optimization. Holmes Pipeline adapts multimodal large language models (MLLMs) for AIGI detection while generating human-verifiable and human-aligned explanations, ultimately yielding our model AIGI-Holmes. During the inference stage, we introduce a collaborative decoding strategy that integrates the model perception of the visual expert with the semantic reasoning of MLLMs, further enhancing the generalization capabilities. Extensive experiments on three benchmarks validate the effectiveness of our AIGI-Holmes. Ziyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun 0016, Jiayi Ji, Shouhong Ding, Xiaoshuai Sun, Yunsheng Wu, Rongrong Ji |
ICCV | 1 |
| 2025 | Towards Rationale-Answer Alignment of LVLMs via Self-Rationale CalibrationabstractLarge Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight “rationale fine-tuning” approach, which modifies the model’s response format to require a rationale before deriving answer without explicit prompts. Next, SRC searches a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs. Yuanchen Wu, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li 0002 |
ICML | 4 |
| 2024 | Variance-Insensitive and Target-Preserving Mask Refinement for Interactive Image SegmentationabstractPoint-based interactive image segmentation can ease the burden of mask annotation in applications such as semantic segmentation and image editing. However, fully extracting the target mask with limited user inputs remains challenging. We introduce a novel method, Variance-Insensitive and Target-Preserving Mask Refinement to enhance segmentation quality with fewer user inputs. Regarding the last segmentation result as the initial mask, an iterative refinement process is commonly employed to continually enhance the initial mask. Nevertheless, conventional techniques suffer from sensitivity to the variance in the initial mask. To circumvent this problem, our proposed method incorporates a mask matching algorithm for ensuring consistent inferences from different types of initial masks. We also introduce a target-aware zooming algorithm to preserve object information during downsampling, balancing efficiency and accuracy. Experiments on GrabCut, Berkeley, SBD, and DAVIS datasets demonstrate our method's state-of-the-art performance in interactive image segmentation. Chaowei Fang, Ziyin Zhou, Junye Chen, Hanjing Su, Qingyao Wu, Guanbin Li |
AAAI | 2 |
| 2024 | StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion ModelabstractThe rapid progress in generative models has given rise to the critical task of AI-Generated Content Stealth (AIGC-S), which aims to create AI-generated images that can evade both forensic detectors and human inspection. This task is crucial for understanding the vulnerabilities of existing detection methods and developing more robust techniques. However, current adversarial attacks often introduce visible noise, have poor transferability, and fail to address spectral differences between AI-generated and genuine images. To address this, we propose StealthDiffusion, a framework based on stable diffusion that modifies AI-generated images into high-quality, imperceptible adversarial examples capable of evading state-of-the-art forensic detectors. StealthDiffusion comprises two main components: Latent Adversarial Optimization, which generates adversarial perturbations in the latent space of stable diffusion, and Control-VAE, a module that reduces spectral differences between the generated adversarial images and genuine images without affecting the original diffusion model's generation process. Extensive experiments show that StealthDiffusion is effective in both white-box and black-box settings, transforming AI-generated images into high-quality adversarial forgeries with frequency spectra similar to genuine images. These forgeries are classified as genuine by advanced forensic classifiers and are difficult for humans to distinguish. Ziyin Zhou, Ke Sun 0016, Zhongxi Chen, Huafeng Kuang, Xiaoshuai Sun, Rongrong Ji |
ACM Multimedia | 1 |
| 2024 | Can't Say Cant? Measuring and Reasoning of Dark Jargons in Large Language Models
Ziyin Zhou, Zhangchi Zhao, Qianqian Qiao, Kaiying Han, Md. Imran Hossen, Xiali Hei 0001 |
SecureComm (4) | 3 |
| 2024 | D2FL: Dimensional Disaster-oriented Backdoor Attack Defense Of Federated LearningabstractDefense algorithms for backdoor attacks in federated learning (FL) commonly rely on model parameter vectorization. However, as neural networks deepen, the exponential growth of model parameters leads to increased dimensionality, exacerbating the curse of dimensionality and reducing the effectiveness of traditional distance-based defenses. To address this, we propose Dimensional Disaster-oriented Backdoor Attack Defense in Federated Learning (D2FL), a method that mitigates attacks by focusing on expressive backdoor modules rather than the entire model. This approach reduces dimensionality and mitigates the challenges posed by large parameter spaces. Our extensive evaluation of D2FL on image classification tasks across various deep neural networks demonstrates its superior efficiency, significantly reducing both defense and aggregation times. Ziyin Zhou, Zezheng Sun, Zeping Li, Jiameng Han, Zhangchi Zhao |
TrustCom | 3 |
| 2024 | Paa-Tee: A Practical Adversarial Attack on Thermal Infrared Detectors with Temperature and Pose AdaptabilityabstractThermal infrared object detectors play an important role in security-related tasks, necessitating feasible adversarial attacks to evaluate their robustness. In many cases, implementing attacks in the physical space by a patch demands intricate and specialized perturbations. However, state-of-the-art adversarial attacks are often impractical, as they require fixed perturbation location and are susceptible to environmental temperature, leading to attack effects overfitting to specific poses and environments. To address this, we propose a practical adversarial attack method named Paa-Tee, with two input transformation strategies. For poses, we continuously alter the patch’s position to mitigate the impact of different poses on the patch’s location. For temperature, leveraging the principles of thermal imaging, we apply various transformations to a single input image to simulate different attack environments. Meanwhile, we utilize hot and cold pastes as low-resolution patches to implement attacks in the physical world. Extensive experiments validate the efficacy of our approach in both the digital and physical worlds. In the digital world, our attacks reduce the average precision of mainstream detectors by 65.44%. In the physical world, we achieve an average attack success rate of 63.77% under various distances, poses, angles, and environmental conditions. Zhangchi Zhao, Liqun Shan, Ziyin Zhou, Kaiying Han, Xiali Hei 0001 |
TrustCom | 4 |
| 2023 | BFMNet: Bilateral feature fusion network with multi-scale context aggregation for real-time semantic segmentation
Fangyu Zhang, Ziyin Zhou |
Neurocomputing | 3 |
| 2021 | Focusing on Persons: Colorizing Old Images Learning from Modern Historical MoviesabstractIn industry, there exist plenty of scenarios where old gray photos need to be automatically colored, such as video sites and archives. In this paper, we present the HistoryNet focusing on historical person's diverse high fidelity clothing colorization based on fine grained semantic understanding and prior. Colorization of historical persons is realistic and practical, however, existing methods do not perform well in the regards. In this paper, a HistoryNet including three parts, namely, classification, fine grained semantic parsing and colorization, is proposed. Classification sub-module supplies classifying of images according to the eras, nationalities and garment types; Parsing sub-network supplies the semantic for person contours, clothing and background in the image to achieve more accurate colorization of clothes and persons and prevent color overflow. In the training process, we integrate classification and semantic parsing features into the coloring generation network to improve colorization. Through the design of classification and parsing subnetwork, the accuracy of image colorization can be improved and the boundary of each part of image can be more clearly. Moreover, we also propose a novel Modern Historical Movies Dataset (MHMD) containing 1,353,166 images and 42 labels of eras, nationalities, and garment types for automatic colorization from 147 historical movies or TV series made in modern time. Various quantitative and qualitative comparisons demonstrate that our method outperforms the state-of-the-art colorization methods, especially on military uniforms, which has correct colors according to the historical literatures. Xin Jin 0015, Zhonglan Li, Dongqing Zou, Xiaodong Li 0013, Xingfan Zhu, Ziyin Zhou, Qilong Sun |
ACM Multimedia | 7 |