Zuopeng Yang

dblp:254/9451 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-2151-1777ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Joint asymmetric discrete hashing for cross-modal retrieval
Jiaxing Li 0009, Zuopeng Yang, Xiaozhao Fang, Shengli Xie 0001, Yong Xu 0001
Pattern Recognit.3
2025 SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection
abstract
Recent open-vocabulary human-object interaction (OV-HOI) detection methods primarily rely on large language model (LLM) for generating auxiliary descriptions and leverage knowledge distilled from CLIP to detect unseen interaction categories. Despite their effectiveness, these methods face two challenges: (1) feature granularity deficiency, due to reliance on last layer visual features for text alignment, leading to the neglect of crucial object-level details from intermediate layers; (2) semantic similarity confusion, resulting from CLIP’s inherent biases toward certain classes, while LLM-generated descriptions based solely on labels fail to adequately capture inter-class similarities. To address these challenges, we propose a stratified granular comparison network. First, we introduce a granularity sensing alignment module that aggregates global semantic features with local details, refining interaction representations and ensuring robust alignment between intermediate visual features and text embeddings. Second, we develop a hierarchical group comparison module that recursively compares and groups classes using LLMs, generating fine-grained and discriminative descriptions for each interaction category. Experimental results on two widely-used benchmark datasets, SWIG-HOI and HICO-DET, demonstrate that our method achieves state-of-the-art results in OV-HOI detection. Codes is available at GitHub.
Chong Shi, Zuopeng Yang, Haojin Tang
CVPR3
2025 Distraction is All You Need for Multimodal Large Language Model Jailbreaking
abstract
Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model’s attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs’ defenses, offering new insights for attack strategies. Our code is available at https://github.com/TeamPigeonLab/CS-DJ.Warning: This paper contains unfiltered content generated by MLLMs that may be offensive to readers
Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Kanghua Mo, Changyu Dong
CVPR1
2025 Reversible generative steganography with distribution-preserving
abstract
Abstract Steganography aims to embed and extract secret information in digital media for enhancing information security, which is widely applied to covert communication, copyright and privacy protection, digital forensics, etc. To resist steganalysis detection, generative steganography is one of the most promising techniques with embedding secret information into a generated image. Although existing generative steganographic methods could perform well with low hiding capacity, most of them encode the secret information in non-distribution-preserving manners, leading to poor security performance against steganalyzers when hiding more secret information. Meanwhile, the secret information tends to be difficult to be extracted with these methods because the secret-to-image transformations are irreversible. To tackle these issues, in this paper, we propose a reversible generative steganography with distribution-preserving scheme, which is mainly composed of a secret message mapping strategy with distribution-preserving and a reversible Glow model. To improve the anti-detectability against steganalyzers, the message mapping strategy with distribution-preserving is customized to encode the secret information into latent vectors which follow the Gaussian distribution as they are usually done in typical image generation models. The Glow model is then trained with reversible transformation to map the latent vectors into the generated stego-images with information hiding. Owing to the distribution-preserving and reversibility of the message mapping and Glow model, the proposed generative steganographic method achieves superior security performance and accurate extraction of secret message. Extensive experimental results demonstrate that the proposed method outperforms several state-of-the-art methods in terms of information extraction accuracy and anti-detectability, especially for high hiding capacity (up to 4.0 bpp).
Weixuan Tang 0002, Yuan Rao 0002, Zuopeng Yang, Fei Peng 0001, Xutong Cui, Peijun Zhu
Cybersecur.3
2024 TD²-Net: Toward Denoising and Debiasing for Video Scene Graph Generation
abstract
Dynamic scene graph generation (SGG) focuses on detecting objects in a video and determining their pairwise relationships. Existing dynamic SGG methods usually suffer from several issues, including 1) Contextual noise, as some frames might contain occluded and blurred objects. 2) Label bias, primarily due to the high imbalance between a few positive relationship samples and numerous negative ones. Additionally, the distribution of relationships exhibits a long-tailed pattern. To address the above problems, in this paper, we introduce a network named TD2-Net that aims at denoising and debiasing for dynamic SGG. Specifically, we first propose a denoising spatio-temporal transformer module that enhances object representation with robust contextual information. This is achieved by designing a differentiable Top-K object selector that utilizes the gumbel-softmax sampling strategy to select the relevant neighborhood for each object. Second, we introduce an asymmetrical reweighting loss to relieve the issue of label bias. This loss function integrates asymmetry focusing factors and the volume of samples to adjust the weights assigned to individual samples. Systematic experimental results demonstrate the superiority of our proposed TD2-Net over existing state-of-the-art approaches on Action Genome databases. In more detail, TD2-Net outperforms the second-best competitors by 12.7% on mean-Recall@10 for predicate classification.
Chong Shi, Yibing Zhan, Zuopeng Yang, Dacheng Tao
AAAI4
2024 MMoT: Mixture-of-Modality-Tokens Transformer for Composed Multimodal Conditional Image Synthesis
Daqing Liu, Minghui Hu 0001, Zuopeng Yang, Changxing Ding, Dacheng Tao
Int. J. Comput. Vis.5
2024 Defending against similarity shift attack for EaaS via adaptive multi-target watermarking
Zuopeng Yang, Kangjun Liu
Inf. Sci.1
2024 Improving the Post-Training Neural Network Quantization by Prepositive Feature Quantization
abstract
Post-training neural network quantization (PTQ) is an effective model compression technology that has revolutionized the deployment of deep neural networks on various edge devices. It provides easy-to-use characteristics and allows for generating a quantized model based on a pre-trained counterpart without re-training. Typical PTQ approaches maintain output consistency through layer-wise calibration. However, these approaches still suffer from performance degradation primarily caused by feature quantization in ultra-low bitwidth conditions. To address this issue, we propose a prepositive feature quantization framework that decouples adjacent layers and calibrates the interaction between feature and parameter quantization perturbations. Additionally, we present a feature-loss-aware optimization strategy to solve the corresponding calibration problem. To validate the effectiveness of our method, we conducted extensive experiments on the ImageNet benchmark dataset. Our approach demonstrates a noticeable improvement in PTQ performance under the 2-bit condition.
Zuopeng Yang, Xiaolin Huang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Eliminating Contextual Prior Bias for Semantic Image Editing via Dual-Cycle Diffusion
abstract
The recent success of text-to-image generation diffusion models has also revolutionized semantic image editing, enabling the manipulation of images based on query/target texts. Despite these advancements, a significant challenge lies in the potential introduction of contextual prior bias in pre-trained models during image editing, e.g., making unexpected modifications to inappropriate regions. To address this issue, we present a novel approach called Dual-Cycle Diffusion, which generates an unbiased mask to guide image editing. The proposed model incorporates a Bias Elimination Cycle that consists of both a forward path and an inverted path, each featuring a Structural Consistency Cycle to ensure the preservation of image content during the editing process. The forward path utilizes the pre-trained model to produce the edited image, while the inverted path converts the result back to the source image. The unbiased mask is generated by comparing differences between the processed source image and the edited image to ensure that both conform to the same distribution. Our experiments demonstrate the effectiveness of the proposed method, as it significantly improves the D-CLIP score from 0.272 to 0.283. The code will be available athttps://github.com/JohnDreamer/DualCycleDiffsion.
Zuopeng Yang, Erdun Gao, Daqing Liu, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2023 Unified Discrete Diffusion for Simultaneous Vision-Language Generation
Minghui Hu 0001, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham, Heliang Zheng, Dacheng Tao, Ponnuthurai N. Suganthan
ICLR3
2022 Modeling Image Composition for Complex Scene Generation
abstract
We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Attention (TwFA) for exploring dependencies of object-to-object, object-to-patch and patch-to-patch. Compared to existing CNN-based and Transformer-based generation models that entangled modeling on pixel-level&patch-level and object-level&patch-level respectively, the proposed focal attention predicts the current patch token by only focusing on its highly-related tokens that specified by the spatial layout, thereby achieving disambiguation during training. Furthermore, the proposed TwFA largely increases the data efficiency during training, therefore we propose the first few-shot complex scene generation strategy based on the well-trained TwFA. Comprehensive experiments show the superiority of our method, which significantly increases both quantitative metrics and qualitative visual realism with respect to state-of-the-art CNN-based and transformer-based methods. Code is available at https://github.com/JohnDreamer/TwFA.
Zuopeng Yang, Daqing Liu, Jie Yang 0002, Dacheng Tao
CVPR1
2022 Human-Centric Image Captioning
Zuopeng Yang, Jie Yang 0002
Pattern Recognit.1
2021 Improving The Robustness Of Convolutional Neural Networks Via Sketch Attention
abstract
The convolutional neural networks (CNNs) is biased towards texture while human eyes relying heavily on the general structure. The inconformity leads to the vulnerability of CNNs. The convolutional results is determined by the local patterns and delicate adversarial perturbation would be amplified layer-wise. Meanwhile the image context and object structure, which can be represented by sketch, stay almost unchanged. Therefore we propose that the sketch information is weak but more robust. In order to transfer the robustness from sketch to image and improve the capture of global structure, a sketch attention guided CNNs (SAG-CNNs) pipeline is constructed. The experiments on CFAR-10 demonstrate that the defensive capability against black-box attacks of SAG-CNNs outperforms other counterparts evidently and achieve preferable trade-off between generalization and robustness.
Zuopeng Yang, Jie Yang 0002, Xiaolin Huang
ICIP2
2019 Feature Fusion Based Deep Spatiotemporal Model for Violence Detection in Videos
Mujtaba Asad, Zuopeng Yang, Zubair Khan, Jie Yang 0002, Xiangjian He
ICONIP (1)2