Zhaoyuan Yang

dblp:227/2718 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
abstract
Prompting has emerged as a practical way to adapt frozen vision-language models (VLMs) for video anomaly detection (VAD). Yet, existing prompts are often overly abstract, overlooking the fine-grained human–object interactions or action semantics that define complex anomalies in surveillance videos. We propose ASK-Hint, a structured prompting framework that leverages action-centric knowledge to elicit more accurate and interpretable reasoning from frozen VLMs. Our approach organizes prompts into semantically coherent groups (e.g. violence, property crimes, public safety) and formulates fine-grained guiding questions that align model predictions with discriminative visual cues. Extensive experiments on UCF-Crime and XD-Violence show that ASK-Hint consistently improves AUC over prior baselines, achieving state-of-the-art performance compared to both fine-tuned and training-free methods. Beyond accuracy, our framework provides interpretable reasoning traces towards anomaly and demonstrates strong generalization across datasets and VLM backbones. These results highlight the critical role of prompt granularity and establish ASK-Hint as a new training-free and generalizable solution for explainable video anomaly detection.
Shu Zou, Lukas Wesemann, Fabian Waschkowski, Zhaoyuan Yang, Jing Zhang 0052
WACV5
2025 Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
abstract
The evolution of Large Vision-Language Models (LVLMs) has progressed from single to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the alteration of image positions. To further explore this issue, we introduce Position-wise Question Answering (PQA), a meticulously designed task to quantify reasoning capabilities at each position. Our analysis reveals a pronounced position bias in LVLMs: open-source models excel in reasoning with images positioned later but underperform with those in the middle or at the beginning, while proprietary models show improved comprehension for images at the beginning and end but struggle with those in the middle. Motivated by this, we propose SoFt Attention (SoFA), a simple, training-free approach that mitigates this bias by employing linear interpolation between inter-image causal attention and bidirectional counterparts. Experimental results demonstrate that SoFA reduces position bias and enhances the reasoning performance of existing LVLMs. The code will be available https://github.com/xytian1008/sofa.
Shu Zou, Zhaoyuan Yang, Jing Zhang 0052
CVPR3
2025 Probability Density Geodesics in Image Diffusion Latent Space
abstract
Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the probability density. In this formulation, a path that traverses a high density (that is, probable) region of image latent space is shorter than the equivalent path through a low density region. We present algorithms for solving the associated initial and boundary value problems and show how to compute the probability density along the path and the geodesic distance between two points. Using these techniques, we analyze how closely video clips approximate geodesics in a pre-trained image diffusion space. Finally, we demonstrate how these techniques can be applied to training-free image sequence interpolation and extrapolation, given a pre-trained image diffusion model.
Qingtao Yu, Zhaoyuan Yang, Peter H. Tu, Jing Zhang 0052, Hongdong Li, Richard I. Hartley, Dylan Campbell
CVPR3
2025 Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
abstract
Few-shot adaptation for Vision-Language Models (VLMs) presents a dilemma: balancing in-distribution accuracy with out-of-distribution generalization. Recent research has utilized low-level concepts such as visual attributes to enhance generalization. However, this study reveals that VLMs overly rely on a small subset of attributes on decision-making, which co-occur with the category but are not inherently part of it, termed spuriously correlated attributes. This biased nature of VLMs results in poor generalization. To address this, 1) we first propose Spurious Attribute Probing (SAP), identifying and filtering out these problematic attributes to significantly enhance the generalization of existing attribute-based methods; 2) We introduce Spurious Attribute Shielding (SAS), a plug-and-play module that mitigates the influence of these attributes on prediction, seamlessly integrating into various Parameter-Efficient Fine-Tuning (PEFT) methods. In experiments, SAP and SAS significantly enhance accuracy on distribution shifts across 11 datasets and 3 generalization tasks without compromising downstream performance, establishing a new state-of-the-art benchmark.
Shu Zou, Zhaoyuan Yang, Mengqi He, Jing Zhang 0052
ICLR3
2025 A Physical Attack for Segmentation-Based Visual Foundation Models
abstract
We formalize a threat model for general patch-based attacks on visual foundation models (VFMs), highlighting the strong assumptions made by previous approaches regarding the proximity of adversarial patches to objects of interest. We demonstrate that these methods often fail to generalize to physical domains. To address these limitations, we propose a patch-based physical attack method that targets VFMs without relying on these assumptions. By operating in the feature space of VFMs and excluding the adversarial patch in the optimization objective, we ensure a more robust attack. Our method achieves the first successful physical adversarial attacks on segmentation-based VFMs, with empirical results showing that transformer-based segmentation VFMs, including SAM, SEEM, and CLIPSeg, are vulnerable to our attack in real-world scenarios.
Mengqi He, Jinhong Ni, Zhaoyuan Yang, Yiwei Fu, John N. Karigiannis
IJCNN3
2024 Adversarial Purification with the Manifold Hypothesis
abstract
In this work, we formulate a novel framework for adversarial robustness using the manifold hypothesis. This framework provides sufficient conditions for defending against adversarial examples. We develop an adversarial purification method with this framework. Our method combines manifold learning with variational inference to provide adversarial robustness without the need for expensive adversarial training. Experimentally, our approach can provide adversarial robustness even if attackers are aware of the existence of the defense. In addition, our method can also serve as a test-time defense mechanism for variational autoencoders.
Zhaoyuan Yang, Jing Zhang 0052, Richard I. Hartley, Peter H. Tu
AAAI1
2024 Grounded Language Acquisition from Object and Action Imagery
abstract
Deep learning approaches to natural language processing have made great strides in recent years. While transformer models have demonstrated impressive knowledge and reasoning capabilities, it is unclear how produced symbols are grounded in data from the world. In this paper, we explore the development of a private language for visual data representation by training emergent language (EL) encoders/decoders in both i) a traditional referential game environment and ii) a contrastive learning environment utilizing a within-class matching training paradigm. An additional classification layer-utilizing neural machine translation and random forest classification-was used to transform symbolic representations (sequences of integer symbols) to class labels. These methods were applied in two experiments focusing on object recognition and action recognition. For object recognition, a set of sketches produced by human participants from real imagery was used and for action recognition, 2D trajectory images were generated from 3D motion capture systems. In order to interpret the symbols produced for data in each experiment, a Gradient-weighted Class Activation Mapping (GradCAM) method was used to identify pixel regions indicating semantic features which contribute evidence towards symbols in learned languages. Results indicate that: i) symbols used to represent images appear to shift focus between different semantic components in an image and ii) this shift occurs gradually over the course of a sentence.
James Kubricht, Zhaoyuan Yang, Jianwei Qiu, Peter H. Tu
AVSS2
2024 ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
abstract
Although soft prompt tuning is effective in efficiently adapting Vision-Language (V&L) models for downstream tasks, it shows limitations in dealing with distribution shifts. We address this issue with Attribute-Guided Prompt Tuning (ArGue), making three key contributions. 1) In contrast to the conventional approach of directly appending soft prompts preceding class names, we align the model with primitive visual attributes generated by Large language Models (LLMs). We posit that a model's ability to express high confidence in these attributes signifies its ca-pacity to discern the correct class rationales. 2) We intro-duce attribute sampling to eliminate disadvantageous at-tributes, thus only semantically meaningful attributes are preserved. 3) We propose negative prompting, explicitly enumerating class-agnostic attributes to activate spurious correlations and encourage the model to generate highly orthogonal probability distributions in relation to these neg-ative features. In experiments, our method significantly out-performs current state-of-the-art prompt tuning methods on both novel class prediction and out-of-distribution general-ization tasks. The code is available https://github.com/Liam-Tian/ArGue.
Shu Zou, Zhaoyuan Yang, Jing Zhang 0052
CVPR3
2024 IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models
abstract
We present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they have significant semantic difference. To bridge this gap, we interpolate in the locally linear and continuous text embedding space and Gaussian latent space. We first optimize the endpoint text embeddings and then map the images to the latent space using a probability flow ODE. Unlike existing work that takes an indirect morphing path, we show that the model adaptation yields a direct path and suppresses ghosting artifacts in the interpolated images. To achieve this, we propose a heuristic bottleneck constraint based on a novel relative perceptual path diversity score that automatically controls the bottleneck size and balances the diversity along the path with its directness. We also propose a perceptually-uniform sampling technique that enables visually smooth changes between the interpolated images. Extensive experiments validate that our IMPUS can achieve smooth, direct, and realistic image morphing and is adaptable to several other generative tasks.
Zhaoyuan Yang, Jing Zhang 0052, Dylan Campbell, Peter H. Tu, Richard I. Hartley
ICLR1
2024 DreamSteerer: Enhancing Source Image Conditioned Editability using Personalized Diffusion Models
abstract
Recent text-to-image (T2I) personalization methods have shown great premise in teaching a diffusion model user-specified concepts given a few images for reusing the acquired concepts in a novel context. With massive efforts being dedicated to personalized generation, a promising extension is personalized editing, namely to edit an image using personalized concepts, which can provide more precise guidance signal than traditional textual guidance. To address this, one straightforward solution is to incorporate a personalized diffusion model with a text-driven editing framework. However, such solution often shows unsatisfactory editability on the source image. To address this, we propose DreamSteerer, a plug-in method for augmenting existing T2I personalization methods. Specifically, we enhance the source image conditioned editability of a personalized diffusion model via a novel Editability Driven Score Distillation (EDSD) objective. Moreover, we identify a mode trapping issue with EDSD, and propose a mode shifting regularization with spatial feature guided sampling to avoid such issue. We further employ two key modifications on the Delta Denoising Score framework that enable high-fidelity local editing with personalized concepts. Extensive experiments validate that DreamSteerer can significantly improve the editability of several T2I personalization baselines while being computationally efficient.
Zhaoyuan Yang, Jing Zhang 0052
NeurIPS2
2020 Justification-Based Reliability in Machine Learning
abstract
With the advent of Deep Learning, the field of machine learning (ML) has surpassed human-level performance on diverse classification tasks. At the same time, there is a stark need to characterize and quantify reliability of a model's prediction on individual samples. This is especially true in applications of such models in safety-critical domains of industrial control and healthcare. To address this need, we link the question of reliability of a model's individual prediction to the epistemic uncertainty of the model's prediction. More specifically, we extend the theory of Justified True Belief (JTB) in epistemology, created to study the validity and limits of human-acquired knowledge, towards characterizing the validity and limits of knowledge in supervised classifiers. We present an analysis of neural network classifiers linking the reliability of its prediction on a test input to characteristics of the support gathered from the input and hidden layers of the network. We hypothesize that the JTB analysis exposes the epistemic uncertainty (or ignorance) of a model with respect to its inference, thereby allowing for the inference to be only as strong as the justification permits. We explore various forms of support (for e.g., k-nearest neighbors (k-NN) and ℓp-norm based) generated for an input, using the training data to construct a justification for the prediction with that input. Through experiments conducted on simulated and real datasets, we demonstrate that our approach can provide reliability for individual predictions and characterize regions where such reliability cannot be ascertained.
Nurali Virani, Naresh Iyer, Zhaoyuan Yang
AAAI3
2020 Variational Encoder-Based Reliable Classification
abstract
Machine learning models provide statistically impressive results which might be individually unreliable. To provide reliability, we propose an Epistemic Classifier (EC) that can provide justification of its belief using support from the training dataset as well as quality of reconstruction. Our approach is based on modified variational auto-encoders that can identify a semantically meaningful low-dimensional space where perceptually similar instances are close in $\ell_{2}-$distance too. Our results demonstrate improved reliability of predictions and robust identification of samples with adversarial attacks as compared to baseline of softmax-based thresholding.
Chitresh Bhushan, Zhaoyuan Yang, Nurali Virani, Naresh Iyer
ICIP2
2020 Adversarial Vulnerability in Doppler-based Human Activity Recognition
abstract
Human activity recognition (HAR) is an important task in many internet of things (IoT) applications. In recent years, significant efforts have been made towards achieving the highest possible recognition performance (accuracy and robustness) by using advanced machine learning techniques, including deep learning. However, to the best of our knowledge, the adversarial vulnerability of the Doppler sensor-based HAR systems has not been studied. In other domains such as computer vision, the vulnerability of deep learning algorithms to adversarial samples has attracted tremendous research interests in the past few years. In this work, we investigate the adversarial vulnerability of the Doppler-based human activity recognition system. Using a case study we demonstrate that the adversarial examples can significantly degrade the performance of the human activity recognition. Specifically, the basic iterative method (BIM) attack can reduce classification accuracy by as much as 85%. We also discuss different types of attacks, e.g., data poisoning attacks and potential strategies of protecting the Doppler-based HAR systems against adversarial attacks.
Zhaoyuan Yang, Yang Zhao 0020, Weizhong Yan
IJCNN1