Jiayi Gao

dblp:315/3227 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reference Attack: A New Cross-Modal Jailbreaking Attack against Multimodal Large Language Models
abstract
Red team testing, an effective proactive method for evaluating the security of multimodal large language models (MLLMs), requires an expanding toolkit alongside the development of MLLM safeguards.We propose the Reference Attack, a powerful tool for red team testing against MLLMs.The Reference Attack is a reference-guided cross-modal jailbreak method that enhances existing prompt-to-image injection attacks by exploiting MLLMs' semantic reconstruction capabilities.Our method embeds malicious prompts in non-text modalities (e.g., images, spreadsheets) and constructs recursive symbolic references in text, enabling MLLMs to gradually recover and generate harmful content through layered reference resolution.The attack introduces a new vector that circumvents conventional content moderation by exploiting MLLMs' lack of security checks during crossmodal reference resolution.We evaluate the Reference Attack on leading MLLMs, including ChatGPT, Gemini, Claude, and the widely used open-source LLaMA model, and achieved an attack success rate of over 93% across all tested models.Compared to state-of-the-art attacks, Reference Attack achieves higher success rates than all baselines under identical evaluation, with a maximum gain of 70.8%.Our study reveals a critical gap in MLLM security and highlights the need for strict security auditing of cross-modal interactions in future content moderation.
Yulong Wang 0001, Yifei Fu, Jiayi Gao
ACL (1)3
2026 UAV 3D Path Planning Based on Multi-strategy Collabo Rative Improved Coati Optimization Algorithm
Jiayi Gao, Lixin Mu
ICIC (13)1
2025 ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
abstract
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce ConMo, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo.
Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng 0001, Kongming Liang, Zhanyu Ma, Jun Guo 0002, Yang Liu 0105
CVPR1
2025 Probabilistic Person-in-Bed Detection Using Accelerometer Signals
abstract
Using accelerometer data in smart bed systems offers a cost-effective solution for person-in-bed detection. In this work, we propose a lightweight probabilistic model for this task. The accelerometer time series is first divided into multiple patches, with high-frequency noise filtered through a combined optimization of 1D convolution and spectral pooling. An LSTM-based feature extraction module is then employed to capture temporal dependencies. Subsequently, a Fourier classification head is applied to generate probabilistic detection outputs. The proposed architecture achieves an accuracy of 1.0 on the segmented detection task and 0.915 on the streaming detection task in the ICASSP 2025 signal Processing Grand Challenge, organized by the Analog Garage. The implementation is available at https://github.com/JiayiGao04/person-in-bed-detection.
Kaite Shi, Jiayi Gao, Xuliang Yu
ICASSP2
2025 A New Method for Detecting Cancer Driver Genes by Constructing a Heterogeneous Network with Test-Time Training from Multi-view
Jiayi Gao, Yuanhao Fan
ICIC (27)2
2025 Benchmarking Graph Foundation Models
abstract
In real-world applications, graph data has garnered significant attention for its representation and analysis using Graph Neural Networks. Recent advancements have led to the development of Graph Foundation Models (GFMs), which aim to enhance cross-domain and cross-task generalization ability. Despite promising results from GFMs, a lack of standardized evaluation processes hinders comparative analysis and cross-domain applicability. To address this gap, we propose GFMBench, an open-source pipeline that standardizes the training, evaluation, and deployment of GFMs across diverse real-world graph applications. GFMBench integrates state-of-the-art GFMs and datasets, providing a modular design for comprehensive support across data preprocessing, model training, and evaluation. The pipeline includes a robust evaluation framework for benchmarking GFM generalization ability, encompassing supervised learning, cross-domain zero-shot and few-shot learning, and in-context learning. To validate the usability of GFMs, we deploy them on the Open Academic Graph, enabling applications such as topic search and author recommendation. This work provides a unified benchmark for GFMs, enabling deeper insights into their generalization ability across various graph tasks and domains. We further open-source GFMBench https://github.com/BUPT-GAMMA/ggfm and related documents https://ggfm.readthedocs.io/en/latest/.
Liangwei Yang, Zeyuan Guo, Jiayi Gao, Tianhao Chai, Cheng Yang 0002, Chuan Shi 0001
KDD (2)4
2025 Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Jiayi Gao, Changcheng Hua, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105
ACM Multimedia1
2025 Evaluating and Mitigating Sycophancy in Large Vision-Language Models
abstract
Large vision-language models (LVLMs) have recently achieved significant advancements, demonstrating powerful capabilities in understanding and reasoning about visual information. However, LVLMs may generate biased responses that reflect the user beliefs rather than the facts, a phenomenon known as sycophancy. Sycophancy can pose serious challenges to the performance, trustworthiness, and security of LVLMs, raising concerns about their practical applications. We note that there is limited work on the evaluation and mitigation of sycophancy in LVLMs. In this paper, we introduce SyEval-VL, a benchmark specifically designed to evaluate sycophancy in LVLMs. SyEval-VL offers a comprehensive evaluation of sycophancy in visual understanding and reasoning across various scenarios with a multi-round dialogue format. We evaluate sycophancy in several popular LVLMs, providing an in-depth analysis of various sycophantic behaviors and their consequential impacts. Additionally, we propose a novel framework, Human Feedback-based Retrieval-Augmented Generation (HFRAG), to mitigate sycophancy in LVLMs by determining the appropriate timing of retrieval, profiling the proper retrieval target, and augmenting the decoding of LVLMs. Extensive experiments demonstrate that the proposed method significantly mitigates sycophancy in LVLMs without requiring additional training. Our code is available at: https://github.com/immc-lab/SyEval-VL
Jiayi Gao, Huaiwen Zhang
ACM Multimedia1
2024 Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection
abstract
Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN.
Jiayi Gao, Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002
AAAI1
2023 Topology Uncertainty Modeling For Imbalanced Node Classification on Graphs
abstract
Most existing graph neural networks work under a class-balanced assumption, while ignoring class-imbalanced scenarios that widely exist in real-world graphs. Although there are many methods in other fields that can alleviate this issue, they do not consider the special topology of the non-Euclidean graph. Hence, we propose Graph Topology Uncertainty (GraphTU), a novel probabilistic class-imbalanced solution specifically for graphs. Firstly, an invisible "uncertain gap" between under-represented minorities in training set and authentic minorities in unseen set is modeled by estimating statistical variances in topology. We extend the training distribution for minorities by sampling in this gap through a non-parametric way. Moreover, a gradient-guided mask is introduced to prevent biased statistics. Extensive experiments demonstrate the superior performance of GraphTU.
Jiayi Gao, Youyong Kong
ICASSP1
2022 An energy-efficiency-adaptive clustering formation mechanism for the wireless sensor networks
abstract
Abstract Energy inequality caused by the process of cluster head election has a large influence on energy efficiency and the network lifetime of wireless sensor networks (WSNs). To this end, a novel concept of EI ec is proposed to evaluate the equality degree of energy consumption. Related theorems for establishing the candidate set of cluster heads are proposed, with the aim of promoting energy equality in each cluster. Subsequently, a novel energy‐efficiency‐adaptive cluster formation mechanism based on economic (ECFE) theory is proposed and detailed. Finally, extensive experiments are carried out to assess its energy efficiency and the network performance by comparisons with the existing classic and latest intelligent clustering algorithms. The results indicate that ECFE improves not only the energy efficiency but also the network performance effectively.
Deyu Lin, Linghe Kong, Chengkun Zhao, Jiayi Gao, Hao Ouyang, Ziyuan Yang 0001, Zhiqiang Zhang 0001
IET Commun.4