Gengyuan Zhang

dblp:305/5662 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Pragmatic Jailbreak on Text-to-image Models
abstract
Tong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang, Shuo Chen, Philip Torr, Vera Demberg, Volker Tresp, Jindong Gu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tong Liu 0019, Zhixin Lai, Gengyuan Zhang, Shuo Chen 0014, Philip Torr 0001, Vera Demberg, Volker Tresp, Jindong Gu
ACL (1)4
2025 FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models
abstract
One-Shot Federated Learning (OSFL), a special decentralized machine learning paradigm, has recently gained significant attention. OSFL requires only a single round of client data or model upload, which reduces communication costs and mitigates privacy threats compared to traditional FL. Despite these promising prospects, existing methods face challenges due to client data heterogeneity and limited data quantity when applied to real-world OSFL systems. Recently, Latent Diffusion Models (LDM) have shown remarkable advancements in synthesizing high-quality images through pretraining on large-scale datasets, thereby presenting a potential solution to overcome these issues. However, directly applying pretrained LDM to heterogeneous OSFL results in significant distribution shifts in synthetic data, leading to performance degradation in classification models trained on such data. This issue is particularly pronounced in rare domains, such as medical imaging, which are underrepresented in LDM’s pretraining data. To address this challenge, we propose Federated Bi-Level Personalization (FedBiP), which personalizes the pretrained LDM at both instance-level and concept-level. Hereby, FedBiP synthesizes images following the client’s local data distribution without compromising the privacy regulations. FedBiP is also the first approach to simultaneously address feature space heterogeneity and client data scarcity in OSFL. Our method is validated through extensive experiments on three OSFL benchmarks with feature space heterogeneity, as well as on challenging medical and satellite image datasets with label heterogeneity. The results demonstrate the effectiveness of FedBiP, which substantially outperforms other OSFL methods. Our code is available at https://github.com/HaokunChen245/FedBiP.
Hang Li 0010, Jinhe Bi, Gengyuan Zhang, Philip Torr 0001, Jindong Gu, Denis Krompass, Volker Tresp
CVPR5
2025 Localizing Events in Videos with Multimodal Queries
abstract
Localizing events in videos based on semantic queries is a pivotal task in video understanding research and user-oriented applications like video search. Yet, current research predominantly relies on natural language queries (NLQs), overlooking the potential of using multimodal queries (MQs) that incorporate images to flexibly represent semantic queries, particularly when it is difficult to express non-verbal or unfamiliar concepts in words. To bridge this gap, we introduce ICQ, a new benchmark designed for localizing events in videos with MQs, alongside an evaluation dataset ICQ-Highlight. To adapt and reevaluate existing video localization models for this new task, we propose 3 Multimodal Query Adaptation methods and a novel Surrogate Fine-Tuning strategy, serving as strong baseline methods. ICQ systematically benchmarks 12 state-of-the-art backbone models, spanning from specialized video localization models to Video Large Language Models. Our extensive experiments highlight the high potential of using MQs in real-world applications. We believe this is a first step toward video event localization with MQs1.
Gengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia 0003, Daniel Cremers, Philip Torr 0001, Volker Tresp, Jindong Gu
CVPR1
2025 Perceive. Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
abstract
Video Question Answering (Video QA) is a challenging video understanding task that requires models to compre-hend entire videos, identify the most relevant information based on contextual cues from a given question, and rea-son accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have trans-formed video QA by leveraging their exceptional common-sense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an ad-ditional space-time alignment poses a considerable chal-lenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA bench-marks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with re-cent advancements in video QA.
Roberto Amoroso, Gengyuan Zhang, Rajat Koner, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp
WACV2
2025 CL-Cross VQA: A Continual Learning Benchmark for Cross-Domain Visual Question Answering
abstract
Visual Question Answering (VQA) systems witnessed a significant advance in recent years due to the development of large-scale Vision-Language Pre-trained Models (VLPMs). As the application scenario and user demand change over time, an advanced VQA system is expected to be capable of continuously expanding its knowledge and capabilities over time, not only to handle new tasks (i.e., new question types or visual scenes) but also to answer questions in new specialized domains without forgetting previously acquired knowledge and skills. Existing works studying CL on VQA tasks primarily consider answer-and question-type incremental learning or sceneand function-incremental learning, whereas how VQA systems perform when they encounter new domains and increasing user demands has not been studied. Motivated by this, we introduce CL-CrossVQA, a rigorous Continual Learning benchmark for Cross-domain Visual Question Answering, through which we conduct extensive experiments on 4 VLPMs, 5 CL approaches, and 5 VQA datasets from different domains. In addition, by probing the forgetting phenomenon of the intermediate layers, we provide insights into how model architecture affects CL performance, why CL approaches can help mitigate forgetting in VLPMs, and how to design CL approaches suitable for VLPMs in this challenging continual learning environment. To facilitate future work on developing an advanced All-in-One VQA system, we will release our datasets and code.
Ahmed Frikha 0002, Denis Krompass, Gengyuan Zhang, Jindong Gu, Volker Tresp
WACV5
2024 RPF-ELD: Regional Prior Fusion using Early and Late Distillation for Breast Cancer Recognition in Ultrasound Images
abstract
Breast cancer is one of the main factors responsible for the deaths of women worldwide. Ultrasound imaging is a key method for early detection of breast cancer, which can help patients gain valuable treatment time and improve their chances of survival. The computer-aided system of breast cancer recognition has started to receive attention due to the lack of experienced sonographers. Presently, most breast cancer recognition methods typically suffer from uncertain locations and proportions of tumor regions in ultrasound images. In this paper, we propose a novel Regional Prior Fusion framework using Early and Late Distillation (RPF-ELD), inspired by the knowledge distillation of the teacher-student framework, for breast cancer recognition in ultrasound images. Firstly, to enhance the concentration of the tumor regions, a high-performing prior-fused model is trained as the teacher model using ultrasound images with the corresponding regional prior information. Next, a diagnostic model is trained as the student model under the prior-fused model distillation using early and late features to implicitly obtain the regional prior knowledge. Finally, the diagnostic model recognizes the categories of breast cancer from only ultrasound images using the experience from distilled prior knowledge. Two publicly released datasets are used to evaluate the proposed RPF-ELD framework. Experimental results demonstrate that the proposed RPF-ELD surpasses current state-of-the-art methods.
Gengyuan Zhang, Fang Lai, Wenwei Cui, Jiexiao Xue, Hao Zhang 0128
BIBM2
2024 Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
abstract
Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is taken based on their knowledge. This makes us wonder if, based on visual cues, Vision-Language Models that are pre-trained with large-scale image-text resources can achieve and even surpass human capability in reasoning times and location. To address this question, we propose a two-stage Recognition & Reasoning probing task applied to discriminative and generative VLMs to uncover whether VLMs can recognize times and location-relevant features and further reason about it. To facilitate the studies, we introduce WikiTiLo, a well-curated image dataset compromising images with rich socio-cultural cues. In extensive evaluation experiments, we find that although VLMs can effectively retain times and location-relevant features in visual encoders, they still fail to make perfect reasoning with context-conditioned visual features. The dataset is available at https://github.com/gengyuanmax/WikiTiLo.
Gengyuan Zhang, Yurui Zhang, Kerui Zhang, Volker Tresp
WACV1
2023 Multi-event Video-Text Retrieval
abstract
Video-Text Retrieval (VTR) is a crucial multi-modal task in an era of massive video-text data on the Internet. A plethora of work characterized by using a two-stream Vision-Language model architecture that learns a joint representation of video-text pairs has become a prominent approach for the VTR task. However, these models operate under the assumption of bijective video-text correspondences and neglect a more practical scenario where video content usually encompasses multiple events, while texts like user queries or webpage metadata tend to be specific and correspond to single events. This establishes a gap between the previous training objective and real-world applications, leading to the potential performance degradation of earlier models during inference. In this study, we introduce the Multi-event Video-Text Retrieval (MeVTR) task, addressing scenarios in which each video contains multiple different events, as a niche scenario of the conventional Video-Text Retrieval Task. We present a simple model, Me-Retriever, which incorporates key event video representation and a new MeVTR loss for the MeVTR task. Comprehensive experiments show that this straightforward framework outperforms other models in the Video-to-Text and Text-to-Video tasks, effectively establishing a robust baseline for the MeVTR task. We believe this work serves as a strong foundation for future studies. Code is available at https://github.com/gengyuanmax/MeVTR.
Gengyuan Zhang, Jisen Ren, Jindong Gu, Volker Tresp
ICCV1
2021 Time-dependent Entity Embedding is not All You Need: A Re-evaluation of Temporal Knowledge Graph Completion Models under a Unified Framework
abstract
Various temporal knowledge graph (KG) completion models have been proposed in the recent literature.The models usually contain two parts, a temporal embedding layer and a score function derived from existing static KG modeling approaches.Since the approaches differ along several dimensions, including different score functions and training strategies, the individual contributions of different temporal embedding techniques to model performance are not always clear.In this work, we systematically study six temporal embedding approaches and empirically quantify their performance across a wide range of configurations with about 4000 experiments and 19000 GPU hours.We classify the temporal embeddings into two classes: (1) timestamp embeddings and (2) time-dependent entity embeddings.Despite the common belief that the latter is more expressive, an extensive experimental study shows that timestamp embeddings can achieve on-par or even better performance with significantly fewer parameters.Moreover, we find that when trained appropriately, the relative performance differences between various temporal embeddings often shrink and sometimes even reverse when compared to prior results.For example, TTransE (Leblay and Chekol, 2018), one of the first temporal KG models, can outperform more recent architectures on ICEWS datasets.To foster further research, we provide the first unified open-source framework for temporal KG completion models with full composability, where temporal embeddings, score functions, loss functions, regularizers, and the explicit modeling of reciprocal relations can be combined arbitrarily.
Zhen Han 0003, Gengyuan Zhang, Yunpu Ma, Volker Tresp
EMNLP (1)2