Qizhou Chen

dblp:376/3942 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-6434-0301ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering
abstract
Taolin Zhang, Dongyang Li, Chen Chen, Qizhou Chen, Jiuheng Wan, Xiaofeng He, Chengyu Wang, Richang Hong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Taolin Zhang 0001, Qizhou Chen, Jiuheng Wan, Chengyu Wang 0001, Richang Hong
ACL (1)4
2026 Taming "Zombie" Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolution
abstract
Taolin Zhang, Pukun Zhao, Qizhou Chen, Jiuheng Wan, Chen Chen, Xiaofeng He, Chengyu Wang, Richang Hong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Taolin Zhang 0001, Pukun Zhao, Qizhou Chen, Jiuheng Wan, Chengyu Wang 0001, Richang Hong
ACL (1)3
2025 Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
abstract
Model editing aims to correct outdated or erroneous knowledge in large models without costly retraining. Recent research discovered that the mid-layer representation of the subject's final token in a prompt has a strong influence on factual predictions, and developed Large Language Model (LLM) editing techniques based on this observation. However, for Vision-LLMs (VLLMs), how visual representations impact the predictions from a decoder-only language model remains largely unexplored. To the best of our knowledge, model editing for VLLMs has not been extensively studied in the literature. In this work, we employ the contribution allocation and noise perturbation methods to measure the contributions of visual representations for token predictions. Our attribution analysis shows that visual representations in mid-to-later layers that are highly relevant to the prompt contribute significantly to predictions. Based on these insights, we propose *VisEdit*, a novel model editor for VLLMs that effectively corrects knowledge by editing intermediate visual representations in regions important to the edit prompt. We evaluated *VisEdit* using multiple VLLM backbones and public VLLM editing benchmark datasets. The results show the superiority of *VisEdit* over the strong baselines adapted from existing state-of-the-art editors for LLMs.
Qizhou Chen, Taolin Zhang 0001, Chengyu Wang 0001, Dakan Wang
AAAI1
2025 BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering
abstract
Multi-hop question answering (QA) involves finding multiple relevant passages and performing step-by-step reasoning to answer complex questions.Previous works on multi-hop QA employ specific methods from different modeling perspectives based on large language models (LLMs), regardless of question types.In this paper, we first conduct an in-depth analysis of public multi-hop QA benchmarks, categorizing questions into four types and evaluating five types of cutting-edge methods: Chainof-Thought (CoT), Single-step, Iterative-step, Sub-step, and Adaptive-step.We find that different types of multi-hop questions exhibit varying degrees of sensitivity to different types of methods.Thus, we propose a Bi-levEL muLti-agEnt reasoning (BELLE) framework to address multi-hop QA by specifically focusing on the correspondence between question types and methods, with each type of method regarded as an "operator" by prompting LLMs differently.The first level of BELLE includes multiple agents that debate to formulate an executable plan of combined "operators" to address the multi-hop QA task comprehensively.During the debate, in addition to the basic roles of affirmative debater, negative debater, and judge, at the second level, we further leverage fast and slow debaters to monitor whether changes in viewpoints are reasonable.Extensive experiments demonstrate that BELLE significantly outperforms strong baselines in various datasets.Additionally, the model consumption of BELLE is higher cost-effectiveness than that of single models in more complex multihop QA scenarios.(B) Single-step [Multi-Hop Question:] What was the former band of the member of Mother Love Bone who died just before the release of Apple?[Answer:] Malfunkshun Multi-Hop Question Retrieval Docs Multi-Hop Answer (C) Iterative-step (D) Sub-step Multi-Hop Question (Intermediate) K times Multi-Hop Answer Multi-Hop Question 1 2 3 A1 A2 A3 Multi-Hop Answer (E) Adaptive-step Multi-Hop Question Classifier (D) Sub-Step (B) Single-step (A) CoT ✔ (A) CoT (C) Iterative-step Let's think step-by-step Multi-Hop Question Multi-Hop Question Retrieval-augmented Reasoning Closed-book Reasoning Our Agent-Based Reasoning Operators Pool …… Multi-Hop Question Inference Comparison Temporal Null 1. Use Sub-Step to decompose query 2. Use Single-Step to retrieve subquery 3. Aggregate the sub-answer Execution Plan Agents Multi-hop QA Task Environ -ment interaction invoke solve
Taolin Zhang 0001, Qizhou Chen, Chengyu Wang 0001
ACL (1)3
2025 Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts
abstract
Model editing aims to correct inaccurate knowledge, update outdated information, and incorporate new data into Large Language Models (LLMs) without the need for retraining. This task poses challenges in lifelong scenarios where edits must be continuously applied for real-world applications. While some editors demonstrate strong robustness for lifelong editing in pure LLMs, Vision LLMs (VLLMs), which incorporate an additional vision modality, are not directly adaptable to existing LLM editors. In this paper, we propose LiveEdit, a Lifelong vision language model Edit to bridge the gap between lifelong LLM editing and VLLMs. We begin by training an editing expert generator to independently produce low-rank experts for each editing instance, with the goal of correcting the relevant responses of the VLLM. A hard filtering mechanism is developed to utilize visual semantic knowledge, thereby coarsely eliminating visually irrelevant experts for input queries during the inference stage of the post-edited model. Finally, to integrate visually relevant experts, we introduce a soft routing mechanism based on textual semantic relevance to achieve multi-expert fusion. For evaluation, we establish a benchmark for lifelong VLLM editing. Extensive experiments demonstrate that LiveEdit offers significant advantages in lifelong VLLM editing scenarios. Further experiments validate the rationality and effectiveness of each module design in LiveEdit.1
Qizhou Chen, Chengyu Wang 0001, Dakan Wang, Taolin Zhang 0001, Wangyue Li
CVPR1
2025 UniEdit: A Unified Knowledge Editing Benchmark for Large Language Models
abstract
Model editing aims to efficiently revise incorrect or outdated knowledge within LLMs without incurring the high cost of full retraining and risking catastrophic forgetting. Currently, most LLM editing datasets are confined to narrow knowledge domains and cover a limited range of editing evaluation. They often overlook the broad scope of editing demands and the diversity of ripple effects resulting from edits. In this context, we introduce \uniedit, a unified benchmark for LLM editing grounded in open-domain knowledge. First, we construct editing samples by selecting entities from 25 common domains across five major categories, utilizing the extensive triple knowledge available in open-domain knowledge graphs to ensure comprehensive coverage of the knowledge domains. To address the issues of generality and locality in editing, we design an Neighborhood Multi-hop Chain Sampling (NMCS) algorithm to sample subgraphs based on a given knowledge piece to entail comprehensive ripple effects to evaluate. Finally, we employ proprietary LLMs to convert the sampled knowledge subgraphs into natural language text, guaranteeing grammatical accuracy and syntactical diversity. Extensive statistical analysis confirms the scale, comprehensiveness, and diversity of our \uniedit benchmark. We conduct comprehensive experiments across multiple LLMs and editors, analyzing their performance to highlight strengths and weaknesses in editing across open knowledge domains and various evaluation criteria, thereby offering valuable insights for future research endeavors.
Qizhou Chen, Dakan Wang, Taolin Zhang 0001, Zaoming Yan, Chengsong You, Chengyu Wang 0001
NeurIPS1
2025 Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large Motion
abstract
Motion in the real world takes place in 3D space. Existing Frame Interpolation methods often estimate global receptive fields in 2D frame space. Due to the limitations of 2D space, these global receptive fields are limited, which makes it difficult to match object correspondences between frames, resulting in sub-optimal performance when handling large-motion scenarios. In this paper, we introduce a novel pipeline for exploring object correspondences based on differential surface theory. The differential surface coordinate system provides a better representation of the real world, enabling effective exploration of object correspondences. Specifically, the pipeline first transforms an input pair of video frames from the image coordinate system to the differential surface coordinate system. Subsequently, within this coordinate system, object correspondences are explored based on surface geometric properties and the surface uniqueness theorem. Experimental findings showcase that our method attains state-of-the-art performance across large motion benchmarks. Our method demonstrates the state-of-the-art performance on these VFI subsets with large motion.
Zaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen, Guixu Zhang, Faming Fang
NeurIPS4
2024 R4: Reinforced Retriever-Reorder-Responder for Retrieval-Augmented Large Language Models
abstract
Retrieval-augmented large language models (LLMs) leverage relevant content retrieved by information retrieval systems to generate correct responses, aiming to alleviate the hallucination problem. However, existing retriever-responder methods typically append relevant documents to the prompt of LLMs to perform text generation tasks without considering the interaction of fine-grained structural semantics between the retrieved documents and the LLMs. This issue is particularly important for accurate response generation as LLMs tend to “lose in the middle” when dealing with input prompts augmented with lengthy documents. In this work, we propose a new pipeline named “Reinforced Retriever-Reorder-Responder” (R4) to learn document orderings for retrieval-augmented LLMs, thereby further enhancing their generation abilities while the large numbers of parameters of LLMs remain frozen. The reordering learning process is divided into two steps according to the quality of the generated responses: document order adjustment and document representation enhancement. Specifically, document order adjustment aims to organize retrieved document orderings into beginning, middle, and end positions based on graph attention learning, which maximizes the reinforced reward of response quality. Document representation enhancement further refines the representations of retrieved documents for responses of poor quality via document-level gradient adversarial learning. Extensive experiments demonstrate that our proposed pipeline achieves better factual question-answering performance on knowledge-intensive tasks compared to strong baselines across various public datasets. The source codes and trained models will be released upon paper acceptance.
Taolin Zhang 0001, Qizhou Chen, Chengyu Wang 0001, Longtao Huang, Hui Xue 0001, Jun Huang 0007
ECAI3
2024 Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt Learning
abstract
Model editing aims to correct outdated or erroneous knowledge in large language models (LLMs) without the need for costly retraining.Lifelong model editing is the most challenging task that caters to the continuous editing requirements of LLMs.Prior works primarily focus on single or batch editing; nevertheless, these methods fall short in lifelong editing scenarios due to catastrophic knowledge forgetting and the degradation of model performance.Although retrieval-based methods alleviate these issues, they are impeded by slow and cumbersome processes of integrating the retrieved knowledge into the model.In this work, we introduce RECIPE, a RetriEval-augmented ContInuous Prompt lEarning method, to boost editing efficacy and inference efficiency in lifelong learning.RECIPE first converts knowledge statements into short and informative continuous prompts, prefixed to the LLM's input query embedding, to efficiently refine the response grounded on the knowledge.It further integrates the Knowledge Sentinel (KS) that acts as an intermediary to calculate a dynamic threshold, determining whether the retrieval repository contains relevant knowledge.Our retriever and prompt encoder are jointly trained to achieve editing properties, i.e., reliability, generality, and locality.In our experiments, RECIPE is assessed extensively across multiple LLMs and editing datasets, where it achieves superior editing performance.RECIPE also demonstrates its capability to maintain the overall performance of LLMs alongside showcasing fast editing and inference speed.
Qizhou Chen, Taolin Zhang 0001, Chengyu Wang 0001, Longtao Huang, Hui Xue 0001
EMNLP1
2024 Single image super-resolution based on trainable feature matching attention network
Qizhou Chen, Qing Shao
Pattern Recognit.1