Qirui Chen

dblp:293/7016 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpecCache: Speculative KV Cache Reuse for Efficient RAG Serving
abstract
Zijian Wen, Tao Zhang, Shuangwu Chen, Shenghao Ye, Yu Guo, Qirui Chen, Jingxian Shuai, Yunpeng Hou, Huasen He, Jianyang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zijian Wen, Tao Zhang 0170, Shuangwu Chen, Shenghao Ye, Qirui Chen, Jingxian Shuai, Yunpeng Hou, Huasen He, Jian Yang 0014
ACL (1)6
2026 ElecThinker: A Three-Stage Framework to Enhance Electronic Diagram Reasoning in Multimodal LLMs
Qirui Chen, Qirui Bai, Dong Jin 0004, Jiangming Li, Shuangwu Chen, Shenghao Ye, Wangming Li
ISCAS1
2026 LogiDiag: Diagnostic Planner-Guided Reasoning With LLMs for Logical Anomaly Diagnosis
abstract
Logical anomalies occur when a product's assembly violates prescribed logical rules, which widely exist in industrial assembly and packaging processes. Due to the difficulty in comprehending such complex logical relationships, a paucity of research has focused on the industrial logical anomaly diagnosis (LAD). Recently, large language models (LLMs) have demonstrated strong semantic understanding and zero-shot reasoning capabilities, making them a promising tool for LAD. However, directly applying LLMs to LAD still face two critical challenges: 1) the scarcity of abnormal samples in real-world settings, and 2) the propensity of LLMs to generate hallucinated or unreliable diagnostic conclusions. To address these challenges, we propose LogiDiag, a novel diagnostic reasoning method for LAD, to pinpoint where and why a product fails to comply with the logical rules, thereby elevating product quality and reducing remedial intervention cost. We design a visual descriptor that identifies product component attributes even in out-of-distribution abnormal images and organizes them into component descriptions for LLMs' comprehension. To mitigate hallucinations, we devise a planner to guide LLMs in diagnostic reasoning through rule orchestration, tool allocation, and chain-of-diagnosis generation. Experimental results on multiple benchmark datasets validate the competitive performance of LogiDiag.
Qirui Bai, Shuangwu Chen, Dong Jin 0004, Qirui Chen, Xiaobin Tan, Jian Yang 0014
IEEE Trans. Ind. Informatics5
2025 Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
abstract
This paper considers the problem of Multi-Hop Video Question Answering (MH-VidQA) in long-form egocentric videos. This task not only requires to answer visual questions, but also to localize multiple relevant time intervals within the video as visual evidences. We develop an automated pipeline to create multi-hop question-answering pairs with associated temporal evidence, enabling to construct a large-scale dataset for instruction-tuning. To monitor the progress of this new task, we further curate a high-quality benchmark, MULTIHOP-EGOQA, with careful manual verification and refinement. Experimental results reveal that existing multimodal systems exhibit inadequate multi-hop grounding and reasoning abilities, resulting in unsatisfactory performance. We then propose a novel architecture, termed as Grounding Scattered Evidence with Large Language Model (GeLM), that enhances multi-modal large language models by incorporating a grounding module to retrieve temporal evidence from videos using flexible grounding tokens. Trained on our visual instruction-tuning data, GeLM demonstrates improved multi-hop grounding and reasoning capabilities, setting a baseline for this new task. Furthermore, when trained on third-person view videos, the same architecture also achieves state-of-the-art performance on the single-hop VidQA benchmark, ActivityNet-RTL, demonstrating its effectiveness.
Qirui Chen, Shangzhe Di, Weidi Xie
AAAI1
2025 Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
abstract
This paper tackles the problem of video question answering (VideoQA), a task that often requires multi-step reasoning and a profound understanding of spatial-temporal dynamics. While large video-language models perform well on benchmarks, they often lack explainability and spatial temporal grounding. In this paper, we propose Agent-of-Thoughts Distillation (AoTD), a method that enhances models by incorporating automatically generated Chain- of-Thoughts (CoTs) into the instruction-tuning process. Specifically, we leverage an agent-based system to decompose complex questions into sub-tasks, and address them with specialized vision models, the intermediate results are then treated as reasoning chains. We also introduce a verification mechanism using a large language model (LLM) to ensure the reliability of generated CoTs. Extensive experiments demonstrate that AoTD improves the performance on multiple-choice and open-ended benchmarks.
Yudi Shi, Shangzhe Di, Qirui Chen, Weidi Xie
CVPR3
2025 Object-Centric Video Question Answering with Visual Grounding and Referring
abstract
Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multiround interactions. In this paper, we make three contributions: (i) we address these limitations by introducing a VideoLLM model, capable of performing both object referring for input and grounding for output in video reasoning tasks, i.e., allowing users to interact with videos using both textual and visual prompts; (ii) we propose STOM (Spatial-Temporal Overlay Module), a novel approach that propagates arbitrary visual prompts input at any single timestamp to the remaining frames within a video; (iii) we present VideoInfer, a manually curated object-centric video instruction dataset featuring questionanswering pairs that require reasoning. We conduct comprehensive experiments on VideoInfer and other existing benchmarks across video question answering and referring object segmentation. The results on 12 benchmarks of 6 tasks show that our proposed model consistently outperforms baselines in both video question answering and segmentation, underscoring its robustness in multimodal, object-centric video and image understanding. Project page: https://qirui-chen.github.io/RGA3-release/.
Qirui Chen, Cilin Yan, Jiayin Cai, Yao Hu 0002, Weidi Xie, Stratis Gavves
ICCV2
2025 Learning Streaming Video Representation via Multitask Training
abstract
Understanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streaming video understanding requires the ability to process video streams frame by frame, preserve historical information, and make low-latency decisions. To address these challenges, our main contributions are three-fold. (i) We develop a novel streaming video backbone, termed as StreamFormer, by incorporating causal temporal attention into a pre-trained vision transformer. This enables efficient streaming video processing while maintaining image representation capability. (ii) To train StreamFormer, we propose to unify diverse spatial-temporal video understanding tasks within a multitask visual-language alignment framework. Hence, StreamFormer learns global semantics, temporal dynamics, and fine-grained spatial relationships simultaneously. (iii) We conduct extensive experiments on online action detection, online video instance segmentation, and video question answering. StreamFormer achieves competitive results while maintaining efficiency, demonstrating its potential for real-time applications.
Yibin Yan, Jilan Xu, Shangzhe Di, Yudi Shi, Qirui Chen, Yifei Huang 0002, Weidi Xie
ICCV6
2024 Multi-sentence Grounding for Long-Term Instructional Video
Qirui Chen, Tengda Han, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ECCV (56)2
2023 Defining and Quantifying the Emergence of Sparse Concepts in DNNs
abstract
This paper aims to illustrate the concept-emerging phenomenon in a trained DNN. Specifically, we find that the inference score of a DNN can be disentangled into the effects of a few interactive concepts. These concepts can be understood as causal patterns in a sparse, symbolic causal graph, which explains the DNN. The faithfulness of using such a causal graph to explain the DNN is theoretically guaranteed, because we prove that the causal graph can well mimic the DNN's outputs on an exponential number of different masked samples. Besides, such a causal graph can be further simplified and re-written as an And-Or graph (AOG), without losing much explanation accuracy. The code is released at https://github.com/sjtu-xai-lab/aog.
Jie Ren 0018, Qirui Chen, Huiqi Deng, Quanshi Zhang
CVPR3
2023 Can We Faithfully Represent Absence States to Compute Shapley Values on a DNN?
Jie Ren 0018, Zhanpeng Zhou, Qirui Chen, Quanshi Zhang
ICLR3