Fanrui Zhang

dblp:321/8347 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-1078-430XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education.
Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang
AAAI3
2026 MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences
abstract
Zizhen Li, Chuanhao Li, Yibin Wang, Jianwen Sun, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Yifei Huang, Kaipeng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zizhen Li, Chuanhao Li 0001, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang
ACL (1)7
2026 Navigating Truth in Multimodal Fact-checking via Retrieval- and Reasoning-Enhanced Large Language Models
abstract
Recent studies show that claims incorporating both text and images spread more effectively than those with text alone, presenting significant challenges for multimodal fact-checking. The rapid development of Multi-modal Large Language Models (MLLMs) has greatly advanced research in this field, enabling stronger performance. However, existing MLLM-based fact-checking methods fail to fully exploit visual evidence, and their reliance on rigid fine-tuning templates limits context-aware explanations and leads to weak deep reasoning. To address these limitations, we propose FACTCOMPASS, a novel framework that combines reasoning-aware fine-tuning with large-scale rule-based reinforcement learning and incorporates a semantic- and knowledge-enhanced retrieval module to strengthen deep reasoning and improve evidence utilization. This framework enhances evidence retrieval by obtaining semantically relevant evidence images, enriching the contextual understanding of claim-related images, and refining textual evidence at the knowledge level. To further enhance reasoning, we introduce a self-refining reinforcement fine-tuning strategy: (1) distilling GPT-4o's reasoning from partially fact-checking data for cold-start Chain-of-Thought learning; (2) activating reasoning across broader datasets using prior knowledge and rejection sampling; (3) applying Group Relative Policy Optimization to explore diverse reasoning paths and optimize factual consistency. Extensive experiments have demonstrated the effectiveness of the proposed framework.
Fanrui Zhang, Qiang Zhang 0051, Chuanhao Li 0001, Jiaxin Ai, Yukang Feng, Zizhen Li, Kaipeng Zhang, Jiawei Liu 0001, Zhengjun Zha
WWW1
2026 CARE: A clinical agentic reasoning engine to enhance real-World diagnostic accuracy via structured medical reasoning
abstract
Recent advancements in large language models (LLMs) have improved performance on standardized medical benchmarks. However, existing benchmarks often rely on truncated context and idealized scenarios. In practice, clinical diagnosis requires synthesizing patient history, physical examination, laboratory tests, and imaging under time constraints, and LLMs can struggle with accuracy and may hallucinate when confronted with authentic cases. To address this challenge, we propose the Clinical Agentic Reasoning Engine (CARE) , a physician-inspired workflow that structures diagnosis into retrieval, preliminary diagnosis, final diagnosis, and confidence-gated recheck, with intermediate outputs serialized in JSON for verifiable, training-free inference. Using CARE as a data-generation pipeline with clinician-prepared diagnostic criteria, we curate 2,000 de-identified real-world cases across 15 abdominal disease categories and produce stepwise CARE annotations under a fixed schema. We adopt Dual-stage Alignment for Reasoning Enhancement (DARE) , which trains on these CARE-annotated trajectories, uses supervised fine-tuning on physician-structured long-form demonstrations and then applies group relative policy optimization to refine policy and promote self-correction. Finally, we introduce CARE-Dx , the resulting diagnosis model obtained by applying DARE to an instruction-tuned backbone, while CARE also remains a training-free inference protocol that can be applied to other LLMs. Experiments show that, under DARE, CARE-Dx achieves strong in-domain and zero-shot out-of-domain performance and approaches leading closed-source accuracy on evaluations, with clinician assessment by 12 experienced physicians from multiple departments confirming that its reasoning aligns with real clinical workflows. Moreover, on the de-identified private cohort Rui-EHR , which covers a broader set of diseases, the CARE pipeline maintains diagnostic quality.
Chenrun Wang, Xingqi He, Angela Lin Wang, Mengzhe Xu, Fanrui Zhang, Kailing Wang, Renhao Yang, Shiyi Yao, Yidong Xu, Carlos Gutiérrez SanRomán
Expert Syst. Appl.8
2025 Multi-granularity and Multi-modal Prompt Learning for Person Re-Identification
Jiawei Liu 0001, Guozhi Zhao, Fanrui Zhang, Zhengjun Zha
CVM (3)5
2025 Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation
abstract
Test-time adaptation using vision- language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single- task test- time adaptation scenarios and have not effectively explored the issue of multi- task adaptation. To address this practical problem, we propose a novel Hierarchical Knowledge Prompt Tuning (HKPT) method, which achieves joint adaptation to multiple target domains by mining more comprehensive source domain discriminative knowledge and hierarchically modeling task- specific and task- shared knowledge. Specifically, HKPT constructs a CLIP prompt distillation framework that utilizes the broader source domain knowledge of large teacher CLIP to guide prompt tuning for lightweight student CLIP from multiple views during testing. Meanwhile, HKPT establishes task- specific dual dynamic knowledge graph to capture fine- grained contextual knowledge from continuous test data. To fully exploit the complementarity among multiple target tasks, HKPT employs an adaptive task grouping strategy for achieving intertask knowledge sharing. Furthermore, HKPT can seamlessly transfer to basic single- task test- time adaptation scenarios while maintaining robust performance. Extensive experimental results in both multi- task and single- task testtime adaptation settings demonstrate that our HKPT significantly outperforms state- of- the- art methods.
Qiang Zhang 0051, Mengsheng Zhao, Jiawei Liu 0001, Fanrui Zhang, Yongchao Xu, Zhengjun Zha
CVPR4
2025 InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
abstract
Zizhen Li, Chuanhao Li, Yibin Wang, Qi Chen, Diping Song, Yukang Feng, Jianwen Sun, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zizhen Li, Chuanhao Li 0001, Diping Song, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang
EMNLP9
2025 ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges
Jiaxin Ai, Zhaopan Xu, Fanrui Zhang, Zizhen Li, Yukang Feng, Baojin Huang, Zhongyuan Wang 0001, Kaipeng Zhang
ICCV5
2025 Sekai: A Video Dataset towards World Exploration
abstract
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications.
Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang
NeurIPS12
2025 Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
abstract
The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100,000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification.
Fanrui Zhang, Qiang Zhang 0051, Jun Chen 0005, Sinbadliu, Junxiong Lin, Jiahong Yan, Jiawei Liu 0001, Zhengjun Zha
NeurIPS1
2025 Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance
abstract
The dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coordinated editing, leading to less reliable outcomes. To tackle these challenges, we propose CARML, a novel retrieval-augmented editing framework that integrates conflict-aware dynamic retrieval with multi-level implicit and explicit guidance for reliable lifelong multimodal editing. Specifically, CARML introduces intra-modal uncertainty and inter-modal conflict quantification to dynamically integrate multi-channel retrieval results, so as to pinpoint the most relevant knowledge to the incoming edit samples. Afterwards, an edit scope classifier discerns whether the edit sample semantically aligns with the edit scope of the retrieved knowledge. If deemed in-scope, CARML refines the retrieved knowledge into information-rich continuous prompt prefixes, serving as the implicit knowledge guide. These prefixes not only include static knowledge prompt that capture key textual semantics but also incorporate token-level, context-aware dynamic prompt to explore fine-grained cross-modal associations between the edit sample and retrieved knowledge. To further enhance reliability, CARML incorporates a "hard correction" mechanism, leveraging explicit label knowledge to adjust the model’s output logits. Extensive experiments across multiple MLLMs and datasets indicate the superior performance of CARML in lifelong multimodal editing scenarios.
Qiang Zhang 0051, Fanrui Zhang, Jiawei Liu 0001, Junjun He, Zhengjun Zha
NeurIPS2
2024 Natural Language-centered Inference Network for Multi-modal Fake News Detection
Qiang Zhang 0051, Jiawei Liu 0001, Fanrui Zhang, Zhengjun Zha
IJCAI3
2024 RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination Corrector
abstract
Visual Spatial Description (VSD) is an emerging image-to-text task which aims at generating descriptions of the spatial relationships between given objects in an image. In this paper, we apply Retrieval-Augmented Generation (RAG) technology in guiding Multimodal Large Language Models (MLLMs) for the task of VSD, complemented by an Adaptive Hallucination Corrector, and further fine-tuning them to bolster semantic understanding and overall model efficacy. We found that our approach demonstrated higher accuracy and fewer hallucination errors in both spatial relationship classification and visual language description tasks within the VSD task, achieving state-of-the-art results.
Jun Yu 0001, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Jianqing Sun, Jiaen Liang
ACM Multimedia7
2024 ESCNet: Entity-enhanced and Stance Checking Network for Multi-modal Fact-Checking
abstract
Recently, misinformation incorporating both texts and images has been disseminated more effectively than those containing text alone on social media, raising significant concerns for multi-modal fact-checking. Existing research makes contributions to multi-modal feature extraction and interaction, but fails to fully enhance the valuable semantic representations or excavate the intricate entity information. Besides, existing multi-modal fact-checking datasets are primarily focused on English and merely concentrate on a single type of misinformation, thereby neglecting a comprehensive summary and coverage of various types of misinformation. Taking these factors into account, we construct the first large-scale Chinese Multi-modal Fact-Checking (CMFC) dataset which encompasses 46,000 claims. The CMFC covers all types of misinformation for fact-checking and is divided into two sub-datasets, Collected Chinese Multi-modal Fact-Checking (CCMF) and Synthetic Chinese Multi-modal Fact-Checking (SCMF). To establish baseline performance, we propose a novel Entity-enhanced and Stance Checking Network (ESCNet), which includes Multi-modal Feature Extraction Module, Stance Transformer, and Entity-enhanced Encoder. The ESCNet jointly models stance semantic reasoning features and knowledge-enhanced entity pair features, in order to simultaneously learn effective semantic-level and knowledge-level claim representations. Our work offers the first step and establishes a benchmark for evidence-based, multi-type, multi-modal fact-checking.
Fanrui Zhang, Jiawei Liu 0001, Qiang Zhang 0051, Yongchao Xu, Zhengjun Zha
WWW1
2024 Event-Driven Heterogeneous Network for Video Deraining
Xueyang Fu, Chengzhi Cao, Senyan Xu, Fanrui Zhang, Zhengjun Zha
Int. J. Comput. Vis.4
2023 ECENet: Explainable and Context-Enhanced Network for Muti-modal Fact verification
abstract
Recently, falsified claims incorporating both text and images have been disseminated more effectively than those containing text alone, raising significant concerns for multi-modal fact verification. Existing research makes contributions to multi-modal feature extraction and interaction, but fails to fully utilize and enhance the valuable and intricate semantic relationships between distinct features. Moreover, most detectors merely provide a single outcome judgment and lack an inference process or explanation. Taking these factors into account, we propose a novel Explainable and Context-Enhanced Network (ECENet) for multi-modal fact verification, making the first attempt to integrate multi-clue feature extraction, multi-level feature reasoning, and justification (explanation) generation within a unified framework. Specifically, we propose an Improved Coarse- and Fine-grained Attention Network, equipped with two types of level-grained attention mechanisms, to facilitate a comprehensive understanding of contextual information. Furthermore, we propose a novel justification generation module via deep reinforcement learning that does not require additional labels. In this module, a sentence extractor agent measures the importance between the query claim and all document sentences at each time step, selecting a suitable amount of high-scoring sentences to be rewritten as the explanation of the model. Extensive experiments demonstrate the effectiveness of the proposed method.
Fanrui Zhang, Jiawei Liu 0001, Qiang Zhang 0051, Esther Sun, Zhengjun Zha
ACM Multimedia1
2023 Hierarchical Semantic Enhancement Network for Multimodal Fake News Detection
abstract
The explosion of multimodal fake news content on social media has sparked widespread concern. Existing multimodal fake news detection methods have made significant contributions to the development of this field, but fail to adequately exploit the potential semantic information of images and ignore the noise embedded in news entities, which severely limits the performance of the models. In this paper, we propose a novel Hierarchical Semantic Enhancement Network (HSEN) for multimodal fake news detection by learning text-related image semantic and precise news high-order knowledge semantic information. Specifically, to complement the image semantic information, HSEN utilizes textual entities as the prompt subject vocabulary and applies reinforcement learning to discover the optimal prompt format for generating image captions specific to the corresponding textual entities, which contain multi-level cross-modal correlation information. Moreover, HSEN extracts visual and textual entities from image and text, and identifies additional visual entities from image captions to extend image semantic knowledge. Based on that, HSEN exploits an adaptive hard attention mechanism to automatically select strongly related news entities and remove irrelevant noise entities to obtain precise high-order knowledge semantic information, while generating attention mask for guiding cross-modal knowledge interaction. Extensive experiments show that our method outperforms state-of-the-art methods.
Qiang Zhang 0051, Jiawei Liu 0001, Fanrui Zhang, Zhengjun Zha
ACM Multimedia3