Peizhi Zhao

dblp:351/5632 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 55% Segmentation and scene understanding · 18% Question answering and dialogue systems · 12%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal benchmark
1.012026
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026
Computer vision › Vision and language
visual grounding
0.912025
Look Around Before Locating: Considering Content and Structure Information for Visual Grounding · AAAI 2025
Computer vision › Segmentation and scene understanding
instance segmentation
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Segmentation and scene understanding
referring image segmentation
0.812024
Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point · AAAI 2024
Computer vision › Face, body and person analysis
person re-identification
0.712023
Linking People across Text and Images Based on Social Relation Reasoning · AAAI 2023
Information retrieval
retrieval-augmented generation
0.312026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026
Computer vision › Vision and language
vision-language pretraining
0.312025
Look Around Before Locating: Considering Content and Structure Information for Visual Grounding · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
social relation reasoning
0.212023
Linking People across Text and Images Based on Social Relation Reasoning · AAAI 2023

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0multimodal large language model · 1.0structure-modulated grounding · 0.9semi-structured reasoning · 0.9cross-modal alignment · 0.9soft ground-truth · 0.8point-based cross-modal comprehension · 0.8binary classification · 0.8social relation reasoning · 0.7
YearPublicationVenuePosition
2026 FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
abstract
We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.
Zichen Tang, Haihong E, Rongjin Li, Linwei Jia, Zhuodi Hao, Zhongjun Yang, Yuanze Li, Haolin Tian, Peizhi Zhao, Xianghe Wang, Xueyuan Lin, Ruofei Bai, Zijian Xie, Ruining Cao, Haocheng Gao
AAAI11
2026 Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
abstract
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Haolin Tian, Pengqi Sun, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Zhongjun Yang, Jintong Chen, Siying Lin
ACL (1)14
2026 Implement Referring Expression Comprehension by Extending Auto-focus Lens to Locked Vision Model
abstract
Referring Expression Comprehension (REC) aims to achieve fine-grained cross-modal content alignment. The traditional two-stage approaches, by decomposing REC into localization (region proposal) and comprehension (expression-based ranking), lead to the isolation of continuous image information and heavily rely on the quality of the proposals. In this article, we propose a point-based two-stage framework for REC to quickly achieve localization by inserting a language-modulated auto-focus module into the locked vision model. Specifically, we redefine REC as two processes: point-based cross-modal comprehension and point-based instance localization. For the comprehension stage, we reconstruct the raw annotations into soft masks at the feature point level as a metric of cross-modal correlation. With this indirect metric, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions. Remarkably, soft masks are shape-independent, which means our method is extremely general. By switching different vision models, different types of predictions (e.g., localization and segmentation) can be obtained. Experiments on multiple benchmarks demonstrate the feasibility and potential of our point-based paradigm. Our code will be public at https://github.com/VILAN-Lab/PBREC-AF .
Shiyi Zheng, Peizhi Zhao, Qingbao Huang, Yi Cai 0001, Haonan Cheng, Qi Wu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Look Around Before Locating: Considering Content and Structure Information for Visual Grounding
abstract
As a long-term challenge and fundamental requirement in vision and language tasks, visual grounding aims to localize a target referred by a natural language query. The regional annotations form a superficial correlation between the subject of expression and some common visual entities, which hinder models from comprehending the linguistic content and structure. However, current one-stage methods struggle to uniformly model the visual and linguistic structure due to the structural gap between continuous image patches and discrete text tokens. In this paper, we propose a semi-structured reasoning framework for visual grounding to gradually comprehend the linguistic content and structure. Specifically, we devise a cross-modal content alignment module to effectively align unlabeled contextual information into a stable semantic space corrected by token-level prior knowledge obtained with CLIP. A multi-branch modulated localization module is also established to obtain modulation grounding by linguistic structure. Through a soft split mechanism, our method can destructure the expression into a fixed semi-structure (i.e., subject and context) while ensuring the completeness of linguistic content. Our method is thus capable of building a semi-structured reasoning system to effectively comprehend the linguistic content and structure by content alignment and structure modulated grounding. Experimental results on five widely-used datasets validate the performance improvements of our proposed method.
Shiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He, Haonan Cheng, Yi Cai 0001, Qingbao Huang
AAAI2
2025 Multi-scale low-frequency enhanced spectral neural operator for reducing low-frequency error in partial differential equations solving
Fengrui Jing, Chuchu Zhai, Peizhi Zhao, Xue Li 0019, Peifu Han, Hongzhen Ding, Yunlong Dong, Long Hao, Tao Song 0001
Eng. Appl. Artif. Intell.3
2025 The Frequency-Domain Corrected Attention Operator for solving PDEs
Qinglong Ma, Xuebin Hu, Peizhi Zhao, Xichen Cao
Inf. Sci.3
2024 Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point
abstract
As a fundamental and challenging task in the vision and language domain, Referring Expression Comprehension (REC) has shown impressive improvements recently. However, for a complex task that couples the comprehension of abstract concepts and the localization of concrete instances, one-stage approaches are bottlenecked by computing and data resources. To obtain a low-cost solution, the prevailing two-stage approaches decouple REC into localization (region proposal) and comprehension (region-expression matching) at region-level, but the solution based on isolated regions cannot sufficiently utilize the context and is usually limited by the quality of proposals. Therefore, it is necessary to rebuild an efficient two-stage solution system. In this paper, we propose a point-based two-stage framework for REC, in which the two stages are redefined as point-based cross-modal comprehension and point-based instance localization. Specifically, we reconstruct the raw bounding box and segmentation mask into center and mass scores as soft ground-truth for measuring point-level cross-modal correlations. With the soft ground-truth, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions on the optimization process. Remarkably, the consistent metrics between center and mass scores allow our system to directly optimize grounding and segmentation by utilizing the same architecture. Experiments on multiple benchmarks show the feasibility and potential of our point-based paradigm. Our code available at https://github.com/VILAN-Lab/PBREC-MT.
Peizhi Zhao, Shiyi Zheng, Wenye Zhao, Dongsheng Xu 0001, Pijian Li, Yi Cai 0001, Qingbao Huang
AAAI1
2023 Linking People across Text and Images Based on Social Relation Reasoning
Peizhi Zhao, Pijian Li, Yi Cai 0001, Qingbao Huang
AAAI2