Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhangquan Chen

dblp:372/5515 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 52% 3D vision · 35% Reinforcement learning · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual reasoning
1.922026
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning · AAAI 2026
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning · ICCV 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning · AAAI 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.012026
Easy for Children, Hard for AI: The Limits of Multimodal LLMs in Early Childhood Learning · AAAI 2026
Computer vision › 3D vision
spatial understanding
1.012026
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning · AAAI 2026
Computer vision › Vision and language
visual grounding
1.012026
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning · AAAI 2026
Computer vision › 3D vision › point cloud registration
non-rigid point cloud registration
0.912025
DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features · CVPR 2025
Computer vision › 3D vision › point cloud registration
point cloud matching
0.912025
DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features · CVPR 2025
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.312026
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning · AAAI 2026
Computer vision › 3D vision › correspondence estimation
dense correspondence
0.312025
DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features · CVPR 2025
Computer vision › 3D vision
point cloud processing
0.312025
DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features · CVPR 2025

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 3.0reverse-expansion-forward-inference · 1.0chain-of-thought · 1.0GRPO · 1.0visual chain-of-thought · 0.9reinforcement learning · 0.9pre-trained visual features · 0.9large multimodal model · 0.9deformation-based module · 0.9
YearPublicationVenuePosition
2026 SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
abstract
Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods have tried to incorporate visual reasoning, however, they fail to leverage attention correction with spatial cues to iteratively refine their focus on prompt-relevant regions. In this paper, we introduce SIFThinker, a spatially-aware “think-with-images” framework that mimics human visual perception. Specifically, SIFThinker enables attention correcting and image region focusing by interleaving depth-enhanced bounding boxes and natural language. Our contributions are twofold: First, we introduce a reverse-expansion-forward-inference strategy that facilitates the generation of interleaved image-text chains of thought for process-level supervision, which in turn leads to the construction of the SIF-50K dataset. Besides, we propose GRPO-SIF, a reinforced training paradigm that integrates depth-informed visual grounding into a unified reasoning pipeline, teaching the model to dynamically correct and focus on prompt-relevant regions. Extensive experiments demonstrate that SIFThinker outperforms state-of-the-art methods in spatial understanding and fine-grained visual perception, while maintaining strong general capabilities, highlighting the effectiveness of our method.
Zhangquan Chen, Ruihui Zhao, Chuwei Luo, Yangyang Kang, Ruqi Huang
AAAI1
2026 Easy for Children, Hard for AI: The Limits of Multimodal LLMs in Early Childhood Learning
abstract
Early childhood is a critical stage for cognitive development, involving core skills such as visual perception and reasoning. While multimodal large language models (MLLMs) have made rapid progress in various general-purpose tasks, their ability to support early education remains largely underexplored. Existing research on child-related AI largely centers on modeling language, emotion, or behavior, with limited focus on evaluating cognitive tasks relevant to early learning. To address this gap, we propose ChildBench, a multimodal benchmark designed to assess models on tasks inspired by early childhood cognitive development. It covers five key domains through ten tasks, including spatial reasoning, visual reasoning, visual discrimination, counting skills, and visual tracking. The benchmark includes 4,890 carefully constructed images and 5,346 manually annotated samples, ensuring both diversity and age-appropriate content. We evaluate a range of state-of-the-art (SoTA) open-source and closed-source MLLMs—including GPT-4o, Gemini, and Qwen2.5-VL—on ChildBench. Despite strong performance on other benchmarks, the best 7B-parameter model with LoRA tuning achieves only 52.01% accuracy, far below the 96% achieved by 5-year-old children. These results reveal critical limitations in fine-grained perception and reasoning. We further analyze failure cases and discuss directions for future model development.
Xueyan Wu, Hanxuan Chen, Zhangquan Chen, Ronghao Chen, Huacan Wang
AAAI5
2025 DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features
abstract
In this paper, we present DV-Matcher, a novel learning-based framework for estimating dense correspondences between non-rigidly deformable point clouds. Learning directly from unstructured point clouds without meshing or manual labelling, our framework delivers high-quality dense correspondences, which is of significant practical utility in point cloud processing. Our key contributions are twofold: First, we propose a scheme to inject prior knowledge from pre-trained vision models into geometric feature learning, which effectively complements the local nature of geometric features with global and semantic information; Second, we propose a novel deformation-based module to promote the extrinsic alignment induced by the learned correspondences, which effectively enhances the feature learning. Experimental results show that our method achieves state-of-the-art results in matching non-rigid point clouds in both near-isometric and heterogeneous shape collection as well as more realistic partial and noisy data. Our code is available at https://github.com/rqhuang88/DV-Matcher.
Zhangquan Chen, Puhua Jiang, Ruqi Huang
CVPR1
2025 VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
abstract
Visual understanding is inherently intention-driven - humans selectively focus on different regions of a scene based on their goals. Recent advances in large multimodal models (LMMs) enable flexible expression of such intentions through natural language, allowing queries to guide visual reasoning processes. Frameworks like Visual Chain-of-Thought have demonstrated the benefit of incorporating explicit reasoning steps, where the model predicts a focus region before answering a query. However, existing approaches rely heavily on supervised training with annotated intermediate bounding boxes, which severely limits scalability due to the combinatorial explosion of intention-region pairs. To overcome this limitation, we propose VisRL, the first framework that applies reinforcement learning (RL) to the problem of intention-driven visual perception. VisRL optimizes the entire visual reasoning process using only reward signals. By treating intermediate focus selection as an internal decision optimized through trial-and-error, our method eliminates the need for costly region annotations while aligning more closely with how humans learn to perceive the world. Extensive experiments across multiple benchmarks show that VisRL consistently outperforms strong baselines, demonstrating both its effectiveness and its strong generalization across different LMMs. Our code is available at https://github.com/zhangquanchen/VisRL.
Zhangquan Chen, Xufang Luo
ICCV1
2025 Bidirectional spatio-temporal generative adversarial network for video super-resolution
Peng Yang 0014, Zhangquan Chen, Yuankang Sun, Zhongjian Hu, Bing Li 0027
Pattern Anal. Appl.2
2024 A Three-Phases-LORA Finetuned Hybrid LLM Integrated with Strong Prior Module in the Education Context
Zhangquan Chen, Chunjiang Liu, Haobin Duan
ICANN (5)1