Peijin Xie

dblp:283/3753 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0004-3021-4576ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 49% Reinforcement learning · 17% 3D vision · 15%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
1.012026
CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.912025
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules · AAAI 2025
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding
0.912025
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules · AAAI 2025
Computer vision › Vision and language › visual reasoning
visual spatial reasoning
0.912025
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules · AAAI 2025
Natural language and speech › Language models and text generation
instruction following
0.312026
CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning · AAAI 2026

Methods — techniques the papers use, named apart from their topics

supervised fine-tuning · 1.0rule-based reinforcement learning · 1.0dynamic task mixing · 1.0diffusion model · 0.9SigLIP · 0.9SAM · 0.9DINO · 0.9CLIP · 0.9
YearPublicationVenuePosition
2026 CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning
abstract
Compositional reasoning is a critical capability for multimodal models, enabling systematic understanding of complex scenes through structured combinations of objects, attributes, and relations. However, existing research on this ability primarily focuses on vision-language models (VLMs, e.g., CLIP and SigLIP), with limited exploration of multimodal large language models (MLLMs). To address this gap, we introduce CR³, a novel framework that enhances compositional reasoning abilities of MLLMs via rule-based reinforcement learning. CR³ leverages rule-based rewards to optimize the MLLM's policy on systematically curated multimodal instruction-following tasks, guided by a model-adaptive dynamic task mixing strategy. Our approach boosts performance by over 19% on three compositional reasoning benchmarks, significantly outperforming supervised fine-tuning (SFT) by at least 12%. Crucially, CR³ demonstrates superior generalization by improving performance on out-of-domain benchmarks where SFT methods degrade, highlighting its effectiveness and data efficiency.
Shun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie, Zhen Xu 0003, Baoxun Wang
AAAI4
2026 Spatial -aware efficient projector for MLLMs via multi-layer feature aggregation
Shun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie, Yunhe Xie, Zhen Xu 0003, Baoxun Wang
Expert Syst. Appl.4
2026 VIP-doc :Visual prompts guide fine-grained document understanding for reader friendly VLLM
Peijin Xie, Lin Sun 0010, Xiangzheng Zhang, Yunhe Xie, Shun Qian, Chengjie Sun, Bingquan Liu
Expert Syst. Appl.1
2025 Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
abstract
Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks like MME, MMBench and SEED comprehensively have evaluated various capabilities which already include visual spatial reasoning(VSR). There is still a lack of sufficient quantity and quality evaluation and optimization datasets for Vision Large Language Models(VLLMs) specifically targeting visual positional reasoning. To handle this, we first diagnosed current VLLMs with the VSR dataset and proposed a unified test set. We found current VLLMs to exhibit a contradiction of over-sensitivity to language instructions and under-sensitivity to visual positional information. By expanding the original benchmark from two aspects of tunning data and model structure, we mitigated this phenomenon. To our knowledge, we expanded spatially positioned image data controllably using diffusion models for the first time and integrated original visual encoding(CLIP) with other 3 powerful visual encoders(SigLIP, SAM and DINO). After conducting combination experiments on scaling data and models, we obtained a VLLM VSR Expert(VSRE) that not only generalizes better to different instructions but also accurately distinguishes differences in visual positional information. VSRE achieved over a 27% increase in accuracy on the VSR test set. It becomes a performant VLLM on the position reasoning of both the VSR dataset and relevant subsets of other evaluation benchmarks. We hope it will accelerate advancements in VLLM on VSR learning.
Peijin Xie, Lin Sun 0010, Bingquan Liu, Xiangzheng Zhang, Chengjie Sun
AAAI1
2020 Gene Ontology aided Compound Protein Binding Affinity Prediction Using BERT Encoding
abstract
The drug-target binding affinity(DTA) indicates the strength of the drug-target interaction; therefore, predicting DTA by computational approaches can considerably benefit drug discovery by narrowing down the searching space and pruning those drug-target pairs with low binding affinity scores. In the computational methods, feature representation of proteins is one of the most important parts due to its strong influence on the following regression task. This paper introduces the BERT-based language representation to embed the gene ontology annotations, combined with the raw sequence to characterize a protein by fusing its physical structure and human knowledge. We exploit CNN network stacked over full connected layers to learn the prediction of DTA scores in a supervised manner. This framework enhances the feature representation ability, leading to the improvement of the DTA prediction precision. The evaluation on the Davis and KIBA datasets compared to the state-of-the-art baselines demonstrates our feature representation's superiority.
Lingling Zhao, Peijin Xie, Lingfeng Hao, Chunyu Wang 0002
BIBM2