Zengwei Yan

dblp:61/8916 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Knowledge representation and reasoning · 36% Vision and language · 28% 3D vision · 28%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Computer vision › 3D vision
spatial understanding
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › abstract reasoning
abstract visual reasoning
0.312026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
0.312026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026

Methods — techniques the papers use, named apart from their topics

visual question answering · 1.0benchmark construction · 1.0
YearPublicationVenuePosition
2026 Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning?
abstract
Multimodal Large Language Models (MLLMs) largely lag human-level performance on abstract visual reasoning (AVR), which requires models to infer latent rules from visual question sets and generalize them to novel scenarios. Most AVR benchmarks are constrained to narrow and repetitive 2D patterns, involving relatively simple spatial relationships and assessing limited dimensions of reasoning ability. Drawing inspiration from real-world paper folding challenges, we propose Paper Folding Puzzles (PFP), a rigorously designed benchmark specifically developed to assess spatial reasoning capabilities. It comprises 150K visual question-answering samples across five diverse tasks, ranging from basic 2D geometric reasoning to 3D spatial understanding. The developed benchmark dataset can be employed to assess core spatial reasoning abilities essential to human cognition, encompassing fundamental symmetry reasoning and 3D spatial comprehension. Furthermore, we conduct a comprehensive evaluation of 18 leading MLLMs (both closed- and open-source variants) on the PFP benchmark to assess their spatial reasoning capabilities. Our findings show that most MLLMs achieve near-chance performance on FPF, exhibiting substantial performance gaps (>30%) relative to human baselines across all tasks. This highlights a critical research gap in improving spatial reasoning capabilities of MLLMs.
Dibin Zhou, Yantao Xu, Zongming Huang, Zengwei Yan, Yongwei Miao, Jianfeng Ren, Fuchang Liu
AAAI4
2022 A novel machine learning based intrusion detection method for 5G empowered CBTC systems
abstract
Wireless communication is the basis for ensuring the safe and efficient operation of rail transit trains. The advantages of 5G technology such as high reliability, high speed and low latency make it a very broad application prospect in the field of rail transit. However, 5G networks also face corresponding security threats in application, such as various network attacks. The communication-based train control (CBTC) systems are vulnerable to network attacks, which can easily lead to emergency braking or collision of trains when it was attacked. In order to protect network security, the intrusion detection systems (IDS) can be used. The IDS currently applied to the CBTC systems have problems such as low detection rate and high false rate. Therefore, this paper proposes an intrusion detection system based on machine learning, and simulates the real network environment by using KDD99 data. First, the data is preprocessed and feature extracted. Then two methods of decision tree and neural network are used to train the classifiers. In addition, the multiple classifiers are trained on the basis of the binary classifiers. Finally, we test the performance of the classifiers. The results show that the classifiers have a good detection accuracy of network traffic.
Zengwei Yan, Renwei He, Liming Zong, Yanjun Zhan
IWCMC2