Dibin Zhou

dblp:80/4873 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Knowledge representation and reasoning · 36% Vision and language · 28% 3D vision · 28%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Computer vision › 3D vision
spatial understanding
1.012026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › abstract reasoning
abstract visual reasoning
0.312026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
0.312026
Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning? · AAAI 2026

Methods — techniques the papers use, named apart from their topics

visual question answering · 1.0benchmark construction · 1.0
YearPublicationVenuePosition
2026 Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning?
abstract
Multimodal Large Language Models (MLLMs) largely lag human-level performance on abstract visual reasoning (AVR), which requires models to infer latent rules from visual question sets and generalize them to novel scenarios. Most AVR benchmarks are constrained to narrow and repetitive 2D patterns, involving relatively simple spatial relationships and assessing limited dimensions of reasoning ability. Drawing inspiration from real-world paper folding challenges, we propose Paper Folding Puzzles (PFP), a rigorously designed benchmark specifically developed to assess spatial reasoning capabilities. It comprises 150K visual question-answering samples across five diverse tasks, ranging from basic 2D geometric reasoning to 3D spatial understanding. The developed benchmark dataset can be employed to assess core spatial reasoning abilities essential to human cognition, encompassing fundamental symmetry reasoning and 3D spatial comprehension. Furthermore, we conduct a comprehensive evaluation of 18 leading MLLMs (both closed- and open-source variants) on the PFP benchmark to assess their spatial reasoning capabilities. Our findings show that most MLLMs achieve near-chance performance on FPF, exhibiting substantial performance gaps (>30%) relative to human baselines across all tasks. This highlights a critical research gap in improving spatial reasoning capabilities of MLLMs.
Dibin Zhou, Yantao Xu, Zongming Huang, Zengwei Yan, Yongwei Miao, Jianfeng Ren, Fuchang Liu
AAAI1
2026 BAFNet: Deep contour-aware features for colorectal polyps segmentation
Dibin Zhou, Ni Chen, Yueping Zhu, Innocent Nyalala, Junfeng Gao
Expert Syst. Appl.1
2025 Robust Vehicle Localization for Spherical Camera Models: Solution, Framework, and Verification
abstract
Vehicle visual localization uses vision sensors to capture environmental information, enabling precise localization of autonomous vehicles within their surroundings. However, current visual localization methods generally have some shortcomings: on one hand, they are limited by the camera’s field of view, on the other hand, their robustness is often inadequate under challenging conditions such as lighting changes, long-term scene changes, or occlusions. To address these issues, we formulate a general spherical camera model for both fisheye and panoramic cameras and propose a minimal solution for pose estimation using this model based on vehicle motion characteristic. The minimal solution cannot filter outliers, so a robust estimation framework is necessary. For outlier-rejection, we introduce two frameworks: a probabilistic optimal RANSAC and a globally optimal graph-based framework. We conduct a probabilistic analysis of the RANSAC to demonstrate its enhanced robustness given by the proposed minimal solution. To achieve robustness to extreme outliers (higher than 90%), we decouple the rotation and translation space through the minimal solution to construct maximum consensus graph for the two sub-problems. We then employs a maximum clique search algorithm to find the optimal solutions, achieving deterministic convergence while maintaining real-time performance. Extensive experiments with synthetic data, real-world fisheye images, and 360∘panoramic images validate the robustness and efficiency of our proposed algorithms.
Yanmei Jiao, Dibin Zhou, Xiumei Li, Rong Xiong, Yue Wang 0020
IEEE Trans. Intell. Transp. Syst.3
2013 Research of virtual Chinese calligraphic learning
abstract
Chinese Calligraphy is one traditional form of Chinese art and was listed in UNESCO's World Culture Heritage in 2009. Traditional method to learn calligraphy is inconvenient as it needs paper, Chinese ink and brush. And learner's especially for foreigners feel difficulty for stroke order of Chinese character is complex and even same character has different stroke orders with different font style. In this paper we introduce a new calligraphic learning system based on mobile device to help people learn calligraphic writing. The system includes the following function: 1) writing process in 2D method based on the extraction of stroke order and stroke thickness is shown; 2) users can copy character and finally evaluation result will be given according the similarity between the original image and user's writing track; 3) a 3D method for recurring is provided to give a virtual writing process.
Yingfei Wu, Zhenming Yuan, Dibin Zhou, Yizhou Cai
ICME3
2007 A Novel Pointcloud-based Isosurface Extraction
abstract
In this paper, we present a novel approach of hardware-accelerated pointcloud-based isosurface extraction on tetrahedral cells. In contrast to previous methods, our method takes advantage of programmable capability on the modern graphics hardware to render the isosurface silhouette and reduces the video memory consumption per tetrahedron. We achieve this by using a GPU-based geometry-generating method called fat-point technique. We classify tetrahedra into three isosurface cases and compute corresponding parameters in vertex processors, and then generate implicit isosurface silhouette in fragment processors. Utilizing the high performance OpenGL vertex buffer objects, our algorithm can achieve a rendering rate of three million tetrahedra per second.
Dibin Zhou, Kangjian Wang, Lijun Xie, Yanni Wang
CAD/Graphics1
2007 High-quality Texture-based Flow Visualization on Surfaces
abstract
In this paper, we present a high-quality texture-based approach for the visualization of vector fields on surface. It adopts a scheme of streamline enhancement by applying a ID high-pass filter to flow textures along the orthogonal flow direction to enhance the intensity contrast among streamlines and improve the image quality of flow textures. The scheme of edge detection and supplement is used in our approach to modulate the characteristic of particles flow. Our algorithm is image space based and can achieve a real-time rendering performance on a personal computer by using parallel processing ability of current graphics hardware.
Dibin Zhou, Kangjian Wang, Yanni Wang
CAD/Graphics1