Zhangyun Tan

dblp:158/9378 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 36% Vision and language · 28% Knowledge representation and reasoning · 28%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Computer vision › 3D vision › multi-view geometry › camera geometry
perspective geometry
0.912025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.912025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.312025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025
Computer vision › 3D vision
spatial consistency
0.312025
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

chain-of-thought prompting · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2025 MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
abstract
Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources are available at https://yunlong10.github.io/MMPerspective/
Yunlong Tang 0002, Pinxin Liu, Mingqian Feng, Zhangyun Tan, Rui Mao 0017, Chao Huang 0033, Jing Bi 0002, Yunzhong Xiao, Susan Liang, Hang Hua, Ali Vosoughi, Luchuan Song, Zeliang Zhang 0001, Chenliang Xu
NeurIPS4
2015 ARFBF model for non stationary random fields and application in HRTEM images
abstract
This paper presents a new model called Autoregressive Fractional Brownian Field (ARFBF) for analyzing textures which contain stationary and non-stationary components. The paper also proposes two estimation methods for the parameter of an isotropic fractional Brownian field based on Wavelet Packet (WP) spectrum: the Log-Regression on Diagonal WP spectrum (Log-RDWP) and the Log-Regression on Polar representation of WP spectrum (Log-RPWP). The Log-RPWP method provides a better estimation performance for small size images. We show the interest of ARFBF model and Log-RPWP for characterizing High-Resolution Transmission Electron Microscopy (HRTEM) images.
Zhangyun Tan, Abdourrahmane M. Atto, Olivier Alata, Maxime Moreaud
ICIP1
2014 Non-stationary texture synthesis from random field modeling
abstract
This paper presents a generalized non-stationary and fractional model for texture synthesis. The model is based on convolution and modulation operations of fractional Brownian fields and its associated spectral representation contains many poles with unit norm. Synthesized textures generated from this model can exhibit several non-trivial fringes which can be visualized in natural textures such those involved in high resolution transmission electron microscopy.
Abdourrahmane M. Atto, Zhangyun Tan, Olivier Alata, Maxime Moreaud
ICIP2