Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mohammadali Banayeeanzade

dblp:400/7778 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 68% Trustworthy machine learning · 27% Image recognition and object detection · 4%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › compositionality
binding problem
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation · CVPR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Machine learning › Trustworthy machine learning › fairness › social bias
social bias in vision-language models
0.912025
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation · CVPR 2025
Computer vision › Vision and language › vision-language model
vision-language model analysis
0.912025
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation · CVPR 2025
Computer vision › Vision and language
visual grounding
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Computer vision › Vision and language
visual reasoning
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Computer vision › Image recognition and object detection
visual search
0.312025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

visual input structuring · 0.9contrastive language-image pretraining · 0.9chain-of-thought prompting · 0.9
YearPublicationVenuePosition
2025 CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
abstract
Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios. This study offers a comprehensive analysis of CLIP’s limitations in these contexts using a specialized dataset, ComCO, designed to evaluate CLIP’s encoders in diverse multi-object scenarios. Our findings reveal significant biases: the text encoder prioritizes first-mentioned objects, and the image encoder favors larger objects. Through retrieval and classification tasks, we quantify these biases across multiple CLIP variants and trace their origins to CLIP’s training process, supported by analyses of the LAION dataset and training progression. Our image-text matching experiments show substantial performance drops when object size or token order changes, underscoring CLIP’s instability with rephrased but semantically similar captions. Extending this to longer captions and text-to-image models like Stable Diffusion, we demonstrate how prompt order influences object prominence in generated images. For more details and access to our dataset and analysis code, visit our project repository: https://clip-oscope.github.io/.
Reza Abbasi, Aminreza Sefid, Mohammadali Banayeeanzade, Mohammad H. Rohban, Mahdieh Soleymani Baghshah
CVPR4
2025 Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs
abstract
Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, visual search, scene description, and spatial relationship understanding. A key factor is that current LVLMs process visual features largely in parallel, lacking mechanisms for spatially grounded, serial attention. This paper introduces Visual Input Structure for Enhanced Reasoning (VISER), a simple, effective method that augments visual inputs with low-level spatial structures and pairs them with a textual prompt that encourages sequential, spatially-aware parsing. We empirically demonstrate substantial performance improvements across core visual reasoning tasks, using only a single-query inference. Specifically, VISER improves GPT-4o performance on visual search, counting, and spatial relationship tasks by 25.0%, 26.8%, and 9.5%, respectively, and reduces edit distance error in scene description by 0.32 on 2D datasets. Furthermore, we find that the visual modification is essential for these gains; purely textual strategies, including Chain-of-Thought prompting, are insufficient and can even degrade performance. VISER underscores the importance of visual input design over purely linguistically based reasoning strategies and suggests that visual structuring is a powerful and general approach for enhancing compositional and spatial reasoning in LVLMs.
Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, Ali Rahimiakbar, Mohammad Mahdi Vahedi, Hosein Hasani, Mahdieh Soleymani Baghshah
NeurIPS2