Hosein Hasani

dblp:217/2122 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 42% Trustworthy machine learning · 31% Learning paradigms · 12%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › compositionality
binding problem
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.912025
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness
spurious correlation
0.912025
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection · NeurIPS 2025
Computer vision › Vision and language
visual grounding
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Computer vision › Vision and language
visual reasoning
0.912025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Machine learning › Learning paradigms
continual learning
0.512021
Generative vs. Discriminative: Rethinking The Meta-Continual Learning · NeurIPS 2021
Machine learning › Learning paradigms › continual learning
meta-continual learning
0.512021
Generative vs. Discriminative: Rethinking The Meta-Continual Learning · NeurIPS 2021
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019
Computer vision › 3D vision › biological vision modeling
visual cortex modeling
0.412019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019
Computer vision › Image recognition and object detection
visual search
0.312025
Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs · NeurIPS 2025
Machine learning › Generative modeling › generative model
generative classifier
0.112021
Generative vs. Discriminative: Rethinking The Meta-Continual Learning · NeurIPS 2021
Machine learning › Representation and self-supervised learning › computational neuroscience
neural coding
0.112019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

visual input structuring · 0.9prototype refinement · 0.9chain-of-thought prompting · 0.9meta-learning · 0.5generative classifier · 0.5surround modulation · 0.4excitatory-inhibitory connections · 0.4
YearPublicationVenuePosition
2025 Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs
abstract
Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This limitation underlies persistent errors in tasks such as counting, visual search, scene description, and spatial relationship understanding. A key factor is that current LVLMs process visual features largely in parallel, lacking mechanisms for spatially grounded, serial attention. This paper introduces Visual Input Structure for Enhanced Reasoning (VISER), a simple, effective method that augments visual inputs with low-level spatial structures and pairs them with a textual prompt that encourages sequential, spatially-aware parsing. We empirically demonstrate substantial performance improvements across core visual reasoning tasks, using only a single-query inference. Specifically, VISER improves GPT-4o performance on visual search, counting, and spatial relationship tasks by 25.0%, 26.8%, and 9.5%, respectively, and reduces edit distance error in scene description by 0.32 on 2D datasets. Furthermore, we find that the visual modification is essential for these gains; purely textual strategies, including Chain-of-Thought prompting, are insufficient and can even degrade performance. VISER underscores the importance of visual input design over purely linguistically based reasoning strategies and suggests that visual structuring is a powerful and general approach for enhancing compositional and spatial reasoning in LVLMs.
Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, Ali Rahimiakbar, Mohammad Mahdi Vahedi, Hosein Hasani, Mahdieh Soleymani Baghshah
NeurIPS6
2025 Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications, where they frequently face data distributions unseen during training. Despite progress, existing methods are often vulnerable to spurious correlations that mislead models and compromise robustness. To address this, we propose SPROD, a novel prototype-based OOD detection approach that explicitly addresses the challenge posed by unknown spurious correlations. Our post-hoc method refines class prototypes to mitigate bias from spurious features without additional data or hyperparameter tuning, and is broadly applicable across diverse backbones and OOD detection settings. We conduct a comprehensive spurious correlation OOD detection benchmarking, comparing our method against existing approaches and demonstrating its superior performance across challenging OOD datasets, such as CelebA, Waterbirds, UrbanCars, Spurious Imagenet, and the newly introduced Animals MetaCoCo. On average, SPROD improves AUROC by 4.8% and FPR@95 by 9.4% over the second best.
Reihaneh Zohrabi, Hosein Hasani, Mahdieh Soleymani Baghshah, Anna Rohrbach, Marcus Rohrbach, Mohammad H. Rohban
NeurIPS2
2021 Generative vs. Discriminative: Rethinking The Meta-Continual Learning
abstract
Deep neural networks have achieved human-level capabilities in various learning tasks. However, they generally lose performance in more realistic scenarios like learning in a continual manner. In contrast, humans can incorporate their prior knowledge to learn new concepts efficiently without forgetting older ones. In this work, we leverage meta-learning to encourage the model to learn how to learn continually. Inspired by human concept learning, we develop a generative classifier that efficiently uses data-driven experience to learn new concepts even from few samples while being immune to forgetting. Along with cognitive and theoretical insights, extensive experiments on standard benchmarks demonstrate the effectiveness of the proposed method. The ability to remember all previous concepts, with negligible computational and structural overheads, suggests that generative models provide a natural way for alleviating catastrophic forgetting, which is a major drawback of discriminative models.
Mohammadamin Banayeeanzade, Rasoul Mirzaiezadeh, Hosein Hasani, Mahdieh Soleymani Baghshah
NeurIPS3
2019 Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks
abstract
Numerous neurophysiological studies have revealed that a large number of the primary visual cortex neurons operate in a regime called surround modulation. Surround modulation has a substantial effect on various perceptual tasks, and it also plays a crucial role in the efficient neural coding of the visual cortex. Inspired by the notion of surround modulation, we designed new excitatory-inhibitory connections between a unit and its surrounding units in the convolutional neural network (CNN) to achieve a more biologically plausible network. Our experiments show that this simple mechanism can considerably improve both the performance and training speed of traditional CNNs in visual tasks. We further explore additional outcomes of the proposed structure. We first evaluate the model under several visual challenges, such as the presence of clutter or change in lighting conditions and show its superior generalization capability in handling these challenging situations. We then study possible changes in the statistics of neural activities such as sparsity and decorrelation and provide further insight into the underlying efficiencies of surround modulation. Experimental results show that importing surround modulation into the convolutional layers ensues various effects analogous to those derived by surround modulation in the visual cortex.
Hosein Hasani, Mahdieh Soleymani Baghshah, Hamid K. Aghajan
NeurIPS1