Zahra Babaiee

dblp:294/9094 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 49% Vision and language · 16% Trustworthy machine learning · 15%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
depthwise separable convolution
2.532025
The Quest for Universal Master Key Filters in DS-CNNs · NeurIPS 2025
The Master Key Filters Hypothesis: Deep Filters Are General · AAAI 2025
Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional Kernels · ICLR 2024
Machine learning › Deep learning architectures and training
convolutional neural network
2.232025
The Quest for Universal Master Key Filters in DS-CNNs · NeurIPS 2025
The Master Key Filters Hypothesis: Deep Filters Are General · AAAI 2025
On-Off Center-Surround Receptive Fields for Accurate and Robust Image Classification · ICML 2021
Machine learning › Trustworthy machine learning
interpretability
1.622025
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models · ICML 2025
Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional Kernels · ICLR 2024
Machine learning › Learning theory
generalization
0.912025
The Quest for Universal Master Key Filters in DS-CNNs · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models · ICML 2025
Computer vision › Vision and language › visual reasoning
visual reasoning evaluation
0.912025
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models · ICML 2025
Computer vision › Image recognition and object detection
image classification
0.822025
On-Off Center-Surround Receptive Fields for Accurate and Robust Image Classification · ICML 2021
The Quest for Universal Master Key Filters in DS-CNNs · NeurIPS 2025
Machine learning › Deep learning architectures and training › convolutional neural network › receptive field
receptive field design
0.512021
On-Off Center-Surround Receptive Fields for Accurate and Robust Image Classification · ICML 2021
Computer vision › Image recognition and object detection › image classification
robust image classification
0.512021
On-Off Center-Surround Receptive Fields for Accurate and Robust Image Classification · ICML 2021

Methods — techniques the papers use, named apart from their topics

unsupervised search · 0.9transfer learning experiments · 0.9graph-based visual tasks · 0.9benchmark evaluation · 0.9unsupervised clustering · 0.8autoencoder · 0.8on-center off-center pathways · 0.5difference of gaussian · 0.5
YearPublicationVenuePosition
2025 The Master Key Filters Hypothesis: Deep Filters Are General
abstract
This paper challenges the prevailing view that convolutional neural network (CNN) filters become increasingly specialized in deeper layers. Motivated by recent observations of clusterable repeating patterns in depthwise separable CNNs (DS-CNNs) trained on ImageNet, we extend this investigation across various domains and datasets. Our analysis of DS-CNNs reveals that deep filters maintain generality, contradicting the expected transition to class-specific features. We demonstrate the generalizability of these filters through transfer learning experiments, showing that frozen filters from models trained on different datasets perform well and can be further improved when sourced from larger, better-performing models. Our findings indicate that spatial features learned by depthwise separable convolutions remain generic across all layers, domains, and architectures. This research provides new insights into the nature of generalization in neural networks, particularly in DS-CNNs, and has significant implications for transfer learning and model design.
Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu
AAAI1
2025 Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
abstract
Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'—the ability to recognize and reason about the same concept despite variations in visual form, a basic ability of human reasoning. To address this challenge, we introduce the Visual Graph Arena (VGA), a dataset featuring six graph-based tasks designed to evaluate and improve AI systems’ capacity for visual abstraction. VGA uses diverse graph layouts (e.g., Kamada-Kawai vs. planar) to test reasoning independent of visual form. Experiments with state-of-the-art vision models and multimodal LLMs reveal a striking divide: humans achieved near-perfect accuracy across tasks, while models totally failed on isomorphism detection and showed limited success in path/cycle tasks. We further identify behavioral anomalies suggesting pseudo-intelligent pattern matching rather than genuine understanding. These findings underscore fundamental limitations in current AI models for visual understanding. By isolating the challenge of representation-invariant reasoning, the VGA provides a framework to drive progress toward human-like conceptualization in AI visual models. The Visual Graph Arena is available at: \href{https://vga.csail.mit.edu/}{vga.csail.mit.edu}.
Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu
ICML1
2025 Scalable Offline Reinforcement Learning for Mean Field Games
Axel Brunnbauer, Julian Lemmel, Zahra Babaiee, Sophie A. Neubauer, Radu Grosu
AAMAS3
2025 The Quest for Universal Master Key Filters in DS-CNNs
abstract
A recent study has proposed the ``Master Key Filters Hypothesis" for convolutional neural network filters. This paper extends this hypothesis by radically constraining its scope to a single set of just 8 universal filters that depthwise separable convolutional networks inherently converge to. While conventional DS-CNNs employ thousands of distinct trained filters, our analysis reveals these filters are predominantly linear shifts (ax+b) of our discovered universal set. Through systematic unsupervised search, we extracted these fundamental patterns across different architectures and datasets. Remarkably, networks initialized with these 8 unique frozen filters achieve over 80\% ImageNet accuracy, and even outperform models with thousands of trainable parameters when applied to smaller datasets. The identified master key filters closely match Difference of Gaussians (DoGs), Gaussians, and their derivatives, structures that are not only fundamental to classical image processing but also strikingly similar to receptive fields in mammalian visual systems. Our findings provide compelling evidence that depthwise convolutional layers naturally gravitate toward this fundamental set of spatial operators regardless of task or architecture. This work offers new insights for understanding generalization and transfer learning through the universal language of these master key filters.
Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu
NeurIPS1
2024 Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional Kernels
abstract
Recent advances in depthwise-separable convolutional neural networks (DS-CNNs) have led to novel architectures, that surpass the performance of classical CNNs, by a considerable scalability and accuracy margin. This paper reveals another striking property of DS-CNN architectures: discernible and explainable patterns emerge in their trained depthwise convolutional kernels in all layers. Through an extensive analysis of millions of trained filters, with different sizes and from various models, we employed unsupervised clustering with autoencoders, to categorize these filters. Astonishingly, the patterns converged into a few main clusters, each resembling the difference of Gaussian (DoG) functions, and their first and second-order derivatives. Notably, we classify over 95\% and 90\% of the filters from state-of-the-art ConvNeXtV2 and ConvNeXt models, respectively. This finding is not merely a technological curiosity; it echoes the foundational models neuroscientists have long proposed for the vision systems of mammals. Our results thus deepen our understanding of the emergent properties of trained DS-CNNs and provide a bridge between artificial and biological visual processing systems. More broadly, they pave the way for more interpretable and biologically-inspired neural network designs in the future.
Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu
ICLR1
2024 Neural Echos: Depthwise Convolutional Filters Replicate Biological Receptive Fields
abstract
In this study, we present evidence suggesting that depthwise convolutional kernels are effectively replicating the structural intricacies of the biological receptive fields observed in the mammalian retina. We provide analytics of trained kernels from various state-of-the-art models substantiating this evidence. Inspired by this intriguing discovery, we propose an initialization scheme that draws inspiration from the biological receptive fields. Experimental analysis of the ImageNet dataset with multiple CNN architectures featuring depthwise convolutions reveals a marked enhancement in the accuracy of the learned model when initialized with biologically derived weights. This underlies the potential for biologically inspired computational models to further our understanding of vision processing systems and to improve the efficacy of convolutional networks.
Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu
WACV1
2021 On-Off Center-Surround Receptive Fields for Accurate and Robust Image Classification
abstract
Robustness to variations in lighting conditions is a key objective for any deep vision system. To this end, our paper extends the receptive field of convolutional neural networks with two residual components, ubiquitous in the visual processing system of vertebrates: On-center and off-center pathways, with an excitatory center and inhibitory surround; OOCS for short. The On-center pathway is excited by the presence of a light stimulus in its center, but not in its surround, whereas the Off-center pathway is excited by the absence of a light stimulus in its center, but not in its surround. We design OOCS pathways via a difference of Gaussians, with their variance computed analytically from the size of the receptive fields. OOCS pathways complement each other in their response to light stimuli, ensuring this way a strong edge-detection capability, and as a result an accurate and robust inference under challenging lighting conditions. We provide extensive empirical evidence showing that networks supplied with OOCS pathways gain accuracy and illumination-robustness from the novel edge representation, compared to other baselines.
Zahra Babaiee, Ramin M. Hasani, Mathias Lechner, Daniela Rus, Radu Grosu
ICML1