Sajad Movahedi

dblp:232/2355 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 35% Efficient and distributed learning · 26% Learning theory · 23%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization
0.912025
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture · ICLR 2025
Machine learning › Learning theory › neural network theory
neural network geometry
0.912025
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture · ICLR 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Fixed-Point RNNs: Interpolating from Diagonal to Dense · NeurIPS 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Fixed-Point RNNs: Interpolating from Diagonal to Dense · NeurIPS 2025
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › one-shot neural architecture search
differentiable architecture search
0.712023
$\Lambda$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells · ICLR 2023
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.712023
$\Lambda$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells · ICLR 2023
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.412019
MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews · AAAI 2019
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect category detection
0.412019
MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews · AAAI 2019
Natural language and speech › Information extraction and text analysis
multilingual NLP
0.412019
MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews · AAAI 2019

Methods — techniques the papers use, named apart from their topics

geometric invariance hypothesis · 0.9fixed-point parameterization · 0.9diagonal linear RNN · 0.9operation selection harmonization · 0.7DARTS · 0.7multilingual word embeddings · 0.4convolutional neural network · 0.4
YearPublicationVenuePosition
2025 Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
abstract
In this paper, we propose the *geometric invariance hypothesis (GIH)*, which argues that the input space curvature of a neural network remains invariant under transformation in certain architecture-dependent directions during training. We investigate a simple, non-linear binary classification problem residing on a plane in a high dimensional space and observe that—unlike MLPs—ResNets fail to generalize depending on the orientation of the plane. Motivated by this example, we define a neural network's **average geometry** and **average geometry evolution** as compact *architecture-dependent* summaries of the model's input-output geometry and its evolution during training. By investigating the average geometry evolution at initialization, we discover that the geometry of a neural network evolves according to the data covariance projected onto its average geometry. This means that the geometry only changes in a subset of the input space when the average geometry is low-rank, such as in ResNets. This causes an architecture-dependent invariance property in the input space curvature, which we dub GIH. Finally, we present extensive experimental results to observe the consequences of GIH and how it relates to generalization in neural networks.
Sajad Movahedi, Antonio Orvieto, Seyed-Mohsen Moosavi-Dezfooli
ICLR1
2025 Fixed-Point RNNs: Interpolating from Diagonal to Dense
abstract
Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e. diagonal) sequence mixing. In this paper, we investigate parameterizations of a large class of dense linear RNNs as fixed-points of parallelizable diagonal linear RNNs. The resulting models can naturally trade expressivity for efficiency at a fixed number of parameters and achieve state-of-the-art results on the state-tracking benchmarks $A_5$ and $S_5$, while matching performance on copying and other tasks.
Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, Antonio Orvieto
NeurIPS1
2023 $\Lambda$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells
Sajad Movahedi, Melika Adabinejad, Ayyoob Imani, Arezou Keshavarz, Mostafa Dehghani 0001, Azadeh Shakery, Babak Nadjar Araabi
ICLR1
2019 MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews
abstract
The advent of the Internet has caused a significant growth in the number of opinions expressed about products or services on e-commerce websites. Aspect category detection, which is one of the challenging subtasks of aspect-based sentiment analysis, deals with categorizing a given review sentence into a set of predefined categories. Most of the research efforts in this field are devoted to English language reviews, while there are a large number of reviews in other languages that are left unexplored. In this paper, we propose a multilingual method to perform aspect category detection on reviews in different languages, which makes use of a deep convolutional neural network with multilingual word embeddings. To the best of our knowledge, our method is the first attempt at performing aspect category detection on multiple languages simultaneously. Empirical results on the multilingual dataset provided by SemEval workshop demonstrate the effectiveness of the proposed method1.
Erfan Ghadery, Sajad Movahedi, Heshaam Faili, Azadeh Shakery
AAAI2
2019 LICD: A Language-Independent Approach for Aspect Category Detection
Erfan Ghadery, Sajad Movahedi, Masoud Jalili Sabet, Heshaam Faili, Azadeh Shakery
ECIR (1)2