Ameen Ali

dblp:280/0176 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 66% Deep learning architectures and training · 32% Representation and self-supervised learning · 2%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.842025
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation · ICLR 2025
The Hidden Attention of Mamba Models · ACL (1) 2025
XAI for Transformers: Better Explanations through Conservative Propagation · ICML 2022
Machine learning › Trustworthy machine learning › interpretability › attention analysis
attention-based explanation
0.912025
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation · ICLR 2025
Machine learning › Deep learning architectures and training › sequence modeling
efficient sequence modeling
0.912025
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation · ICLR 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
The Hidden Attention of Mamba Models · ACL (1) 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.612022
XAI for Transformers: Better Explanations through Conservative Propagation · ICML 2022
Machine learning › Trustworthy machine learning › interpretability › attribution methods
layer-wise relevance propagation
0.612022
XAI for Transformers: Better Explanations through Conservative Propagation · ICML 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
XAI for Transformers: Better Explanations through Conservative Propagation · ICML 2022

Methods — techniques the papers use, named apart from their topics

explainability methods · 0.9causal self-attention · 0.9attribution methods · 0.9attention mechanism · 0.9layer-wise relevance propagation · 0.6gradient propagation · 0.6gradient-based attribution · 0.5back-projection · 0.5
YearPublicationVenuePosition
2025 The Hidden Attention of Mamba Models
abstract
The Mamba layer offers an efficient selective state-space model (SSM) that is highly effective in modeling multiple domains, includingNLP, long-range sequence processing, and computer vision. Selective SSMs are viewed as dual models, in which one trains in parallel on the entire sequence via an IO-aware parallel scan, and deploys in an autoregressive manner. We add a third view and show that such models can be viewed as attention-driven models. This new perspective enables us to empirically and theoretically compare the underlying mechanisms to that of the attention in transformers and allows us to peer inside the inner workings of the Mamba model with explainability methods. Our code is publicly available.
Ameen Ali, Itamar Zimerman, Lior Wolf
ACL (1)1
2025 Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
abstract
Recent advances in efficient sequence modeling have led to attention-free layers, such as Mamba, RWKV, and various gated RNNs, all featuring sub-quadratic complexity in sequence length and excellent scaling properties, enabling the construction of a new type of foundation models. In this paper, we present a unified view of these models, formulating such layers as implicit causal self-attention layers. The formulation includes most of their sub-components and is not limited to a specific part of the architecture. The framework compares the underlying mechanisms on similar grounds for different layers and provides a direct means for applying explainability methods. Our experiments show that our attention matrices and attribution method outperform an alternative and a more limited formulation that was recently proposed for Mamba. For the other architectures for which our method is the first to provide such a view, our method is effective and competitive in the relevant metrics compared to the results obtained by state-of-the-art Transformer explainability methods. Our code is publicly available.
Itamar Zimerman, Ameen Ali, Lior Wolf
ICLR2
2023 Degree-based stratification of nodes in Graph Neural Networks
Ameen Ali, Lior Wolf, Hakan Çevikalp
ACML1
2022 XAI for Transformers: Better Explanations through Conservative Propagation
abstract
Transformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets.
Ameen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon, Klaus-Robert Müller, Lior Wolf
ICML1
2022 Video and Text Matching with Conditioned Embeddings
abstract
We present a method for matching a text sentence from a given corpus to a given video clip and vice versa. Traditionally video and text matching is done by learning a shared embedding space and the encoding of one modality is independent of the other. In this work, we encode the dataset data in a way that takes into account the query’s relevant information. The power of the method is demonstrated to arise from pooling the interaction data between words and frames. Since the encoding of the video clip depends on the sentence compared to it, the representation needs to be recomputed for each potential match. To this end, we propose an efficient shallow neural network. Its training employs a hierarchical triplet loss that is extendable to paragraph/video matching. The method is simple, provides explainability, and achieves state-of-the-art results for both sentence-clip and video-text by a sizable margin across five different datasets: ActivityNet, DiDeMo, YouCook2, MSR-VTT, and LSMDC. We also show that our conditioned representation can be transferred to video-guided machine translation, where we improved the current results on VATEX. Source code is available at https://github.com/AmeenAli/VideoMatch.
Ameen Ali, Idan Schwartz, Tamir Hazan, Lior Wolf
WACV1
2021 Visualization of Supervised and Self-Supervised Neural Networks via Attribution Guided Factorization
abstract
Neural network visualization techniques mark image locations by their relevancy to the network's classification. Existing methods are effective in highlighting the regions that affect the resulting classification the most. However, as we show, these methods are limited in their ability to identify the support for alternative classifications, an effect we name the saliency bias hypothesis. In this work, we integrate two lines of research: gradient-based methods and attribution-based methods, and develop an algorithm that provides per-class explainability. The algorithm back-projects the per pixel local influence, in a manner that is guided by the local attributions, while correcting for salient features that would otherwise bias the explanation. In an extensive battery of experiments, we demonstrate the ability of our methods to class-specific visualization, and not just the predicted label. Remarkably, the method obtains state of the art results in benchmarks that are commonly applied to gradient-based methods as well as in those that are employed mostly for evaluating attribution methods. Using a new unsupervised procedure, our method is also successful in demonstrating that self-supervised methods learn semantic information. Our code is available at: https://github.com/shirgur/AGFVisualization.
Shir Gur, Ameen Ali, Lior Wolf
AAAI2