Fawaz Sammani

dblp:248/8242 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-7659-4410ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 62% Transfer learning and domain adaptation · 16% Vision and language · 12%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
3.042025
Zero-Shot Natural Language Explanations · ICLR 2025
Visualizing and Understanding Contrastive Learning · IEEE Trans. Image Process. 2024
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
1.622025
Zero-Shot Natural Language Explanations · ICLR 2025
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
1.422025
Zero-Shot Natural Language Explanations · ICLR 2025
NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks · CVPR 2022
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation
0.812024
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge · NeurIPS 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Visualizing and Understanding Contrastive Learning · IEEE Trans. Image Process. 2024
Computer vision › Vision and language
vision-language model
0.812024
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
visual explanation
0.812024
Visualizing and Understanding Contrastive Learning · IEEE Trans. Image Process. 2024
Computer vision › Vision and language
image captioning
0.412020
Show, Edit and Tell: A Framework for Editing Image Captions · CVPR 2020
Machine learning › Trustworthy machine learning › interpretability › concept-based explanation
concept discovery
0.312025
Zero-Shot Natural Language Explanations · ICLR 2025
Computer vision › Image recognition and object detection
image classification
0.212024
Visualizing and Understanding Contrastive Learning · IEEE Trans. Image Process. 2024

Methods — techniques the papers use, named apart from their topics

multi-layer perceptron · 0.9class embedding mapping · 0.9textual concept-based explanation · 0.8mutual knowledge analysis · 0.8contrastive learning · 0.8pre-training · 0.6language model · 0.6denoising autoencoder · 0.4copy mechanism · 0.4LSTM · 0.4
YearPublicationVenuePosition
2025 Zero-Shot Natural Language Explanations
abstract
Natural Language Explanations (NLEs) interpret the decision-making process of a given model through textual sentences. Current NLEs suffer from a severe limitation; they are unfaithful to the model’s actual reasoning process, as a separate textual decoder is explicitly trained to generate those explanations using annotated datasets for a specific task, leading them to reflect what annotators desire. In this work, we take the first step towards generating faithful NLEs for any visual classification model without any training data. Our approach models the relationship between class embeddings from the classifier of the vision model and their corresponding class names via a simple MLP which trains in seconds. After training, we can map any new text to the classifier space and measure its association with the visual features. We conduct experiments on 38 vision models, including both CNNs and Transformers. In addition to NLEs, our method offers other advantages such as zero-shot image classification and fine-grained concept discovery.
Fawaz Sammani, Nikos Deligiannis
ICLR1
2024 Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
abstract
Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retrieving the class closest to the image. This work provides a new approach for interpreting CLIP models for image classification from the lens of mutual knowledge between the two modalities. Specifically, we ask: what concepts do both vision and language CLIP encoders learn in common that influence the joint embedding space, causing points to be closer or further apart? We answer this question via an approach of textual concept-based explanations, showing their effectiveness, and perform an analysis encompassing a pool of 13 CLIP models varying in architecture, size and pretraining datasets. We explore those different aspects in relation to mutual knowledge, and analyze zero-shot predictions. Our approach demonstrates an effective and human-friendly way of understanding zero-shot classification decisions with CLIP.
Fawaz Sammani, Nikos Deligiannis
NeurIPS1
2024 Visualizing and Understanding Contrastive Learning
abstract
Contrastive learning has revolutionized the field of computer vision, learning rich representations from unlabeled data, which generalize well to diverse vision tasks. Consequently, it has become increasingly important to explain these approaches and understand their inner workings mechanisms. Given that contrastive models are trained with interdependent and interacting inputs and aim to learn invariance through data augmentation, the existing methods for explaining single-image systems (e.g., image classification models) are inadequate as they fail to account for these factors and typically assume independent inputs. Additionally, there is a lack of evaluation metrics designed to assess pairs of explanations, and no analytical studies have been conducted to investigate the effectiveness of different techniques used to explaining contrastive learning. In this work, we design visual explanation methods that contribute towards understanding similarity learning tasks from pairs of images. We further adapt existing metrics, used to evaluate visual explanations of image classification systems, to suit pairs of explanations and evaluate our proposed methods with these metrics. Finally, we present a thorough analysis of visual explainability methods for contrastive learning, establish their correlation with downstream tasks and demonstrate the potential of our approaches to investigate their merits and drawbacks.
Fawaz Sammani, Boris Joukovsky, Nikos Deligiannis
IEEE Trans. Image Process.1
2023 Model-Agnostic Visual Explanations via Approximate Bilinear Models
abstract
This paper proposes InteractionLIME: a model-agnostic attribution technique to explain deep models predictions in terms of feature interactions. Specifically, we regress a bilinear form to approximate the output of two-input models, by sampling perturbations of both inputs simultaneously. Upon training, we retrieve a global explanation and a set of feature partitioning maps via the singular value decomposition of the learned interaction matrix of the bilinear model. We demonstrate InteractionLIME on vision and text-vision contrastive models, using visual examples and quantitative evaluation metrics. Our results show that the bilinear model successfully retrieves important interacting features from both inputs, while strongly reducing the occurrence of incomplete or asymmetric explanations produced by a linear model.
Boris Joukovsky, Fawaz Sammani, Nikos Deligiannis
ICIP2
2022 NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks
abstract
Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, high-level and fine-grained. Current NLE models11Throughout this paper, we refer to NLE models as Natural Language Explanation models aimed for vision and vision-language tasks. explain the decision-making process of a vision or vision-language model (a.k.a., task model), e.g., a VQA model, via a language model (a.k.a., explanation model), e.g., GPT. Other than the additional memory resources and inference time required by the task model, the task and explanation models are completely independent, which disassociates the explanation from the reasoning process made to predict the answer. We introduce NLX-GPT, a general, compact and faithful language model that can simultaneously predict an answer and explain it. We first conduct pre-training on large scale data of image-caption pairs for general understanding of images, and then formulate the answer as a text prediction task along with the explanation. Without region proposals nor a task model, our resulting overall framework attains better evaluation scores, contains much less parameters and is 15× faster than the current SoA model. We then address the problem of evaluating the explanations which can be in many times generic, data-biased and can come in several forms. We therefore design 2 new evaluation measures: (1) explain-predict and (2) retrieval-based attack, a selfevaluation framework that requires no labels. Code is at: https://github.com/fawazsammani/nlxgpt.
Fawaz Sammani, Tanmoy Mukherjee, Nikos Deligiannis
CVPR1
2020 Show, Edit and Tell: A Framework for Editing Image Captions
abstract
Most image captioning frameworks generate captions directly from images, learning a mapping from visual features to natural language. However, editing existing captions can be easier than generating new ones from scratch. Intuitively, when editing captions, a model is not required to learn information that is already present in the caption (i.e. sentence structure), enabling it to focus on fixing details (e.g. replacing repetitive words). This paper proposes a novel approach to image captioning based on iterative adaptive refinement of an existing caption. Specifically, our caption-editing model consisting of two sub-modules: (1) EditNet, a language module with an adaptive copy mechanism (Copy-LSTM) and a Selective Copy Memory Attention mechanism (SCMA), and (2) DCNet, an LSTM-based denoising auto-encoder. These components enable our model to directly copy from and modify existing captions. Experiments demonstrate that our new approach achieves state of-art performance on the MS COCO dataset both with and without sequence-level training.
Fawaz Sammani, Luke Melas-Kyriazi
CVPR1
2019 Look and Modify: Modification Networks for Image Captioning
Fawaz Sammani, Mahmoud Elsayed
BMVC1