Mohamed Ashraf Abdelsalam

dblp:332/0937 · also Mohamed A. Abdelsalam · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 41% Learning paradigms · 21% Video understanding and tracking · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › image captioning
controllable image captioning
0.812024
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation · ECCV (66) 2024
Computer vision › Vision and language
image captioning
0.812024
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation · ECCV (66) 2024
Machine learning › Generative modeling › multimodal generation
multimodal generative model
0.712023
GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023
Computer vision › Video understanding and tracking › activity recognition › procedural activity understanding
procedural video understanding
0.712023
GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.512021
IIRC: Incremental Implicitly-Refined Classification · CVPR 2021
Computer vision › Image recognition and object detection › image classification › hierarchical classification
coarse-to-fine classification
0.512021
IIRC: Incremental Implicitly-Refined Classification · CVPR 2021
Machine learning › Learning paradigms
continual learning
0.512021
IIRC: Incremental Implicitly-Refined Classification · CVPR 2021
Computer vision › Vision and language
video-language modeling
0.212023
GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023
Performance modeling and evaluation
benchmarking
0.112021
IIRC: Incremental Implicitly-Refined Classification · CVPR 2021

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.0structured semantic augmentation · 0.8zero-shot transfer · 0.7pretraining on text corpus · 0.7generative model · 0.7
YearPublicationVenuePosition
2024 CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Kalliopi Basioti, Mohamed Ashraf Abdelsalam, Federico Fancellu, Vladimir Pavlovic 0001, Afsaneh Fazly
ECCV (66)2
2023 GePSAn: Generative Procedure Step Anticipation in Cooking Videos
abstract
We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focuses on the problem of data scarcity in procedural video datasets, another core challenge of future anticipation is how to account for multiple plausible future realizations in natural settings. This problem has been largely overlooked in previous work. To address this challenge, we frame future step prediction as modelling the distribution of all possible candidates for the next step. Specifically, we design a generative model that takes a series of video clips as input, and generates multiple plausible and diverse candidates (in natural language) for the next step. Following previous work, we side-step the video annotation scarcity by pretraining our model on a large text-based corpus of procedural activities, and then transfer the model to the video domain. Our experiments, both in textual and video domains, show that our model captures diversity in the next step prediction and generates multiple plausible future predictions. Moreover, our model establishes new state-of-the-art results on YouCookII, where it outperforms existing baselines on the next step anticipation. Finally, we also show that our model can successfully transfer from text to the video domain zero-shot, i.e., without fine-tuning or adaptation, and produces good-quality future step predictions from video.
Mohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis, Afsaneh Fazly
ICCV1
2022 Visual Semantic Parsing: From Images to Abstract Meaning Representation
abstract
Mohamed Ashraf Abdelsalam, Zhan Shi, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic, Afsaneh Fazly. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022.
Mohamed Ashraf Abdelsalam, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic 0001, Afsaneh Fazly
CoNLL1
2021 IIRC: Incremental Implicitly-Refined Classification
abstract
We introduce the "Incremental Implicitly-Refined Classification (IIRC)" setup, an extension to the class incremental learning setup where the incoming batches of classes have two granularity levels. i.e., each sample could have a high-level (coarse) label like "bear" and a low-level (fine) label like "polar bear". Only one label is provided at a time, and the model has to figure out the other label if it has already learned it. This setup is more aligned with real-life scenarios, where a learner usually interacts with the same family of entities multiple times, discovers more granularity about them, while still trying not to forget previous knowledge. Moreover, this setup enables evaluating models for some important lifelong learning challenges that cannot be easily addressed under the existing setups. These challenges can be motivated by the example "if a model was trained on the class bear in one task and on polar bear in another task, will it forget the concept of bear, will it rightfully infer that a polar bear is still a bear? and will it wrongfully associate the label of polar bear to other breeds of bear?". We develop a standardized benchmark that enables evaluating models on the IIRC setup. We evaluate several state-of-the-art lifelong learning algorithms and highlight their strengths and limitations. For example, distillation-based methods perform relatively well but are prone to incorrectly predicting too many labels per image. We hope that the proposed setup, along with the benchmark, would provide a meaningful problem setting to the practitioners.
Mohamed Ashraf Abdelsalam, Mojtaba Faramarzi, Shagun Sodhani, Sarath Chandar
CVPR1
2021 A Brief Study on the Effects of Training Generative Dialogue Models with a Semantic loss
abstract
Neural models trained for next utterance generation in dialogue task learn to mimic the n-gram sequences in the training set with training objectives like negative log-likelihood (NLL) or cross-entropy.Such commonly used training objectives do not foster generating alternate responses to a context.But, the effects of minimizing an alternate training objective that fosters a model to generate alternate response and score it on semantic similarity has not been well studied.We hypothesize that a language generation model can improve on its diversity by learning to generate alternate text during training and minimizing a semantic loss as an auxiliary objective.We explore this idea on two different sized data sets on the task of next utterance generation in goal oriented dialogues.We make two observations (1) minimizing a semantic objective improved diversity in responses in the smaller data set (Frames) but only as-good-as minimizing the NLL in the larger data set (Mul-tiWoZ) (2) large language model embeddings can be more useful as a semantic loss objective than as initialization for token embeddings.
Prasanna Parthasarathi, Mohamed Ashraf Abdelsalam, Sarath Chandar, Joelle Pineau
SIGDIAL2