VLDB 2026 Research / reviewers in the wild / expert
Mohamed Ashraf Abdelsalam
dblp:332/0937 · also Mohamed A. Abdelsalam
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 41% Learning paradigms · 21% Video understanding and tracking · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › image captioning
controllable image captioning |
0.8 | 1 | 2024 | CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation · ECCV (66) 2024 |
Computer vision › Vision and language
image captioning |
0.8 | 1 | 2024 | CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation · ECCV (66) 2024 |
Machine learning › Generative modeling › multimodal generation
multimodal generative model |
0.7 | 1 | 2023 | GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023 |
Computer vision › Video understanding and tracking › activity recognition › procedural activity understanding
procedural video understanding |
0.7 | 1 | 2023 | GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.5 | 1 | 2021 | IIRC: Incremental Implicitly-Refined Classification · CVPR 2021 |
Computer vision › Image recognition and object detection › image classification › hierarchical classification
coarse-to-fine classification |
0.5 | 1 | 2021 | IIRC: Incremental Implicitly-Refined Classification · CVPR 2021 |
Machine learning › Learning paradigms
continual learning |
0.5 | 1 | 2021 | IIRC: Incremental Implicitly-Refined Classification · CVPR 2021 |
Computer vision › Vision and language
video-language modeling |
0.2 | 1 | 2023 | GePSAn: Generative Procedure Step Anticipation in Cooking Videos · ICCV 2023 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2021 | IIRC: Incremental Implicitly-Refined Classification · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 1.0structured semantic augmentation · 0.8zero-shot transfer · 0.7pretraining on text corpus · 0.7generative model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Kalliopi Basioti, Mohamed Ashraf Abdelsalam, Federico Fancellu, Vladimir Pavlovic 0001, Afsaneh Fazly |
ECCV (66) | 2 |
| 2023 | GePSAn: Generative Procedure Step Anticipation in Cooking VideosabstractWe study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focuses on the problem of data scarcity in procedural video datasets, another core challenge of future anticipation is how to account for multiple plausible future realizations in natural settings. This problem has been largely overlooked in previous work. To address this challenge, we frame future step prediction as modelling the distribution of all possible candidates for the next step. Specifically, we design a generative model that takes a series of video clips as input, and generates multiple plausible and diverse candidates (in natural language) for the next step. Following previous work, we side-step the video annotation scarcity by pretraining our model on a large text-based corpus of procedural activities, and then transfer the model to the video domain. Our experiments, both in textual and video domains, show that our model captures diversity in the next step prediction and generates multiple plausible future predictions. Moreover, our model establishes new state-of-the-art results on YouCookII, where it outperforms existing baselines on the next step anticipation. Finally, we also show that our model can successfully transfer from text to the video domain zero-shot, i.e., without fine-tuning or adaptation, and produces good-quality future step predictions from video. Mohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis, Afsaneh Fazly |
ICCV | 1 |
| 2022 | Visual Semantic Parsing: From Images to Abstract Meaning RepresentationabstractMohamed Ashraf Abdelsalam, Zhan Shi, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic, Afsaneh Fazly. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022. Mohamed Ashraf Abdelsalam, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic 0001, Afsaneh Fazly |
CoNLL | 1 |
| 2021 | IIRC: Incremental Implicitly-Refined ClassificationabstractWe introduce the "Incremental Implicitly-Refined Classification (IIRC)" setup, an extension to the class incremental learning setup where the incoming batches of classes have two granularity levels. i.e., each sample could have a high-level (coarse) label like "bear" and a low-level (fine) label like "polar bear". Only one label is provided at a time, and the model has to figure out the other label if it has already learned it. This setup is more aligned with real-life scenarios, where a learner usually interacts with the same family of entities multiple times, discovers more granularity about them, while still trying not to forget previous knowledge. Moreover, this setup enables evaluating models for some important lifelong learning challenges that cannot be easily addressed under the existing setups. These challenges can be motivated by the example "if a model was trained on the class bear in one task and on polar bear in another task, will it forget the concept of bear, will it rightfully infer that a polar bear is still a bear? and will it wrongfully associate the label of polar bear to other breeds of bear?". We develop a standardized benchmark that enables evaluating models on the IIRC setup. We evaluate several state-of-the-art lifelong learning algorithms and highlight their strengths and limitations. For example, distillation-based methods perform relatively well but are prone to incorrectly predicting too many labels per image. We hope that the proposed setup, along with the benchmark, would provide a meaningful problem setting to the practitioners. Mohamed Ashraf Abdelsalam, Mojtaba Faramarzi, Shagun Sodhani, Sarath Chandar |
CVPR | 1 |
| 2021 | A Brief Study on the Effects of Training Generative Dialogue Models with a Semantic lossabstractNeural models trained for next utterance generation in dialogue task learn to mimic the n-gram sequences in the training set with training objectives like negative log-likelihood (NLL) or cross-entropy.Such commonly used training objectives do not foster generating alternate responses to a context.But, the effects of minimizing an alternate training objective that fosters a model to generate alternate response and score it on semantic similarity has not been well studied.We hypothesize that a language generation model can improve on its diversity by learning to generate alternate text during training and minimizing a semantic loss as an auxiliary objective.We explore this idea on two different sized data sets on the task of next utterance generation in goal oriented dialogues.We make two observations (1) minimizing a semantic objective improved diversity in responses in the smaller data set (Frames) but only as-good-as minimizing the NLL in the larger data set (Mul-tiWoZ) (2) large language model embeddings can be more useful as a semantic loss objective than as initialization for token embeddings. Prasanna Parthasarathi, Mohamed Ashraf Abdelsalam, Sarath Chandar, Joelle Pineau |
SIGDIAL | 2 |