VLDB 2026 Research / reviewers in the wild / expert
Mennatullah Siam
dblp:163/9048
· DBLP profile ↗
20ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-1854-3698ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TAM-VT: Transformation-Aware Multi-Scale Video Transformer for Segmentation and TrackingabstractVideo Object Segmentation (VOS) has emerged as an increasingly important problem with availability of larger datasets and more complex and realistic settings, which involve long videos with global motion (e.g., in egocentric settings), depicting small objects undergoing both rigid and non-rigid (including state) deformations. While a number of recent approaches have been explored for this task, these data characteristics still present challenges. In this work we propose a novel, clip-based DETR-style encoder-decoder architecture, which focuses on systematically analyzing and addressing aforementioned challenges. Specifically, we propose a novel transformation-aware loss thatfocuses learning on portions of the video where an object undergoes significant deformations - a form of “soft” hard examples mining. Further, we propose a multiplicative time-coded memory, beyond vanilla additive positional encoding, which helps propagate context across long videos. Finally, we incorporate these in our proposed holistic multiscale video transformer for tracking via multi-scale memory matching and decoding to ensure sensitivity and accuracy for long videos and small objects. Our model enables on-line inference with long videos in a windowed fashion, by breaking the video into clips and propagating context among them. We illustrate that short clip length and longer memory with learned time-coding are important design choices for improved performance. Collectively, these technical contributions enable our model to achieve new state-of-the-art (SoTA) performance on two complex egocentric datasets - VISOR [13] and VOST [44], while achieving comparable to SoTA results on the conventional VOS benchmark, DAVIS’ 17 [38]. Detailed ablations validate our design choices and provide insights into the importance of parameter choices and impact on performance.11Code is available at the project page. Raghav Goyal, Wan-Cyuan Fan, Mennatullah Siam, Leonid Sigal |
WACV | 3 |
| 2025 | Temporal Transductive Inference for Few-Shot Video Object Segmentation
Mennatullah Siam |
Int. J. Comput. Vis. | 1 |
| 2025 | Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal NetworksabstractThere is limited understanding of the information captured by deep spatiotemporal models in their intermediate representations. For example, while evidence suggests that action recognition algorithms are heavily influenced by visual appearance in single frames, no quantitative methodology exists for evaluating such static bias in the latent representation compared to bias toward dynamics. We tackle this challenge by proposing an approach for quantifying the static and dynamic biases of any spatiotemporal model, and apply our approach to three tasks, action recognition, automatic video object segmentation (AVOS) and video instance segmentation (VIS). Our key findings are: (i) Most examined models are biased toward static information. (ii) Some datasets that are assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual channels in an architecture can be biased toward static, dynamic or jointly encode a combination static and dynamic information. (iv) Most models converge to their culminating biases in the first half of training. We then explore how these biases affect performance on dynamically biased datasets. For action recognition, we propose StaticDropout, a semantically guided dropout that debiases a model from static information toward dynamics. For AVOS, we design a better combination of fusion and cross connection layers compared with previous architectures. Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale ApproachabstractThe emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research has demonstrated that such models, when prompted appropriately, are excellent for few-shot inference. However, such techniques are under-explored for dense prediction tasks like semantic segmentation. In this work, we examine the effectiveness of prompting a transformer-decoder with learned visual prompts for the generalized few-shot segmentation (GFSS) task. Our goal is to achieve strong performance not only on novel categories with limited examples, but also to retain performance on base cat-egories. We propose an approach to learn visual prompts with limited examples. These learned visual prompts are used to prompt a multiscale transformer decoder to facilitate accurate dense predictions. Additionally, we introduce a unidirectional causal attention mechanism between the novel prompts, learned with limited examples, and the base prompts, learned with abundant data. This mecha-nism enriches the novel prompts without deteriorating the base class performance. Overall, this form of prompting helps us achieve state-of-the-art performance for GFSS on two different benchmark datasets: COCO-20i and Pascal-Si, without the need for test-time optimization (or transduction). Furthermore, test-time optimization leveraging unlabelled test data can be used to improve the prompts, which we refer to as transductive prompt tuning. Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little |
CVPR | 2 |
| 2024 | The State of Computer Vision Research in AfricaabstractDespite significant efforts to democratize artificial intelligence (AI), computer vision which is a sub-field of AI, still lags in Africa. A significant factor to this, is the limited access to computing resources, datasets, and collaborations. As a result, Africa’s contribution to top-tier publications in this field has only been 0.06% over the past decade. Towards improving the computer vision field and making it more accessible and inclusive, this study analyzes 63,000 Scopus-indexed computer vision publications from Africa. We utilize large language models to automatically parse their abstracts, to identify and categorize topics and datasets. This resulted in listing more than 100 African datasets. Our objective is to provide a comprehensive taxonomy of dataset categories to facilitate better understanding and utilization of these resources. We also analyze collaboration trends of researchers within and outside the continent. Additionally, we conduct a large-scale questionnaire among African computer vision researchers to identify the structural barriers they believe require urgent attention. In conclusion, our study offers a comprehensive overview of the current state of computer vision research in Africa, to empower marginalized communities to participate in the design and development of computer vision systems. Abdul-Hakeem Omotayo, Ashery Mbilinyi, Lukman Ismaila, Houcemeddine Turki, Mahmod Abdien, Karim Gamal, Idriss Tondji, Yvan Pimi, Naome A. Etori, Marwa M. Matar, Clifford Broni-Bediako, Abigail Oppong, Mai Gamal, Eman Ehab, Gbètondji J.-S. Dovonon, Zainab Akinjobi, Daniel Ajisafe, Oluwabukola Grace Adegboro, Mennatullah Siam |
J. Artif. Intell. Res. | 19 |
| 2024 | Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and BenchmarkabstractLearning with limited labeled data is a challenging problem in various applications, including remote sensing. Few-shot semantic segmentation is one approach that can encourage deep learning models to learn from few labeled examples for novel classes not seen during the training. The generalized few-shot segmentation setting has an additional challenge which encourages models not only to adapt to the novel classes but also to maintain strong performance on the training base classes. While previous datasets and benchmarks discussed the few-shot segmentation setting in remote sensing, we are the first to propose a generalized few-shot segmentation benchmark for remote sensing. The generalized setting is more realistic and challenging, which necessitates exploring it within the remote sensing context. We release the dataset augmenting OpenEarthMap (OEM) with additional classes labeled for the generalized few-shot evaluation setting. The dataset is released during the OEM land cover mapping generalized few-shot challenge in the learning with limited labeled data for image and video understanding (L3D-IVU) workshop in conjunction with computer vision and pattern recognition (CVPR) 2024. In this work, we summarize the dataset and challenge details in addition to providing the benchmark results on the two phases of the challenge for the validation and test sets. Clifford Broni-Bediako, Junshi Xia, Jian Song 0010, Hongruixuan Chen, Mennatullah Siam, Naoto Yokoya |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | MED-VT: Multiscale Encoder-Decoder Video Transformer with Application to Object SegmentationabstractMultiscale video transformers have been explored in a wide variety of vision tasks. To date, however, the multiscale processing has been confined to the encoder or decoder alone. We present a unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in videos. Multiscale representation at both encoder and decoder yields key benefits of implicit extraction of spatiotemporal features (i.e. without reliance on input optical flow) as well as temporal consistency at encoding and coarse-to-fine detection for high-level (e.g. object) semantics to guide precise localization at decoding. Moreover, we propose a transductive learning scheme through many-to-many label propagation to provide temporally consistent predictions. We showcase our Multiscale Encoder-Decoder Video Transformer (MED-VT) on Automatic Video Object Segmentation (AVOS) and actor/action segmentation, where we outperform state-of-the-art approaches on multiple benchmarks using only raw images, without using optical flow. Rezaul Karim, He Zhao 0004, Richard P. Wildes, Mennatullah Siam |
CVPR | 4 |
| 2023 | AI-Assisted Tool for Early Diagnosis and Prevention of Colorectal Cancer in AfricaabstractColorectal cancer (CRC) is considered the third most common cancer worldwide and is recently increasing in Africa. It is mostly diagnosed at an advanced state causing high fatality rates, which highlights the importance of CRC early diagnosis. There are various methods used to enable early diagnosis of CRC, which are vital to increase survival rates such as colonoscopy. Recently, there are calls to start an early detection program in Egypt using colonoscopy. It can be used for diagnosis and prevention purposes to detect and remove polyps, which are benign growths that have the risk of turning into cancer. However, there tends to be a high miss rate of polyps from physicians, which motivates machine learning guided polyp segmentation methods in colonoscopy videos to aid physicians. To date, there are no large-scale video polyp segmentation dataset that is focused on African countries. It was shown in AI-assisted systems that under-served populations such as patients with African origin can be misdiagnosed. There is also a potential need in other African countries beyond Egypt to provide a cost efficient tool to record colonoscopy videos using smart phones without relying on video recording equipment. Since most of the equipment used in Africa are old and refurbished, and video recording equipment can get defective. Hence, why we propose to curate a colonoscopy video dataset focused on African patients, provide expert annotations for video polyp segmentation and provide an AI-assisted tool to record colonoscopy videos using smart phones. Our project is based on our core belief in developing research by Africans and increasing the computer vision research capacity in Africa. Bushra Ibnauf, Mohammed Aboul Ezz, Ayman Abdel Aziz, Khalid Elgazzar, Mennatullah Siam |
IJCAI | 5 |
| 2022 | A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic InformationabstractDeep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been observed that action recognition algorithms are heavily influenced by visual appearance in single static frames, there is no quantitative methodology for evaluating such static bias in the latent representation compared to bias toward dynamic information (e.g. motion). We tackle this challenge by proposing a novel approach for quantifying the static and dynamic biases of any spatiotemporal model. To show the efficacy of our approach, we analyse two widely studied tasks, action recognition and video object segmentation. Our key findings are threefold: (i) Most examined spatiotemporal models are biased toward static information; although, certain two-stream architectures with cross-connections show a better balance between the static and dynamic information captured. (ii) Some datasets that are commonly assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual units (channels) in an architecture can be biased toward static, dynamic or a combination of the two.11Project page and code Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis |
CVPR | 2 |
| 2021 | Monocular Instance Motion Segmentation for Autonomous Driving: KITTI InstanceMotSeg Dataset and Multi-Task BaselineabstractMoving object segmentation is a crucial task for autonomous vehicles as it can be used to segment objects in a class agnostic manner based on their motion cues. It enables the detection of unseen objects during training (e.g., moose or a construction truck) based on their motion and independent of their appearance. Although pixel-wise motion segmentation has been studied in autonomous driving literature, it has been rarely addressed at the instance level, which would help separate connected segments of moving objects leading to better trajectory planning. As the main issue is the lack of large public datasets, we create a new InstanceMotSeg dataset comprising of 12.9K samples improving upon our KITTIMoSeg dataset. In addition to providing instance level annotations, we have added 4 additional classes which is crucial for studying class agnostic motion segmentation. We adapt YOLACT and implement a motion-based class agnostic instance segmentation model which would act as a baseline for the dataset. We also extend it to an efficient multi-task model which additionally provides semantic instance segmentation sharing the encoder. The model then learns separate prototype coefficients within the class agnostic and semantic heads providing two independent paths of object detection for redundant safety. To obtain real-time performance, we study different efficient encoders and obtain 39 fps on a Titan Xp GPU using MobileNetV2 with an improvement of 10% mAP relative to the baseline. Our model improves the previous state of the art motion segmentation method by 3.3%. The dataset and qualitative results video are shared in our website at https://sites.google.com/view/instancemotseg. Eslam Mohamed, Mahmoud Ewaisha, Mennatullah Siam, Hazem Rashed, Senthil Kumar Yogamani, Waleed Hamdy, Mohamed El-Dakdouky, Ahmad El Sallab |
IV | 3 |
| 2020 | Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic EmbeddingsabstractSignificant progress has been made recently in developing few-shot object segmentation methods. Learning is shown to be successful in few-shot segmentation settings, using pixel-level, scribbles and bounding box supervision. This paper takes another approach, i.e., only requiring image-level label for few-shot object segmentation. We propose a novel multi-modal interaction module for few-shot object segmentation that utilizes a co-attention mechanism using both visual and word embedding. Our model using image-level labels achieves 4.8% improvement over previously proposed image-level few-shot object segmentation. It also outperforms state-of-the-art methods that use weak bounding box supervision on PASCAL-5^i. Our results show that few-shot segmentation benefits from utilizing word embeddings, and that we are able to perform few-shot segmentation using stacked joint visual semantic processing with weak image-level labels. We further propose a novel setup, Temporal Object Segmentation for Few-shot Learning (TOSFL) for videos. TOSFL can be used on a variety of public video data such as Youtube-VOS, as demonstrated in both instance-level and category-level TOSFL experiments. Mennatullah Siam, Naren Doraiswamy, Boris N. Oreshkin, Hengshuai Yao, Martin Jägersand |
IJCAI | 1 |
| 2019 | AMP: Adaptive Masked Proxies for Few-Shot SegmentationabstractDeep learning has thrived by training on large-scale datasets. However, in robotics applications sample efficiency is critical. We propose a novel adaptive masked proxies method that constructs the final segmentation layer weights from few labelled samples. It utilizes multiresolution average pooling on base embeddings masked with the label to act as a positive proxy for the new class, while fusing it with the previously learned class signatures. Our method is evaluated on PASCAL-5idataset and outperforms the state-of-the-art in the few-shot semantic segmentation. Unlike previous methods, our approach does not require a second branch to estimate parameters or prototypes, which enables it to be used with 2-stream motion and appearance based segmentation networks. We further propose a novel setup for evaluating continual learning of object segmentation which we name incremental PASCAL (iPASCAL) where our method outperforms the baseline method. Our code is publicly available at https://github. com/MSiam/AdaptiveMaskedProxies. Mennatullah Siam, Boris N. Oreshkin, Martin Jägersand |
ICCV | 1 |
| 2019 | Online Object and Task Learning via Human Robot InteractionabstractThis work describes the development of a robotic system that acquires knowledge incrementally through human interaction where new objects and motions are taught on the fly. The robotic system developed was one of the five finalists in the KUKA Innovation Award competition and demonstrated during the Hanover Messe 2018 in Germany. The main contributions of the system are i) a novel incremental object learning module - a deep learning based localization and recognition system - that allows a human to teach new objects to the robot, ii) an intuitive user interface for specifying 3D motion task associated with the new object, and iii) a hybrid force-vision control module for performing compliant motion on an unstructured surface. This paper describes the implementation and integration of the main modules of the system and summarizes the lessons learned from the competition. Masood Dehghan, Zichen Zhang 0001, Mennatullah Siam, Jun Jin 0001, Laura Petrich, Martin Jägersand |
ICRA | 3 |
| 2019 | Video Object Segmentation using Teacher-Student Adaptation in a Human Robot Interaction (HRI) SettingabstractVideo object segmentation is an essential task in robot manipulation to facilitate grasping and learning affordances. Incremental learning is important for robotics in unstructured environments. Inspired by the children learning process, human robot interaction (HRI) can be utilized to teach robots about the world guided by humans similar to how children learn from a parent or a teacher. A human teacher can show potential objects of interest to the robot, which is able to self adapt to the teaching signal without providing manual segmentation labels. We propose a novel teacher-student learning paradigm to teach robots about their surrounding environment. A two-stream motion and appearance “teacher” network provides pseudo-labels to adapt an appearance “student” network. The student network is able to segment the newly learned objects in other scenes, whether they are static or in motion. We also introduce a carefully designed dataset that serves the proposed HRI setup, denoted as (I)nteractive (V)ideo (O)bject (S)egmentation. Our IVOS dataset contains teaching videos of different objects, and manipulation tasks. Our proposed adaptation method outperforms the state-of-theart on DAVIS and FBMS with 6.8% and 1.2% in F-measure respectively. It improves over the baseline on IVOS dataset with 46.1% and 25.9% in mIoU. Mennatullah Siam, Steven Weikai Lu, Laura Petrich, Mahmoud Gamal, Martin Jägersand |
ICRA | 1 |
| 2018 | RTSeg: Real-Time Semantic Segmentation Comparative StudyabstractSemantic segmentation benefits robotics related applications, especially autonomous driving. Most of the research on semantic segmentation only focuses on increasing the accuracy of segmentation models with little attention to computationally efficient solutions. The few work conducted in this direction does not provide principled methods to evaluate the different design choices for segmentation. In this paper, we address this gap by presenting a real-time semantic segmentation benchmarking framework with a decoupled design for feature extraction and decoding methods. The framework is comprised of different network architectures for feature extraction such as VGG16, Resnet18, MobileNet, and ShuffleNet. It is also comprised of multiple meta-architectures for segmentation that define the decoding methodology. These include SkipNet, UNet, and Dilation Frontend. Experimental results are presented on the Cityscapes dataset for urban scenes. The modular design allows novel architectures to emerge, that lead to 143x GFLOPs reduction in comparison to SegNet. This benchmarking framework is publicly available at11https://github.com/MSiam/TFSegmentation. Mennatullah Siam, Mostafa Gamal, Moemen Abdel-Razek, Senthil Kumar Yogamani, Martin Jägersand |
ICIP | 1 |
| 2018 | Real-Time Segmentation with Appearance, Motion and GeometryabstractReal-time Segmentation is of crucial importance to robotics related applications such as autonomous driving, driving assisted systems, and traffic monitoring from unmanned aerial vehicles imagery. We propose a novel two-stream convolutional network for motion segmentation, which exploits flow and geometric cues to balance the accuracy and computational efficiency trade-offs. The geometric cues take advantage of the domain knowledge of the application. In case of mostly planar scenes from high altitude unmanned aerial vehicles (UAVs), homography compensated flow is used. While in the case of urban scenes in autonomous driving, with GPS/IMU sensory data available, sparse projected depth estimates and odometry information are used. The network provides 4.7× speedup over the state of the art networks in motion segmentation from 153ms to 36ms, at the expense of a reduction in the segmentation accuracy in terms of pixel boundaries. This enables the network to perform real-time on a Jetson T×2. In order to recuperate some of the accuracy loss, geometric priors is used while still achieving a much improved computational efficiency with respect to the state-of-the-art. The usage of geometric priors improved the segmentation in UAV imagery by 5.2 % using the metric of IoU over the baseline network. While on KITTI-MoSeg the sparse depth estimates improved the segmentation by 12.5 % over the baseline. Our proposed motion segmentation solution is verified on the popular KITTI and VIVID datasets, with additional labels we have produced. The code for our work is publicly available at1. Mennatullah Siam, Sara Elkerdawy, Mostafa Gamal, Moemen Abdel-Razek, Martin Jägersand, Hong Zhang 0013 |
IROS | 1 |
| 2017 | Convolutional gated recurrent networks for video segmentationabstractSemantic segmentation has recently witnessed major progress, but most of the previous work focused on improving single image segmentation. In this paper, we introduce a novel approach to implicitly utilize temporal data in videos for online segmentation. This design receives a sequence of consecutive video frames and outputs the segmentation of the last frame. Convolutional gated recurrent networks are used for the recurrent part to preserve spatial connectivities in the image. This architecture is tested for both binary and semantic video segmentation tasks. Experiments are conducted on the recent benchmarks in SegTrack V2, Davis, Camvid, and Synthia. Using recurrent fully convolutional networks improved the baseline network performance in all of our experiments. Namely, 5% and 3% improvement of F-measure in SegTrack2 and Davis respectively, 5.7% and 1.6% improvement in mean IoU in Synthia and Camvid. Thus, RFCN networks can be seen as a method to improve any baseline segmentation network by embedding them into a recurrent module that utilizes temporal data. Mennatullah Siam, Sepehr Valipour, Martin Jägersand, Nilanjan Ray |
ICIP | 1 |
| 2017 | Unifying Registration Based Tracking: A Case Study with Structural SimilarityabstractThis paper adapts a popular image quality measure called structural similarity for high precision registration based tracking while also introducing a simpler and faster variant of the same. Further, these are evaluated comprehensively against existing measures using a unified approach to study registration based trackers that decomposes them into three constituent sub modules - appearance model, state space model and search method. Several popular trackers in literature are broken down using this method so that their contributions - as of this paper - are shown to be limited to only one or two of these submodules. An open source tracking framework is made available that follows this decomposition closely through extensive use of generic programming. It is used to perform all experiments on four publicly available datasets so the results are easily reproducible. This framework provides a convenient interface to plug in a new methodfor any sub module and combine it with existing methods for the other two. It can also serve as a fast and flexible solution for practical tracking needs due to its highly efficient implementation. Abhineet Singh, Mennatullah Siam, Martin Jägersand |
WACV | 2 |
| 2017 | Recurrent Fully Convolutional Networks for Video SegmentationabstractImage segmentation is an important step in most visual tasks. While convolutional neural networks have shown to perform well on single image segmentation, to our knowledge, no study has been done on leveraging recurrent gated architectures for video segmentation. Accordingly, we propose and implement a novel method for online segmentation of video sequences that incorporates temporal data. The network is built from a fully convolutional network and a recurrent unit that works on a sliding window over the temporal data. We use convolutional gated recurrent unit that preserves the spatial information and reduces the parameters learned. Our method has the advantage that it can work in an online fashion instead of operating over the whole input batch of video frames. The network is tested on video segmentation benchmarks in Segtrack V2 and Davis. It proved to have 5% improvement in Segtrack and 3% improvement in Davis in F-measure over a plain fully convolutional network. Sepehr Valipour, Mennatullah Siam, Martin Jägersand, Nilanjan Ray |
WACV | 2 |
| 2015 | Comparative Evaluation of 3D Pose Estimation of Industrial Objects in RGB Pointclouds
Bjarne Großmann, Mennatullah Siam, Volker Krüger |
ICVS | 2 |