Mohamed Abdelfattah

dblp:18/8169 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0002-5732-4340ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 25% Representation and self-supervised learning · 25% 3D vision · 19%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
1.522024
S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition · ECCV (32) 2024
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024
Computer vision › 3D vision
depth estimation
0.912025
FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution · ICCV 2025
Computer vision › 3D vision › depth estimation
video depth estimation
0.912025
FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution · ICCV 2025
Computer vision › Video understanding and tracking
action recognition
0.812024
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › predictive learning
joint-embedding predictive architecture
0.812024
S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition · ECCV (32) 2024
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
masked contrastive learning
0.812024
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Computer vision › Image recognition and object detection
object detection
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Computer vision › Segmentation and scene understanding
object segmentation
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Machine learning › Trustworthy machine learning
robustness
0.212024
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024
Computer vision › Vision and language
art image understanding
0.212022
ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

semantic segmentation · 1.1instance segmentation · 1.1transfer learning · 0.9pretrained single-image depth models · 0.9transformer · 0.8self-supervised learning · 0.8multi-level contrastive learning · 0.8attention-guided masking · 0.8emotion annotation · 0.6cross-cultural analysis · 0.6
YearPublicationVenuePosition
2025 FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution
abstract
A versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three requirements, performing depth estimation on a 2044x1148 streaming video at 24 FPS. We show that, with careful modifications to pretrained single-image depth models, these capabilities are enabled with relatively little data and training. We evaluate our approach across multiple unseen datasets against state-of-the-art depth models, and find that ours outperforms them in terms of boundary sharpness and speed by a significant margin, while maintaining competitive accuracy. We hope our model will enable various applications that require high-resolution depth, such as video editing, and online decision-making, such as robotics. We release all code and model weights at https://github.com/Eyeline-Research/FlashDepth
Gene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah, Bharath Hariharan, Noah Snavely, Ning Yu 0006, Paul E. Debevec
ICCV4
2025 OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation
abstract
We present OSKAR, the first multimodal foundation model based on bootstrapped latent feature prediction. Unlike generative or contrastive methods, it avoids memorizing unnecessary details (e.g., pixels), and does not require negative pairs, large memory banks, or hand-crafted augmentations. We propose a novel pretraining strategy: given masked tokens from multiple modalities, predict a subset of missing tokens per modality, supervised by momentum-updated uni-modal target encoders. This design efficiently utilizes the model capacity in learning high-level representations while retaining modality-specific information. Further, we propose a scalable design which decouples the compute cost from the number of modalities using a fixed representative token budget—in both input and target tokens—and introduces a parameter-efficient cross-attention predictor that grounds each prediction in the full multimodal context. We instantiate OSKAR on video, skeleton, and text modalities. Extensive experimental results show that OSKAR's unified pretrained encoder outperforms models with specialized architectures of similar size in action recognition (rgb, skeleton, frozen, low-shot) and localization, video-text retrieval, and video question answering. Project website: https://multimodal-oskar.github.io
Mohamed Abdelfattah, Kaouther Messaoud, Alexandre Alahi
NeurIPS1
2024 MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning
abstract
Current transformer-based skeletal action recognition models tend to focus on a limited set of joints and low-level motion patterns to predict action classes. This results in significant performance degradation under small skeleton perturbations or changing the pose estimator between training and testing. In this work, we introduce MaskCLR, a new Masked Contrastive Learning approach for Robust skeletal action recognition. We propose an Attention-Guided Proba-bilistic Masking strategy to occlude the most important joints and encourage the model to explore a larger set of discrimi-native joints. Furthermore, we propose a Multi-Level Contrastive Learning paradigm to enforce the representations of standard and occluded skeletons to be class-discriminative, i.e., more compact within each class and more dispersed across different classes. Our approach helps the model capture the high-level action semantics instead of low-level joint variations, and can be conveniently incorporated into transformer-based models. Without loss of generality, we combine MaskCLR with three transformer backbones: the vanilla transformer, DSTFormer, and STTFormer. Extensive experiments on NTU60, NTU120, and Kinetics400 show that MaskCLR consistently outperforms previous state-of-the-art methods on standard and perturbed skeletons from different pose estimators, showing improved accuracy, generalization, and robustness. Project website: https://maskclr.github.io.
Mohamed Abdelfattah, Mariam Hassan, Alexandre Alahi
CVPR1
2024 S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition
Mohamed Abdelfattah, Alexandre Alahi
ECCV (32)1
2022 ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
abstract
Less than 35% of recyclable waste is being actually recycled in the US [2], which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/
Dina Bashkirova, Mohamed Abdelfattah, Ziliang Zhu, James Akl, Fadi M. Alladkani, Ping Hu 0001, Vitaly Ablavsky, Berk Çalli, Sarah Adel Bargal, Kate Saenko
CVPR2
2022 ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture
abstract
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Xiangliang Zhang 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001
EMNLP2
2012 Transparent structural online test for reconfigurable systems
abstract
FPGA-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. However, the reliability of modern FPGAs is threatened by latent defects and aging effects. Hence, it is mandatory to ensure the reliable operation of the FPGA's reconfigurable fabric. This can be achieved by periodic or on-demand online testing. In this paper, a system-integrated, transparent structural online test method for runtime reconfigurable systems is proposed. The required tests are scheduled like functional workloads, and thorough optimizations of the test overhead reduce the performance impact. The proposed scheme has been implemented on a reconfigurable system. The results demonstrate that thorough testing of the reconfigurable fabric can be achieved at negligible performance impact on the application.
Mohamed Abdelfattah, Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich
IOLTS1