EDBT 2026 Demo / reviewers in the wild / expert
Mohamed Abdelfattah
dblp:18/8169
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0002-5732-4340ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 25% Representation and self-supervised learning · 25% 3D vision · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
1.5 | 2 | 2024 | S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition · ECCV (32) 2024 MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution · ICCV 2025 |
Computer vision › 3D vision › depth estimation
video depth estimation |
0.9 | 1 | 2025 | FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution · ICCV 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.8 | 1 | 2024 | MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › predictive learning
joint-embedding predictive architecture |
0.8 | 1 | 2024 | S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition · ECCV (32) 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
masked contrastive learning |
0.8 | 1 | 2024 | MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024 |
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation |
0.6 | 1 | 2022 | ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022 |
Computer vision › Image recognition and object detection
object detection |
0.6 | 1 | 2022 | ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.6 | 1 | 2022 | ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2024 | MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning · CVPR 2024 |
Computer vision › Vision and language
art image understanding |
0.2 | 1 | 2022 | ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
semantic segmentation · 1.1instance segmentation · 1.1transfer learning · 0.9pretrained single-image depth models · 0.9transformer · 0.8self-supervised learning · 0.8multi-level contrastive learning · 0.8attention-guided masking · 0.8emotion annotation · 0.6cross-cultural analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionabstractA versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three requirements, performing depth estimation on a 2044x1148 streaming video at 24 FPS. We show that, with careful modifications to pretrained single-image depth models, these capabilities are enabled with relatively little data and training. We evaluate our approach across multiple unseen datasets against state-of-the-art depth models, and find that ours outperforms them in terms of boundary sharpness and speed by a significant margin, while maintaining competitive accuracy. We hope our model will enable various applications that require high-resolution depth, such as video editing, and online decision-making, such as robotics. We release all code and model weights at https://github.com/Eyeline-Research/FlashDepth Gene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah, Bharath Hariharan, Noah Snavely, Ning Yu 0006, Paul E. Debevec |
ICCV | 4 |
| 2025 | OSKAR: Omnimodal Self-supervised Knowledge Abstraction and RepresentationabstractWe present OSKAR, the first multimodal foundation model based on bootstrapped latent feature prediction. Unlike generative or contrastive methods, it avoids memorizing unnecessary details (e.g., pixels), and does not require negative pairs, large memory banks, or hand-crafted augmentations. We propose a novel pretraining strategy: given masked tokens from multiple modalities, predict a subset of missing tokens per modality, supervised by momentum-updated uni-modal target encoders. This design efficiently utilizes the model capacity in learning high-level representations while retaining modality-specific information. Further, we propose a scalable design which decouples the compute cost from the number of modalities using a fixed representative token budget—in both input and target tokens—and introduces a parameter-efficient cross-attention predictor that grounds each prediction in the full multimodal context. We instantiate OSKAR on video, skeleton, and text modalities. Extensive experimental results show that OSKAR's unified pretrained encoder outperforms models with specialized architectures of similar size in action recognition (rgb, skeleton, frozen, low-shot) and localization, video-text retrieval, and video question answering. Project website: https://multimodal-oskar.github.io Mohamed Abdelfattah, Kaouther Messaoud, Alexandre Alahi |
NeurIPS | 1 |
| 2024 | MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation LearningabstractCurrent transformer-based skeletal action recognition models tend to focus on a limited set of joints and low-level motion patterns to predict action classes. This results in significant performance degradation under small skeleton perturbations or changing the pose estimator between training and testing. In this work, we introduce MaskCLR, a new Masked Contrastive Learning approach for Robust skeletal action recognition. We propose an Attention-Guided Proba-bilistic Masking strategy to occlude the most important joints and encourage the model to explore a larger set of discrimi-native joints. Furthermore, we propose a Multi-Level Contrastive Learning paradigm to enforce the representations of standard and occluded skeletons to be class-discriminative, i.e., more compact within each class and more dispersed across different classes. Our approach helps the model capture the high-level action semantics instead of low-level joint variations, and can be conveniently incorporated into transformer-based models. Without loss of generality, we combine MaskCLR with three transformer backbones: the vanilla transformer, DSTFormer, and STTFormer. Extensive experiments on NTU60, NTU120, and Kinetics400 show that MaskCLR consistently outperforms previous state-of-the-art methods on standard and perturbed skeletons from different pose estimators, showing improved accuracy, generalization, and robustness. Project website: https://maskclr.github.io. Mohamed Abdelfattah, Mariam Hassan, Alexandre Alahi |
CVPR | 1 |
| 2024 | S-JEPA: A Joint Embedding Predictive Architecture for Skeletal Action Recognition
Mohamed Abdelfattah, Alexandre Alahi |
ECCV (32) | 1 |
| 2022 | ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered ScenesabstractLess than 35% of recyclable waste is being actually recycled in the US [2], which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/ Dina Bashkirova, Mohamed Abdelfattah, Ziliang Zhu, James Akl, Fadi M. Alladkani, Ping Hu 0001, Vitaly Ablavsky, Berk Çalli, Sarah Adel Bargal, Kate Saenko |
CVPR | 2 |
| 2022 | ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureabstractYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Xiangliang Zhang 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 2 |
| 2012 | Transparent structural online test for reconfigurable systemsabstractFPGA-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. However, the reliability of modern FPGAs is threatened by latent defects and aging effects. Hence, it is mandatory to ensure the reliable operation of the FPGA's reconfigurable fabric. This can be achieved by periodic or on-demand online testing. In this paper, a system-integrated, transparent structural online test method for runtime reconfigurable systems is proposed. The required tests are scheduled like functional workloads, and thorough optimizations of the test overhead reduce the performance impact. The proposed scheme has been implemented on a reconfigurable system. The results demonstrate that thorough testing of the reconfigurable fabric can be achieved at negligible performance impact on the application. Mohamed Abdelfattah, Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich |
IOLTS | 1 |