VLDB 2026 Research / reviewers in the wild / expert
Mohamed Omar
dblp:50/6339
· DBLP profile ↗
15ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | R-SIEL: a physics-informed learning algorithm for discovering dynamics of serial manipulators
Mohamed Omar, Ruifeng Li 0001, Ke Wang 0028, Ahmed Asker |
Expert Syst. Appl. | 1 |
| 2025 | NL-WCS: A Novel Data-Driven Algorithm for Extracting the Dynamics of Serial Robots Considering Non-Linear FrictionabstractRecently, developed data-driven SINDy-based techniques can identify the dynamic model of serial robots without simplifying assumptions nor pre-knowledge of all kinematics and geometric details. However, these techniques cannot handle non-linear friction models, which significantly affects the precision of the dynamic model identification. This study proposes a novel data-driven approach for dynamic model identification considering the non-linear friction model along with SINDy concept. This approach is termed as non-linear-weighted-constrained SINDy (NL-WCS). The SINDy concept is extended to accommodate any non-linear friction model by efficiently incorporating the Levenberg-Marquardt (LM) algorithm. Weighted L1 regularization is combined with the physics and the data constraints, to promote sparsity in the recovery of the dynamic equations. Moreover, this combination makes the approach robust against the regression matrix’s ill-conditionality and noise. LM is integrated with the robot’s SIMULINK model to get initial values for the non-linear friction empirical parameters. NL-WCS is experimentally evaluated by utilizing three distinct trajectories to validate the extracted dynamic model of 6-DOF UR10 and 7-DOF KUKA robots. In addition, five non-linear friction models are compared. NL-WCS outperforms all the previous SINDy-data-driven methods since it reduced the RMSE significantly up to 60.65%. NL-WCS also demonstrated robustness when tested against different levels of noise. Note to Practitioners—High-performance serial robot model-based controllers need precise dynamic model identification. The classical analytical methods for modeling serial robots rely on deriving the dynamic equations with simplified assumptions and then identifying the inertial parameters. Specifically, they assume that all the robot’s geometric details are known. Furthermore, there are uncertainties in geometric parameter values due to manufacturing/assembly errors. Data-driven-based SINDy methods can derive the dynamic model without pre-knowledge of the geometric parameters of the robot. However, it can’t incorporate non-linear friction models. Thus, this paper proposed a novel data-driven technique to extract the dynamic model of any serial manipulator by coupling SINDy approach with the nonlinear friction models which is the realistic case. The practitioners can benefit from a technique like that, in such a way of applying NL-WCS to different types of industrial robots, particularly, those that have missing manufacturer data sets or sheets. Simple and reliable dynamic models can be derived only by running the robot with various trajectories. Also, this helps eliminate any uncertainties in the dynamic model discovery/building process. Furthermore, Incorporating the non-linear friction model in the derivation process enhances the dynamic modeling accuracy as it effectively takes into account all friction characteristics. The proposed approach can be converted into a software package that the practitioners can use to derive the dynamics without getting their heads around the overwhelming robot’s kinematic details. Mohamed Omar, Ke Wang 0028, Ruifeng Li 0001, Ossama B. Abouelatta, Tao Xie 0010, Mohamed Gouda Alkalla |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | SOFW: A Synergistic Optimization Framework for Indoor 3D Object DetectionabstractIn this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar |
IEEE Trans. Multim. | 9 |
| 2024 | A Multimodal Benchmark and Improved Architecture for Zero Shot LearningabstractIn this work, we demonstrate that due to the inadequacies in the existing evaluation protocols and datasets, there is a need to revisit and comprehensively examine the multimodal Zero-Shot Learning (MZSL) problem formulation. Specifically, we address two major challenges faced by current MZSL approaches; (1) Established baselines are frequently incomparable and occasionally even flawed since existing evaluation datasets often have some overlap with the training dataset, thus violating the zero-shot paradigm; (2) Most existing methods are biased towards seen classes, which significantly reduces the performance when evaluated on both seen and unseen classes. To address these challenges, we first introduce a new multimodal dataset for zero-shot evaluation called MZSL-50 with 4462 videos from 50 widely diversified classes and no overlap with the training data. Further, we propose a novel multimodal zero-shot transformer (MZST) architecture that leverages attention bottlenecks for multimodal fusion. Our model directly predicts the semantic representation and is superior at reducing the bias towards seen classes. We conduct extensive ablation studies, and achieve state-of-the-art results on three benchmark datasets and our novel MZSL-50 dataset. Specifically, we improve the conventional MZSL performance by a margin of 2.1%, 9.81% and 8.68% on VGG-Sound, UCF-101 and ActivityNet, respectively. Finally, we expect the introduction of the MZSL-50 dataset will promote the future in-depth research on multimodal zero-shot learning in the community.1 Keval Doshi, Amanmeet Garg, Burak Uzkent, Mohamed Omar |
WACV | 5 |
| 2023 | Dynamic Inference with Grounding Based Vision and Language ModelsabstractTransformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks. On the other hand, there exists a large amount of computational redundancy in these large models which skips their run-time efficiency. To address this problem, we propose dynamic inference for grounding based vision and language models conditioned on the input image-text pair. We first design an approach to dynamically skip multihead self-attention and feed forward network layers across two backbones and multimodal network. Additionally, we propose dynamic token pruning and fusion for two backbones. In particular, we remove redundant tokens at different levels of the backbones and fuse the image tokens with the language tokens in an adaptive manner. To learn policies for dynamic inference, we train agents using reinforcement learning. In this direction, we replace the CNN backbone in a recent grounding-based vision and language model, MDETR, with a vision transformer and call it ViTMDETR. Then, we apply our dynamic inference method to ViTMDETR, called D-ViTDMETR, and perform experiments on image-language tasks. Our results show that we can improve the run-time efficiency of the state-of-the-art models MDETR and GLIP by up to ~ 50% on Referring Expression Comprehension and Segmentation, and VQA with only maximum ~ 0.3% accuracy drop. Burak Uzkent, Amanmeet Garg, Keval Doshi, Jingru Yi, Mohamed Omar |
CVPR | 7 |
| 2023 | Selective Structured State-Spaces for Long-Form Video UnderstandingabstractEffective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence ($S4$) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all imagetokens equally as done by$S4$model can adversely affect its efficiency and accuracy. To address this limitation, we present a novel Selective$S4$(i.e.,$S5)$model that employs a lightweight mask generator to adaptively select informative image tokens resulting in more efficient and accurate modeling of long-term spatiotemporal dependencies in videos. Unlike previous mask-based token reduction methods used in transformers, our$S5$model avoids the dense self-attention calculation by making use of the guidance of the momentum-updated$S4$model. This enables our model to efficiently discard less informative tokens and adapt to various long-form video understanding tasks more effectively. However, as is the case for most token reduction methods, the informative image tokens could be dropped incorrectly. To improve the robustness and the temporal horizon of our model, we propose a novel long-short masked contrastive learning (LSMCL) approach that enables our model to predict longer temporal context using shorter input videos. We present extensive comparative results using three challenging long-form video understanding datasets (LVU, COIN and Breakfast), demonstrating that our approach consistently outperforms the previous state-of-the-art S4 model by up to 9.6% accuracy while reducing its memory footprint by 23%. Pichao Wang, Linda Liu, Mohamed Omar, Raffay Hamid |
CVPR | 6 |
| 2023 | Multiscale Audio Spectrogram Transformer for Efficient Audio ClassificationabstractAudio event has a hierarchical architecture in both time and frequency and can be grouped together to construct more abstract semantic audio classes. In this work, we develop a multiscale audio spectrogram Transformer (MAST) that employs hierarchical representation learning for efficient audio classification. Specifically, MAST employs one-dimensional (and two-dimensional) pooling operators along the time (and frequency domains) in different stages, and progressively reduces the number of tokens and increases the feature dimensions. MAST significantly outperforms AST [1] by 22.2%, 4.4% and 4.7% on Kinetics-Sounds, Epic-Kitchens-100 and VGGSound in terms of the top-1 accuracy without external training data. On the downloaded AudioSet dataset, which has over 20% missing audios, MAST also achieves slightly better accuracy than AST. In addition, MAST is 5× more efficient in terms of multiply-accumulates (MACs) with 42% reduction in the number of parameters compared to AST. Through clustering metrics and visualizations, we demonstrate that the proposed MAST can learn semantically more separable feature representations from audio signals. Mohamed Omar |
ICASSP | 2 |
| 2023 | Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature AlignmentabstractText-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advancement by ECLIPSE has improved long-range text-to-video retrieval by developing an audiovisual video representation. Nonetheless, the objective of the text-to-video retrieval task is to capture the complementary audio and video information that is pertinent to the text query rather than simply achieving better audio and video alignment. To address this issue, we introduce TEFAL, a TExt-conditioned Feature ALignment method that produces both audio and video representations conditioned on the text query. Instead of using only an audiovisual attention block, which could suppress the audio information relevant to the text query, our approach employs two independent cross-modal attention blocks that enable the text to attend to the audio and video representations separately. Our proposed method’s efficacy is demonstrated on four benchmark datasets that include audio: MSR-VTT, LSMDC, VATEX, and Charades, and achieves better than state-of-the-art performance consistently across the four datasets. This is attributed to the additional text-query-conditioned audio representation and the complementary information it adds to the text-query-conditioned video representation. Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, Mohamed Omar |
ICCV | 6 |
| 2023 | FAQT-2: A customer-oriented method for MCDM with statistical verification applied to industrial robot selection
Hassan Soltan, Khaled Janada, Mohamed Omar |
Expert Syst. Appl. | 3 |
| 2021 | Switched Reluctance Motor Design for an EV Propulsion ApplicationabstractThis paper introduces a design methodology for a Switched Reluctance Motor (SRM) for an 80 kW Battery Electric Vehicle (BEV) propulsion application. The methodology aims to satisfy the high power density requirement targeted for an electric motor for a BEV application, while maintaining improved efficiency and torque quality. Iterative modeling effort has been employed for the design, including finite element analysis for the electromagnetic characteristics of the motor and dynamic modeling for performance analyses. The design approach starts with determining the motor geometry followed by sensitivity analysis for the motor performance considering various motor parameters. Then the SRM conduction angles are optimized with multi-objective genetic algorithm to improve the torque density and reduce torque ripple. The performance is further analyzed and compared to the target motor. Omar Zayed, Mohamed Omar, Mohamed H. Bakr, Mehdi Narimani, Ali Emadi, Berker Bilgin |
IECON | 2 |
| 2016 | HAUCA Curves for the Evaluation of Biomarker Pilot Studies with Small Sample Sizes and Large Numbers of Features
Frank Klawonn, Ina Koch, Jörg Eberhard, Mohamed Omar |
IDA | 5 |
| 2012 | Speech Activity Detection for Noisy Data Using Adaptation Techniques
Mohamed Omar |
INTERSPEECH | 1 |
| 2012 | On the Hardness of Counting and Sampling Center StringsabstractGiven a set S of n strings, each of length l, and a nonnegative value d, we define a center string as a string of length l that has Hamming distance at most d from each string in S. The #CLOSEST STRING problem aims to determine the number of center strings for a given set of strings S and input parameters n, l, and d. We show #CLOSEST STRING is impossible to solve exactly or even approximately in polynomial time, and that restricting #CLOSEST STRING so that any one of the parameters n, ‘, or d is fixed leads to a fully polynomial-time randomized approximation scheme (FPRAS). We show equivalent results for the problem of efficiently sampling center strings uniformly at random (u.a.r.). Christina Boucher 0001, Mohamed Omar |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | On the Hardness of Counting and Sampling Center Strings
Christina Boucher 0001, Mohamed Omar |
SPIRE | 2 |
| 2006 | Asymptotics of Largest Components in Combinatorial Structures
Mohamed Omar, Daniel Panario, L. Bruce Richmond, Jacki Whitely |
Algorithmica | 1 |