VLDB 2026 Research / reviewers in the wild / expert
Madjid Maidi
dblp:50/919
· DBLP profile ↗
14ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0001-7070-006XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid CNN-Transformer Architecture for Object Detection and Multimodal Captioning in Educational Contexts
Leila Habibi, Madjid Maidi, Larbi Boubchir, Boubaker Daachi |
IEEE Big Data | 2 |
| 2025 | Dual Focus Multiscale Attention for Object Detection in Mixed Reality: Leveraging Customizable Synthetic DatasetsabstractWe propose a novel object detection framework tailored for mixed reality (MR), combining a customizable synthetic dataset with a lightweight attention-enhanced detection model. Our dataset generation pipeline synthesizes planetary and telescope foregrounds with hybrid real-synthetic backgrounds, enabling robust learning across variable lighting and occlusion scenarios—challenges common in educational MR environments. At the core of our architecture is the Dual Focus Multiscale Attention (DFMA) module, which simultaneously refines spatial and channel-wise features at multiple scales. Integrated into a YOLO-based (You Only Look Once) backbone and FPN, DFMA significantly improves feature discrimination while preserving real-time efficiency. On MS COCO our model improves mean Average Precision (mAP) across Intersection over Union (IoU) thresholds from 0.5 to 0.95 ([email protected]:0.95) over state-of-the-art nano detectors from 39.3% to 41.3% (± 2%) at only +6% params and +3% GFLOPs, with a notable reduction in false positives on visually similar, low-textured objects. We further demonstrate real-time deployment in a Unity-based MR application, highlighting the system's effectiveness in immersive astronomy-focused educational scenarios. Our results underscore the potential of synthetic data and multiscale attention to bridge accuracy, speed, and realism in next generation MR systems. Salah-eddine Laidoudi, Madjid Maidi, Samir Otmane |
ISMAR | 2 |
| 2025 | Gated Temporal Shifts with Depth-Efficient Channel Attention for Real-Time Hand-Gesture InteractionabstractWe introduce a compact video-classification pipeline for real-time dynamic hand-gesture recognition in mixed-reality (MR) settings. The network marries a MobileNetV3 backbone with two purpose-built temporal components: (1) a Gated Discriminative Temporal Shift Module (G-DiTSM) that inserts first-order motion differences and learns channel-wise gates to fuse them adaptively, and (2) a lightweight Depth-Efficient Channel Attention (DepthECA) block that recalibrates spatial features on the fly. Operating on eight sparsely sampled frames per clip (Temporal Segment Network paradigm), the resulting model contains 2.65 M parameters and requires only 0.084 GFLOPs per inference. Evaluated on the RGB-only 20BN Jester benchmark (148k clips spanning 27 gesture classes) recorded from front-facing viewpoints. The system reaches 95.34% Top-1 and 99.80% Top-5 accuracy, surpassing recent 3D CNNs and transformer baselines while being an order of magnitude lighter. Ablations confirm that DepthECA and G-DiTSM provide complementary gains (+18.78% and +0.93% Top-1, respectively, over the MobileNetV3 baseline). Because all components are plug-and-play and introduce minimal overhead, the architecture is well suited to the tight latency and power budgets of standalone MR headsets, paving the way for natural grab, rotate, and command interactions using only on-board RGB cameras Salah-eddine Laidoudi, Madjid Maidi, Samir Otmane |
VRST | 2 |
| 2024 | Real-Time Indoor Object Detection Based on Hybrid CNN-Transformer ApproachabstractReal-time object detection in indoor settings is a challenging area of computer vision, faced with unique obstacles such as variable lighting and complex backgrounds. This field holds significant potential to revolutionize applications like augmented and mixed realities by enabling more seamless interactions between digital content and the physical world. However, the scarcity of research specifically fitted to the intricacies of indoor environments has highlighted a clear gap in the literature. To address this, our study delves into the evaluation of existing datasets and computational models, leading to the creation of a refined dataset. This new dataset is derived from OpenImages v7[14], focusing exclusively on 32 indoor categories selected for their relevance to real-world applications. Alongside this, we present an adaptation of a CNN detection model, incorporating an attention mechanism to enhance the model’s ability to discern and prioritize critical features within cluttered indoor scenes. Our findings demonstrate that this approach is not just competitive to existing state-of-the-art models in accuracy and speed but also opens new avenues for research and application in the field of real-time indoor object detection. Salah-eddine Laidoudi, Madjid Maidi, Samir Otmane |
ECAI | 2 |
| 2023 | Real-Time Detection of Low-Textured Objects Based on Deep LearningabstractIn this paper, a custom Single Shot Multi-box Detector (SSD) [1] is proposed for object detection on difficult scenes. The fruit 360 dataset [2], with low-textured images of different fruits and vegetables, is used as a training and validation data set. The purpose of this research is to implement the detector on mobile devices for mixed and augmented reality experiences, so a lighter weight SSD [1] model was designed while retaining its performance. The custom model is 4 times faster than the original SSD [1] model and the tests showed that it is even more accurate on the designated data set. The model is implemented in Python using Tensorflow and will soon be available on GitHub for public use. Salah-eddine Laidoudi, Madjid Maidi, Samir Otmane |
MMSP | 2 |
| 2022 | Multimodal 2D/3D Registration for Open Augmented Reality ApplicationsabstractThis work aims to build a multimodal and open Augmented Reality application to enable overlaying videos or 3D graphics on natural objects in a real-time tracking process. The target object is detected within the sequence of images through its local points of interest and their corresponding descriptors. The similarity between the query and the reference image is computed and the homography relating 2D features is determined to produce the projective transformation used for the video registration mode. To add a 3D model into the scene, we estimate the real camera pose and we resolve the transformations connecting the virtual and the real camera reference frames. The obtained results from experiments are accurate and computationally effective. The developed system detects and tracks in real-time markerless objects, then overlays accurately videos or 3D graphics on targets. Madjid Maidi, Samir Otmane |
ISM | 1 |
| 2020 | Open Augmented Reality System For Mobile Markerless TrackingabstractThe aim of this work is to present an open solution for building an Augmented Reality (AR) system without using any existing SDK. The proposed approach relies upon 2D planar object recognition for mobile real-time tracking applications. The transformation relating the world and the camera coordinate systems is determined using pose estimation. Once the projective transform relating 3D and 2D features is computed, a virtual 3D graphic is registered on the image. Many tests have been performed to show the efficiency of the proposed approach and to prove its relevance in terms of accuracy and time computation. The final application enabled real-time mobile tracking of markerless images augmented with 3D models to enrich the visual perception of the user. Madjid Maidi, Yassine Lehiani, Marius Preda |
ICIP | 1 |
| 2014 | Markerless identification and tracking for scalable image databaseabstractIn this paper we present a novel approach for object identification and tracking in large image datasets. Objects of interest are represented by feature points and descriptors extracted and compared to a set of reference data. An optimized matching paradigm is designed to deal with scalable image databases while keeping a good recognition rate in real-life environment conditions. Experiments are conducted to evaluate the effectiveness of the method and the obtained results demonstrate a true interest of the proposed approach. Madjid Maidi, Marius Preda, Yassine Lehiani |
ICIP | 1 |
| 2014 | Vision-based tracking in large image database for real-time mobile augmented realityabstractThis paper presents an approach for tracking natural objects in augmented reality applications. The targets are detected and identified using a markerless approach relying upon the extraction of image salient features and descriptors. The method deals with large image databases using a novel strategy for feature retrieval and pairwise matching. Further-more, the developed method integrates a real-time solution for 3D pose estimation using an analytical technique based on camera perspective transformations. The algorithm associates 2D feature samples coming from the identification part with 3D mapped points of the object space. Next, a sampling scheme for ordering correspondences is carried out to establishing the 2D/3D projective relationship. The tracker performs localization using the feature images and 3D models to enhance the scene view with overlaid graphics by computing the camera motion parameters. The modules built within this architecture are deployed on a mobile platform to provide an intuitive interface for interacting with the surrounding real world. The system is experimented and evaluated on challenging scalable image dataset and the obtained results demonstrate the effectiveness of the approach towards versatile augmented reality applications. Madjid Maidi, Marius Preda, Yassine Lehiani, Traian Lavric |
MMSP | 1 |
| 2013 | Interactive media control using natural interaction-based KinectabstractIn this work we present a novel interaction approach based on a gesture recognition system using a Microsoft Kinect sensor. Gestures are defined and interpreted in order to activate controls on a media device. This natural interface enables intuitive interaction with the multimedia content. The depth sensor observes the scene to detect a request for control in the form of a gesture. Then, the application assigns a control for the media system. The application is tested under various scenarios and proved the reliability and the effectiveness of the proposed approach. Madjid Maidi, Marius Preda |
ICASSP | 1 |
| 2011 | Characters Identification in TV SeriesabstractThis work aims to realize a recognition system for a software engine that will automatically generate a quiz starting from a video content and reinsert it into the video, turning thus any available foreign-language video (such as news or TV series) into a remarkable learning tool. Our system includes a face tracking application which integrates the eigen face method with a temporal tracking approach. The main part of our work is to detect and identify faces from movies and to associate specific quizzes for each recognized character. The proposed approach allows to label the detected faces and maintains face tracking along the video stream. This task is challenging since characters present significant variation in their appearance. Therefore, we employed eigen faces to reconstruct the original image from training models and we developed a new technique based on frames buffering for continuous tracking in unfavorable environment conditions. Many tests were conducted and proved that our system is able to identify multiple characters. The obtained results showed the performance and the effectiveness of the proposed method. Madjid Maidi, Veronica Scurtu, Marius Preda |
ISM | 1 |
| 2010 | A performance study for camera pose estimation using visual marker based tracking
Madjid Maidi, Jean-Yves Didier, Fakhreddine Ababsa, Malik Mallem |
Mach. Vis. Appl. | 1 |
| 2009 | Vision-inertial tracking system for robust fiducials registration in augmented realityabstractThis paper describes a multimodal tracking system to resolve occlusions in augmented reality applications. The first module of the proposed architecture is composed of a vision based system and allows identification and tracking of visible targets. When targets are partially occluded by scene elements, a second module relieves the vision based module and tracks feature points using a robust algorithm. Finally, a multi-sensors tracking approach is implemented to handle total occlusion of targets and maintains registration even if all markers are not visible. Experimental results and many evaluations have been performed to show the efficiency and robustness of the proposed multimodal approach of tracking and occlusion handling in augmented reality. Madjid Maidi, Fakhreddine Ababsa, Malik Mallem |
CIMSIVP | 1 |
| 2005 | Vision-inertial system calibration for tracking in augmented reality
Madjid Maidi, Fakhreddine Ababsa, Malik Mallem |
ICINCO | 1 |