Abdulmotaleb El Saddik

dblp:06/1887 · also Abdulmotaleb El-Saddik, Abdulmotaleb Elsaddik · DBLP profile ↗
← Back
300ranked-venue papers
13as first author
98since 2021 · last 2026
0000-0002-7690-8547ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 173 · 7 first-author · 52 since 2021Artificial intelligence and machine learning · 60 · 1 first-author · 31 since 2021Computer networks · 48 · 2 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 34 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 12 · 1 since 2021Systems, architecture and hardware · 9 · 2 first-author · 2 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking Feature Conditioning for Robust Forged Media Detection in Edge AI Sensing Systems
Izaldein Al-Zyoud, Abdulmotaleb El Saddik
IWCMC2
2026 Person search with deep learning
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik
Eng. Appl. Artif. Intell.4
2026 Enhancing multimodal emotion recognition with dynamic fuzzy membership and attention fusion
Nhut Minh Nguyen, Trung Minh Nguyen, Thanh Trung Nguyen, Phuong-Nam Tran 0001, Nhat Truong Pham, Linh Le, Alice Othmani, Abdulmotaleb El Saddik, Duc Ngoc Minh Dang
Eng. Appl. Artif. Intell.8
2026 Leveraging model explainability and fine-grained cutmix augmentation for robust detection of apricot diseases in UAV images
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
Expert Syst. Appl.3
2026 Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
Pattern Recognit.7
2026 SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detection
abstract
3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments.
Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Intell. Transp. Syst.7
2026 CP-Diffusion: Conditional Prompt-Based Diffusion Models for Video Generation
abstract
Motion customization plays a pivotal role in video generation by preserving the original appearance and context while adhering to specific motion patterns. In contrast, video generation techniques often lack coherence and realism due to difficulties in capturing and transferring motion patterns. Building upon the Video Motion Customization (VMC) framework, we proposed a few-shot learning approach using our unified Multi-Head Temporal Attention (MHTA) module for motion customization in text-to-video diffusion models. This significantly reduces computational requirements while maintaining and improving motion quality. Our model provides a streamlined mechanism for motion distillation while maintaining separate self-, cross-, and temporal attention. Moreover, the temporal attention layer is adapted through a simplified mechanism with efficient Q/K/V projections, while maintaining fixed spatial self- and cross-attention. The model distills a ground-truth motion vector from consecutive frames to align the predicted and ground-truth motion. Our proposed MHTA model outperforms the baseline in video generation using motion customization while being significantly more resource-efficient. Moreover, our approach can easily be applied to generate conditional prompt-based videos in the gaming industry.
Mustaqeem Khan 0001, Muhammad Saad 0005, Nasir Rahim, Wail Gueaieb, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2026 Graph-Based Event and Sub-Event Grouping in User-Generated Videos with Distortion-Aware Keyframe Clustering
abstract
User-generated content (UGC) videos recorded in uncontrolled environments often exhibit blur, camera shake, lighting fluctuations, and large viewpoint differences, making it difficult to organize multiple recordings of the same event. This work proposes an integrated pipeline that groups UGC videos into coherent events and sub-events by combining distortion-aware keyframe selection, adaptive audio–visual fusion, and confidence-weighted graph construction. The method first filters and clusters segment-level representations to obtain reliable keyframes, then fuses audio and visual cues through a lightweight gating module to produce robust multimodal descriptors. These descriptors populate a similarity graph whose strong and weak edges reveal sub-event and event structure without requiring shot boundaries or manual segmentation. Although distortion modeling, keyframe extraction, and multimodal similarity have been studied separately, existing approaches do not integrate them for hierarchical UGC video grouping. Experiments on the JIKU dataset and a curated YouTube dataset show consistent improvements in fidelity, diversity, and clustering metrics, demonstrating the applicability of the approach to video summarization, multi-view organization, and other UGC analysis tasks.
Malya Singh, Wei Tsang Ooi, Abdulmotaleb El Saddik, Mukesh Saini
ACM Trans. Multim. Comput. Commun. Appl.3
2026 A Step Closer Towards the Digital Twin of the Plant
abstract
Digital twins can provide vital insights into agricultural products and processes. There have been a lot of documented attempts at digital twins in agriculture. However, majority of these attempts build synthetic models and ignore the temporal dimension of the plant growth. Therefore, the existing models fail to depict actual plant details and growth. Our work replicates the actual growth of a real plant in the digital world by acquiring 3D meshes of the plant at various instants. It focuses on the transition between those acquired meshes by approximating all the consecutive pairs into approximate mesh pairs that have a common topology. The quality of these common approximate mesh pairs is quantitatively measured by an Energy term, which is minimized during the optimization process. Later, the meshes with the common topology are interpolated (morphing) to build the final digital twin of the plant. Experimental results show that the proposed methodology to attain the final morph has the potential to be a vital module, which could be responsible for the visual updates in the digital replica of the digital twin of the plant.
Karanvir Singh, Abdulmotaleb El Saddik, Mukesh Saini
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Mobiflip: Information-Bottleneck-Guided Minimal Federated Adaptation for Cross-Modal Models
abstract
Cross-modal federated learning is constrained by bandwidth and on-device compute. We present Mobiflip: a minimalist strategy that freezes a lightweight backbone and communicates only a channel-wise \(1\times 1\) scaling adapter appended to the image branch. Guided by the Information Bottleneck, we prove that under common distributional and linear-encoder surrogates, per-channel scaling attains the linear optimum; coupled with the directional geometry of (Mobile)CLIP, the adapter is, in first-order approximation, an optimal preconditioner of the cosine-similarity space—preserving discriminative directions while compressing redundancy and suppressing inter-client drift. We adopt MobileCLIP as a mobile-friendly backbone to jointly minimize compute and communication. On CIFAR-10/100 and medical imaging, a single aggregation already yields stable Bacc; each round transmits only about 0.7% of backbone parameters with \(>\!\!92\%\) reduction in communication. Compared with recent federated multimodal/large-model methods, Mobiflip maintains—or even improves—accuracy under ultra-low communication.
Zishan Xu, Jiansen Zhang, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.7
2025 ASimp: Automatic High-Poly 3D Mesh Simplification for Preprocessing Based on QoE
abstract
Mesh simplification of 3D models can accelerate rendering, reduce storage space, and improve performance. However, for high-poly 3D models, there are ongoing concerns about potentially compromising the Quality of Experience (QoE), the need to set simplification ratios or parameters, and the time-consuming nature of the simplification process. To address these issues, we proposed a new mesh simplification for the preprocessing step. Based on the Quadratic Error Metric (QEM) simplification algorithm, we conducted human-centered 3D model comparison experiments to determine the optimal simplification ratio for high-poly 3D models in full body shots. From experimental data, we proposed and implemented ASimp, an automatic 3D mesh simplification scheme. In evaluation experiments, ASimp demonstrated rapid preprocessing speeds while ensuring QoE and the effectiveness of its simplification products. We hope that ASimp will contribute to the optimization of 3D models and find applications in fields such as cultural heritage, archaeology, visual effects, video games, medicine, metaverse, and beyond.
Lehao Lin, Hong Kang, Yuqi Shi, Haihan Duan, Abdulmotaleb El Saddik, Wei Cai 0002
ICME5
2025 ASTAnet: Transformer-based Siamese Network for Robust Audio-to-Audio Alignment in Amateur User Generated Audio Clips
abstract
Audio alignment involves synchronizing two or more audio recordings. Existing methods depend on handcrafted features and struggle with precision in lengthy or noisy recordings. Deep learning techniques have proven effective across various domains; however, their application in audio-to-audio alignment is still in its infancy. We propose ASTAnet, a framework that integrates the Vision Transformer for feature extraction with the Siamese network for similarity estimation. With timestamp positional encoding, ASTAnet improves temporal precision and reduces alignment errors using a contrastive learning objective based on Euclidean distance. Our experiments achieved an overall mean absolute error value of 0.005, a 1.8X improvement compared to the previous works. Extensive evaluations demonstrate its effectiveness, particularly for varying and longer audio recordings.
Malya Singh, Priyankar Choudhary, Abdulmotaleb El Saddik, Mukesh Saini
ICME3
2025 GroMo25: ACM Multimedia 2025 Grand Challenge for Plant Growth Modeling with Multiview Images
Shreya Bansal, Ruchi Bhatt, Amanpreet Chander, Malya Singh, Mohan Kankanhalli, Abdulmotaleb El Saddik, Mukesh Saini
ACM Multimedia7
2025 Agent-to-Agent (A2A) Protocol Integrated Digital Twin System with AgentIQ for Multimodal AI Fitness Coaching and Personalized Well-Being
Kamran Gholizadeh HamlAbadi, Monica (Monireh) Vahdati, Fedwa Laamarti, Abdulmotaleb El Saddik
ACM Multimedia4
2025 A Multi-Agent Digital Twin Framework for AI-Driven Fitness Coaching
abstract
Figure 1: The architecture of the proposed Digital Twin AI Fitness Coaching System (DTAIFC).The system integrates multimodal input sensing, short and long-term memory, a Crew-inspired multi-agent AI coordination core, and a user-facing interaction interface.Orchestration Agent operates as the central brain coordinating the Recommendation and Feedback Agents for intelligent workout planning and real-time adjustment.
Monica (Monireh) Vahdati, Kamran Gholizadeh HamlAbadi, Fedwa Laamarti, Abdulmotaleb El Saddik
IMX4
2025 RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
Expert Syst. Appl.7
2025 Multi-object tracking with scale-aware transformer and enhanced association strategy
Xuezhi Xiang, Xiankun Zhou, Mingliang Zhai, Abdulmotaleb El Saddik
Multim. Syst.5
2025 Decentralized Web3 Non-Fungible Token Community for Societal Prosperity? A Social Capital Perspective
abstract
In the rapidly evolving Web3 world, non-fungible token (NFT) communities are reshaping the formation, distribution, and activation of social capital in ways distinct from traditional models. However, despite their growing impact on societal prosperity, a comprehensive understanding of social capital dynamics within Web3 NFT communities remains limited. This study explores the Mfers community, a key example within Web3 NFT ecosystems. By analyzing social media and blockchain data and using a Delphi method-based human-large language model (LLM) collaboration, we uncovered unique social capital patterns across six dimensions. Our findings highlight a compelling blend of decentralization, inclusion, trust, and empowerment but also raise critical questions about wealth inequality, content quality, and ethical challenges. Based on the findings, we discussed the uniqueness of social capital in Web3 NFT communities, the tension between technical and power decentralization, and the multidimensional nature of societal prosperity. We also suggested directions for future research on decentralized online communities in the CSCW field. This study provides a systematic perspective on social capital in Web3 NFT communities and introduces an innovative human-LLM collaborative analysis, offering insights into the design and governance of benign decentralized online communities.
Hongzhou Chen, Chenyu Zhou 0009, Abdulmotaleb El Saddik, Wei Cai 0002
Proc. ACM Hum. Comput. Interact.3
2025 S2Match: Self-paced sampling for data-limited semi-supervised learning
Dayan Guan, Yun Xing 0001, Jiaxing Huang 0001, Aoran Xiao, Abdulmotaleb El Saddik, Shijian Lu
Pattern Recognit.5
2025 DC-Net: Divide-and-conquer for salient object detection
Xuebin Qin, Abdulmotaleb El Saddik
Pattern Recognit.3
2025 Adaptive Social Metaverse Streaming Based on Federated Multiagent Deep Reinforcement Learning
abstract
The social metaverse is a growing digital ecosystem that blends virtual and physical worlds. It allows users to interact socially, work, shop, and enjoy entertainment. However, privacy remains a major challenge, as immersive interactions require continuous collection of biometric and behavioral data. At the same time, ensuring high-quality, low-latency streaming is difficult due to the demands of real-time interaction, immersive rendering, and bandwidth optimization. To address these issues, we propose adaptive social metaverse streaming (ASMS), a novel streaming system based on federated multiagent proximal policy optimization (F-MAPPO). ASMS leverages F-MAPPO, which integrates federated learning (FL) and deep reinforcement learning (DRL) to dynamically adjust streaming bit rates while preserving user privacy. Experimental results show that ASMS improves user experience by at least 14% compared to existing streaming methods across various network conditions. Therefore, ASMS enhances the social metaverse experience by providing seamless and immersive streaming, even in dynamic and resource-constrained networks, while ensuring that sensitive user data remain on local devices.
Zijian Long, Haiwei Dong 0001, Abdulmotaleb El Saddik
IEEE Trans. Comput. Soc. Syst.4
2025 A Multidimensional Contract Design for Smart Contract-as-a-Service
abstract
Empowered by blockchain technology, smart contracts have attracted considerable interest from Web3 users due to their distinct advantages. Nevertheless, it is challenging to address problems caused by the dramatic expansion of the Web3 ecosystem. This article introduces the smart contract-as-a-service (SCaaS) paradigm to mitigate smart contracts’ redundant deployment via their composability and reusability. Moreover, we design trust and incentive schemes to ensure project security and developer engagement in SCaaS. Specifically, we first introduce a reputation filter by leveraging the authentic on-chain data, aiming to eliminate high-risk contracts. We then design a contract-based incentive mechanism to help the foundation attract heterogeneous developers with multidimensional private information, and maximize the foundation’s utility by inducing developers to undertake projects of differing complexities based on their ability. We further differentiate between veteran and newcome developers and examine their influences on foundational strategies. Finally, extensive experimental results demonstrate that our proposed contracts can efficiently remove high-risk smart contracts, maximize the foundation’s utility, and ensure that developers select contracts honestly and participate in the SCaaS ecosystem actively.
Jinghan Sun, Hou-Wan Long, Hong Kang, Zhixuan Fang, Abdulmotaleb El Saddik, Wei Cai 0002
IEEE Trans. Comput. Soc. Syst.5
2025 DDCI: Unsupervised Domain Adaptation for Remote Sensing Images Based on Diffusion Causal Distillation
abstract
The distribution of remote sensing (RS) images can vary significantly due to seasonal changes and lighting conditions, making it difficult for deep learning models to generalize effectively across different RS datasets. This variation leads to a domain gap that hampers model performance when applied to new, unseen data. To tackle this challenge, we introduce DDCI, a novel unsupervised domain adaptation (UDA) framework designed to bridge the domain gap in RS image perception. Our framework consists of two key components, i.e., the adaptation diffusion distillation (ADD) module and the consistent causal intervention (CCI) module. The ADD module addresses the domain gap by aligning the source and target domains. It enhances the representation of the target domain by distilling semantic knowledge from the teacher model of the source domain. This process allows the target domain to benefit from the rich features of the source domain, leading to improved model generalization. The CCI module focuses on removing spurious correlations between domain-agnostic knowledge and domain-specific knowledge. By carefully considering the distinct characteristics of the target domain while preserving the specificity of the source domain, the CCI module ensures that only relevant, causal information is transferred between domains. This prevents overfitting to irrelevant domain-specific features and enhances model robustness. We demonstrate the effectiveness of the DDCI framework on RS scene classification tasks, utilizing four widely recognized RS datasets. Our results show significant performance improvements, underscoring the potential of this approach to boost the adaptability of deep learning models across diverse RS image datasets.
Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.7
2025 GLFRNet: Global-Local Feature Refusion Network for Remote Sensing Image Instance Segmentation
abstract
Instance segmentation is a significant way for remote sensing image (RSI) interpretation. The large number, sharp variation of sizes, and complex background of objects raise higher demands for instance segmentation models. The synergistic usage of global and local features has drawn great attention due to its superior performance but has not been fully explored in mainstream instance segmentation methods. In this work, a global-local feature refusion network (GLFRNet) with two fusion procedures is proposed to fully utilize coarse-grained and fine-grained features for RSI instance segmentation. In this model, the backbone integrates both convolutional neural network (CNN)-based and VMamba-based branches to extract local and global features, respectively. Three novel models are proposed to leverage the features adaptively, i.e., the cross-dim feature fusion (CDFF) module, the semantic complementary feature fusion (SCFF) module, and the guided feature refusion module (GFRM). The CDFF module is designed to aggregate features flexibly by fusing features from two backbones with different attention modules in the first fusion procedure. The GFRM and SCFF module are proposed in the refusion procedure to generate accurate segmentation results. Inspired by agent attention, the GFRM dynamically assembles detailed features for mask generation by refusing local and global features with the guidance of fusion results from CDFF. The SCFF module complements the significant features by enhancing and integrating global, local, and detailed features, and finally generates masks of instances. Extensive experiments demonstrate that GLFRNet outperforms the second-best model by 1.9, 1.3, and 0.3 in mask average precisions (APs) on NWPU VHR-10, WHU Building, and iSAID datasets.
Jiaqi Zhao 0001, Yari Wang, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.6
2025 ST-Mamba: Spatio-Temporal Synergistic Model for Remote Sensing Change Detection
abstract
The advancement of remote sensing and deep learning has spurred interest in high-resolution image change detection (CD). However, pseudo-changes in multi-temporal images, due to complex scenes and variable imaging conditions, often lead to significant misdetection in current methods. To address this problem, we propose a new CD framework: Spatio-Temporal Mamba (ST-Mamba), which consists of three key components. Firstly, a Mamba-based Feature Extraction Module (MFEM) is designed as the encoder to extract essential features from multi-temporal images by leveraging Mamba’s capability to capture inherent information in long data sequences. Secondly, a Spatio-Temporal Synergy Module (STSM) is developed to unify the background features of multi-temporal feature maps into a common domain by employing the state-space model for spatio-temporal modeling. Finally, a Spatio-Temporal Fusion Module (STFM) is created to guide the fusion of image features at different scales and across channels by utilizing a feature map of the unified background features. Experimental results on five widely used change detection datasets show significant improvements over current state-of-the-art methods.
Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.6
2025 Adversarial Geometric Attacks for 3D Point Cloud Object Tracking
abstract
3D point cloud object tracking (3D PCOT) plays a vital role in applications such as autonomous driving and robotics. Adversarial attacks offer a promising approach to enhance the robustness and security of tracking models. However, existing adversarial attack methods for 3D PCOT seldom leverage the geometric structure of point clouds and often overlook the transferability of attack strategies. To address these limitations, this paper proposes an adversarial geometric attack method tailored for 3D PCOT, which includes a point perturbation attack module (non-isometric transformation) and a rotation attack module (isometric transformation). First, we introduce a curvature-aware point perturbation attack module that enhances local transformations by applying normal perturbations to critical points identified through geometric features such as curvature and entropy. Second, we design a Thompson sampling-based rotation attack module that applies subtle global rotations to the point cloud, introducing tracking errors while maintaining imperceptibility. Additionally, we design a fused loss function to iteratively optimize the point cloud within the search region, generating adversarially perturbed samples. The proposed method is evaluated on multiple 3D PCOT models and validated through black-box tracking experiments on benchmarks. For P2B, white-box attacks on KITTI reduce the success rate from 53.3% to 29.6% and precision from 68.4% to 37.1%. On NuScenes, the success rate drops from 39.0% to 27.6%, and precision from 39.9 to 26.8%. Black-box attacks show a transferability, with BAT showing a maximum 47.0% drop in success rate and 47.2% in precision on KITTI, and a maximum 22.5% and 27.0% on NuScenes.
Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
IEEE Trans. Multim.6
2025 Intrinsic Consistency Preservation With Adaptively Reliable Samples for Source-Free Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to alleviate the domain shift by transferring knowledge learned from a labeled source dataset to an unlabeled target domain. Although UDA has seen promising progress recently, it requires access to data from both domains, making it problematic in source data-absent scenarios. In this article, we investigate a practical task source-free domain adaptation (SFDA) that alleviates the limitations of the widely studied UDA in simultaneously acquiring source and target data. In addition, we further study the imbalanced SFDA (ISFDA) problem, which addresses the intra-domain class imbalance and inter-domain label shift in SFDA. We observe two key issues in SFDA that: 1) target data form clusters in the representation space regardless of whether the target data points are aligned with the source classifier and 2) target samples with higher classification confidence are more reliable and have less variation in their classification confidence during adaptation. Motivated by these observations, we propose a unified method, named intrinsic consistency preservation with adaptively reliable samples (ICPR), to jointly cope with SFDA and ISFDA. Specifically, ICPR first encourages the intrinsic consistency in the predictions of neighbors for unlabeled samples with weak augmentation (standard flip-and-shift), regardless of their reliability. ICPR then generates strongly augmented views specifically for adaptively selected reliable samples and is trained to fix the intrinsic consistency between weakly and strongly augmented views of the same image concerning predictions of neighbors and their own. Additionally, we propose to use a prototype-like classifier to avoid the classification confusion caused by severe intra-domain class imbalance and inter-domain label shift. We demonstrate the effectiveness and general applicability of ICPR on six benchmarks of both SFDA and ISFDA tasks. The reproducible code of our proposed ICPR method is available at https://github.com/CFM-MSG/Code_ICPR.
Abdulmotaleb El Saddik, Xing Xu 0001, Dongshuai Li, Zuo Cao, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.2
2025 Unleashing Creativity in the Metaverse: Generative AI and Multimodal Content
abstract
The metaverse presents an emerging creative expression and collaboration frontier where generative artificial intelligence (GenAI) can play a pivotal role with its ability to generate multimodal content from simple prompts. These prompts allow the metaverse to interact with GenAI, where context information, instructions, input data, or even output indications constituting the prompt can come from within the metaverse. However, their integration poses challenges regarding interoperability, lack of standards, scalability, and maintaining a high-quality user experience. This article explores how GenAI can productively assist in enhancing creativity within the contexts of the metaverse and unlock new opportunities. We provide a technical, in-depth overview of the different generative models for image, video, audio, and 3D content within the metaverse environments. We also explore the bottlenecks, opportunities, and innovative applications of GenAI from the perspectives of end users, developers, service providers, and AI researchers. This survey commences by highlighting the potential of GenAI for enhancing the metaverse experience through dynamic content generation to populate massive virtual worlds. Subsequently, we shed light on the ongoing research practices and trends in multimodal content generation, enhancing realism and creativity and alleviating bottlenecks related to standardization, computational cost, privacy, and safety. Last, we share insights into promising research directions toward the integration of GenAI with the metaverse for creative enhancement, improved immersion, and innovative interactive applications.
Abdulmotaleb El Saddik, Jamil Ahmad 0003, Mustaqeem Khan 0001, Saad Abouzahir, Wail Gueaieb
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Haptic Network Protocols: A Comprehensive Review and Directions for Next-Gen Metaverse Applications
abstract
This article presents a systematic review of haptic network protocols in the context of the Metaverse. With the increasing integration of haptic technologies into applications like remote collaboration and robotic surgery, the need for reliable, low-latency data transmission has intensified. This work provides a comprehensive analysis of existing haptic protocols and frameworks, focusing on their development, implementation, and the methods employed to optimize Quality of Service (QoS) parameters such as latency, delay, packet loss, jitter, throughput, and bandwidth. By examining the strengths and limitations of these protocols in real-time applications, this article identifies critical areas for improvement and suggests future directions, including the potential for incorporating machine learning (ML) and artificial intelligence (AI) to enable next-generation haptic communication suited for high-demand environments like the Metaverse.
Mohammed Faisal, Roberto Alejandro Martinez Velazquez, Fedwa Laamarti, Hussein Al Osman, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-Identification
abstract
Traditional text-based person re-identification relies on identity labels. However, it is impossible to annotate large datasets, since identity annotation is expensive and time-consuming. Weakly supervised text-based person re-identification, where only text–image pairs are available without annotation of identities, is very practical in real life. While dealing with the weakly supervised person re-identification, two issues should be strengthed, i.e., alignment caused by different modal, and cross-modal matching ambiguity caused by the lack of identity labels. In this article, we propose a similarity regulation and calibration alignment (SRCA) framework, which consists of two unimodal encoders for images and text, respectively, and a multi-modal encoder for the masked language modeling task. First, a similarity regulation (SR) strategy is proposed to relax the strict one-to-one constraints for the local similarities between different pairs by introducing a novel soft objective. The soft objective can adjust hard objectives to achieve soft cross-modal alignment by establishing a many-to-many relationship between two modalities. Second, the calibration alignment (CA) module is proposed to improve intra-class compactness by modeling pseudo-label assignment as optimal transport. The ambiguity of cross-modal matching can be reduced by aligning features and pseudo-labels of different modalities and gradually calibrating the distribution of pseudo-labels. Experimental results show that our method has achieved obvious advantages compared with existing methods and also demonstrated competitive performance compared with fully supervised methods.
Ao Fu, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2025 Meta-Review of Wearable Devices for Healthcare in the Metaverse
abstract
In recent years, there has been a growing interest in leveraging the metaverse to enhance community engagement and healthcare. This article provides a comprehensive examination of wearable devices and sensors utilized within immersive environments to improve well-being and healthcare outcomes. We categorize the healthcare application domains that employ wearable devices and identify commonly used devices and sensors based on a thorough review of the literature. Our study offers a detailed summary of these applications, highlighting their potential to enhance overall quality of life through remote monitoring, rehabilitation, and chronic disease management. Furthermore, we address existing research gaps and challenges in this field, offering insights for future research directions. This meta-review emphasizes the need for further exploration in the rapidly evolving domain of wearable healthcare technologies within the metaverse, presenting an overview of the current state of wearable devices in healthcare and underscoring their significance in advancing healthcare delivery and outcomes.
Monireh Vahdati, Fedwa Laamarti, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Immersive Multimedia Communication: State-of-the-Art on Extended Reality Streaming
abstract
Extended reality (XR) is rapidly advancing and poised to revolutionize content creation and consumption. In XR, users integrate various sensory inputs to form a cohesive perception of the virtual environment. This survey reviews the state-of-the-art in XR streaming, focusing on multiple paradigms. To begin, we define XR and introduce various XR headsets along with their multimodal interaction methods to provide a foundational understanding. We then analyze XR traffic characteristics to highlight the unique data transmission requirements. We also explore factors that influence the quality of experience in XR systems, aiming to identify key elements for enhancing user satisfaction. Following this, we present visual attention-based optimization methods for XR streaming to improve efficiency and performance. Finally, we examine current applications and highlight challenges to provide insights into ongoing and future developments of XR.
Haiwei Dong 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MuralAgent: Enhancing Ancient Mural Outpainting with RAG-Based Texts and Multimodal Integration
abstract
In the context of the digital age, utilizing cutting-edge technology for the digitization and creative expansion of ancient murals is crucial, aimed at preserving and passing on cultural heritage. Existing image outpainting techniques suffer from a lack of semantic guidance. This article introduces MuralAgent, a multimodal model based on Retrieval-Augmented Generation (RAG) technology. It precisely extracts key information from mural images and integrates it with a constructed ancient texts knowledge base to ensure the cultural and semantic consistency of the expanded images. Moreover, fine-tuning the Stable Diffusion model ensures the fidelity of the generated image styles. Specifically, this study involves constructing an ancient texts knowledge base for accurate matching, designing specific prompts for GPT-4V(ision) to extract key information, and innovatively expanding artworks through Stable Diffusion, providing a novel way for the public to reinterpret ancient murals.
Zishan Xu, Xiaofeng Zhang 0006, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.8
2025 Wakeup-Darkness: When Multimodal Meets Unsupervised Low-Light Image Enhancement
abstract
Low-light image enhancement is a crucial visual task, and many unsupervised methods overlook the degradation of visible information in low-light scenes, adversely affecting the fusion of complementary information and hindering the generation of satisfactory results. To address this, we introduce Wakeup-Darkness, a multimodal enhancement framework that innovatively enriches user interaction through voice and textual commands. This approach signifies a technical leap and represents a paradigm shift in user engagement. We introduce a Cross-Modal Feature Fusion (CMFF) that synergizes semantic and depth context with low-light enhancement operations. Moreover, we propose a Gated Residual Block (GRB) and a channel-aware Look-Up Table (LUT) to adjust the intensity distribution of each channel. Crucially, the proposed Wakeup-Darkness scheme demonstrates remarkable generalization in unsupervised scenarios. The source code can be accessed from https://github.com/zhangbaijin/Wakeup-Dakness .
Xiaofeng Zhang 0006, Zishan Xu, Hao Tang 0005, Chaochen Gu, Wei Chen 0036, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Efficient Test-Time Adaptation of Vision-Language Models
abstract
Test-time adaptation with pre-trained vision-language models has attracted increasing attention for tackling distribution shifts during the test time. Though prior studies have achieved very promising performance, they in-volve intensive computation which is severely unaligned with test-time adaptation. We design TDA, a training-free dynamic adapter that enables effective and efficient test-time adaptation with vision-language models. TDA works with a lightweight key-value cache that maintains a dy-namic queue with few-shot pseudo labels as values and the corresponding test-sample features as keys. Leveraging the key-value cache, TDA allows adapting to test data gradually via progressive pseudo label refinement which is super-efficient without incurring any backpropagation. In addition, we introduce negative pseudo labeling that alleviates the adverse impact of pseudo label noises by assigning pseudo labels to certain negative classes when the model is uncertain about its pseudo label predictions. Extensive experiments over two benchmarks demonstrate TDA's superior effectiveness and efficiency as compared with the state-of-the-art. The code has been released in https://kdiaaa.github.io/tda/
Adilbek Karmanov, Dayan Guan, Shijian Lu, Abdulmotaleb El Saddik, Eric P. Xing
CVPR4
2024 Contrastive Learning with Counterfactual Explanations for Radiology Report Generation
Mingjie Li 0006, Haokun Lin, Xiaodan Liang, Ling Chen 0006, Abdulmotaleb El Saddik, Xiaojun Chang
ECCV (43)6
2024 Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and Depth
abstract
Recurrently refining the optical flow based on a single high-resolution feature demonstrates high performance. We exploit the strength of this strategy to build a novel architecture for the joint learning of optical flow and depth. Our pro-posed architecture is improved to work in the case of training on unlabeled data, which is extremely challenging. The loss is computed for the iterations carried out over a single high-resolution feature, where the reconstruction loss fails to optimize the accuracy particularity in occluded regions. Therefore, we propose to hierarchically refine the optical flow across multiple scales while feeding the rigid flow calculated from depth and camera pose to provide more refinement. We further propose a self-supervised patch-based similarity loss to be optimized with the reconstruction loss to improve accuracy in the occluded regions. Our proposed method demonstrates efficient performance on the KITTI 2015 dataset, with more improvement in the occluded regions.
Rokia Abdein, Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik
ICASSP5
2024 Knowledge-Infused Learning for Fine-Grained Plant Disease Recognition
abstract
Domain knowledge exists in various forms, including text, ontologies, graphs, images, audio, and videos. In plant disease detection, most works solely utilize images with disease labels, neglecting textual descriptions of visual disease symptoms used by human experts for diagnosis. These text descriptions and sample images aid expert identification of visual symptoms. We propose a novel method that leverages text descriptions and image data by modeling domain-specific knowledge about visual symptoms in leaf images as separate feature channels. Each channel corresponds to specific features whose absence or presence in the image influences model predictions. We introduce a channel attention-guided fusion module for weighting each channel based on the input and corresponding output. The combined feature channels are transformed into a standardized 3-channel input format, which can then be processed by any pre-trained convolutional neural network (CNN) as input for feature extraction and subsequent classification. Furthermore, intermediate activations of the channel attention layer combined with the weights from the fusion layer make model predictions explainable. Experimental results on three publicly available datasets of apple and cucumber leaf diseases demonstrate improvements of up to 5% utilizing various state-of-the-art CNN architectures, indicating the efficacy of incorporating textual disease descriptions using the proposed approach.
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
ICIP3
2024 Deepskinformer: Skin Lesion Segmentation Using Hierarchical Transformers And Edge Enhancement
abstract
Segmentation of skin lesions from dermatological images is critical in diagnosing and treating skin cancer. Despite this, the diversity of lesion shapes, sizes, and textures against a similar-toned skin backdrop makes these images challenging to analyze. Current segmentation methods are often less precise in delineating boundaries and more susceptible to interference from background noise. To address this issue, we introduce an end-to-end framework called DeepSkinFormer (DSF) for skin lesion segmentation using the Skin Edge Enhancement Module (SEEM) to enhance boundaries for efficient detection. We evaluate the proposed model on standard benchmarks, HAM10000, ISIC2017, and PH2 datasets. Our model outperforms existing methods and achieves stateof-the-art results using the Dice and mean Intersection Over Union (mIOU) scores. Furthermore, we conduct an ablation study to confirm the significant contributions of DSFspecialized modules to their effectiveness.
Ufaq Khan, Umair Nawaz, Mustaqeem Khan 0001, Wail Gueaieb, Abdulmotaleb El Saddik
ICIP5
2024 HuBERT-CLAP: Contrastive Learning-Based Multimodal Emotion Recognition using Self-Alignment Approach
Long H. Nguyen, Nhat Truong Pham, Mustaqeem Khan 0001, Alice Othmani, Abdulmotaleb El Saddik
MMAsia5
2024 Duopoly Competition in Blockchain Game with Interoperability
abstract
As a bridge connecting the Web3 financial ecosystem and digital games, smart contracts empowered blockchain games have attracted significant attention from the Web3 community in recent years. By providing players ownership over assets and interoperable Non-Fungible Tokens (NFTs), blockchain games enable the reuse of in-game assets beyond the original games, thereby overturning the “walled garden” among traditional games. Nonetheless, blockchain games diminish the monopolistic edge previously held by traditional game providers, forcing them to compete with players by token distribution. Therefore, this paper explores the duopoly competition within the blockchain game market, emphasizing the role of interoperable NFTs together with NFT wear and tear. Specifically, we propose a three-stage game to formulate the interactions between game providers and players. Besides, we revealed the relationship between game providers' code disclosure strategies for NFT interoperability and token retention strategies. Finally, the experimental results demonstrate how the token distribution, players' preferences, and the NFT wear level affect the profits of game providers.
Jinghan Sun, Abdulmotaleb El Saddik, Wei Cai 0002
SMC3
2024 CamoFocus: Enhancing Camouflage Object Detection with Split-Feature Focal Modulation and Context Refinement
abstract
Camouflage Object Detection (COD) involves the challenge of isolating a target object from a visually similar background, presenting a formidable challenge for learning algorithms. Drawing inspiration from state-of-the-art (SOTA) Focal Modulation Networks, our objective is to proficiently modulate the foreground and background components, thereby capturing the distinct features of each. We introduce a Feature Split and Modulation (FSM) module to attain this goal. This module efficiently separates the object from the background by utilizing foreground and background modulators guided by a supervisory mask. For enhanced feature refinement, we propose a Context Refinement Module (CRM), which considers features acquired from FSM across various spatial scales, leading to comprehensive enrichment and highly accurate prediction maps. Through extensive experimentation, we showcase the superiority of CamoFocus over recent SOTA COD methods. Our evaluations encompass diverse benchmark datasets, including CAMO, COD10K, CHAMELEON, and NC4K. The findings underscore the potential and significance of the proposed CamoFocus model and establish its efficacy in addressing the critical challenges of camouflage object detection.
Mustaqeem Khan 0001, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
WACV4
2024 A reusable AI-enabled defect detection system for railway using ensembled CNN
Rahatara Ferdousi, Fedwa Laamarti, Chunsheng Yang, Abdulmotaleb El Saddik
Appl. Intell.4
2024 Learning feature contexts by transformer and CNN hybrid deep network for weakly supervised person search
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.5
2024 DBMHT: A double-branch multi-hypothesis transformer for 3D human pose estimation in video
Xuezhi Xiang, Xiaoheng Li, Weijie Bao, Yulong Qiao, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.5
2024 A GCN and Transformer complementary network for skeleton-based action recognition
Xuezhi Xiang, Xiaoheng Li, Xuzhao Liu, Yulong Qiao, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.5
2024 Self-supervised monocular depth estimation with self-distillation and dense skip connection
Xuezhi Xiang, Wei Li 0109, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.4
2024 Temporal adaptive feature pyramid network for action detection
Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.4
2024 Yield estimation and health assessment of temperate fruits: A modular framework
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
Eng. Appl. Artif. Intell.3
2024 MSER: Multimodal speech emotion recognition using cross-attention with deep fusion
Mustaqeem Khan 0001, Wail Gueaieb, Abdulmotaleb El Saddik, Soonil Kwon
Expert Syst. Appl.3
2024 Open-vocabulary object detection via debiased curriculum self-training
Hanlue Zhang, Dayan Guan, Xiangrui Ke, Abdulmotaleb El Saddik, Shijian Lu
Expert Syst. Appl.4
2024 MADRL-Based Rate Adaptation for 360° Video Streaming With Multiviewpoint Prediction
abstract
Over the last few years, 360video traffic on the network has grown significantly. A key challenge of 360video playback is ensuring a high quality of experience (QoE) with limited network bandwidth. Currently, most studies focus on tile-based adaptive bitrate (ABR) streaming based on single viewport prediction to reduce bandwidth consumption. However, the performance of models for single-viewpoint prediction is severely limited by the inherent uncertainty in head movement, which can not cope with the sudden movement of users very well. This paper first presents a multimodal spatial-temporal attention transformer to generate multiple viewpoint trajectories with their probabilities given a historical trajectory. The proposed method models viewpoint prediction as a classification problem and uses attention mechanisms to capture the spatial and temporal characteristics of input video frames and viewpoint trajectories for multi-viewpoint prediction. After that, a multi-agent deep reinforcement learning (MADRL)-based ABR algorithm utilizing multi-viewpoint prediction for 360video streaming is proposed for maximizing different QoE objectives under various network conditions. We formulate the ABR problem as a decentralized partially observable Markov decision process (Dec-POMDP) problem and present a MAPPO algorithm based on centralized training and decentralized execution (CTDE) framework to solve the problem. The experimental results show that our proposed method improves the defined QoE metric by up to 85.5% compared to existing ABR methods.
Zijian Long, Haiwei Dong 0001, Abdulmotaleb El Saddik
IEEE Internet Things J.4
2024 3D hand pose estimation and reconstruction based on multi-feature fusion
Jiye Wang, Xuezhi Xiang, Abdulmotaleb El Saddik
J. Vis. Commun. Image Represent.4
2024 GloFP-MSF: monocular scene flow estimation with global feature perception
Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik
Multim. Syst.5
2024 Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik
Multim. Tools Appl.7
2024 Simple Primitives With Feasibility- and Contextuality-Dependence for Open-World Compositional Zero-Shot Learning
abstract
The task of Open-World Compositional Zero-Shot Learning (OW-CZSL) is to recognize novel state-object compositions in images from all possible compositions, where the novel compositions are absent during the training stage. The performance of conventional methods degrades significantly due to the large cardinality of possible compositions. Some recent works consider simple primitives (i.e., states and objects) independent and separately predict them to reduce cardinality. However, it ignores the heavy dependence between states, objects, and compositions. In this paper, we model the dependence via feasibility and contextuality. Feasibility-dependence refers to the unequal feasibility of compositions, e.g., hairy is more feasible with cat than with building in the real world. Contextuality-dependence represents the contextual variance in images, e.g., cat shows diverse appearances when it is dry or wet. We design Semantic Attention (SA) to capture the feasibility semantics to alleviate impossible predictions, driven by the visual similarity between simple primitives. We also propose a generative Knowledge Disentanglement (KD) to disentangle images into unbiased representations, easing the contextual bias. Moreover, we complement the independent compositional probability model with the learned feasibility and contextuality compatibly. In the experiments, we demonstrate our superior or competitive performance, SA-and-kD-guided Simple Primitives (SAD-SP), on three benchmark datasets.
Zhe Liu 0023, Lina Yao 0001, Xiaojun Chang, Wei Fang 0001, Xiaojun Wu 0001, Abdulmotaleb El Saddik
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Deep Scene Flow Learning: From 2D Images to 3D Point Clouds
abstract
Scene flow describes the 3D motion in a scene. It can be modeled as a single task or as a composite of the auxiliary tasks of depth, camera motion, and optical flow estimation. Deep learning's emergence in recent years has broadened the horizons for new methodologies in estimating these tasks, either as separate tasks or as joint tasks to reconstruct the scene flow. The sequence of images that are either synthesized or captured by a camera is used as input for these methods, which face the challenge of dealing with various situations in images to provide the most accurate motion, such as image quality. Nowadays, images have been superseded by point clouds, which provide 3D information, thereby expediting and enhancing the estimated motion. In this paper, we dig deeply into scene flow estimation in the deep learning era. We provide a comprehensive overview of the important topics regarding both image-based and point-cloud-based methods. In addition, we cover the methodologies for each category, highlighting the network architecture. Furthermore, we provide a comparison between these methods in terms of performance and efficiency. Finally, we conclude this survey with insights and discussions on the open issues and future research directions.
Xuezhi Xiang, Rokia Abdein, Wei Li 0109, Abdulmotaleb El Saddik
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 InvFlow: Involution and multi-scale interaction for unsupervised learning of optical flow
Xuezhi Xiang, Rokia Abdein, Ning Lv 0001, Abdulmotaleb El Saddik
Pattern Recognit.4
2024 Incentive Mechanism Design Toward a Win-Win Situation for Generative Art Trainers and Artists
abstract
The recent development of generative art, a typical category of artificial intelligence-generated content (AIGC), is essentially beneficial for social good, which can help amateurs to create artwork and improve experts’ efficiency. However, some artists are opposed to generative art technologies due to the copyright infringement and influence of the artists’ way of earning a living, which makes the artists protest against generative art technologies, causing a lose–lose situation. Adversarial attacks against generative model training are potential solutions to address this issue, while the lose–lose situation cannot be improved. To build a win–win situation, a feasible method is to incentivize the artists to actively contribute their artworks to generative model training without influencing their living or infringing copyright, such as data crowdsourcing, but traditional data crowdsourcing methods cannot well fit the generative art area. Therefore, this article builds a blockchain-based trading system for generative model training data collection and generated artwork circulation. Specifically, this article formulates a social welfare maximization problem based on the reverse auction and designs a corresponding incentive mechanism. The conducted theoretical analysis and numerical evaluation demonstrate the effectiveness of the proposed incentive mechanism toward a win–win situation for generative art model trainers and artists.
Haihan Duan, Abdulmotaleb El Saddik, Wei Cai 0002
IEEE Trans. Comput. Soc. Syst.2
2024 Dual-Branch Hybrid Learning Network for Unbiased Scene Graph Generation
abstract
The current studies of Scene Graph Generation (SGG) focus on solving the long-tailed problem for generating unbiased scene graphs. However, most de-biasing methods over-emphasize the tail predicates and underestimate head ones throughout training, thereby wrecking the representation ability of head predicate features. Furthermore, these impaired features from head predicates harm the learning of tail predicates. In fact, the inference of tail predicates heavily depends on the general patterns learned from head ones, e.g., “standing on” depends on “on”. Thus, these de-biasing SGG methods can neither achieve excellent performance on tail predicates nor satisfying behaviors on head ones. To address this issue, we propose a Dual-branch Hybrid Learning network (DHL) to take care of both head predicates and tail ones for SGG, including a Coarse-grained Learning Branch (CLB) and a Fine-grained Learning Branch (FLB). Specifically, the CLB is responsible for learning expertise and robust features of head predicates, while the FLB is expected to predict informative tail predicates. Furthermore, DHL is equipped with a Branch Curriculum Schedule (BCS) to make the two branches work well together. Experiments show that our approach achieves a new state-of-the-art performance on VG and GQA datasets and makes a trade-off between the performance of tail predicates and head ones. Moreover, extensive experiments on two downstream tasks (i.e., Image Captioning and Sentence-to-Graph Retrieval) further verify the generalization and practicability of our method. Our code is available athttps://github.com/aa200647963/SGG-DHL/.
Chaofan Zheng, Lianli Gao, Xinyu Lyu, Pengpeng Zeng, Abdulmotaleb El Saddik, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.5
2024 OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images
abstract
Oriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer.
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.7
2024 How to Cache Important Contents for Multi-Modal Service in Dynamic Networks: A DRL-Based Caching Scheme
abstract
With the continuous evolution of networking technologies, multi-modal services that involve video, audio, and haptic contents are expected to become the dominant multimedia service in the near future. Edge caching is a key technology that can significantly reduce network load and content transmission latency, which is critical for the delivery of multi-modal contents. However, existing caching approaches only rely on a limited number of factors, e.g., popularity, to evaluate their importance for caching, which is inefficient for caching multi-modal contents, especially in dynamic network environments. To overcome this issue, we propose a content importance-based caching scheme which consists of a content importance evaluation model and a caching model. By leveraging dueling double deep Q networks (D3QN) model, the content importance evaluation model can adaptively evaluate contents' importance in dynamic networks. Based on the evaluated contents' importance, the caching model can easily cache and evict proper contents to improve caching efficiency. The simulation results show that the proposed content importance-based caching scheme outperforms existing caching schemes in terms of caching hit ratio (at least 15% higher), reduced network load (up to 22% reduction), average number of hops (up to 27% lower), and unsatisfied requests ratio (more than 47% reduction).
Zhe Zhang 0010, Marc St-Hilaire, Xin Wei 0001, Haiwei Dong 0001, Abdulmotaleb El Saddik
IEEE Trans. Multim.5
2024 Web3 Metaverse: State-of-the-Art and Vision
abstract
The metaverse, as a rapidly evolving socio-technical phenomenon, exhibits significant potential across diverse domains by leveraging Web3 (a.k.a. Web 3.0) technologies such as blockchain, smart contracts, and non-fungible tokens (NFTs). This survey aims to provide a comprehensive overview of the Web3 metaverse from a human-centered perspective. We (i) systematically review the development of the metaverse over the past 30 years, highlighting the balanced contributions from its core components: Web3, immersive convergence, and crowd intelligence communities, (ii) define the metaverse that integrates the Web3 community as the Web3 metaverse and propose an analysis framework from the community, society, and human layers to describe the features, missions, and relationships for each community and their overlapping sections, (iii) survey the state-of-the-art of the Web3 metaverse from a human-centered perspective, namely, the identity, field, and behavior aspects, and (iv) provide supplementary technical reviews. To the best of our knowledge, this work represents the first systematic, interdisciplinary survey on the Web3 metaverse. Specifically, we commence by discussing the potential for establishing decentralized identities (DID) utilizing mechanisms such as profile picture (PFP) NFTs, domain name NFTs, and soulbound tokens (SBTs). Subsequently, we examine land, utility, and equipment NFTs within the Web3 metaverse, highlighting interoperable and full on-chain solutions for existing centralization challenges. Lastly, we spotlight current research and practices about individual, intra-group, and inter-group behaviors within the Web3 metaverse, such as Creative Commons Zero license (CC0) NFTs, decentralized education, decentralized science (DeSci), and decentralized autonomous organizations (DAO). Furthermore, we share our insights into several promising directions, encompassing three key socio-technical facets of Web3 metaverse development.
Hongzhou Chen, Haihan Duan, Maha Abdallah, Yufeng Zhu, Yonggang Wen 0001, Abdulmotaleb El Saddik, Wei Cai 0002
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Black-box Attack against Self-supervised Video Object Segmentation Models with Contrastive Loss
abstract
Deep learning models have been proven to be susceptible to malicious adversarial attacks, which manipulate input images to deceive the model into making erroneous decisions. Consequently, the threat posed to these models serves as a poignant reminder of the necessity to focus on the model security of object segmentation algorithms based on deep learning. However, the current landscape of research on adversarial attacks primarily centers around static images, resulting in a dearth of studies on adversarial attacks targeting Video Object Segmentation (VOS) models. Given that a majority of self-supervised VOS models rely on affinity matrices to learn feature representations of video sequences and achieve robust pixel correspondence, our investigation has delved into the impact of adversarial attacks on self-supervised VOS models. In response, we propose an innovative black-box attack method incorporating contrastive loss. This method induces segmentation errors in the model through perturbations in the feature space and the application of a pixel-level loss function. Diverging from conventional gradient-based attack techniques, we adopt an iterative black-box attack strategy that incorporates contrastive loss across the current frame, any two consecutive frames, and multiple frames. Through extensive experimentation conducted on the DAVIS 2016 and DAVIS 2017 datasets using three self-supervised VOS models and one unsupervised VOS model, we unequivocally demonstrate the potent attack efficiency of the black-box approach. Remarkably, theJ&Fmetric value experiences a significant decline of up to 50.08% post-attack.
Ying Chen 0005, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Meetor: A Human-Centered Automatic Video Editing System for Meeting Recordings
abstract
Widely adopted digital cameras and smartphones have generated a large number of videos, which have brought a tremendous workload to video editors. Recently, a variety of automatic/semi-automatic video editing methods have been proposed to tackle these issues in some specific areas. However, for the production of meeting recordings, the existing studies highly depend on extra equipment in the conference venues, such as the infrared camera or special microphone, which are not practical. In this article, we design and implement Meetor, a human-centered automatic video editing system for meeting recordings. The Meetor mainly contains three parts: an audio-based video synchronization algorithm, human-centered video content flaw detection algorithms, and an automatic video editing algorithm. Two main experiments are conducted from both objective and subjective aspects to evaluate the performance of the Meetor. The experimental results on a testbed illustrate that the proposed algorithms could achieve state-of-the-art (SOTA) performance in video content flaw detection. However, the conducted user study demonstrates that Meetor could generate meeting recordings with a satisfactory quality compared with professional video editors. Moreover, we also present a practical application of the Meetor in a university campus prototype, in which the Meetor is applied in the automatic editing of lecture recordings. All in all, the proposed Meetor can be utilized in practical applications to release the workload of professional video editors.
Haihan Duan, Junhua Liao, Lehao Lin, Abdulmotaleb El Saddik, Wei Cai 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Meta-Review on Brain-Computer Interface (BCI) in the Metaverse
abstract
This article presents a comprehensive meta-review of the intersection between Brain-Computer Interface (BCI) technologies and the Metaverse, emphasizing the enhancement of immersive experiences through VR, AR, MR, XR, Digital Twin, and haptic interfaces. The study classifies BCI devices into wearable and non-wearable categories, with a focus on their applications in robotics. It explores BCI user feedback mechanisms and their impact on medical and non-medical settings, including personalized rehabilitation and immersive gaming. The review introduces two frameworks for leveraging the Metaverse to navigate multisensory integration between BCI and assistive devices. Applications such as VR therapies for stroke patients and neuro-responsive multiplayer gaming environments showcase the potential of BCIs to enhance Metaverse interactions. To the best of our knowledge, this is the first meta-review on the integration of BCI and the Metaverse, identifying key challenges and research gaps, and serves as a foundational reference for future research and development in this interdisciplinary field.
Kamran Gholizadeh HamlAbadi, Fedwa Laamarti, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Seventeen Years of the ACM Transactions on Multimedia Computing, Communications and Applications: A Bibliometric Overview
abstract
ACM Transactions on Multimedia Computing, Communications, and Applications has been dedicated to advancing multimedia research, fostering discoveries, innovations, and practical applications since 2005. The journal consistently publishes top-notch, original research in emerging fields through open submissions, calls for articles, special issues, rigorous review processes, and diverse research topics. This study aims to delve into an extensive bibliometric analysis of the journal, utilising various bibliometric indicators. The article seeks to unveil the latent implications within the journal’s scholarly landscape from 2005 to 2022. The data primarily draws from the Web of Science Core Collection database. The analysis encompasses diverse viewpoints, including yearly publication rates and citations, identifying highly cited articles, and assessing the most prolific authors, institutions, and countries. The article employs VOSviewer-generated graphical maps, effectively illustrating networks of co-citations, keyword co-occurrences, and institutional and national bibliographic couplings. Furthermore, the study conducts a comprehensive global and temporal examination of co-occurrences of the author’s keywords. This investigation reveals the emergence of numerous novel keywords over the past decades.
Walayat Hussain, Honghao Gao, Rafiul Karim, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical Transformers
abstract
Supervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer .
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Multi-Modal LiDAR Point Cloud Semantic Segmentation with Salience Refinement and Boundary Perception
abstract
Point cloud segmentation is essential for scene understanding, which provides advanced information for many applications, such as autonomous driving, robots, and virtual reality. To improve the accuracy and robustness of point cloud segmentation, many researchers have attempted to fuze camera images to complement the color and texture information. The common fusion strategy is the combination of convolutional operations with concatenation, element-wise addition or element-wise multiplication. However, conventional convolutional operators tend to confine the fusion of modal features within their receptive fields, which can be incomplete and limited. In addition, the inability of encoder–decoder segmentation networks to explicitly perceive segmentation boundary information results in semantic ambiguity and classification errors at object edges. These errors are further amplified in point cloud segmentation tasks, significantly affecting the accuracy of point cloud segmentation. To address the above issues, we propose a novel self-attention multi-modal fusion semantic segmentation network for point cloud semantic segmentation. Firstly, to effectively fuze different modal features, we propose a self-cross fusion module (SCF), which models long-range modality dependencies and transfers complementary image information to the point cloud to fully leverage the modality-specific advantages. Secondly, we design the salience refinement module (SR), which calculates the importance of channels in the feature maps and global descriptors to enhance the representation capability of salient modal features. Finally, we propose the local-aware anisotropy loss measure the element-level importance in the data and explicitly provide boundary information for the model, which alleviates the inherent semantic ambiguity problem in segmentation networks. Extensive experiments on two benchmark datasets demonstrate that our proposed method surpasses current state-of-the-art methods.
Yong Zhou 0003, Zeming Xie, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2023 StyleRF: Zero-Shot 3D Style Transfer of Neural Radiance Fields
abstract
3D style transfer aims to render stylized novel views of a 3D scene with multiview consistency. However, most existing work suffers from a three-way dilemma over accurate geometry reconstruction, high-quality stylization, and being generalizable to arbitrary new styles. We propose StyleRF (Style Radiance Fields), an innovative 3D style transfer technique that resolves the three-way dilemma by performing style transformation within the feature space of a radiance field. StyleRF employs an explicit grid of high-level features to represent 3D scenes, with which highfidelity geometry can be reliably restored via volume rendering. In addition, it transforms the grid features according to the reference style which directly leads to high-quality zero-shot style transfer. StyleRF consists of two innovative designs. The first is sampling-invariant content transformation that makes the transformation invariant to the holistic statistics of the sampled 3D points and accordingly ensures multi-view consistency. The second is deferred style transformation of 2D feature maps which is equivalent to the transformation of 3D points but greatly reduces memory footprint without degrading multi-view consistency. Extensive experiments show that StyleRF achieves superior 3D stylization quality with precise geometry reconstruction and it can generalize to various new styles in a zero-shot manner. Project website: https://kunhao-liu.github.io/StyleRF/
Kunhao Liu, Fangneng Zhan, Yingchen Yu, Abdulmotaleb El Saddik, Shijian Lu, Eric P. Xing
CVPR6
2023 3D Semantic Segmentation in the Wild: Learning Generalized Models for Adverse-Condition Point Clouds
abstract
Robust point cloud parsing under all-weather conditions is crucial to level-5 autonomy in autonomous driving. However, how to learn a universal 3D semantic segmentation (3DSS) model is largely neglected as most existing benchmarks are dominated by point clouds captured under normal weather. We introduce SemanticSTF, an adverse-weather point cloud dataset that provides dense point-level annotations and allows to study 3DSS under various adverse weather conditions. We study all-weather 3DSS modeling under two setups: 1) domain adaptive 3DSS that adapts from normal-weather data to adverse-weather data; 2) domain generalizable 3DSS that learns all-weather 3DSS models from normal-weather data. Our studies reveal the challenge while existing 3DSS methods encounter adverse-weather data, showing the great value of SemanticSTF in steering the future endeavor along this very meaningful research direction. In addition, we design a domain randomization technique that alternatively randomizes the geometry styles of point clouds and aggregates their embeddings, ultimately leading to a generalizable model that can improve 3DSS under various adverse weather effectively. The SemanticSTF and related codes are available at https://github.com/xiaoaoran/SemanticSTF.
Aoran Xiao, Jiaxing Huang 0001, Weihao Xuan, Ruijie Ren, Kangcheng Liu, Dayan Guan, Abdulmotaleb El Saddik, Shijian Lu, Eric P. Xing
CVPR7
2023 Deformable Cross Attention for Learning Optical Flow
abstract
Optical flow is the process of estimating motion in scenes. Each object in the scene has a homogeneous motion, i.e., moves in the same direction with the same velocity. Therefore, connecting the parts of an image globally provides an essential cue for learning accurate motion. Convolution-based methods estimate the motion features from the local regions, which miss this important cue. Recently, some methods used Transformer to model global dependencies to improve optical flow. However, Transformer suffers from excessive attention computations and still brings irrelevant parts into the region of interest. Therefore, we propose a deformable cross-attention for optical flow estimation, which provides two important advantages: connecting the parts of the image globally while deforming the attention to the objects’ shapes in the image and reducing the memory consumption. Our proposed method achieved competitive performance on Sintel and KITTI 2015 datasets in terms of accuracy and efficiency.
Rokia Abdein, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik
ICASSP4
2023 Transpointflow: Learning Scene Flow from Point Clouds with Transformer
abstract
Scene flow estimation is the task of obtaining 3D motion from a dynamic scene. Due to the sparseness of point clouds, extracting features for a local group of points separately may result in different features that may all belong to the same object. This difference makes global correlation prone to producing an unacceptable flow. Local correlation restricts the algorithm to capturing limited movements and fails when fast movement or large deformation of an object occurs. Therefore, we propose a transformer-based scene flow method that can perform global feature modeling through a self-attention layer. Moreover, we propose a cross-attention-based flow embedding layer for global feature matching. We further propose a learnable attention-based up-sampling layer to up-sample the estimated flow to higher resolution based on a single feature scale, eliminating the need to model global dependencies at all scales. Experimental results show that our model produces competitive results on Flyingthings3D and KITTI datasets with efficient performance.
Rokia Abdein, Xuezhi Xiang, Abdulmotaleb El Saddik
ICIP3
2023 CEAFFOD: Cross-Ensemble Attention-based Feature Fusion Architecture Towards a Robust and Real-time UAV-based Object Detection in Complex Scenarios
abstract
Deploying object detectors in embedded devices such as unmanned aerial vehicles (UAVs) comes with many challenges. This is due to both the UAV itself having low embedded resources in terms of computation and memory, and also due to the nature of the captured visual data with the variations in objects' scale, orientation, density, viewpoint, distribution, shape, context and others. It is crucial for the object detector to be robust with high accuracy, real-time with fast inference and light-weight to be applicable. Inspired by YOLO architecture, we propose a novel single-stage detection architecture. Our contributions are, first, feature fusion spatial pyramid pooling (FFSPP) block that applies attention-based feature fusion across both time and space utilizing the information of subsequent frames and scales in an efficient manner. Secondly, we introduce a multi-dilated attention-based cross-stage partial connection (MDACSP) block that helps in increasing the receptive field and producing per-channel modulation weights after aggregating the feature maps across their spatial domain. Third, scaled feature fusion head (SFFH) fuses both the FFSPP block features and the connected MDACSP block features specific for this head. For a more robust result across different scenarios, we perform cross-ensembling with three of the top UAV/traffic surveillance datasets: UAVDT, UA-DETRAC and VisDrone. Our ablation study shows how every contribution improves over the baseline. Our approach yielded the state-of-the-art results in all the aforementioned datasets achieving 89.3% mAP, 93.5% mAP, and 42.9% mAP respectively. Testing the model performance on NVIDIA Jetson Xavier NX board shows a desirable balance between the inference time and the memory cost. We also show qualitatively the model robustness and efficiency across the diverse complex scenarios of these datasets. We hope this work facilitates the advancement of the UAV-based perception in such crucial industrial applications.
Ahmed Elhagry, Hang Dai, Abdulmotaleb El Saddik, Wail Gueaieb, Giulia De Masi
ICRA3
2023 Metaverse Services: The Way of Services Towards the Future
abstract
With the emergence of new generation of digital technologies, e.g., artificial intelligence, blockchain, cloud computing, big data, edge computing, 5G/6G, VR/AR/MR, and the Internet of Things, an exciting era of metaverse is coming. Interacted and linked with the physical world, metaverse offers a platform of a new social ecosystem, dealing with digital twins and empowering virtual-reality symbiosis. In metaverse, social activities and business processes are performed based on the sequences of workflow or service processes. Bridging both the virtual space and the real world, such metaverse services are more complicated and present many new challenges and research topics. In this paper, the concept and characteristics of metaverse services are presented, the key technologies and typical use cases are reviewed, and the future challenges and opportunities of metaverse services are also discussed.
Xiaofei Xu 0001, Quan Z. Sheng, Boualem Benatallah, Zhong Chen 0001, Robert Gazda, Abdulmotaleb El Saddik, Munindar P. Singh
ICWS6
2023 Pseudo-Stereo++: Cycled Generative Pseudo-Stereo for Monocular 3D Object Detection in Autonomous Driving
abstract
Recently, the feature-level generation has demonstrated the effectiveness of pseudo-stereo synthesis in Monocular 3D Detection (M3D). In this paper, we aim to further bridge the gap between the stereo and the monocular 3D object detectors in autonomous driving through direct image-level pseudo-stereo generation. We propose a novel Cycled Generative Pseudo-Stereo (CGPS) architecture to generate the right-view image from the left-view for constructing a pseudo-stereo pair to stereo 3D object detectors while maintaining the natural of M3D with the left-view image as the only input. Moreover, we use a triplet consistency loss to focus on the detected objects in the pseudo-stereo generation. Besides, we demonstrate that the proposed CGPS is an ad-hoc module to adapt top stereo 3D object detectors into monocular 3D object detectors. The proposed framework with CGPS achieves 74.80%, 55.28%, and 46.70% 3DAP for easy, moderate, and hard difficulty levels in monocular 3D detection on the KITTI benchmark with comparable performance to the stereo 3D object detectors but using a monocular image as the only input. Till the submission, the proposed M3D framework ranks 1stwith dramatic improvements against the existing monocular 3D detectors on the KITTI benchmark.
Ahmed Elhagry, Hang Dai, Abdulmotaleb El Saddik
IROS3
2023 Weakly Supervised 3D Open-vocabulary Segmentation
abstract
Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impeded by the lack of large-scale and diverse 3D open-vocabulary segmentation datasets for training robust and generalizable models. Distilling knowledge from pre-trained 2D open-vocabulary segmentation models helps but it compromises the open-vocabulary feature as the 2D models are mostly finetuned with close-vocabulary datasets. We tackle the challenges in 3D open-vocabulary segmentation by exploiting pre-trained foundation models CLIP and DINO in a weakly supervised manner. Specifically, given only the open-vocabulary text descriptions of the objects in a scene, we distill the open-vocabulary multimodal knowledge and object reasoning capability of CLIP and DINO into a neural radiance field (NeRF), which effectively lifts 2D features into view-consistent 3D segmentation. A notable aspect of our approach is that it does not require any manual segmentation annotations for either the foundation models or the distillation process. Extensive experiments show that our method even outperforms fully supervised models trained with segmentation annotations in certain scenes, suggesting that 3D open-vocabulary segmentation can be effectively learned from 2D images and text-image pairs. Code is available at https://github.com/Kunhao-Liu/3D-OVS.
Kunhao Liu, Fangneng Zhan, Muyu Xu, Yingchen Yu, Abdulmotaleb El Saddik, Christian Theobalt, Eric P. Xing, Shijian Lu
NeurIPS6
2023 Global-aware and local-aware enhancement network for person search
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik
Comput. Vis. Image Underst.5
2023 Human-Centric Resource Allocation for the Metaverse With Multiaccess Edge Computing
abstract
Multiaccess edge computing (MEC) is a promising solution to the computation-intensive, low-latency rendering tasks of the metaverse. However, how to optimally allocate limited communication and computation resources at the edge to a large number of users in the metaverse is quite challenging. In this article, we propose an adaptive edge resource allocation method based on multiagent soft actor–critic with graph convolutional networks (SAC-GCN). Specifically, SAC-GCN models the multiuser metaverse environment as a graph where each agent is denoted by a node. Each agent learns the interplay between agents by graph convolutional networks with a self-attention mechanism to further determine the resource usage for one user in the metaverse. The effectiveness of SAC-GCN is demonstrated through the analysis of user experience, balance of resource allocation, and resource utilization rate by taking a virtual city park metaverse as an example. Experimental results indicate that SAC-GCN outperforms other resource allocation methods in improving overall user experience, balancing resource allocation, and increasing resource utilization rate by at least 27%, 11%, and 8%, respectively.
Zijian Long, Haiwei Dong 0001, Abdulmotaleb El Saddik
IEEE Internet Things J.3
2023 Context-aware and part alignment for visible-infrared person re-identification
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Abdulmotaleb El Saddik
Image Vis. Comput.6
2023 Guest Editorial Digital Twins for Mobile Networks - Part I
abstract
Digital twins (DTs), defined as the virtual representation of a real-world entity or system, act as a mirror to provide a way to simulate, predict physical behaviors, and possibly control the real-world entity where applicable. Originating in the industry, advances in computing capacity and recent progress in artificial intelligence (AI)-based analytics make DTs attractive to a broader set of use cases including mobile networks.
Shahid Mumtaz, Soumaya Cherkaoui, Mohsen Guizani, Joel J. P. C. Rodrigues, Abdulmotaleb El Saddik, Sabita Maharjan, Yang Xiao 0001, Muhammad Ikram Ashraf
IEEE J. Sel. Areas Commun.5
2023 Guest Editorial Digital Twins for Mobile Networks - Part II
abstract
6G communication networks are expected to become an integral part of the infrastructure needed for developing a smart society in the future. Addressing the challenges on the road towards realizing 6G network requirements in terms of quality of service, user experience, and security, is therefore of utmost importance. The digital twin (DT) technology can potentially improve the efficiency, reliability, and security of 6G networks. Digital twins for mobile networks (DTMNs) are seen as a key factor in harnessing the full benefits of 6G. Using digital twins can help address several problems, including network optimization, fault diagnosis, and fault management. Furthermore, DTMNs can characterize the physical entities in a 6G network and their relationships to each other, build their virtual models, and use simulation, learning, and reasoning capabilities to make predictions and support informed decision-making,
Shahid Mumtaz, Soumaya Cherkaoui, Mohsen Guizani, Joel J. P. C. Rodrigues, Abdulmotaleb El Saddik, Sabita Maharjan, Yang Xiao 0001, Muhammad Ikram Ashraf
IEEE J. Sel. Areas Commun.5
2023 EMHIFormer: An Enhanced Multi-Hypothesis Interaction Transformer for 3D human pose estimation in video
Xuezhi Xiang, Kaixu Zhang, Yulong Qiao, Abdulmotaleb El Saddik
J. Vis. Commun. Image Represent.4
2023 AAD-Net: Advanced end-to-end signal processing system for human emotion detection & recognition using attention-based deep echo state network
Mustaqeem Khan 0001, Abdulmotaleb El Saddik, Fahd Alotaibi 0001, Nhat Truong Pham
Knowl. Based Syst.2
2023 A Multimodal Coupled Graph Attention Network for Joint Traffic Event Detection and Sentiment Classification
abstract
Traffic events are one of the main causes of traffic accidents, leading to traffic event detection being a challenging research problem in traffic management and intelligent transportation systems (ITSs). The main gap in this task lies in how to extract and represent the valuable information from various kinds of traffic data. Considering the important role that social networks play in traffic data analysis, we argue that sentiment classification and traffic event detection are two closely related tasks in ITSs, where event and sentiment can reveal both explicit and implicit traffic accidents, respectively. Unfortunately, none of the recent approaches in traffic event detection have taken sentiment knowledge into view. This paper proposes a multimodal coupled graph attention network (MCGAT). It aims to construct a multimodal multitask interactive graphical structure where terms (sucha as words, and pixels) are treated as nodes, and their contextual and cross-modal correlations are formalized as edges. The key components are cross-modal and cross-task graph connection layers. The cross-modal graph connection layer captures the multimodal representation, where each node in one modality connects all nodes in another modality. The cross-task graph connection layer is designed by connecting the multimodal node in one task to two single nodes in another task. Empirical evaluation of two benchmarking datasets, such as MGTES and Twitter, shows the effectiveness of the proposed model over state-of-the-art baselines in terms of F1 and accuracy, with significant improvements of 2.4%, 2.4%, 2.7%, and 2.7%.
Yazhou Zhang 0001, Prayag Tiwari, Abdulmotaleb El Saddik, M. Shamim Hossain
IEEE Trans. Intell. Transp. Syst.4
2023 Spatial-Channel Enhanced Transformer for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task in computer vision, aiming at matching people across images from visible and infrared modalities. The widely used VI-ReID framework consists of a convolution neural backbone network that extracts the visual features, and a feature embedding network to project heterogeneous features to the same feature space. However, many studies based on the existing pre-trained models neglect potential correlations between different locations and channels within a single sample during the feature extraction. Inspired by the success of the Transformer in computer vision, we extend it to enhance feature representation for VI-ReID. In this paper, we propose a discriminative feature learning network based on a visual Transformer (DFLN-ViT) for VI-ReID. Firstly, to capture long-term dependencies between different locations, we propose a spatial feature awareness module (SAM), which utilizes a single-layer Transformer with a novel patch-embedding strategy to encode location information. Secondly, to refine the representation at each channel, we design a channel feature enhancement module (CEM). The CEM treats the features of each channel as a sequence of Transformer inputs, taking advantage of the Transformer's ability to model long-term dependencies. Finally, we propose a Triplet-aided Hetero-Center (THC) loss to learn more discriminative feature representation by balancing the cross-modality distance and intra-modality distance of the center. The experimental results on two datasets show that our method can significantly improve the VI-ReID performance, outperforming most state-of-the-art methods.
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Silin Chen, Abdulmotaleb El Saddik
IEEE Trans. Multim.6
2023 Distilled Meta-learning for Multi-Class Incremental Learning
abstract
Meta-learning approaches have recently achieved promising performance in multi-class incremental learning. However, meta-learners still suffer from catastrophic forgetting, i.e., they tend to forget the learned knowledge from the old tasks when they focus on rapidly adapting to the new classes of the current task. To solve this problem, we propose a novel distilled meta-learning (DML) framework for multi-class incremental learning that integrates seamlessly meta-learning with knowledge distillation in each incremental stage. Specifically, during inner-loop training, knowledge distillation is incorporated into the DML to overcome catastrophic forgetting. During outer-loop training, a meta-update rule is designed for the meta-learner to learn across tasks and quickly adapt to new tasks. By virtue of the bilevel optimization, our model is encouraged to reach a balance between the retention of old knowledge and the learning of new knowledge. Experimental results on four benchmark datasets demonstrate the effectiveness of our proposal and show that our method significantly outperforms other state-of-the-art incremental learning methods.
Hao Liu 0065, Zhaoyu Yan, Bing Liu 0016, Jiaqi Zhao 0001, Yong Zhou 0003, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2022 Special issue deep learning for multimedia healthcare
abstract
Text, radiological pictures, audio notes, video, and other types of multimedia healthcare data are all generated by today's smart healthcare system [1].The evolution of COVID-19 has resulted in an incremental rise in current healthcare data.The study of multimodal healthcare data on such a big scale has revealed both obstacles and potential.Thanks to artificial intelligence (AI) and, more specifically, deep learning (DL) algorithms, which have been widely used by researchers for handling massive amounts of epidemic data, predicting live epidemic crises, and initiating new research directions in the analysis of healthcare multimedia data [2].As a result, deep learning for multimedia healthcare data analysis is becoming a hot topic in multimedia and computer vision research.The call for papers attracted 54 submissions and after a rigorous review, 20 papers have been accepted for this special issue.A brief summary of papers in this special issue is presented in the following:The paper titled "A Novel Study for Automatic Twoclass Covid-19 Diagnosis (between Covid-19 and Healthy, Pneumonia) on X-ray Images using Texture Analysis and 2-D/3-D Convolutional Neural Networks" aims to diagnose COVID-19 early using X-ray images, automatic two-class classification was carried out in four different titles: COVID-19/Healthy, COVID-19 Pneumonia/Bacterial Pneumonia, COVID-19 Pneumonia/Viral Pneumonia, and COVID-19 Pneumonia/Other Pneumonia.In the study, besides using
M. Shamim Hossain, Josu Bilbao, Diana P. Tobón, Muhammad Ghulam, Abdulmotaleb El Saddik
Multim. Syst.5
2022 Deep learning in multimedia healthcare applications: a review
Diana P. Tobón, M. Shamim Hossain, Muhammad Ghulam, Josu Bilbao, Abdulmotaleb El Saddik
Multim. Syst.5
2022 CLT-Det: Correlation Learning Based on Transformer for Detecting Dense Objects in Remote Sensing Images
abstract
Challenges still exist in the task of object detection in remote sensing images with densely distributed objects due to large variation in scale and neglect of the relative position and correlation. To address these issues, a Correlation Learning Detector based on Transformer (CLT-Det) is proposed for detecting dense objects in remote sensing images. A Transformer Attention Module (TAM) is designed to improve the densely packed objects’ model representation ability by learning pixel-wise attention with Transformer. To alleviate the semantic gap caused by variations in scale, a Feature Refinement Module (FRM) is proposed by improving the multi-scale feature pyramid. A Correlation Transformer Module (CTM) is proposed to extract correlation information and encodes position information of dense objects’ features on the classification branch for fully utilizing the position information and correlation among objects. Extensive experiments compared with several state-of-art methods on two challenging remote sensing datasets, namely DOTA and HRSC2016, demonstrate that the proposed CLT-Det achieves promising and competitive performance.
Yong Zhou 0003, Silin Chen, Jiaqi Zhao 0001, Rui Yao 0006, Yong Xue, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.6
2022 Special Section on Edge-AI for Connected Living
abstract
introduction Share on Special Section on Edge-AI for Connected Living Editors: M. Shamim Hossain King Saud University, Saudi Arabia King Saud University, Saudi ArabiaView Profile , Changsheng Xu Chinese Academy of Sciences, China Chinese Academy of Sciences, ChinaView Profile , Josu Bilbao IKERLAN, Spain IKERLAN, SpainView Profile , Md. Abdur Rahman University of Prince Mugrin, KSA University of Prince Mugrin, KSAView Profile , Abdulmotaleb El Saddik University of Ottawa, Canada University of Ottawa, CanadaView Profile , Mohamed Bin Zayed University of Artificial Intelligence, UAE & University of Ottawa, Canada University of Artificial Intelligence, UAE & University of Ottawa, CanadaView Profile Authors Info & Claims ACM Transactions on Internet TechnologyVolume 22Issue 3August 2022 Article No.: 55epp 1–3https://doi.org/10.1145/3514196Published:14 March 2022Publication History 0citation176DownloadsMetricsTotal Citations0Total Downloads176Last 12 Months176Last 6 weeks24 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
M. Shamim Hossain, Changsheng Xu, Josu Bilbao, Mohamed Abdur Rahman 0001, Abdulmotaleb El Saddik, Mohamed Bin Zayed
ACM Trans. Internet Techn.5
2022 MMSUM Digital Twins: A Multi-view Multi-modality Summarization Framework for Sporting Events
abstract
Sporting events generate a massive amount of traffic on social media with live moment-to-moment accounts as any given situation unfolds. The generated data are intensified by fans feelings, reactions, and subjective opinions towards what happens during the event, all of which are based on their individual points of view. Analyzing and summarizing this data will generate a comprehensive overview of the event in terms of how the event evolves and how fans react and view the event based on their perspectives. Previously, most of the summarization works ignore fan reactions and subjective opinions, and focus primarily on generating an objective-view summary. We believe that an effective and useful summary should consider human reactions, sentiment, and point of view, as opposed to simply describing what happens during the event. Accordingly, in this work, we propose MMSUM Digital Twins: a summarization framework that is capable of generating a multi-view multi-modal summary for sporting events in real-time. The proposed digital twins-based framework consists of four main components: sub-event recognition which detects the event’s key moments, tweet categorization, which determines which team the tweets’ writers support and assigns tweets to their teams, sentiment analysis to track fans’ state of mind, and image popularity prediction for selecting representative images. Furthermore, the MMSUM employs a visual-filtering model to address the issue of noisy images that inundate social media, compromising the summarization quality. We leverage the knowledge of sport fans to evaluate the generated multi-view summarization through an online user study. The experiment results confirm the effectiveness of our proposed approach for summarizing sporting events by considering multimedia data, sentiment, and subjective views of the event.
Samah Al-Oufi, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Special Section on AI-empowered Multimedia Data Analytics for Smart Healthcare
abstract
No abstract available.
M. Shamim Hossain, Rita Cucchiara, Muhammad Ghulam, Diana P. Tobón, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identification
abstract
In person re-identification (Re-ID) , the data annotation cost of supervised learning, is huge and it cannot adapt well to complex situations. Therefore, compared with supervised deep learning methods, unsupervised methods are more in line with actual needs. In unsupervised learning, a key to solving Re-ID is to find a standard that can effectively distinguish the difference (distance) between the features of images belonging to different pedestrian identities. However, there are some differences in the images captured by different cameras (such as brightness, angle, etc.). It is well known that the training of neural networks is mainly based on the distance between features, while in unsupervised learning, especially in unsupervised learning methods based on hierarchical clustering, the distance between features plays a more important role in the clustering phase. We improve the accuracy of a deep learning method based on hierarchical clustering under fully unsupervised conditions, starting from both feature and distance metrics. First, we propose to use spherical features, by normalizing the images in the feature space, to weaken the structural differences (length) between features, while saving the feature differences (direction) between different identities. Then, we use the sum of squared errors (SSE) as a regularization term to balance different cluster states. We evaluate our method on four large-scale Re-ID datasets, and experiments show that our method achieves better results than the state-of-the-art unsupervised methods.
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.7
2022 Building the Metaverse by Digital Twins at All Scales, State, Relation
abstract
The new-generation information technology development enables Digital Twins to reshape the physical world into the virtual digital space and provide technical support for Metaverse construction. The Metaverse objects can be mesoscale or macro-micro-scales. Metaverse is a complex collection of both solid substances and liquid, gaseous, plasma, and other uncertain states. Additionally, Metaverse integrates the tangibles with social relations, such as interpersonal relations (friendship, love, and blood relations) and the overall social relations (ethics, morality, and law). This work also introduces some principles or laws to construct the Digital Twins model for social relations, such as broken windows theory, small-world phenomenon, survivor bias, and herd behavior. Thus, from multiple angles, it reviews mapping the tangible and intangible real-world objects to the Metaverse using the Digital Twins model.
Zhihan Lyu, Shuxuan Xie, Yuxi Li 0005, M. Shamim Hossain, Abdulmotaleb El Saddik
Virtual Real. Intell. Hardw.5
2021 Stable and Effective One-Step Method for Person Search
abstract
Person search, which requires both pedestrian detection and person re-identification, is a challenging computer vision task applied to real-world scenarios. The challenges faced by detection and re-identification, such as occlusion, poor illumination, confusing background, are still urgent for person search. In addition, one-step methods for person search need to deal with the divergence between two tasks. In this work, we propose an end-to-end model containing the feature extractor, the region proposal network, and the multi-task learning module. In order to process divergence between detection and re-identification, we introduce switchable normalization and gradient centralization to improve the stability of the model. To solve the imbalance problem of hard examples, we introduce focal loss as a classification loss in the multi-task learning module. The experimental results on two bench-marks, i.e., CUHK-SYSU and PRW, well demonstrate that our method outperforms the state-of-the-art one-step methods.
Ning Lv 0001, Xuezhi Xiang, Rokia Abdein, Abdulmotaleb El Saddik
ICASSP6
2021 Unsupervised cross-domain person re-identification with self-attention and joint-flexible optimization
Haopeng Hou, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Ying Chen 0005, Abdulmotaleb El Saddik
Image Vis. Comput.7
2021 An Explainable Deep Learning Ensemble Model for Robust Diagnosis of Diabetic Retinopathy Grading
abstract
Diabetic retinopathy (DR) is one of the most common causes of vision loss in people who have diabetes for a prolonged period. Convolutional neural networks (CNNs) have become increasingly popular for computer-aided DR diagnosis using retinal fundus images. While these CNNs are highly reliable, their lack of sufficient explainability prevents them from being widely used in medical practice. In this article, we propose a novel explainable deep learning ensemble model where weights from different models are fused into a single model to extract salient features from various retinal lesions found on fundus images. The extracted features are then fed to a custom classifier for the final diagnosis of DR severity level. The model is trained on an APTOS dataset containing retinal fundus images of various DR grades using a cyclical learning rates strategy with an automatic learning rate finder for decaying the learning rate to improve model accuracy. We develop an explainability approach by leveraging gradient-weighted class activation mapping and shapely adaptive explanations to highlight the areas of fundus images that are most indicative of different DR stages. This allows ophthalmologists to view our model's decision in a way that they can understand. Evaluation results using three different datasets (APTOS, MESSIDOR, IDRiD) show the effectiveness of our model, achieving superior classification rates with a high degree of precision (0.970), sensitivity (0.980), and AUC (0.978). We believe that the proposed model, which jointly offers state-of-the-art diagnosis performance and explainability, will address the black-box nature of deep CNN models in robust detection of DR grading.
Mohammad Shorfuzzaman, M. Shamim Hossain, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement Networks
abstract
Many deep learning applications benefit from multi-task learning with several related objectives. In autonomous driving scenarios, being able to accurately infer motion and spatial information is essential for scene understanding. In this paper, we combine an adaptive feature refinement module and a unified framework for joint learning of optical flow, depth and camera pose estimation in an unsupervised manner. The feature refinement module is embedded into motion estimation and depth prediction sub-networks, which can exploit more channel-wise relationships and contextual information for feature learning. Given a monocular video, our network firstly estimates depth and camera motion, and calculates rigid optical flow. Then, we design an auxiliary flow network for inferring non-rigid flow fields. In addition, a forward-backward consistency check is adopted for occlusion reasoning. Extensive experiments on KITTI dataset demonstrate that the proposed method achieves potential results comparing to recent deep learning networks.
Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik
ICASSP4
2020 Pavement crack detection network based on pyramid structure and attention mechanism
abstract
Automatic detection of pavement crack is an important task for conducting road maintenance. However, as an important part of the intelligent transportation system, automatic pavement crack detection is challenging due to the poor continuity of cracks, the different width of cracks, and the low contrast between cracks and the surrounding pavement. This study proposes a novel pavement crack detection method based on an end‐to‐end trainable deep convolution neural network. The authors build the network using the encoder–decoder architecture and adopt a pyramid module to exploit global context information for the complex topology structures of cracks. Moreover, they introduce a spatial‐channel combinational attention module into the encoder–decoder network for refining crack features. Further, the dilated convolution is used to reduce the loss of crack details due to the pooling operation in the encoder network. In addition, they introduce a lovász hinge loss function, which is suitable for small objects. They train the authors' network on the CRACK500 dataset and evaluate it on three pavement crack datasets. Among the methods they compare, their method can achieve the best experimental results.
Xuezhi Xiang, Abdulmotaleb El Saddik
IET Image Process.3
2020 A CNNs-based method for optical flow estimation with prior constraints and stacked U-Nets
Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Yulong Qiao, Abdulmotaleb El Saddik
Neural Comput. Appl.5
2020 Dual-Path Part-Level Method for Visible-Infrared Person Re-identification
Xuezhi Xiang, Ning Lv 0001, Mingliang Zhai, Rokia Abdein, Abdulmotaleb El Saddik
Neural Process. Lett.5
2020 Attention-Based Generative Adversarial Network for Semi-supervised Image Classification
Xuezhi Xiang, Zeting Yu, Ning Lv 0001, Abdulmotaleb El Saddik
Neural Process. Lett.5
2020 Optical Flow Estimation Using Dual Self-Attention Pyramid Networks
abstract
Recently, optical flow estimation benefits greatly from deep learning based techniques. Most approaches use encoder-decoder architecture (U-Net) or spatial pyramid network (SPN) to learn optical flow. Both U-Net and SPN can extract multi-scale features and can predict optical flow directly. However, existing networks ignore to exploit the global information among channel features and inter-spatial relationship of features. In this paper, we propose a dual self-attention pyramid network, which adaptively integrates local features with their global dependencies and focuses on important features and suppresses unimportant features. Specifically, we introduce two types of attention modules into SPN, which emphasizes meaningful features along channel and spatial axes. The channel attention can adaptively re-weight channel-wise features by considering interdependencies among channels. Moreover, the spatial attention can utilize global contextual information to emphasize or suppress features in different spatial locations. In addition, two attention modules are embedded into each pyramidal level, which can refine features at different scale. We evaluate our method on MPI-Sintel and KITTI. The experimental results show that using the dual self-attention module can improve the representation power of network and further increase the accuracy of optical flow estimation.
Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik
IEEE Trans. Circuits Syst. Video Technol.5
2020 An Object Context Integrated Network for Joint Learning of Depth and Optical Flow
abstract
Supervised depth prediction and optical flow estimation have achieved promising performance due to the advanced deep network architectures. Since the ground truths are difficult to be collected, many recent works try to learn the depth and flow in an unsupervised manner. However, existing methods only use features from convolutional layers or a simple aggregation of multi-level features to predict the depth and flow maps, which is insufficient to exploit context information. In this paper, we attempt to exploit object contextual information and investigate the effect of the object context for joint learning of depth and optical flow. Specifically, we present a novel combination of object context and the framework of joint learning depth and optical flow. Our proposed network can exploit and integrate the object context for both tasks by aggregating the context according to pair-wise similarities. Furthermore, we adopt the existing spatial pyramid network (SPN) to estimate the depth and flow in a coarse-to-fine strategy effectively. Given temporally adjacent stereo pairs, our network can be trained end-to-end in an unsupervised manner and can predict the depth and flow maps simultaneously. We conduct experiments on two publicly available datasets, KITTI2012 and KITTI2015. Our proposed approach yields comparable performance on both depth and flow tasks, compared to the recent deep learning-based approaches. Experimental results demonstrate that exploiting object contextual information is useful and beneficial for depth and optical flow estimation.
Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik
IEEE Trans. Image Process.5
2019 Ad-net: Attention Guided Network for Optical Flow Estimation Using Dilated Convolution
abstract
Variational models for optical flow estimation usually define an energy function that contains prior assumptions to explore rudimentary statistics of images. However, such methods cannot learn motion knowledge from the pre-prepared data and have many parameters that need to be set manually. Nowadays, convolutional neural networks (CNNs) have been used in optical flow estimation successfully, which can learn weights from the training dataset and can predict optical flow end-to-end. In this paper, we propose an attention guided network for learning optical flow, named AD-Net, which contains several attention units for modelling the relativities between the channels. Further, we introduce dilated convolution into supervised network for reducing the loss of motion details. In addition, some prior auxiliary constraints are embedded in the supervised network as auxiliary loss terms. Our proposed approach is tested on MPI-Sintel and KITTI2012 datasets and can preserve motion edges and details effectively.
Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik
ICASSP5
2019 Optical Flow Estimation Using Spatial-Channel Combinational Attention-Based Pyramid Networks
abstract
Recently, learning to estimate optical flow via deep convolutional networks is attracting significant attention. In this paper, we introduce a spatial-channel attention module into optical flow estimation, which infers attention maps along two separated dimensions, channel and spatial, and then integrates these separated attention maps into a fusion attention map for feature refinement. We embed this module into spatial pyramid network, which can adaptively learn the channel and spatial attention maps at each level for modifying the different scaled features and can further improve the accuracy of optical flow estimation. Our network is trained on FlyingChairs and FlyingThings3D datasets with a supervised manner, and is further tested on MPI-Sintel benchmark. The experimental results show that using the spatial-channel attention unit is beneficial for dense flow estimation and our approach is comparable with the state-of-the-art methods.
Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik
ICIP5
2019 Robust Adaptive Tracking Synchronization Protocols for Leader-follower Multirotor Aerial Vehicles with Uncertainty
abstract
This paper develops robust adaptive tracking synchronization protocol for a group of cloud connected leader-follower multirotor aerial vehicles (MAVs) with uncertainty. The design combines adaptive learning mechanism with sliding mode control vectors to solve consensus tracking synchronization problem for both attitude and position dynamics. The protocols are constructed by using local and neighboring states of the vehicles provided that the vehicles can share states information with neighboring vehicles via local area network. Adaptive learning algorithms are used locally for each vehicle to deal with uncertainty associated with nonlinear dynamics and uncertain flying environment. Lyapunov and sliding mode control method is employed to design and analyze asymptotic convergence of the consensus tracking error functions. The convergence analysis shows that the states of the follower vehicles can reach an agreement and synchronize to the leader vehicle achieving ensuring tracking synchronization property. The protocols design and implementation is simple as they do not require exact knowledge of the dynamical model and uncertainty.
Abdulmotaleb El Saddik, Anderson Sunda-Meya
SMC2
2019 Robust Cooperative Load-Frequency Tracking Protocols for Leader-Follower Smart Power Grid Networks With Uncertainty
abstract
This paper introduces consensus based leader-follower robust cooperative load frequency tracking control(LFTC) protocols for multi-area smart power grid networks. The LFC protocols combines local states with the states of the neighboring areas with directed communication topology. We first develop Robust LFTC protocols by assuming that the bounds of the uncertainty associated with power network dynamics and disturbance are known a priori. Lyapunov and graph theory used to show the finite-time convergence of the states of the follower control area power grid networks to the states of the leader control area. Then, we relax the demand of the bound on the uncertainty from LFTC protocols by integrating a robust adaptive learning algorithms. Robust adaptive learning control uses to deal with the presence of uncertainty associated with the power grid networks and disturbance. Convergence analysis of the closed-loop multi-area power grid networks are shown by using Lyapunov and graph theory. Analysis shows that the states of the follower control areas can reach an agreement and track the states of the leader control area asymptotically. Evaluation results on a four-area interconnected power grid networks are presented to show the effectiveness of the proposed consensus based distributed robust LFTC protocols for real-time applications.
Abdulmotaleb El Saddik, Anderson Sunda-Meya
SMC2
2019 Robust Load Frequency Control for Smart Power Grid Over Open Distributed Communication Network with Uncertainty
abstract
This work develops delay dependent load frequency control scheme for multi-area smart power grid over open communication networks with the presence of uncertainty and unsymmetrical time varying delays. First, the design employs direct method using differential inequalities and matrix measures to derive stability conditions for power grid network systems. The design assumed that the uncertainty associated with modeling errors and external fault disturbances are bounded. The stability conditions are given together with the upper bound of the delays and the convergence rate of the solution trajectory of the closed loop system. Second, we introduce Lyapunov based indirect method to establish stability criterion for LFC systems. The stability condition is established for both symmetrical and unsymmetrical time varying delays in measurement and control channel. The design analyzes the upper bound of the delay and solution trajectory in the presence of uncertainty varying with the state and constant. Compared with the existing designs, the proposed design and analysis uses time varying delays both in measurement channel from RTU to control center and control channel from control center to power generation unit. Unlike the existing LFC schemes, the design employs the uncertainty appearing into multi-area power system networks from the modeling errors, variation of loads and other external disturbances.
Abdulmotaleb El Saddik, Anderson Sunda-Meya
SMC2
2019 Optical flow estimation using channel attention mechanism and dilated convolutional neural networks
Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik
Neurocomputing5
2019 CASP: context-aware stress prediction system
Raneem Alharthi, Rajwa Alharthi, Benjamin Guthier, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2019 Haptic Codecs for the Tactile Internet
abstract
The Tactile Internet will enable users to physically explore remote environments and to make their skills available across distances. An important technological aspect in this context is the acquisition, compression, transmission, and display of haptic information. In this paper, we present the fundamentals and state of the art in haptic codec design for the Tactile Internet. The discussion covers both kinesthetic data reduction and tactile signal compression approaches. We put a special focus on how limitations of the human haptic perception system can be exploited for efficient perceptual coding of kinesthetic and tactile information. Further aspects addressed in this paper are the multiplexing of audio and video with haptic information and the quality evaluation of haptic communication solutions. Finally, we describe the current status of the ongoing IEEE standardization activity P1918.1.1 which has the ambition to standardize the first set of codecs for kinesthetic and tactile information exchange across communication networks.
Eckehard G. Steinbach, Matti Strese, Mohamad A. Eid, Amit Bhardwaj, Qian Liu 0001, Mohammad Al Ja'afreh, Toktam Mahmoodi, Rania Hassen, Abdulmotaleb El Saddik, Oliver Holland
Proc. IEEE10
2019 Toward citation recommender systems considering the article impact in the extended nearby citation network
Abdulrhman Alshareef, Mohammed F. Alhamid, Abdulmotaleb El Saddik
Peer-to-Peer Netw. Appl.3
2019 EVM-CNN: Real-Time Contactless Heart Rate Estimation From Facial Video
abstract
With the increase in health consciousness, noninvasive body monitoring has aroused interest among researchers. As one of the most important pieces of physiological information, researchers have remotely estimated the heart rate (HR) from facial videos in recent years. Although progress has been made over the past few years, there are still some limitations, like the processing time increasing with accuracy and the lack of comprehensive and challenging datasets for use and comparison. Recently, it was shown that HR information can be extracted from facial videos by spatial decomposition and temporal filtering. Inspired by this, a new framework is introduced in this paper to remotely estimate the HR under realistic conditions by combining spatial and temporal filtering and a convolutional neural network. Our proposed approach shows better performance compared with the benchmark on the MMSE-HR dataset in terms of both the average HR estimation and short-time HR estimation. High consistency in short-time HR estimation is observed between our method and the ground truth.
Juan Sebastian Arteaga-Falconi, Haiwei Dong 0001, Abdulmotaleb El Saddik
IEEE Trans. Multim.5
2019 A Deep Learning System for Recognizing Facial Expression in Real-Time
abstract
This article presents an image-based real-time facial expression recognition system that is able to recognize the facial expressions of several subjects on a webcam at the same time. Our proposed methodology combines a supervised transfer learning strategy and a joint supervision method with center loss, which is crucial for facial tasks. A newly proposed Convolutional Neural Network (CNN) model, MobileNet, which has both accuracy and speed, is deployed in both offline and in a real-time framework that enables fast and accurate real-time output. Evaluations towards two publicly available datasets, JAFFE and CK+, are carried out respectively. The JAFFE dataset reaches an accuracy of 95.24%, while an accuracy of 96.92% is achieved on the 6-class CK+ dataset, which contains only the last frames of image sequences. At last, the average run-time cost for the recognition of the real-time implementation is around 3.57ms/frame on a NVIDIA Quadro K4200 GPU.
Haiwei Dong 0001, Jihad Mohamad Jaam, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Editorial to Special Issue on Deep Learning for Intelligent Multimedia Analytics
Wei Zhang 0031, Ting Yao 0003, Shiai Zhu, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Deep Learning-Based Multimedia Analytics: A Review
abstract
The multimedia community has witnessed the rise of deep learning–based techniques in analyzing multimedia content more effectively. In the past decade, the convergence of deep-learning and multimedia analytics has boosted the performance of several traditional tasks, such as classification, detection, and regression, and has also fundamentally changed the landscape of several relatively new areas, such as semantic segmentation, captioning, and content generation. This article aims to review the development path of major tasks in multimedia analytics and take a look into future directions. We start by summarizing the fundamental deep techniques related to multimedia analytics, especially in the visual domain, and then review representative high-level tasks powered by recent advances. Moreover, the performance review of popular benchmarks gives a pathway to technology advancement and helps identify both milestone works and future directions.
Wei Zhang 0031, Ting Yao 0003, Shiai Zhu, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2018 AI + Multimedia Make Better Life?
abstract
No abstract available.
Wen-Huang Cheng, Jiaying Liu 0001, Mohan Kankanhalli, Abdulmotaleb El Saddik, Benoit Huet
ACM Multimedia4
2018 Multimedia analysis with collective intelligence
Meng Wang 0001, Qi Tian 0001, Abdulmotaleb El Saddik, Mathias Lux, Tie Yun
J. Vis. Commun. Image Represent.3
2018 Predicting muscle forces measurements from kinematics data using kinect in stroke rehabilitation
Mohamad Hoda, Yehya Hoda, Basim Hafidh, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2017 The Solar System as a 3D Metaphor to Visualize User Interactions in a Social Network
abstract
With the popularity of social networks, users can communicate with each other in a more convenient way. However, the increasing amount of data poses new challenges for the analysis of the social activities of the users. In this paper, we propose to visualize the heterogeneous information of user interactions in a social network in a three-dimensional way using the concept of solar systems. The target user represents the center of the solar system. We determine physical variables from interaction frequencies between the user and their friends to be visualized. This gives the user a better insight about their relationships on Facebook. To show the interactions between the user and their friends, we choose four variables related to the solar system: the size of each planet, its angular velocity, and the semi-major and the semi-minor axis of its elliptical orbit. Our system measures the interaction frequencies between a user and their friends based on a linear model. Its coefficients are derived from an online survey we performed. The experimental results indicate that the accuracy of our estimated interaction frequency is better than the accuracies for each individual interaction feature. The average accuracy improvement is 15.93%.
Wanjun Pei, Benjamin Guthier, Abdulmotaleb El Saddik
ISM3
2017 Shall IoT User Interfaces Start Recommending Multimedia Devices as Well?
abstract
Recommendation of media objects, such as audio and video clips, has been there for a while. Most user interfaces include a list of favourites or most popular media objects. On the other hand, recommendation of multimedia devices is limited to shopping websites. With the evolution of IoT, however, users these days are surrounded by many interconnected devices that can be used to accomplish the same task at any given time. For example, while at home, a user can choose to play a media le on a smartphone, tablet, laptop, TV, or home theatre. In this article we investigate the question of whether or not users are ready to accept automatic recommendation of physical things with a case study of media playback devices. We further investigate various factors that a ect user's choice of media playback device with a user study. The analysis shows that users like device recommendation in general. In addition, many users even prefer the playback to be automatically transferred directly to the most appropriate device, while other users just want a noti cation. We also found that user's gender, profession, age, and the duration of media le a ect the choice of playback device.
Mukesh Saini, Ali Danesh, Abdulmotaleb El Saddik
ISM3
2017 Language-independent data set annotation for machine learning-based sentiment analysis
abstract
Social media platforms provide large amounts of user-generated text which can be utilized for text-based Sentiment Analysis in order to obtain insights about opinions on many aspects of life. Current approaches that are based on supervised learning require manually annotated data sets that are time-consuming to create and are specific to a single language. In this work, we present an approach to generate ground truth sentiment values for a data set from Twitter. We use a sentiment emoji lexicon and distribute known polarities of hashtags over neighbors in a graph we build on them. This approach is language-independent. Native speakers of five different languages evaluate the accuracy of the sentiment values assigned by our method on a corpus of Tweets. Our experiments show that the quality of our automatically assigned sentiment values is sufficiently high to be used for training of machine learning-based sentiment analysis.
Benjamin Guthier, Kalun Ho, Abdulmotaleb El Saddik
SMC3
2017 Distributed robust adaptive finite-time voltage control for microgrids with uncertainty
abstract
Consensus based distributed robust adaptive finite-time secondary voltage control is designed for inverter-based islanded AC microgrids. The design combines decentralized local states information with the states of the neighboring distributed generators with directed communication topology. Robust control algorithms are used locally for each distributed generator to deal with uncertainty. Lyapunov and terminal sliding mode theory uses to guarantee that the proposed distributed control design can restore voltage to the reference value in finite-time. Analysis shows that the finite-time robust consensus can force the voltage of the distributed generators to reach the designed terminal sliding surface in finite-time and remain there. The proposed distributed secondary controller does not require a priori knowledge of the nonlinear dynamical model and uncertainty associated with microgrids.
Peter Xiaoping Liu, Abdulmotaleb El Saddik
SMC3
2017 Haptics based bilateral shared telemanipulation of aerial vehicle over open communication network
abstract
In this paper, we develop haptic based force reflecting interaction interface for bilateral telemanipulation of miniature aerial vehicle. The human-master interface combines shared control term with the reflected force fields mapped by using artificial force field and virtual impedance force field. The shared control for the human-master comprises velocity signals of the with the scaled position of the master haptic manipulator. The shared input interface for the slave-flying environment is developed by combining scaled position of the master manipulator with the velocity of the remote MAV. The data transmission between ground station and remote vehicle are carried out by open internet communication network. Evaluation results on a qudrotor MAV system are presented to demonstrate the effectiveness for real-time applications.
Peter Xiaoping Liu, Abdulmotaleb El Saddik
SMC3
2017 Consensus based distributed cooperative control for multiple miniature aerial vehicles with uncertainty
abstract
In this paper, we investigate distributed consensus problems for multiple miniature aerial vehicles (MAVs) with nonlinear dynamics and uncertainty. We develop distributed consensus protocol to solve regulation synchronization problem for leaderless MAVs with directed interaction topology. Adaptive control algorithms are used locally for each vehicle to deal with nonlinear dynamics and uncertainty associated with flying environment, such as, wind gust, payload mass, aerodynamic friction and other external disturbances. The resulting protocol for synchronization problem combines simple decentralized proportional-plus-derivative like term and robust adaptive control term with position signal based consensus protocol. Lyapunov method uses to show the asymptotic convergence of the consensus errors of the closed loop systems formulated by multiple MAVs. It is shown in our analysis that all MAVs reach an agreement and synchronize to a common value which is not a priori defined. The convergence of the asymptotic consensus error is shown by using Lyapunov method and sliding mode control theory. The proposed design is simple as it does not require exact knowledge of the dynamical model and uncertainty.
Peter Xiaoping Liu, Abdulmotaleb El Saddik
SMC3
2017 Mobile cloud-based physical activity advisory system using biofeedback sensors
Hawazin Badawi, Haiwei Dong 0001, Abdulmotaleb El Saddik
Future Gener. Comput. Syst.3
2017 City digital pulse: a cloud based heterogeneous data analysis platform
Zhongli Li, Shiai Zhu, Huiwen Hong, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2017 InCloud: a cloud-based middleware for vehicular infotainment systems
Mukesh Saini, Kazi Masudul Alam, Haolin Guo, Abdulhameed Alelaiwi, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2017 Development of an automatic 3D human head scanning-printing system
Longyu Zhang, Bote Han, Haiwei Dong 0001, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2017 DST: days spent together using soft sensory information on OSNs - a case study on Facebook
Fatimah Al-Zamzami, Mukesh Saini, Abdulmotaleb El Saddik
Soft Comput.3
2016 Promoting active participation of the learners in an authoring based learning movie system
abstract
Netflix, Hulu, etc are some of the most popular video content streaming services that are increasingly being accessed through many popular consumer devices such as Apple TV, XBox, Wii, etc. It has now become possible to conveniently interact with the video contents by using the input hardwares that these devices provide. We emulate the setups that many of these popular platforms provide in order to develop learning based video interaction games. The games leverage the user interaction feature with the video contents. In the award winning learning television series such as Mickey Mouse ClubHouse, WordWorld, Super Why etc., the protagonists present learning based questions by using various scenarios and viewers learn the answers passively as they wait. In order to foster active participation of the viewers, we author the movie with learning questions at particular timelines of the video and provide interaction options. In those specified timelines, the learners interact with the presented questions by using Wii's pointMe or XBox Kinect's gesture based interactions and input answers. The interactions assist the learners to engage with the video contents and make it possible to actively participate in the learning process. In order to examine the suitability of the proposed approach, we perform usability experiments in a technology-augmented learning space and report our findings.
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
AICCSA2
2016 Extreme Learning Machines for approximating nonlinear dimensionality reduction mappings: Application to Haptic handwritten signatures
abstract
The abundance of computing and mobile devices makes the problem of user identification and verification an essential requirement for many applications. Haptics devices include the sense of touch in the form of kinesthetic and tactile feedback which provide additional features within handwritten signatures. However, they generate high dimensional data and dimensionality reduction techniques become useful for data mining, machine learning and visualization. Nonlinear transformations have been used for this, but in present day scenarios (Big Data, the Internet of Things, massive data streams, etc.) the computation becomes more complex, time consuming or impractical. Moreover, the relationships between the features of the original and the target spaces are more difficult to uncover. Extreme Learning Machines (ELM) are used for approximating nonlinear manifold learning methods in two ways: as a functional representation for implicit methods, and as simpler surrogate models for explicit mapping techniques. In the context of Haptic handwritten signatures, five implicit and explicit nonlinear transformation methods are investigated. In all cases it was found that ELM approximations to the mappings obtained with the original methods exhibit very good behavior and can be used either as functional representations for the implicit methods or as simpler surrogate models for explicit techniques.
Julio J. Valdés, Fawaz A. Alsulaiman, Abdulmotaleb El Saddik
IJCNN3
2016 E-Tourism: Mobile Dynamic Trip Planner
abstract
In this paper, we propose an algorithm called the Balanced Orienteering Problem, to design trips for tourists. This algorithm, combined with a recommender system for tourism suggestions, create the infrastructure for the mobile application of the tourism guide we developed. A comparison study between some of the current algorithms and our proposed one were performed and the initial results illustrate that our proposed algorithm yields comparable results to existing system, yet it outperforms them in the average execution time.
Hamzah Alghamdi, Shiai Zhu, Abdulmotaleb El Saddik
ISM3
2016 Sentiment Analysis on Multi-View Social Data
Teng Niu, Shiai Zhu, Abdulmotaleb El Saddik
MMM (2)4
2016 MUDVA: A multi-sensory dataset for the vehicular CPS applications
abstract
Vehicular Cyber-Physical System (VCPS) is a new trend in the research of the intelligent transport systems (ITS). In VCPS, vehicles work as a hub of sensors to collect interior and exterior information about the vehicle. Vehicles can use ad-hoc networking or 3G/LTE communication technology to share useful information with their neighboring vehicles or with the infrastructures to accomplish user safety, comfort, and entertainment tasks. In order to facilitate efficient sensor-services fusion in the VCPS applications, we need real life vehicular sensory datasets. While there has been many datasets containing vehicle mobility traces, there is hardly any that contains sensory information to be shared on the network. In this paper, we present a scenario specific modular dataset architecture along with some multi-sensory dataset modules. One of the dataset modules provides time synchronized multi-vehicle data including multi-view video, multi-directional sound, GPS, accelerometer, gyroscope, and magnetic field sensors. Each of the three vehicles recorded front, back, left, and right videos while moving closely in the suburban areas to let explore vehicular cooperative applications. Another module presents necessary tools and datasets to identify vehicular events such as acceleration, deceleration, turn, and no-turn events. We also present development details of a safety application using the presented datasets along with a list of other possible applications.
Kazi Masudul Alam, Mohammed Bin Hariz, Seyed Vahid Hosseinioun, Mukesh Saini, Abdulmotaleb El Saddik
MMSP5
2016 Social media analytics and learning
Zhengjun Zha, Tao Mei 0001, Abdulmotaleb El Saddik
Neurocomputing3
2016 On the learning of image social relevance from heterogeneous social network
Shiai Zhu, Samah Al-Oufi, Abdulmotaleb El Saddik
Neurocomputing4
2016 Towards context-aware media recommendation based on social tagging
Mohammed F. Alhamid, Majdi Rawashdeh, M. Anwar Hossain 0001, Abdulhameed Alelaiwi, Abdulmotaleb El Saddik
J. Intell. Inf. Syst.5
2016 RecAm: a collaborative context-aware framework for multimedia recommendations in an ambient intelligence environment
Mohammed F. Alhamid, Majdi Rawashdeh, Haiwei Dong 0001, M. Anwar Hossain 0001, Abdulhameed Alelaiwi, Abdulmotaleb El Saddik
Multim. Syst.6
2016 Special issue on collaborative haptic audio-visual environments and systems
Xiaohu Guo, B. Prabhakaran 0001, Abdulmotaleb El Saddik
Multim. Syst.3
2016 Personality assessment using multiple online social networks
Shally Bhardwaj, Pradeep K. Atrey, Mukesh Saini, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2016 Tag-based personalized recommendation in social media services
Majdi Rawashdeh, Mohammed F. Alhamid, Jihad Mohamad Jaam, Awny Alnusair, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2016 See in 3D: state of the art of 3D display technologies
Haiwei Dong 0001, Abdulhameed Alelaiwi, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2016 Exploring Latent Preferences for Context-Aware Personalized Recommendation Systems
abstract
Context-aware recommendations offer the potential of exploiting social contents and utilize related tags and rating information to personalize the search for content considering a given context. Recommendation systems tackle the problem of trying to identify relevant resources from the vast number of choices available online. In this study, we propose a new recommendation model that personalizes recommendations and improves the user experience by analyzing the context when a user wishes to access multimedia content. We conducted empirical analysis on a dataset from last.fm to demonstrate the use of latent preferences for ranking items under a given context. Additionally, we use an optimization function to maximize the mean average precision measure of the resulted recommendation. Experimental results show a potential improvement to the quality of the recommendation in terms of accuracy when compared with state-of-the-art algorithms.
Mohammed F. Alhamid, Majdi Rawashdeh, Haiwei Dong 0001, M. Anwar Hossain 0001, Abdulmotaleb El Saddik
IEEE Trans. Hum. Mach. Syst.5
2016 From 3D Sensing to Printing: A Survey
abstract
Three-dimensional (3D) sensing and printing technologies have reshaped our world in recent years. In this article, a comprehensive overview of techniques related to the pipeline from 3D sensing to printing is provided. We compare the latest 3D sensors and 3D printers and introduce several sensing, postprocessing, and printing techniques available from both commercial deployments and published research. In addition, we demonstrate several devices, software, and experimental results of our related projects to further elaborate details of this process. A case study is conducted to further illustrate the possible tradeoffs during the process of this pipeline. Current progress, future research trends, and potential risks of 3D technologies are also discussed.
Longyu Zhang, Haiwei Dong 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Estimating two-dimensional blood flow velocities from videos
abstract
We estimate the velocity field of the blood flow in a human face from videos. Our approach first performs spatial preprocessing to improve the signal-to-noise ratio (SNR) and the computational efficiency. The discrete Fourier transform (DFT) and a temporal band-pass filter are then applied to extract the frequency corresponding to the subject's heart rate. We propose two techniques for reducing the noise from the resulting phase and amplitude maps. The 2D blood flow field is then estimated from the relative phase shift between the pixels. We evaluate our approach on real and synthetic face videos using two different metrics. Our method produces a velocity field with an angular error of 20° and an error in magnitude of 29% on the average.
Benjamin Guthier, Abdulmotaleb El Saddik
ICIP3
2015 Utilizing image social clues for automated image tagging
abstract
Social tags have been successfully utilized for image search and recommendation, yet the tags may be bias and noisy. Assisting users to annotate their images with tags that meet their preferences and efficiently describe the visual content is a fundamental objective in multimedia. In this work, we propose to leverage the image social information, such as tagging preferences of an image owner and social groups that an image has been shared with, by adopting the well-known neighbor voting approach for automated image tagging. In specific, we assign more contributions of neighborhood images which are socially closer to the target image in the voting procedure. Meanwhile, the social strength of reference images with respect to the target image is jointly considered. The experiments on a large scale image dataset for tag recommendation and image search show the advantages of considering image social clues.
Shiai Zhu, Samah Al-Oufi, Abdulmotaleb El Saddik
ICME3
2015 Design and Development of a Cloud Based Cyber-Physical Architecture for the Internet-of-Things
abstract
Internet-of-Things (IoT) is considered as the next big disruptive technology field which main goal is to achieve social good by enabling collaboration among physical things or sensors. We present a cloud based cyber-physical architecture to leverage the Sensing as-a-Service (SenAS) model, where every physical thing is complemented by a cloud based twin cyber process. In this model, things can communicate using direct physical connections or through the cyber layer using peer-to-peer inter process communications. The proposed model offers simultaneous communication channels among groups of things by uniquely tagging each group with a relationship ID. An intelligent service layer ensures custom privacy and access rights management for the sensor owners. We also present the implementation details of an IoT platform and demonstrate its practicality by developing case study applications for the Internet-of-Vehicles (IoV) and the connected smart home.
Kazi Masudul Alam, Alex Sopena, Abdulmotaleb El Saddik
ISM3
2015 Development of a Web-Based Haptic Authoring Tool for Multimedia Applications
abstract
In this paper, we introduce an MPEG-V based haptic authoring tool intended for simplifying the development process of haptics-enabled multimedia applications. The developed tool provides a web-based interface for users to create haptic environments by importing 3D models and adding haptic properties to them. The user can then export the resulting environment to a standard MPEG-V format. The latter can be imported to a haptic player that renders the described haptics-enabled 3D scene. The proposed tool can support many haptic devices, including Geomagic Devices, Force Dimension Devices, Novint Falcon Devices, and Moog FCS HapticMaster Devices. We conduct a proof of concept HTML5 haptic game project and user studies on haptic effects, development process and user interface, which shows our tool's effectiveness in simplifying the development process of haptics-enabled multimedia applications.
Haiwei Dong 0001, Hussein Al Osman, Abdulmotaleb El Saddik
ISM4
2015 Employing Sensors and Services Fusion to Detect and Assess Driving Events
abstract
With the remarkable increase in use of sensors in our daily lives, various methods have been devised to detect events in a driving environment using smart-phones as they provide two main advantages: they eliminate the need to have dedicated hardware in vehicles and they are widely accessible. Since rewarding safe driving is an important issue for insurance companies, some companies are implementing Usage-Based Insurance (UBI) as opposed to traditional History-Based plans. The collection of driving events, such as acceleration and turning, is a prerequisite requirement for the adoption of such plans. Mobile phone sensors are capable of detecting whether a car is accelerating or braking, while through service fusion we can detect other events like speeding or instances of severe weather. We propose a new and robust hybrid classification algorithm that detects acceleration-based events with an F1-score of 0.9304 and turn events with an F1-score of 0.9038. We further propose a method for measuring the driving performance index using the detected events.
Seyed Vahid Hosseinioun, Hussein Al Osman, Abdulmotaleb El Saddik
ISM3
2015 Boosting Prediction of Geo-location for Web Images Through Integrating Multiple Knowledge Sources
abstract
Estimating geographical information of a given photo is a challenging task due to the massive spread of candidate locations on the earth. With the help of freely available geo-tagged Web images, the problem can be addressed by propagating geo-coordinates (latitude and longitude) of geo-related training data, which is obtained using document retrieval techniques. The state-of-the-art approach adopts language modeling technique to estimate the probability distribution of image associated tags in a local region. Under this framework, we propose to differentiate the tags based on the knowledge explored from multiple sources. Finally, a set of geo-informative tags are identified and further emphasized during the model learning and geo-location prediction. In addition, accurate geo-coordinates are estimated by incorporating the image visual information. Experiments on a large-scale geo-tagged Flickr image dataset demonstrate the effectiveness of proposed method at different levels of evaluation granularity.
Hao Kuang, Shiai Zhu, Abdulmotaleb El Saddik
ICMR3
2015 An Elicitation Study on Gesture Attitudes and Preferences Towards an Interactive Hand-Gesture Vocabulary
abstract
With the introduction of new depth sensing technologies, interactive hand-gesture devices are rapidly emerging. However, the hand-gestures used in these devices do not follow a common vocabulary, making certain control command device-specific. In this paper we present an initial effort to create a standardized interactive hand-gesture vocabulary for the next generation of television applications. We conduct a user-elicitation study using a survey in order to define a common vocabulary for specific control commands, such as Volume up/down, Menu open/close, etc. This survey is entirely user-oriented and thus it has two phases. In the first phase, we ask open questions about specific commands. In the second phase, we use the answers suggested from the first phase to create a multiple choice questionnaire. Based on the results from the survey, we study the gesture attitudes and preferences between gender groups, and between age groups with a quantitative and qualitative statistical analysis. Finally, the hand-gesture vocabulary is derived after applying an agreement analysis on the user-elicited gestures. The proposed methodology for gesture set design is comparable with existing methodologies and yields higher agreement levels than relevant user-elicited studies in the field.
Haiwei Dong 0001, Nadia Figueroa, Abdulmotaleb El Saddik
ACM Multimedia3
2015 A Proxemic Multimedia Interaction over the Internet of Things
Ali Danesh, Mukesh Saini, Abdulmotaleb El Saddik
MMM (2)3
2015 A stochastic approach to group recommendations in social media systems
Heung-Nam Kim, Abdulmotaleb El Saddik
Inf. Syst.2
2015 Towards context-sensitive collaborative media recommender system
Mohammed F. Alhamid, Majdi Rawashdeh, Hussein Al Osman, M. Shamim Hossain, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2015 Design and development of a user centric affective haptic jacket
Faisal Arafsha, Kazi Masudul Alam, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2015 Guest editorial: advances in multimedia for health
M. Shamim Hossain, Stefan Göbel 0001, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2015 Development of a haptic video chat system
Longyu Zhang, Jamal Saboune, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2015 A Combined Approach Toward Consistent Reconstructions of Indoor Spaces Based on 6D RGB-D Odometry and KinectFusion
abstract
We propose a 6D RGB-D odometry approach that finds the relative camera pose between consecutive RGB-D frames by keypoint extraction and feature matching both on the RGB and depth image planes. Furthermore, we feed the estimated pose to the highly accurate KinectFusion algorithm, which uses a fast ICP (Iterative Closest Point) to fine-tune the frame-to-frame relative pose and fuse the depth data into a global implicit surface. We evaluate our method on a publicly available RGB-D SLAM benchmark dataset by Sturm et al. The experimental results show that our proposed reconstruction method solely based on visual odometry and KinectFusion outperforms the state-of-the-art RGB-D SLAM system accuracy. Moreover, our algorithm outputs a ready-to-use polygon mesh (highly suitable for creating 3D virtual worlds) without any postprocessing steps.
Nadia Figueroa, Haiwei Dong 0001, Abdulmotaleb El Saddik
ACM Trans. Intell. Syst. Technol.3
2015 Modeling and Stability Analysis of Automatic Generation Control Over Cognitive Radio Networks in Smart Grids
abstract
Due to its great potential to improve the overall performance of data transmission with its dynamic and adaptive spectrum allocation capability in comparison with many other networking technologies, cognitive radio (CR) networking technology has been increasingly employed in networking and communication infrastructures for smart grids. However, a secondary user (SU) of a CR network has to be squeezed out from a channel when a primary user reclaims the channel, which may occur in a randomized fashion. The random interruption of SU traffic may cause packet losses and delays for SU data, and it will in turn affect the stability of the monitoring and control of smart grids. In this paper, we address this problem and investigate the modeling and stability analysis of the automatic generation control (AGC) of a smart grid for which CR networks are used as the infrastructure for the aggregation and communication of both system-wide information and local measurement data. For this purpose, a randomly switched power system model is proposed for the AGC of the smart grid. By modeling the CR network as an On–Off switch with sojourn times, the stability of the AGC of the smart grid is analyzed. In particular, we investigate the smart grid with two main types of CR networks: 1) the sojourn times are arbitrary but bounded and 2) the sojourn times follow an independent and identical distribution process. The sufficient conditions are obtained for the stability of the AGC of the smart grid with these two CR networks, respectively. Simulation results show the effects of the CR networks on the dynamic performance of the AGC of the smart grid and illustrate the usefulness of the developed sufficient conditions in the design of CR networks in order to ensure the stability of the AGC of the smart grid.
Shichao Liu 0001, Peter Xiaoping Liu, Abdulmotaleb El Saddik
IEEE Trans. Syst. Man Cybern. Syst.3
2014 Towards consistent reconstructions of indoor spaces based on 6D RGB-D odometry and KinectFusion
abstract
We focus on generating consistent reconstructions of indoor spaces from a freely moving handheld RGB-D sensor, with the aim of creating virtual models that can be used for measuring and remodeling. We propose a novel 6D RGBD odometry approach that finds the relative camera pose between consecutive RGB-D frames by keypoint extraction and feature matching both on the RGB and depth image planes. Furthermore, we feed the estimated pose to the highly accurate KinectFusion algorithm, which uses a fast ICP (Iterative-Closest-Point) to fine-tune the frame-to-frame relative pose and fuse the Depth data into a global implicit surface. We evaluate our method on a publicly available RGB-D SLAM benchmark dataset by Sturm et al. The experimental results show that our proposed reconstruction method solely based on visual odometry and KinectFusion outperforms the state-of-the-art RGB-D SLAM system accuracy. Our algorithm outputs a ready-to-use polygon mesh (highly suitable for creating 3D virtual worlds) without any post-processing steps.
Haiwei Dong 0001, Nadia Figueroa, Abdulmotaleb El Saddik
IROS3
2014 A Low-cost Serious Game Therapy Environment with Inverse Kinematic Feedback for Children Having Physical Disability
abstract
Recently, the use of non-invasive ways to track joint motion in the human body has drawn significant attention in the therapy domain. One of the reasons for this popularity is due to availability of economically priced 3D motion sensors. In this paper, we present a web-based 3D interactive serious game interface that uses non-invasive methods to recognize the movements of the body. Motion data of a subject is collected through two motion sensors, a Kinect and a LEAP, in a non-invasive manner. Joint motion along with inverse kinematic joint information is displayed in 3D environment in a live manner or recorded for offline replaying and data analysis. To facilitate the complex therapy authoring process, the system incorporates an authoring tool that allows a therapist to design a complex therapy in terms of primitive therapies and assign it to a patient. The subject as well as other members of the community of interest such as therapists, parents and caregivers can view the results at any time and can follow up with patient's progress.
Mohamed Abdur Rahman 0001, Delwar Hossain, Ahmad M. Qamar, Faizan Ur Rehman, Asad H. Toonsi, Mohamed A. Ahmed, Abdulmotaleb El Saddik, Saleh M. Basalamah
ICMR7
2014 EMASC14: 1st International Workshop on Emerging Multimedia Applications and Services for Smart Cities
abstract
Smart city is the vision of future city - with increasingly instrumented, inter-connected and intelligent urban systems - to improve the quality of life in many aspects including public safety, healthcare, transportation, or energy. With the ever-increasing presence of multimodal sensors in the smart city infrastructure, multimedia plays an indispensable role. The proliferation of multimedia, sensors, pervasive devices, and infrastructures for realizing smart city has brought many challenges that are the core focus of EMASC workshop.
M. Anwar Hossain 0001, Abdulmotaleb El Saddik
ACM Multimedia2
2014 A Real-Time Smart Assistant for Video Surveillance Through Handheld Devices
abstract
In a remote surveillance system, a high resolution surveillance camera streams its video to a user's handheld device. Such devices are unable to make use of the high resolution video due to their limited display size and bandwidth. In this paper, we propose a method to assist the mobile operator of the surveillance camera in focusing on sensitive regions of the video. Our system automatically identifies relevant regions. We introduce a pan and zoom strategy to ensure that the operator is able to see fine details in these areas while maintaining contextual knowledge. Regions of interest are identified using foreground detection as well as face and body detection. The efficacy of the proposed method is demonstrated through a user study. Our proposed method was reported to be more useful than two comparable approaches for getting an understanding of the activities in a surveillance scene while maintaining context.
Hao Kuang, Benjamin Guthier, Mukesh Saini, Dwarikanath Mahapatra, Abdulmotaleb El Saddik
ACM Multimedia5
2014 New stability and tracking criteria for a class of bilateral teleoperation systems
Peter Xiaoping Liu, Abdulmotaleb El Saddik
Inf. Sci.3
2014 Utility based decision support engine for camera view selection in multimedia surveillance systems
Dewan Tanvir Ahmed, M. Anwar Hossain 0001, Shervin Shirmohammadi, Abdullah Sharaf Alghamdi, Pradeep K. Atrey, Abdulmotaleb El Saddik
Multim. Tools Appl.6
2014 Collective control over sensitive video data using secret sharing
Pradeep K. Atrey, Saeed Alharthi, M. Anwar Hossain 0001, Abdullah Sharaf Alghamdi, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2014 Target-shooting exergame with a hand gesture control
Nasser H. Dardas, Juan M. Silva, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2014 Slingshot 3D: A synchronous haptic-audio-video game
Mohamad A. Eid, Ahmad El Issawi, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2014 SmartPads: a plug-N-play configurable tangible user interface
Basim Hafidh, Hussein Al Osman, Ali Karime, Jihad Mohamad Jaam, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2014 U-biofeedback: a multimedia-based reference model for ubiquitous biofeedback systems
Hussein Al Osman, Mohamad A. Eid, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2014 Context-aware multimedia services modeling: an e-Health perspective
Mohamed Abdur Rahman 0001, M. Shamim Hossain, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2014 A context-aware multimedia framework toward personal social network services
Mohamed Abdur Rahman 0001, Heung-Nam Kim, Abdulmotaleb El Saddik, Wail Gueaieb
Multim. Tools Appl.3
2014 A Quality of Experience Model for Haptic Virtual Environments
abstract
Haptic-based Virtual Reality (VR) applications have many merits. What is still obscure, from the designer's perspective of these applications, is the experience the users will undergo when they use the VR system. Quality of Experience (QoE) is an evaluation metric from the user's perspective that unfortunately has received limited attention from the research community. Assessing the QoE of VR applications reflects the amount of overall satisfaction and benefits gained from the application in addition to laying the foundation for ideal user-centric design in the future. In this article, we propose a taxonomy for the evaluation of QoE for multimedia applications and in particular VR applications. We model this taxonomy using a Fuzzy logic Inference System (FIS) to quantitatively measure the QoE of haptic virtual environments. We build and test our FIS by conducting a users' study analysis to evaluate the QoE of a haptic game application. Our results demonstrate that the proposed FIS model reflects the user's estimation of the application's quality significantly with low error and hence is suited for QoE evaluation.
Abdelwahab Hamam, Abdulmotaleb El Saddik, Jihad Mohamad Jaam
ACM Trans. Multim. Comput. Commun. Appl.2
2013 An edutainment system for assisting qatari children with moderate intellectual and learning disability through exerting physical activities
abstract
Children with Moderate Intellectual Disability (MID) and those with Moderate Learning Disability (MLD) are growing up with extensive exposure to computer technology. Computers and computer-related devices have the potential to help these children in education, career development, and independent living. However, most of the software, games, and web sites that MID and MLD children interact with are designed without consideration of their special needs, making the applications less effective or completely inaccessible. This paper introduces an edutainment system specifically designed to help these children have an enhanced and enjoyable learning process, while addresses the need for integrating physical activity into their daily lives. The proposed system consists of a padded floor mat that includes sixteen square tiles supported by sensors, which are used to interact with a number of software games specifically designed to suit the mental needs of children with ID. The system aims for enhancing both MID and MLD children learning capabilities, understanding, communications, thinking, memorization, and obesity problems.
Moutaz Saleh Mustafa Saleh, Jihad Mohamad Jaam, Ali Karime, Abdulmotaleb El Saddik
EDUCON4
2013 Plenary talks: From whiskers to fingertips - A biomimetic approach to active touch sensing
abstract
How do animals understand the physical world they live in? One answer, due to Gibson, is that their sensory systems are tuned to pick up relevant affordances for behavior, but how is it that the brain and the sensory apparatus become suitably adapted to perform this feat? To cast light on this question we have been investigating active touch sensing in mammals, including humans, and developing biomimetic robots that can help us understand these biological systems whilst also developing useful haptic technologies. An important focus has been on the vibrissal (whisker) system of rodents, and its emergence through evolution and development, which we have investigated through a combination of (i) ethological studies of behaving animals, (ii) computational neuroscience models of the neural circuits involved in vibrissal processing, and (iii) biomimetic robots embodying many of the characteristics of whiskered animals in their design and control. This work has resulted in a series of whiskered robots, the most recent of which, Shrewbot, is able to construct tactile maps of its environment and recognize and track moving objects. We are also studying humanoid touch, focusing on the development of Bayesian strategies for active tactile sensing with robot hands. Here our results have provided the first demonstration of hyperacuity in robot touch whilst also indicating that tactile perception is improved in unstructured environments by appropriate active control. The active sensing framework can also be applied to the development of haptic interfaces for human users that can augment our existing sensory capability. For instance, we are developing a head-mounted “remote touch” system that links distance sensors (ultrasound arrays) with vibrotactile displays. Here an interesting question is how the signals that are delivered through the displays should be modulated to take into account the intentional head and body movements of the user and in order to provide a meaningful and intuitive experience. The talk will present converging lines of evidence, from these different research strands, for the importance of active control in haptics. Our results will also be used to illustrate how experimental, computational, and robotic approaches can operate together to advance our understanding of sensorimotor cognition in behaving systems.
Tony J. Prescott, Masahiko Inami, Abdulmotaleb El Saddik, Wayne J. Book
World Haptics3
2013 A framework toward detecting and visualizing kinematic data for children with Hemiplegia
abstract
In this paper we propose a multimedia environment that can capture kinematic data from live gestures of a child having Hemiplegia disability and generate live analytical results to be useful for decision making system of a therapist. The kinematic data is obtained from some clinically suggested therapy modules that are used to monitor quality of improvement of a disabled child, which includes exercises involving the affected joints and muscles. The proposed environment uses the 3D depth sensing Microsoft Kinect device to detect, recognize and track the movement of different key joints of the body and deduce kinematic data from these movements. The method is non-invasive as the child does not need to wear any external devices in the body. The proposed environment incorporates Second Life serious game environment where the live therapeutic movement of child, therapist and one's community of interest is synchronized between physical and virtual world. Finally, we share our preliminary test data, which is validated by the therapists from three different disability hospitals that treat children with Hemiplegia.
Mohamed Abdur Rahman 0001, Saleh M. Basalamah, Asad H. Toonsi, Abdulmotaleb El Saddik
Healthcom4
2013 "Anti-fatigue" control for over-actuated bionic arm with muscle force constraints
abstract
In this paper, we propose an “anti-fatigue” control method for bionic actuated systems. Specifically, the proposed method is illustrated on an over-actuated bionic arm. Our control method consists of two steps. In the first step, a set of linear equations is derived by connecting the acceleration description in both joint and muscle space. The pseudo inverse solution to these equations provides an initial optimal muscle force distribution. As a second step, we derive a gradient direction for muscle force redistribution. This allows the muscles to satisfy force constraints and generate an even distribution of forces throughout all the muscles (i.e. towards "anti-fatigue"). The overall proposed method is tested for a bending-stretching movement. We used two models (bionic arm with 6 and 10 muscles) to verify the method. The force distribution analysis verifies the “anti-fatigue” property of the computed muscle force. The efficiency comparison shows that the computational time does not increase significantly with the increase of muscle number. The tracking error statistics of the two models show the validity of the method.
Haiwei Dong 0001, Setareh Yazdkhasti, Nadia Figueroa, Abdulmotaleb El Saddik
IROS4
2013 Towards Context-Aware Recommendations of Multimedia in an Ambient Intelligence Environment
abstract
Given today's mobile and smart devices, and the ability to access different multimedia contents in real-time, it is difficult for users to find the right multimedia content from such a large number of choices. Users also consume diverse multimedia based on many contexts, with different personal preferences and settings. For these reasons, there is a need to reinforce recommendation process with context-adaptive information that can be used to select the right multimedia content and deliver the recommendations in preferred mechanisms. This paper proposes a framework to establish a bridge between the multimedia content, the user and joint preferences, contextual information including the physiological parameters, and the Ambient Intelligent (AmI) environment, using multi-modal recommendation interfaces.
Mohammed F. Alhamid, Majdi Rawashdeh, Abdulmotaleb El Saddik
ISM3
2013 Evaluating Player Experience in Cycling Exergames
abstract
Obesity has become a worldwide problem which most countries are trying to fight. It affects many people, irrespective of age, race, gender, or religion, anyone can suffer from obesity that leads to serious problems for individuals and for society as a whole. In this study we have selected two groups of people: the basic people who rarely exercise on a weekly basis, and the average people who exercise regularly every week. We have explored the attitude of the two groups towards mixing exercises with games in order to motivate the people with basic activity levels to exercise more frequently. We have used a qualitative standard online questionnaire from AttrakDiff and we have done a quantitative study of some important factors during exercises. The results of the qualitative and quantitative studies were very encouraging, as they reveal that mixing games with exercises can transform boring exercises into entertaining ones. It can also motivate players to continue and repeat the exercises. The ANOVA test has been applied and it shows that combining games with the bike has a significant effect on the speed and the average rotation per minute of the participants.
Mohamad Hoda, Rana Alattas, Abdulmotaleb El Saddik
ISM3
2013 Tailoring recommendations to groups of users: a graph walk-based approach
abstract
With the rapid popularity of smart devices, users are easily and conveniently accessing rich multimedia content. Consequentially, the increasing need for recommender services, from both individual users and groups of users, has arisen. In this paper, we present a graph-based approach to a recommender system that can make recommendations most notably to groups of users. From rating information, we first model a signed graph that contains both positive and negative links between users and items. On this graph we examine two distinct random walks to separately quantify the degree to which a group of users would like or dislike items. We then employ a differential ranking approach for tailoring recommendations to the group. Our empirical evaluations on the MovieLens dataset demonstrate that the proposed group recommendation method performs better than existing alternatives. We also demonstrate the feasibility of Folkommender for smartphones.
Heung-Nam Kim, Majdi Rawashdeh, Abdulmotaleb El Saddik
IUI3
2013 Knowing Who You Are and Who You Know: Harnessing Social Networks to Identify People via Mobile Devices
Mark Bloess, Heung-Nam Kim, Majdi Rawashdeh, Abdulmotaleb El Saddik
MMM (1)4
2013 Social Media Annotation and Tagging Based on Folksonomy Link Prediction in a Tripartite Graph
Majdi Rawashdeh, Heung-Nam Kim, Abdulmotaleb El Saddik
MMM (1)3
2013 Muscle Force Control of a Kinematically Redundant Bionic Arm with Real-Time Parameter Update
abstract
Redundant muscle-driven arms have numerous advantages, such as increased robustness, ability for load distribution, impedance change etc. However, controlling such a muscle-driven arm is a difficult task. This is mainly due to its redundancy, specially when the muscle force is required to follow certain output constraints and fulfill optimization objectives. In this paper, a new method for controlling such muscle-like systems is proposed. By considering both joint and muscle acceleration contributions, a set of linear equations was constructed. Driving muscle activation is thus framed as the only unknown vector. To solve this linear equation set, a pseudo-inverse solution was used. The null space within this solution represents the internal force, which was used to evenly distribute the muscle forces, which is considered as "anti-fatigue" way. Moreover, to make the proposed method adaptive to modeling errors, an estimated system model was added to represent the real model. By updating the parameters of the estimated model based on prediction error, the estimated model approaches the real model gradually in real time. The overall method was tested for the case of a bending-stretching movement. The presented results verify the validity of the method, and illustrate its useful features and advantages.
Haiwei Dong 0001, Nadia Figueroa, Abdulmotaleb El Saddik
SMC3
2013 From Sense to Print: Towards Automatic 3D Printing from 3D Sensing Devices
abstract
In this paper, we introduce the From Sense to Print system. It is a system where a 3D sensing device connected to the cloud is used to reconstruct an object or a human and generate 3D CAD models which are sent automatically to a 3D printer. In other words, we generate ready-to-print 3D models of objects without manual intervention in the processing pipeline. Our proposed system is validated with an experimental prototype using the Kinect sensor as the 3D sensing device, the KinectFusion algorithm as our reconstruction algorithm and a fused deposition modeling (FDM) 3D printer. In order for the pipeline to be automatic, we propose a semantic segmentation algorithm applied to the 3D reconstructed object, based on the tracked camera poses obtained from the reconstruction phase. The segmentation algorithm works with both inanimate objects lying on a table/floor or with humans. Furthermore, we automatically scale the model to fit in the maximum volume of the 3D printer at hand. Finally, we present initial results from our experimental prototype and discuss the current limitations.
Nadia Figueroa, Haiwei Dong 0001, Abdulmotaleb El Saddik
SMC3
2013 Folksonomy link prediction based on a tripartite graph for tag recommendation
Majdi Rawashdeh, Heung-Nam Kim, Jihad Mohamad Jaam, Abdulmotaleb El Saddik
J. Intell. Inf. Syst.4
2013 Folkommender: a group recommender system based on a graph-based ranking algorithm
Heung-Nam Kim, Mark Bloess, Abdulmotaleb El Saddik
Multim. Syst.3
2013 Guest editorial: selected papers from ICIMCS 2011
Chong-Wah Ngo, Changsheng Xu, Xiao Wu 0001, Abdulmotaleb El Saddik
Multim. Syst.4
2013 Mobile PointMe-based spatial haptic interaction with annotated media for learning purposes
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
Multim. Syst.2
2013 Exertion interfaces for computer videogames using smartphones as input controllers
Juan M. Silva, Abdulmotaleb El Saddik
Multim. Syst.2
2013 Effect of kinesthetic and tactile haptic feedback on the quality of experience of edutainment applications
Abdelwahab Hamam, Mohamad A. Eid, Abdulmotaleb El Saddik
Multim. Tools Appl.3
2013 Adaptive interaction support in ambient-aware environments based on quality of context information
M. Anwar Hossain 0001, Ali A. Nazari Shirehjini, Abdullah Sharaf Alghamdi, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2013 Mobile haptic e-book system to support 3D immersive reading in ubiquitous environments
abstract
In order to leverage the use of various modalities such as audio-visual materials in instilling effective learning behavior we present an intuitive approach of annotation based hapto-audio-visual interaction with the traditional digital learning materials such as e-books. By integrating the home entertainment system in the user's reading experience combined with haptic interfaces we want to examine whether such augmentation of modalities influence the user's learning patterns. The proposed Haptic E--Book (HE-Book) system leverages the haptic jacket, haptic arm band as well as haptic sofa interfaces to receive haptic emotive signals wirelessly in the form of patterned vibrations of the actuators and expresses the learning material by incorporating image, video, 3D environment based augmented display in order to pave ways for intimate reading experience in the popular mobile e-book platform.
Kazi Masudul Alam, Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2013 Identity verification based on handwritten signatures with haptic information using genetic programming
abstract
In this article, haptic-based handwritten signature verification using Genetic Programming (GP) classification is presented. A comparison of GP-based classification with classical classifiers including support vector machine, k -nearest neighbors, naïve Bayes, and random forest is conducted. In addition, the use of GP in discovering small knowledge-preserving subsets of features in high-dimensional datasets of haptic-based signatures is investigated and several approaches are explored. Subsets of features extracted from GP-generated models (analytic functions) are also exploited to determine the importance and relevance of different haptic data types (e.g., force, position, torque, and orientation) in user identity verification. The results revealed that GP classifiers compare favorably with the classical methods and use a much fewer number of attributes (with simple function sets).
Fawaz A. Alsulaiman, Nizar Sakr, Julio J. Valdés, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2013 Human perception of haptic-to-video and haptic-to-audio skew in multimedia applications
abstract
The purpose of this research is to assess the sensitivity of humans to perceive asynchrony among media signals coming from a computer application. Particularly we examine haptic-to-video and haptic-to-audio skew. For this purpose we have designed an experimental setup, where users are exposed to a basic multimedia presentation resembling a ping-pong game. For every collision between a ball and a racket, the user is able to perceive auditory, visual, and haptic cues about the collision event. We artificially introduce negative and positive delay to the auditory and visual cues with respect to the haptic stream. We subjectively evaluate the perception of inter-stream asynchrony perceived by the users using two types of haptic devices. The statistical results of our evaluation show perception rates of around 100 ms regardless of modality and type of device.
Juan M. Silva, Mauricio Orozco Trujillo, Jongeun Cha, Abdulmotaleb El Saddik, Emil M. Petriu
ACM Trans. Multim. Comput. Commun. Appl.4
2013 Exploring social tagging for personalized community recommendations
Heung-Nam Kim, Abdulmotaleb El Saddik
User Model. User Adapt. Interact.2
2012 Identity verification based on haptic handwritten signatures: Genetic programming with unbalanced data
abstract
In this paper, haptic-based handwritten signature verification using Genetic Programming (GP) classification is presented. The relevance of different haptic data types (e.g., force, position, torque, and orientation) in user identity verification is investigated. In particular, several fitness functions are used and their comparative performance is investigated. They take into account the unbalance dataset problem (large disparities within the class distribution), which is present in identity verification scenarios. GP classifiers using such fitness functions compare favorably with classical methods. In addition, they lead to simple equations using a much smaller number of attributes. It was found that collectively, haptic features were approximately as equally important as visual features from the point of view of their contribution to the identity verification process.
Fawaz A. Alsulaiman, Julio J. Valdés, Abdulmotaleb El Saddik
CISDA3
2012 Admux Communication Protrocol for Real-Time Multimodal Intreaction
abstract
In our previous work [1], we proposed an adaptive application layer communication framework, named Admux, for multimedia applications incorporating hap tic, video, auditory, and graphics information for non-dedicated networks such as the Internet. In this paper, the contribution is two-fold: first, we present a thorough description of Admux communication protocol and content access/communication management. Second, the evaluation of Admux, using an interactive multimodal 3D Office Slingshot game, is described. The 3D Office Slingshot game involves the communication of synchronous hap tic-audio-video media â" with both tactile and kinesthetic hap tic feedback. The performance evaluation shows that Admux is capable of delivering synchronous hap tic-video data while adapting to network conditions by allocating proportional resources to various media channels. The usability testing with 20 subjects has shown that players have expressed positive feedback about the game.
Mohamad A. Eid, Abdulmotaleb El Saddik
DS-RT2
2012 PhacePhinder: harnessing social networks to build social face databases for mobile devices
abstract
This demo presents a client-server application which collects images and personal information from social networks to build face recognition databases. The client runs on a mobile phone allowing any face captured by a mobile phone's camera to be identified. Using the personal information collected we give back the most meaningful social connection between the user and the recognized individual.
Mark Bloess, Heung-Nam Kim, Abdulmotaleb El Saddik
ACM Multimedia3
2012 ACM international workshop on cloud-based multimedia applications and services for e-health(CBMAS-EH 2012)
abstract
Cloud-Based Multimedia services and technologies are emerging as an innovative means of accessing and delivering e-health resources and services in the years to come. The research in multimedia cloud for E-health is still in its infancy, and several technical issues remain open. Prior to its general use and adoption, careful consideration and evaluation are required. This article provides a summary and overview of the First International ACM workshop on Cloud-Based Multimedia Applications and Services for E-Health.
M. Shamim Hossain, Abdulmotaleb El Saddik
ACM Multimedia2
2012 A group trust metric for identifying people of trust in online social networks
Samah Al-Oufi, Heung-Nam Kim, Abdulmotaleb El Saddik
Expert Syst. Appl.3
2012 Leveraging personal photos to inferring friendships in social network services
Heung-Nam Kim, Abdulmotaleb El Saddik, Jin-Guk Jung
Expert Syst. Appl.2
2012 Folksonomy-based personalized search and ranking in social media services
Heung-Nam Kim, Majdi Rawashdeh, Abdullah Sharaf Alghamdi, Abdulmotaleb El Saddik
Inf. Syst.4
2012 Tableaux-based optimization of schema mappings for data integration
Md. Anisur Rahman, Mehedi Masud, Iluju Kiringa, Abdulmotaleb El Saddik
J. Intell. Inf. Syst.4
2012 Chaos-cryptography based privacy preservation technique for video surveillance
Sk. Md. Mizanur Rahman, M. Anwar Hossain 0001, Hussein T. Mouftah, Abdulmotaleb El Saddik, Eiji Okamoto
Multim. Syst.4
2012 Determining trust in media-rich websites using semantic similarity
Pradeep K. Atrey, Hicham Ibrahim, M. Anwar Hossain 0001, Sheela Ramanna, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2012 RFID-based interactive multimedia system for the children
Ali Karime, M. Anwar Hossain 0001, Abu Saleh Md. Mahfujur Rahman, Wail Gueaieb, Jihad Mohamad Jaam, Abdulmotaleb El Saddik
Multim. Tools Appl.6
2012 Social media filtering based on collaborative tagging in semantic space
Heung-Nam Kim, Andrew Roczniak, Pierre Lévy, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2012 Guest EditorialMultimedia Services and Technologies for E-Health (MUST-EH)
abstract
The 11 papers in this special section focus on multimedia services and technologies for E-Health (MUST-EH).
M. Shamim Hossain, Stefan Göbel 0001, Abdulmotaleb El Saddik
IEEE Trans. Inf. Technol. Biomed.3
2011 HE-book: A prototype haptic interface for immersive e-book reading experience
abstract
This paper presents an intuitive approach of annotation based haptic interaction with traditional digital reading materials such as eBooks. The research targets to bring a multi-sensory interface consisting of perceptual, cognitive and vibrotactile interactions with the digital reading contents. It leverages a previously developed haptic jacket to receive haptic emotive signals wirelessly in the form of patterned vibrations of the actuators in order to pave ways for intimate reading experience for the readers in various eBook platforms.
Kazi Masudul Alam, Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
World Haptics3
2011 A context-aware e-health framework for students with moderate intellectual and learning disabilities
abstract
In this paper we address the challenges of adding context-awareness to e-Health systems for students with moderate learning and intellectual disability. Our proposed framework provides a personalized e-Health environment containing context aware learning media, services and user interface to deal with individuals with disability. The framework adapts accessibility as well as user interfaces dynamically based on user disability. We present the detailed design and implementation of the framework.
Rajwa Alharthi, Rania Albalawi, Mahmud Abdo, Abdulmotaleb El Saddik
ICME4
2011 Fusion of face networks through the surveillance of public spaces to address sociological security recommendations
abstract
Researchers around the world are trying to address the ever increasing security requirements by bringing new approaches to surveillance specifically in public places like school, rail way, subway station, air port etc. To establish and sustain security in public spaces, surveillance plays a key role in technology-dependent governance common to many countries in the world. Traditionally, through the routine surveillance, an automated security system gains knowledge about people and their activities in a certain space. In this paper we are proposing a fusion algorithm to aggregate surveillance parameters from more than one such spaces. Inspired by existing works on social network analysis based on human photos, we propose a new face network structure model. These face network structures are later fused to obtain sociological parameters of a person of interest and gather recommendations about the circle of associates of that individual. We believe these type of recommendations are helpful in comprehensive investigation purposes.
S. K. Alamgir Hossain, Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ICME3
2011 Learn-pads: A mathematical exergaming system for children's physical and mental well-being
abstract
Child obesity is one of the major challenges facing modern societies, especially in developed countries. Exergaming tools are considered as effective means to reduce obesity among kids because they require the children to exert physical strength while playing the games. However, most of the existing exergaming tools focus more on the physical well-being of its users and almost neglect the mental aspect. In this paper, we present an exergaming system that combines both aspects by promoting not only entertainment, but also learning through physical activity. The system consists of a set of footpads that allow the user to interact with video games enriched with multimedia and aimed at enhancing the math knowledge of children. Our study shows that the system have created an atmosphere of fun among the children and engaged them in learning.
Ali Karime, Hussein Al Osman, Wail Gueaieb, Jihad Mohamad Jaam, Abdulmotaleb El Saddik
ICME5
2011 E-Glove: An electronic glove with vibro-tactile feedback for wrist rehabilitation of post-stroke patients
abstract
Arm paresis is a very common disability among post-stroke survivors. It is characterized by the inability of a person to perform some specific movements in the arm. A Long term Rehabilitation process plays a key role in the recovery of this kind of disabilities, but such treatment might not be easily accessible to people living away from the cities where most of the rehabilitation centers are located. In this paper, we present our interactive rehabilitation system called "E-Glove" that is aimed to help patients with wrist impairments to perform some daily exercises in a joyful and interactive manner. A 2D golf game that could be played with the glove was developed for this purpose.
Ali Karime, Hussein Al Osman, Wail Gueaieb, Abdulmotaleb El Saddik
ICME4
2011 Mobile pointme based pervasive gaming interaction with learning objects annotated physical atlas
abstract
The prevalent visions of ambient intelligence leverage natural interaction between user and available services in a learning space. In this pursuit, we propose a framework to facilitate handheld device based PointMe interaction with annotated media content, where the user points his/her handheld device to the annotated physical atlas for interacting with a world map. The proposed system performs annotations by specifying spatial location of the atlas and mapping related learning information to them. Each annotated data is encoded in customized Learning Object Metadata (LOM) format and they provide access points for available information about the specific countries in the map. This real world interaction technique with the physical environment and seamless virtual learning information acquisition make the system transparent from the young learners and help them to become engaged in their learning activities.
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ICME2
2011 HKiss: Real world based haptic interaction with virtual 3D avatars
abstract
Many researchers around the world are aiming to leverage the sense of touch in the communication medium between multiuser 3D virtual world and real environment. In this paper we propose a system that brings 3D avatar centric virtual interpersonal communication events as a form of haptic stimulation to the real world users. In order to render the haptic stimulations, we considered a neck piece (tactile haptic device) that the real users can wear in a scarf-like suit. Further, we enhanced the Linden Lab's multiuser online 3D virtual world, Second Life in order to facilitate the haptic communications. In our model when one of the virtual avatars kisses the other in Second Life, an event is triggered. The event is decoded by our system to send haptic based kiss to the real user via the Bluetooth-enabled neck piece hardware. Some of the potential applications of the proposed approach includes distant lover's communication, remote child caring, and stress recovery.
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ICME2
2011 Gesture recognition on a mobile device for remote event generation
abstract
This paper describes an application created on an Android mobile phone that recognizes gestures using the smartphone's orientation sensor. These gestures can be used to trigger events in another program running on a remote computer. We present a prototype application that generates different events for controlling a PowerPoint presentation, namely starting and stopping, as well as displaying the next and previous slide.
Eric Torunski, Abdulmotaleb El Saddik, Emil M. Petriu
ICME2
2011 Photo search in a personal photo diary by drawing face position with people tagging
abstract
In recent years, people tend to maintain personal photos in digital spaces not only to share their experiences with social friends but also to jog their own memory. Therefore, an effective solution is crucial to the growth of the needs for recording one's daily life. In this study, we have developed a complete system for personal photo diary system, namely MePTroy. With a friendly user interface, users can easily maintain personal episodes and memories with photos. In addition, we also support a flexible method for photo search based on the position of facial appearance that enables users to access episodes quickly. By integrating face detection and recognition technologies, as well as a friendly UI, MePTory offers diverse functionalities to annotate and search photos.
Heung-Nam Kim, Abdulmotaleb El Saddik, Kee-Sung Lee, Yeon-Ho Lee
IUI2
2011 Folksonomy-boosted social media search and ranking
abstract
With the rapid proliferation of social media services, users on the social Web are overwhelmed by the huge amount of social media available. In this paper, we look into the potential of social tagging in social media services to help users in retrieving social media. By leveraging social tagging, we propose a new personalized search method to enhance not only retrieval accuracy but also retrieval coverage. Our approach first determines the similarities between resources and between tags. Thereafter, we build two models: a user-tag relation model that reflects how a certain user has assigned tags similar to a given tag and a tag-item relation model that captures how a certain tag has been tagged to resources similar to a given resource. We then seamlessly map the tags on the items depending on a particular user's query in order to find the most attractive media content relevant to the user needs. The experimental evaluations have shown the proposed method achieves better search results than state-of-the art algorithms in terms of accuracy and coverage.
Majdi Rawashdeh, Heung-Nam Kim, Abdulmotaleb El Saddik
ICMR3
2011 Serious games
abstract
No abstract available.
Abdulmotaleb El Saddik
ACM Multimedia1
2011 Transparent non-intrusive multimodal biometric system for video conference using the fusion of face and ear recognition
abstract
Mono-modal biometric systems face many limitations such as noisy data, intra-class variations, distinctiveness, spoof attacks, non-universality, and unacceptable error rates. Working on enhancing the performance of a mono-modal biometric system may not be highly efficient and effective. A multimodal biometric system combines two or more biometric features into a single identification system. It aims to improve several of the mono-modal biometric systems drawbacks and improve the recognition coverage and performance. In this paper, a transparent non-intrusive multimodal biometric system based on the fusion of the face and ear biometrics is proposed to identify individuals during a video conference environment with minimal explicit user involvement and hassle. The results of the experiment show that the performance of the transparent non-intrusive multimodal biometric system of the face and ear is higher than that of the mono-modal face or ear.
Abbas Javadtalab, Laith Abbadi, Mona Omidyeganeh, Shervin Shirmohammadi, Carlisle M. Adams, Abdulmotaleb El Saddik
PST6
2011 Personalized PageRank vectors for tag recommendations: inside FolkRank
abstract
This paper looks inside FolkRank, one of the well-known folksonomy-based algorithms, to present its fundamental properties and promising possibilities for improving performance in tag recommendations. Moreover, we introduce a new way to compute a differential approach in FolkRank by representing it as a linear combination of the personalized PageRank vectors. By the linear combination, we present FolkRank's probabilistic interpretation that grasps how FolkRank works on a folksonomy graph in terms of the random surfer model. We also propose new FolkRank-like methods for tag recommendations to efficiently compute tags' rankings and thus reduce expensive computational cost of FolkRank. We show that the FolkRank approaches are feasible to recommend tags in real-time scenarios as well. The experimental evaluations show that the proposed methods provide fast tag recommendations with reasonable quality, as compared to FolkRank. Additionally, we discuss the diversity of the top n tags recommended by FolkRank and its variants.
Heung-Nam Kim, Abdulmotaleb El Saddik
RecSys2
2011 Leveraging Collaborative Filtering to Tag-Based Personalized Search
Heung-Nam Kim, Majdi Rawashdeh, Abdulmotaleb El Saddik
UMAP3
2011 Collaborative error-reflected models for cold-start recommender systems
Heung-Nam Kim, Abdulmotaleb El Saddik
Decis. Support Syst.2
2011 Collaborative user modeling for enhanced content filtering in recommender systems
Heung-Nam Kim, Inay Ha, Kee-Sung Lee, Abdulmotaleb El Saddik
Decis. Support Syst.5
2011 Collaborative user modeling with user-generated tags for social recommender systems
Heung-Nam Kim, Abdulmajeed Alkhaldi, Abdulmotaleb El Saddik
Expert Syst. Appl.3
2011 Effective multimedia surveillance using a human-centric approach
Pradeep K. Atrey, Abdulmotaleb El Saddik, Mohan Kankanhalli
Multim. Tools Appl.2
2011 Modeling and assessing quality of information in multisensor multimedia monitoring systems
abstract
Current sensor-based monitoring systems use multiple sensors in order to identify high-level information based on the events that take place in the monitored environment. This information is obtained through low-level processing of sensory media streams, which are usually noisy and imprecise, leading to many undesired consequences such as false alarms, service interruptions, and often violation of privacy. Therefore, we need a mechanism to compute the quality of sensor-driven information that would help a user or a system in making an informed decision and improve the automated monitoring process. In this article, we propose a model to characterize such quality of information in a multisensor multimedia monitoring system in terms of certainty, accuracy/confidence and timeliness. Our model adopts a multimodal fusion approach to obtain the target information and dynamically compute these attributes based on the observations of the participating sensors. We consider the environment context, the agreement/disagreement among the sensors, and their prior confidence in the fusion process in determining the information of interest. The proposed method is demonstrated by developing and deploying a real-time monitoring system in a simulated smart environment. The effectiveness and suitability of the method has been demonstrated by dynamically assessing the value of the three quality attributes with respect to the detection and identification of human presence in the environment.
M. Anwar Hossain 0001, Pradeep K. Atrey, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2011 Introduction to ACM multimedia 2010 best paper candidates
abstract
No abstract available.
Shervin Shirmohammadi, Jiebo Luo 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2010 Location Aware Question Answering Based Product Searching in Mobile Handheld Devices
abstract
In this research we present a question answering based searching technique for location based shopping. A user may ask a question to the system as they naturally ask to a human while the system retrieves the search results by analyzing the given question and the current GPS location. Based on the retrieved results, the system carry out conversation with the user to explicitly understand his/her needs and accordingly filters search results for display. The conversation between the system and the user is based on word co-occurrence keyword extraction and Artificial Intelligence Markup Language (AIML) technique. As per initial experimentation, we found out that the proposed approach of conversation based shopping is appealing and useful to the user.
S. K. Alamgir Hossain, Abu Saleh Md. Mahfujur Rahman, Thomas T. Tran, Abdulmotaleb El Saddik
DS-RT4
2010 Augmented Rendering of Makeup Features in a Smart Interactive Mirror System for Decision Support in Cosmetic Products Selection
abstract
We propose a smart mirror system to display an augmented 3D representation of the user with makeup features. In this approach the user is able to view the possible outcomes of different makeup applications in the smart mirror without affecting the real face appearance in the process. The system incorporates 3D face construction, IR based face tracking and OpenGL material extensive rendering approach to deliver the augmented made-up face. We argue that by viewing the augmented grooming features the users will be able to flexibly decide the makeup products of their choice.
Abu Saleh Md. Mahfujur Rahman, Thomas T. Tran, S. K. Alamgir Hossain, Abdulmotaleb El Saddik
DS-RT4
2010 Feature selection in haptic-based handwritten signatures using rough sets
abstract
This paper explores the use of rough set theory for feature selection in high dimensional haptic-based handwritten signatures (exploited for user identification). Two rough set-based methods for feature selection are analyzed, the first is a greedy approach while the second relies on genetic algorithms to find minimal subsets of attributes. Also, to further reduce the haptic feature space while maximizing user identification accuracy, a method is proposed where feature vectors are subsampled prior to the feature selection procedure. Rough set-generated minimal subsets are initially exploited to determine the importance of different haptic data types (e.g. force, position, torque and orientation) in discriminating between different users. In addition, a comparison between rough set-based methods and classical machine learning techniques in the selection of minimal information-preserving subsets of features in high dimensional haptic datasets, is provided. The criteria for comparison are the length of the selected subsets of features and their corresponding discrimination power. Support Vector Machine classifiers are used to evaluate the accuracy of the selected minimal feature vectors. The results demonstrated that the combination of rough set and genetic algorithm techniques can outperform well-established machine learning methods in the selection of minimal subsets of features present in haptic-based handwritten signatures.
Nizar Sakr, Fawaz A. Alsulaiman, Julio J. Valdés, Abdulmotaleb El Saddik, Nicolas D. Georganas
FUZZ-IEEE4
2010 Touch me interaction paradigm for physically browsing personal learning spaces
abstract
The prevalent visions of ambient intelligence leverage natural interaction between user and available services in a learning space. In this pursuit, we present a framework for augmenting physical objects with annotated information in order to improve physical browsing. The proposed system incorporates an intuitive camera based annotation scheme to catalog and author information about the physical learning objects in an environment. Each annotated data is encoded in such a way that they provide access points for available information or services about the physical objects. The system uses adequate visual cues and provides tactile feedback in order to leverage the touch-based interactions with the objects. These real world interaction techniques with the physical environment and seamless virtual learning information acquisition make the system transparent from the young learners and help them to become engaged in their learning activities.
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ICME2
2010 A real-time privacy-sensitive data hiding approach based on chaos cryptography
abstract
A multimedia surveillance system aims to provide security and safety of people in a monitored space. However, due to the nature of surveillance, privacy-sensitive information, such as face, gait and other physical parameters based on the captured media from multiple sensors, can be revealed without the concern of the people. This is a major concern in recent days. Therefore, it is desirable to have such mechanism that can hide privacy-sensitive information as much as possible, yet supporting effective surveillance tasks. In this paper, we propose a chaos cryptography based data hiding approach that can be applied on selected regions of interest (ROIs) in video camera footage, which contains privacy-sensitive data. Our approach also supports multiple levels of abstraction of data hiding depending on the role of the authorized user. In order to evaluate the suitability of this approach, we applied our algorithm on some video camera footage and observed that our approach is computationally efficient and applicable for real-time video surveillance tasks.
Sk. Md. Mizanur Rahman, M. Anwar Hossain 0001, Hussein T. Mouftah, Abdulmotaleb El Saddik, Eiji Okamoto
ICME4
2010 Bridging the Gap between Virtual and Real World by Bringing an Interpersonal Haptic Communication System in Second Life
abstract
The sense of touch has much importance in technology-mediated human emotion communication and interaction. Many researchers around the world are aiming to leverage the sense of touch in the communication medium between multi-user 3D virtual world and real environment. Driven by the motivation, we explored the possibilities of integrating haptic interactions with Linden Lab’s multi-user online virtual world, Second Life. We enhanced the open source Second Life viewer client in order to facilitate the communications of emotional feedbacks such as human touch, encouraging pat and comforting hug to the participating users through real-world haptic stimulation. These emotional feedbacks that are fundamental to physical and emotional development in turn can enhance the users interactive and immersive experiences with the virtual social communities in the Second Life. In this paper, we describe the development of a prototype that realizes the aforementioned virtual-real communication through a haptic-jacket system. Some of the potential applications of the proposed approach includes distant lover’s communication, remote child caring, and stress recovery.
Abu Saleh Md. Mahfujur Rahman, S. K. Alamgir Hossain, Abdulmotaleb El Saddik
ISM3
2010 Deducing user's fatigue from haptic data
abstract
Undesired physical fatigue reduces the overall Quality of Experience (QoE) of virtual reality haptics applications. Detecting fatigue is the first step in rectifying this problem. Fatigue in usability analysis is usually detected through conducting questionnaires and observations. This paper introduces an objective indirect discovery of user's fatigue through analyzing data of a haptic writing application. Our results show that if users are feeling tired their kinetic energy would decrease. We can compute this kinetic energy from the velocity of the arm movement during the usage of the haptic device.
Abdelwahab Hamam, Nicolas D. Georganas, Fawaz A. Alsulaiman, Abdulmotaleb El Saddik
ACM Multimedia4
2010 Adding haptic feature to YouTube
abstract
In this paper, we present a web-based framework in which users can annotate tactile feeling to a YouTube video and experience the tactile feeling by wearing a tactile device while watching\annotating the video. The tactile device is embedded into a wearable garment, a haptic jacket and a haptic arm band in this paper, and has a rectangular layout like a video screen. Therefore, the tactile information is represented as a sequence of rectangular arrays with time stamps and stored in XML format. Each element of the array represents a tactile intensity, a magnitude of actuation. In the framework we provide a web-based authoring tool to add tactile feeling while navigating a video and setting tactile intensity in a time line. We also introduce a web browser in which a tactile device driver is embedded to activate the tactile device based on the annotated tactile information.
Mohamed Abdur Rahman 0001, Abdulmajeed Alkhaldi, Jongeun Cha, Abdulmotaleb El Saddik
ACM Multimedia4
2010 Multimodal fusion for multimedia analysis: a survey
Pradeep K. Atrey, M. Anwar Hossain 0001, Abdulmotaleb El Saddik, Mohan Kankanhalli
Multim. Syst.3
2010 Spatial-geometric approach to physical mobile interaction based on accelerometer and IR sensory data fusion
abstract
Interaction with the physical environment using mobile phones has become increasingly desirable and feasible. Nowadays mobile phones are being used to control different devices and access information/services related to those devices. To facilitate such interaction, devices are usually marked with RFID tags or visual markers, which are read by a mobile phone equipped with an integrated RFID reader or camera to fetch related information about those objects and initiate further actions. This article contributes in this domain of mobile physical interaction; however, using a spatial-geometric approach for interacting with indoor physical objects and artifacts instead of RFID based solutions. Using this approach, a mobile phone can point from a distance to an annotated object or a spatial subregion of that object for the purpose of interaction. The pointing direction and location is determined based on the fusion of IR camera and accelerometer data, where the IR cameras are used to calculate the 3D position of the mobile phone users and the accelerometer in the phone provides its tilting and orientation information. The annotation of objects and their subregions with which the mobile phone interacts is performed by specifying their geometric coordinates and associating related information or services with them. We perform experiment in a technology-augmented smart space and show the applicability and potential of the proposed approach.
Abu Saleh Md. Mahfujur Rahman, M. Anwar Hossain 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2010 Introduction to the best papers of ACM multimedia 2009
abstract
No abstract available.
Changsheng Xu, Eckehard G. Steinbach, Abdulmotaleb El Saddik, Michelle X. Zhou
ACM Trans. Multim. Comput. Commun. Appl.3
2009 Feature selection and classification in genetic programming: Application to haptic-based biometric data
abstract
In this paper, a study is conducted in order to explore the use of genetic programming, in particular gene expression programming (GEP), in finding analytic functions that can behave as classifiers in high-dimensional haptic feature spaces. More importantly, the determined explicit functions are used in discovering minimal knowledge-preserving subsets of features from very high dimensional haptic datasets, thus acting as general dimensionality reducers. This approach is applied to the haptic-based biometrics problem; namely, in user identity verification. GEP models are initially generated using the original haptic biometric datatset, which is imbalanced in terms of the number of representative instances of each class. This procedure was repeated while considering an under-sampled (balanced) version of the datasets. The results demonstrated that for all datasets, whether imbalanced or under-sampled, a certain number (on average) of perfect classification models were determined. In addition, using GEP, great feature reduction was achieved as the generated analytic functions (classifiers) exploited only a small fraction of the available features.
Fawaz A. Alsulaiman, Nizar Sakr, Julio J. Valdés, Abdulmotaleb El Saddik, Nicolas D. Georganas
CISDA4
2009 Magic stick: A tangible interface for the edutainment of young children
abstract
Recently, there has been a high demand for developing tools that promote education through learning. We introduce our edutainment tool called Magic Stick that helps children learn about new objects by providing their names associated by visual representations regarding these objects. Children's parents or teachers can pick the entities they would like their children to learn about by simply attaching RFID tags to these entities. Afterwards, they can customize the type of information and visualizations related to these entities through the use of a friendly GUI designed for this purpose. In our study with young children, we found that the Magic Stick created an entertaining atmosphere among children and greatly engaged them in learning.
Ali Karime, M. Anwar Hossain 0001, Wail Gueaieb, Abdulmotaleb El Saddik
ICME4
2009 A Framework to bridge social network and body sensor network: An e-Health perspective
abstract
Body sensor networks (BSN) can capture physical phenomena from a human body, contextual information from the environment and high level events of a person. Associating contextual information and events with the captured raw sensory data can serve as a crucial input for many applications such as e- Health. For example, to accurately and timely monitor an elderly person with several physical disabilities while he is at home or outdoors, the context and event information along with raw sensory data needs to be reached to an e-Health service provider to assist in taking time critical decision. Such process includes receiving the sensory data, analyzing it to trigger necessary services such as sending an alert message to the family physician, hospital, emergency service, his immediate caregiver, family members, friends and so on. A BSN also allows members of one's community of interest, referred to as a social network, to query real-time sensory, contextual and event data. Combining the social network with BSN is envisioned to enhance the current state of the art in e-Health applications. In this paper, we propose a framework, called SenseFace, that can dynamically pass sensory data from one's BSN to his/her social network and vice versa. Finally, we illustrate the design and implementation of the framework.
Mohamed Abdur Rahman 0001, Mohammed F. Alhamid, Abdulmotaleb El Saddik, Wail Gueaieb
ICME3
2009 HugMe: synchronous haptic teleconferencing
abstract
Traditional teleconferencing systems have enabled remote communications via audiovisual modalities. However, in real life, human touch such as encouraging pat plays a fundamental role to physical and emotional communication between persons. This paper presents a synchronous haptic teleconferencing system with touch interaction to convey affection and intimacy. We present a preliminary prototype called HugMe. In this system, two remote users could see as well as touch each other.
Jongeun Cha, Mohamad A. Eid, Ahmad Barghout, Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ACM Multimedia5
2009 ACM 2009 workshop on ambient media computing (AMC'09) overview
abstract
No abstract available.
Howard Leung, Cha Zhang, Qing Li 0001, Rynson W. H. Lau, Benjamin W. Wah, Abdulmotaleb El Saddik, K. Selçuk Candan, Irene Cheng 0001
ACM Multimedia6
2009 Motion-path based gesture interaction with smart home services
abstract
In this paper, we propose a motion-path based gesture recognition technique and show its application in a smart home environment. Users hand gestures are recognized by capturing the motion-path while they draw different symbols in the air. In order to capture the motion-path, we use infra-red camera's IR sensing capability. The IR camera tracks the infra-red emitter attached to the user's hand gloves and produces a sequence of motion-points, which are then analyzed syntactically to recognize the intended hand gesture. The recognized gesture is used to interact with the intelligent environment for accessing various services. Toggling a lamp switch, changing the light intensity, and playing/pausing a movie are few examples where we have integrated the gesture-based interaction. Our experiment shows that the proposed gesture recognition technique is robust and its use in the smart home environment is interesting and appealing to the people.
Abu Saleh Md. Mahfujur Rahman, M. Anwar Hossain 0001, Jorge Parra, Abdulmotaleb El Saddik
ACM Multimedia4
2009 Guest Editorial: Distributed Simulation, Virtual Environments and Real-time Applications
Abdulmotaleb El Saddik, Mirela Sechi Moretti Annoni Notare
Concurr. Comput. Pract. Exp.1
2009 A biologically inspired framework for multimedia service management in a ubiquitous environment
abstract
Abstract This paper addresses several key issues in distributed multimedia services management and composition such as scalability, heterogeneity, and quality of service (QoS). The proposed framework introduces biologically inspired multimedia service management through the composition of basic multimedia services such as streaming services and different transcoding services. The biologically inspired approach is used for collecting the QoS requirements from individual transcoding services in order to select the most suitable services for the desired composition process. A prototype of the proposed framework is designed, implemented, and evaluated in terms of scalability and load balancing. Copyright © 2009 John Wiley & Sons, Ltd.
M. Shamim Hossain, Atif Alamri, Abdulmotaleb El Saddik
Concurr. Comput. Pract. Exp.3
2009 A framework for human-centered provisioning of ambient media services
M. Anwar Hossain 0001, Jorge Parra, Pradeep K. Atrey, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2009 Improving robustness of P2P applications in mobile environments
Andrew Roczniak, Abdulmotaleb El Saddik
Peer-to-Peer Netw. Appl.2
2009 Touchable 3D video system
abstract
Multimedia technologies are reaching the limits of providing audio-visual media that viewers consume passively. An important factor, which will ultimately enhance the user's experience in terms of impressiveness and immersion, is interaction. Among daily life interactions, haptic interaction plays a prominent role in enhancing the quality of experience of users, and in promoting physical and emotional development. Therefore, a critical step in multimedia research is expected to bring the sense of touch, or haptics, into multimedia systems and applications. This article proposes a touchable 3D video system where viewers can actively touch a video scene through a force-feedback device, and presents the underlying technologies in three functional components: (1) contents generation, (2) contents transmission, and (3) viewing and interaction. First of all, we introduce a depth image-based haptic representation (DIBHR) method that adds haptic and heightmap images, in addition to the traditional depth image-based representation (DIBR), to encode the haptic surface properties of the video media. In this representation, the haptic image contains the stiffness, static friction, and dynamic friction, whereas the heightmap image contains roughness of the video contents. Based on this representation method, we discuss how to generate synthetic and natural (real) video media through a 3D modeling tool and a depth camera, respectively. Next, we introduce a transmission mechanism based on the MPEG-4 framework where new MPEG-4 BIFS nodes are designed to describe the haptic scene. Finally, a haptic rendering algorithm to compute the interaction force between the scene and the viewer is described. As a result, the performance of the haptic rendering algorithm is evaluated in terms of computational time and smooth contact force. It operates marginally within a 1 kHz update rate that is required to provide stable interaction force and provide smoother contact force with the depth image that has high frequency geometrical noise using a median filter.
Jongeun Cha, Mohamad A. Eid, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2008 Evaluating ALPHAN with Multi-user Collaboration
abstract
In our previous work we introduced a novel application layer protocol, named ALPHAN, for haptic data communication. In this paper, we present a thorough evaluation of the protocol using a multi-user collaborative haptic application. The benchmark application consists of a simple game where three users attempt to lift a 3D triangular shape and place it in a triangular hole. The performance metrics and the test bed of the protocol evaluation are also discussed. It is found that a delay of 150 ms or higher caused the participating users not even to feel the existence of each other. Also a comparison between the two users and three users scenarios is considered. Finally, we comment on our findings and provide directions for prospective research.
Hussein Al Osman, Mohamad A. Eid, Abdulmotaleb El Saddik
DS-RT3
2008 Design of Application-Specific Incentives in P2P Networks
abstract
A rational P2P node may decide not to provide a particular resource or to provide it with degraded quality. If nodes are very likely to behave this way, or if the failure of an P2P-based application is associated with serious consequences, then usage of P2P infrastructure to support this application becomes questionable. Timely and extensive research works address this issue. These are presented and discussed in a broader survey of available strategies, tools and techniques to analyze and mitigate effects of rational behavior. We are designing an incentive mechanism for a P2P-based sharing of multimedia files. Our contribution is to incorporate application-specific characteristics into the incentive mechanism in order to improve its performance. We illustrate this approach by analyzing a collaborative file sharing protocol when various incentive mechanisms are implemented. For each such mechanism, we consider the case when it uses the application-specific characteristic, and when it does not. In this publication, we report on our preliminary findings.
Andrew Roczniak, Abdulmotaleb El Saddik, Ross Kouhi
DS-RT2
2008 Communication Complexity Evaluation for Longest-Lived Directional Multicasting in WANETs
abstract
We consider the problem of maximizing the multicast lifetime in wireless ad hoc networks with directional antennas. By a simulation study, we have discovered that the existing distributed algorithm for such optimization problem may generate considerable number of control messages to build up a longest-lived multicast tree in large-scale networks. This may prohibit them from being used directly in ad hoc networks with limited energy and bandwidth. In this paper, we would like to investigate some mechanisms to improve the communication complexity of the distributed algorithm. We explore some important properties of this optimization problem from a graph theory perspective and derive several localized operations that are especially beneficial to the resource-constrained (e.g. limited energy, memory, and computation capabilities) wireless ad hoc networks. The localized operations have low complexity for both memory and computation requirements at each node. Our simulation results show that the proposed localized operations would allow our distributed algorithms to achieve an expected linear communication complexity.
Song Guo 0001, Abdulmotaleb El Saddik
ICC3
2008 A Message Complexity Oriented Design of Distributed Algorithm for Long-Lived Multicasting in Wireless Sensor Networks
abstract
We consider an optimization problem in wireless sensor networks (WSNs) that is to find a multicast tree rooted at the source node and including all the destination nodes such that the lifetime of the tree is maximized. While a recently proposed distributed algorithm for this problem guarantees to obtain optimal solutions, we show that its high message complexity may prevent such contribution from being practically used in resource-constrained WSNs. In this paper, we proposed a new distributed suboptimal algorithm that achieves a good balance on the algorithm-optimality and message complexity. In particular, we prove that it has a linear-message complexity. The tradeoff between algorithm sub-optimality and message complexity is also studied by simulations.
Song Guo 0001, Abdulmotaleb El Saddik
ICCCN3
2008 Automatic scheduling of CCTV camera views using a human-centric approach
abstract
In large scale surveillance systems, a number of CCTV cameras are installed in distributed premises and are connected to a central control station, where human operators observe the different camera views for identifying a probable security breach. In such situations, it is particularly difficult for the operator to pay attention to all camera views. Studies have shown that a human operator can effectively monitor only four camera views at a time. This paper attempts to solve the problem of dynamically selecting and scheduling the four best CCTV views. We adopt a human-centric approach in which the system computes the operatorpsilas attention in the CCTV views to automatically determine the importance of events captured by the respective cameras. The experiments show that the proposed method helps a human operator in identifying important events occuring in the environment.
Pradeep K. Atrey, M. Anwar Hossain 0001, Abdulmotaleb El Saddik
ICME3
2008 Context-aware QoI computation in multi-sensor systems
abstract
Multi-sensor systems are increasingly being deployed in many application scenarios due to the enormous potential they can offer. However, as the processing of sensory data often results in imprecise outcome, measuring the quality of information (QoI) in these systems has become an important issue. The measurement of QoI is usually performed by processing the elementary data provided by the heterogeneous sensors, which is also influenced by the techniques involved in sensor management. However, the effect of context, such as environmental geometry, sensor placement, orientation, time, and other parameters in computing QoI has not yet been explored extensively in the literature. This paper proposes a context evolution model and studies its impact in the QoI computation. In particular, we show that the dynamic context information can be utilized to manage a multi-sensor system to improve its QoI.
M. Anwar Hossain 0001, Pradeep K. Atrey, Abdulmotaleb El Saddik
MASS3
2008 A glimpse of multimedia ambient intelligence
abstract
Ambient Intelligence (AmI) characterizes a new paradigm for novel interactions between a person and his/her everyday environment. This will impose major challenges on multimedia research to support human-centered information, communication, services and entertainment everyday and everywhere.
Rosa Iglesias, Abdulmotaleb El Saddik
ACM Multimedia2
2008 Haptics technologies: theory and applications from a multimedia perspective
abstract
The desire for natural intuitive modes of interactions with digital media has led to development of multimodal interfaces that aim to engage the users through a confluence of modalities such as audio, video etc. Haptic systems that enable touch based interactions with digital environments are a recent addition to multimodal systems and have widespread applications; there is a need to develop a sound design, development and evaluation strategy to leverage the availability of the haptic modality. This tutorial aims to provide an initial impetus in the direction of enabling multimedia researchers to conduct research in the area of haptic user interfaces. The tutorial will present the audience with an introduction to the field of haptics. The presented material will be made accessible to the multimedia community, relating material from the haptics domain to multimedia algorithms, systems etc. At the completion of the tutorial, the students will be able to analyze and apply algorithms and strategies for design, development, and evaluation of touch based user interfaces.
Abdulmotaleb El Saddik, Jongeun Cha, Kanav Kahol
ACM Multimedia1
2008 A biologically inspired multimedia content repurposing system in heterogeneous environments
M. Shamim Hossain, Abdulmotaleb El Saddik
Multim. Syst.2
2008 C-HAVE: Collaborative Haptic Audio Visual Environments and Systems
Abdulmotaleb El Saddik
Multim. Syst.1
2008 ACM/Springer Mobile Networks and Applications (MONET)
Abdulmotaleb El Saddik, Klaus Moessner, K. Selçuk Candan, Ben Liang 0001, Jiangchuan Liu
Mob. Networks Appl.1
2008 Gain-based Selection of Ambient Media Services in Pervasive Environments
M. Anwar Hossain 0001, Pradeep K. Atrey, Abdulmotaleb El Saddik
Mob. Networks Appl.3
2008 PECOLE: P2P multimedia collaborative environment
Abdulmotaleb El Saddik, Abu Saleh Md. Mahfujur Rahman, Souhail Abdala, Bogdan Solomon
Multim. Tools Appl.1
2008 Touching beyond audio and video
Abdulmotaleb El Saddik, Shervin Shirmohammadi
Multim. Tools Appl.1
2008 Experiments in haptic-based authentication of humans
Mauricio Orozco Trujillo, Matthew Graydon, Shervin Shirmohammadi, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2008 Experiments in haptic-based authentication of humans
Mauricio Orozco Trujillo, Matthew Graydon, Shervin Shirmohammadi, Abdulmotaleb El Saddik
Multim. Tools Appl.4
2008 Confidence Evolution in Multimedia Systems
abstract
Multimedia systems utilize multiple media streams, each of which have different confidence levels in accomplishing various detection tasks. For example, in a multimedia surveillance system, one would usually have higher confidence in an audio stream compared to a video stream for detecting human shouting events. The pre-computation of these confidence levels is cumbersome especially when new media streams are dynamically added to the system. This paper proposes a novel method, which dynamically computes the confidence levels of new streams based on the past history of their agreement/disagreement with the already trusted streams. To demonstrate the utility of the proposed method, we provide the experimental results for detecting events in a multimedia surveillance scenario.
Pradeep K. Atrey, Abdulmotaleb El Saddik
IEEE Trans. Multim.2
2007 Collision Detection and Force Response in Highly-Detailed Point-Based Hapto-Visual Virtual Environments
abstract
In this paper, we present a collision detection algorithm and a force response algorithm both for use in dynamic, rigid-bodied, highly-detailed, hapto-visual virtual environments in which the models' geometry is point-based. Our collision detection algorithm partitions the virtual space into a modified octree in a preprocessing step. At runtime, collision detection involves querying the octree for the octant where the end-effector currently is, as well as the indices of neighboring octants. After the world space is narrowed down to a volume of interest, the algorithm checks to see if the end-effector falls inside any axes-aligned bounding box that is centered at the model surface points that reside in the aforementioned volume of interest. A collision is defined as the haptic endeffector being found inside an axes-aligned bounding box centered at a model surface point. After a collision is detected, the force response algorithm calculates a force vector that starts at the end-effector's current position and ends at the model surface point closest to the end-effector. This is an adaptation of the common god/proxy-object approach but for use in point-based models.
Naim R. El-Far, Nicolas D. Georganas, Abdulmotaleb El Saddik
DS-RT3
2007 SimSITE: The HLA/RTI Based Emergency Preparedness and Response Training Simulation
abstract
Many wake-up calls have been received for emergency response, due to natural disasters such as hurricanes, fires or man-made incidents. The emergency responders need to work in a coordinated, well-planned manner to best mitigate the impact of an emergency incident. Simulation systems provide a wider range of training at a much lower expense for emergency preparedness and response. It is identified as the only feasible approach when it is difficult to emulate real-life experiments. The presented research demonstrates the emergency training simulation for the SITE building at the University of Ottawa, where the real-time interaction and collaboration are achieved over HLA/RTI, the IEEE standard for distributed simulation and modeling.
Abdulmotaleb El Saddik, Azzedine Boukerche, Nicolas D. Georganas
DS-RT3
2007 Design and Implementation of Haptic Tele-mentoring over the Internet
abstract
Haptic Tele-mentoring refers to an educational technique in which the mentor can teach the mentee, in a hand-by-hand manner over communication networks, through the coupling of two haptic devices. Essentially, the realization of tele-mentoring relies on the efficient transmission of haptic information, such that either end of the network can sense and/or impart forces. A few of obstacles in developing the tele-mentoring applications over the internet are network delay, jitter, and packet loss. These impairments potentially affect the stability of the tele-mentoring system and degrade the feeling of guiding. This paper analyzes the design and implementation constraints of the system, as well as the simulation and experimental results. To compensate for the network latency, a novel approach based on the behaviors of the human arm trajectory is proposed to lower the overshoot so as to improve the overall system stability. The experimental results show the effectiveness of the anti-overshoot algorithm.
Jilin Zhou, Abdulmotaleb El Saddik, Nicolas D. Georganas
DS-RT3
2007 A Haptic Enabled UML Case Tool
abstract
This paper describes a haptic enabled UML CASE tool that enables software engineering developers to physically manipulate and touch the UML modeling elements and feel the force feedback. We propose an architecture and a software design for the tool. The current implementation of the tool uses the Omni Phantom device, a quite common haptic interface among the haptic research community. The CASE tool, from a user perspective, consists of three parts: a drawing area, a palette, and a tool bar. Our preliminary usability study demonstrated the potential of adding the haptic modality to UML development tools.
Atif Alamri, Mohamad A. Eid, Abdulmotaleb El Saddik
ICME3
2007 Confidence Building Among Correlated Streams in Multimedia Surveillance Systems
Pradeep K. Atrey, Mohan Kankanhalli, Abdulmotaleb El Saddik
MMM (2)3
2007 Tools for transparent synchronous collaborative environments
Abdulmotaleb El Saddik, Nicolas D. Georganas
Multim. Tools Appl.1
2006 Haptic Applications Meta-Language
abstract
A wide range of haptic devices exist that possess the potential to offer users a rich experience in a virtual reality environment. This however depends on the haptic device to be used. Surely, removing the burden of users having to 'adjust to' operating haptic devices is welcomed. In this paper, we propose the creation of the Haptic Applications Meta Language - hereon HAML - which is an XML-based language created to describe Haptic-enabled frameworks to a high degree of detail. The envisioned goal of HAML is to allow for the creation of plug-and-play environments in which a wide array of supported haptic devices can be used in a multitude of virtual environments, with the compatibility issues being handled by automated engines instead of programmatically by the user. Therefore, we introduce the HAML framework and discuss its tentative structure, proof-of-concept implementation, and avenues for future work.
Fayez R. El-Far, Mohamad A. Eid, Mauricio Orozco Trujillo, Abdulmotaleb El Saddik
DS-RT4
2006 Design of Distributed Collaborative Application through Service Aggregation
abstract
The Service Oriented Architecture (SOA) is used to support loosely-coupled integration of existing applications. We are investigating the possibility of creating entirely new applications based on SOA. We present a case study for a collaborative authoring application targeting groups of around five users collaborating over the Internet. We highlight the basic requirements of the application and show how these can be fulfilled by utilizing certain services accessed through standard HTTP, Jabber and JXTA set of protocols, and off-the-shelf techniques. By measuring the performance of the application in a heterogeneous environment and by providing details of an alternate service fulfilling the application's requirement, we show that SOA-based collaborative applications can be quickly designed and deployed.
Andrew Roczniak, Jamil Melhem, Pierre Lévy, Abdulmotaleb El Saddik
DS-RT4
2006 Secured MPEG-21 Digital Item Adaptation for H.264 Video
abstract
Seamless adaptation and transcoding techniques to adapt the digital content have achieved significant focus to serve the consumers with the desired content in a feasible way. With the succession of time we sense that secured adaptation should also be taken care of for not only serving sensitive digital contents but also to offer security as an embedded feature of the adaptation practice to ensure digital right management and confidentiality. In this paper, we propose an encryption framework for a transcoder while adapting H.264 video conforming to MPEG-21 DIA. Encryption mechanism is applied on the adapted video content thus reducing computational overhead compared to that on the original content
Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik
ICME3
2006 MeTaMaF: Metadata Tagging and Mapping Framework for Managing Multimedia Content
abstract
Metadata comes into forefront as a savior of multimedia search and management. However, the existence of the diverse set of metadata standards and the different vocabularies used by these standards has made that task especially challenging in recent days. In this paper, we propose a framework for managing multimedia content by effectively dealing with the heterogeneous multimedia metadata in a transparent fashion. The cornerstone of our approach is to leverage the existing metadata standards and provide schema and data level mapping among them. The proposed framework also allows adding new tags/vocabularies with the given metadata vocabularies, defining equivalency relationship among those vocabularies, and separately managing the metadata vocabularies outside of the actual media files. We have developed a prototype of the framework and evaluated its performance in terms of some basic query execution times
M. Anwar Hossain 0001, Md. Anisur Rahman, Iluju Kiringa, Abdulmotaleb El Saddik
ISM4
2006 Compressed-Domain Encryption of Adapted H.264 Video
abstract
Commercial service providers and secret services yearn to employ the available environment for conveyance of their data in a secured way. In order to encrypt or to ensure personalized security of the video contents in an intermediary node, it is necessary to have the content structure conforming to an international standard. Moreover, pressure to satisfy user preferences and device requirements seamlessly are raising the need for content to be customized providing the best possible experience. In this paper, we present perceptual encryption scheme for video encryption that is incorporated with a dynamic temporal adaptation technique of the H.264 video conforming ISO/IEC MPEG-21 Digital Item Adaptation. Encryption is performed on demand directly from the adapted bitstream and its generic Bitstream Syntax Description (gBSD).
Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik
ISM3
2006 Eye & Why: A Prototype for Learning Objects Visualization in Virtual Environment
abstract
In this paper, we have introduced a 3D Car gaming metaphor, an interactive interface for searching learning objects from distributed search repositories. By analyzing the semantic information of the search results, the metaphor creates groups and presents them as virtual roads, where each rendered traffic sign corresponds to the visual representation of a learning object. The road network based learning object visualization scheme assists the learner in realizing the relationships (if one exists) among the displayed results. Hence, the learner would be able to discard or explore a group of information very easily from the rendered visual relationship constructs, thus resulting in an effective information navigation process. Additionally, the game motivated car metaphor employs customized and interactive 3D visual schemes to make the learning and navigational process more entertaining.
Abu Saleh Md. Mahfujur Rahman, Abdulmotaleb El Saddik
ISM2
2006 Security Considerations for SOA-Based Multimedia Applications
abstract
Growing levels of digitalization and broadband access drives extremely fast progress in multimedia and networking technologies and allows consumers to create requirements at an accelerating rate. Producers response is to emphasize the speed of delivery and upgradability of applications. Development of Web services and service-oriented architecture puts the emphasis on creation of specific services which then can be aggregated to achieve a particular goal. By substituting or adding new services, a particular application can be adapted faster to changing requirements. Any application thus composed must address security concerns, including security guarantees after a service substitution. Based on our framework for creating multimedia collaborative authoring applications from a set of standard services, we provide a security analysis of the resulting application and present a novel framework for ensuring uniform access control guarantees across different constituent services
Andrew Roczniak, Alexandre Miège, Abdulmotaleb El Saddik
ISM3
2006 User-credential based role mapping in multi-domain environment
abstract
No abstract available.
Ajith Kamath, Ramiro Liscano, Abdulmotaleb El Saddik
PST3
2005 JADE: jabber-based authoring in distributed environments
abstract
We present our initial results in developing a framework for collaborative multimedia authoring tools. This research is motivated by the lack of tools that take into account consumers' quality of experience. By mapping factors that have an impact on the quality of experience into requirements, we are developing a framework for tools that allow retrieval and manipulation of multimedia objects, and collaborative authoring of multimedia documents based on Jabber set of protocols.
Andrew Roczniak, Abdulmotaleb El Saddik
ACM Multimedia2
2005 Impact of incentive mechanisms on quality of experience
abstract
Since entities participating in P2P networks are usually autonomous and therefore free to decide on their level of participation, mechanisms to resolve conflicts between individual and collective rationality are needed. How can implementations of such mechanisms be compared? This paper introduces a qualitative reference framework, highlighting essential elements and major design decisions in any implementation of incentive mechanisms. In the context of multimedia applications built on top of P2P architectures, the reference framework can be used in assessing the impact on the quality of experience (QoE) when incentive mechanisms are included.
Andrew Roczniak, Abdulmotaleb El Saddik
ACM Multimedia2
2005 Haptic: the new biometrics-embedded media to recognizing and quantifying human patterns
abstract
Authentication for the purposes of security has taken giant strides since the introduction of Biometrics to help identify people by their behavioral and physiological features. From organizations and corporations to educational institutes, electronic resources, and even crime scenes, Biometrics offers a wide application scope to detect fraud attempts. This paper proposes a research path to achieve the task of authenticating users that are working in a haptic-based environment. The field of Biometrics can be divided into two main classes of human features. Birth-given characteristics like fingerprints and facial features cannot be developed or altered by humans. Behavioral characteristics such as hand signature and voice fall into the second class [1]. The work presented in this paper pursues the latter class and specifically studies how a person reacts to using daily devices or tools. The fact that we can exploit people's habits in handling devices to detect identity was the hypothesis that motivated this work.
Mauricio Orozco Trujillo, Ismail Shakra, Abdulmotaleb El Saddik
ACM Multimedia3
2004 LORNAV: A Demo of a Virtual Reality Tool for Navigation and Authoring of Learning Object Repositories
abstract
Navigation in 3D world has reached its pinnacle with the advent of several technologies like JAVA, XML, WEB SERVICES, VRML, X3D etc. A lot of efforts have been given to visualize and navigate in virtual mall, cities, digital libraries etc. Most of them visualize only static objects. We designed a Virtual Reality (VR) Tool, called Learning Object Repository Navigation and Authoring in Virtual environment (LORNAV) that extracts Learning Object Metadata (LOM) dynamically from repositories and creates 3D representation of these objects. The proposed tool provides several facilities such as personalized navigation allowing a user to view content that is of interest to him. It also helps users to create new aggregated Learning Objects (LOs) from existing ones in a 3D Environment to be presented in SMIL format.
Mohamed Abdur Rahman 0001, M. Anwar Hossain 0001, Abdulmotaleb El Saddik
DS-RT3
2004 Architecture and Evaluation of Tele-Haptic Environments
abstract
A collaborative, haptic, audio and visual environment (C-HAVE) consists of a network of nodes. Each node in the C-HAVE world contributes to the shared environment with some virtual objects. These can be static, e.g., a sculpture or the ground, or dynamic, e.g., an object that can be virtually manipulated. We aim at developing a heterogeneous scalable architecture for large collaborative haptics environments where a number of potential users participate with different kinds of haptic devices. The main objective of the presented research is the development of three prototypes to demonstrate quantitatively the effects of adding haptics to a task. The experimental results reveal the effects of the different implementations on the performance and time delay of a particular task through objective measurement results.
Jilin Zhou, Abdulmotaleb El Saddik, Nicolas D. Georganas
DS-RT3
2004 A QoS-Based Framework for Distributed Content Adaptation
abstract
The tremendous growth of the Internet has introduced a number of interoperability problems for distributed multimedia applications. These problems are related to the heterogeneity of client devices, network connectivity, content formats, and user's preferences. The purpose of this paper is to present a framework for transcoding multimedia streams. The proposed infrastructure takes into consideration the profile of communicating devices, network connectivity, exchanged content format, context description, and available customization services to find a chain of services that could be applied to adapt the content to the required needed format. Part of the framework is a QoS-based selection algorithm that finds the best sequence of adaptation services which can maximize users' satisfaction with the delivered content.
Khalil El-Khatib, Gregor von Bochmann, Abdulmotaleb El Saddik
QSHINE3
2003 Peer-to-Peer Suitability for Collaborative Multiplayer Games
abstract
Peer-to-peer communication is emerging as one of the most potentially disruptive technologies in the networking sector. If the interest in such technologies as Napster, Morpheus and Gnutella is any indication, peer-to-peer networks will be a major component in the future of computer communications and the Internet. In this work, we conducted a case study by implementing the Xiangqi game using a P2P networking technology. In so doing, we gained a deeper understanding of P2P communication in general and a more in-depth knowledge of the specific technology used. We encountered a number of issues while implementing our multimedia application using the JXTA P2P framework, but we were nevertheless able to create a working prototype that functioned satisfactorily. We thus concluded that P2P was a feasible alternative to the current server-based multi-player gaming paradigm.
Abdulmotaleb El Saddik, Andre Dufour
DS-RT1
2003 Topic Introduction
Abdulmotaleb El Saddik, László Böszörményi
Euro-Par2
2003 Peer-to-Peer Communication through the Design and Implementation of Xiangqi
Abdulmotaleb El Saddik, Andre Dufour
Euro-Par1
2003 JASMINE: A Java Tool for Multimedia Collaboration on the Internet
Shervin Shirmohammadi, Abdulmotaleb El Saddik, Nicolas D. Georganas, Ralf Steinmetz
Multim. Tools Appl.2
2003 RST-invariant digital image watermarking based on log-polar mapping and phase correlation
abstract
Based on log-polar mapping (LPM) and phase correlation, the paper presents a novel digital image watermarking scheme that is invariant to rotation, scaling, and translation (RST). We embed a watermark in the LPMs of the Fourier magnitude spectrum of an original image, and use the phase correlation between the LPM of the original image and the LPM of the watermarked image to calculate the displacement of watermark positions in the LPM domain. The scheme preserves the image quality by avoiding computing the inverse log-polar mapping (ILPM), and produces smaller correlation coefficients for unwatermarked images by using phase correlation to avoid exhaustive search. The evaluations demonstrate that the scheme is invariant to rotation and translation, invariant to scaling when the scale is in a reasonable range, and very robust to JPEG compression.
Dong Zheng 0002, Jiying Zhao, Abdulmotaleb El Saddik
IEEE Trans. Circuits Syst. Video Technol.3
2001 Keep It Small and Smart
abstract
The production of interactive multimedia content is in most cases an expensive task in terms of time and cost. It must hence be the goal to optimize the production by exploiting the reusability of interactive multimedia elements. Reusability can be triggered by a combination of reusable multimedia components, together with the appropriate use of metadata to control the components as well as their combination. In this article, we discuss reusability and adaptability aspects of interactive multimedia content in web-based learning systems. In contrast to existing approaches, we extend a component-based architecture to build up interactive multimedia visualization units by the use of metadata for reusability and customizability issues.
Abdulmotaleb El Saddik, Ralf Steinmetz
AICCSA1
2001 Perceived Consistency
abstract
Quality of service guarantees for multimedia communication systems have been considered on several abstraction levels. In the multimedia networking field it is typical to identify the minimal QoS requirements of an application to save resources by guaranteeing its functionality. Many of these applications can operate in spite of an imperfect delivery of media data, while other applications such as distributed databases or distributed file systems consider perfect QoS necessary but accept delay. The basic problems of the latter is the consistency of their data, while the former require a consistent perception of the content. More generically, both QoS requirements can be interpreted as a problem of maintaining a consistent system state. Consequently we assume that many distributed applications, including most distributed multimedia applications, can fulfil their tasks in spite of imperfect consistency. Since the application requirements differ widely, the elements that make up "consistency" must be separated and classified. This paper introduces Consistency QoS and proposes a classification of elements that determine an application's consistency requirements. The low level QoS requirements that these separate parameters rely on are shown, and example parameter sets for application classes are given.
Carsten Griwodz, Michael Liepert, Abdulmotaleb El Saddik, Giwon On, Michael Zink, Ralf Steinmetz
AICCSA3
2001 Reusability and adaptability of interactive resources in Web-based educational systems
abstract
The production of interactive multimedia content is in most cases an expensive task in terms of time and cost. Hence, optimizing production by exploiting the reusability of interactive multimedia elements is mandatory. Reusability can be triggered by a combination of resuable multimedia components and the appropriate use of metadata to control the components as well as their combination. In this article, we discuss the reusability aspects of interactive multimedia content in web-based learning systems. In contrast to existing approaches, we extend a component-based architecture to build interactive multimedia visualization units with the use of metadata for reusability and customizability. In the three-tier model, the lowest layer of the paradigm corresponds to the programmer (code reusability). The user interface of an educational visualization is located at the top layer where the interaction with the end-user (student) takes place. The educator is located between the top and the bottom layers. This medium layer allows adapting interactive multimedia content according to the needs of the user, applying a predefined set of metadata. The teacher can both adjust the level of explanation and the level of interactivity of an animation, and influence the presentation and the results of the algorithms being illustrated (program reusability). After a theoretical overview, we explain our architecture by giving an example of an application.
Abdulmotaleb El Saddik, Stephan Fischer 0001, Ralf Steinmetz
ACM J. Educ. Resour. Comput.1
2001 Web-based multimedia tools for sharing educational resources
abstract
Many educational resources and objects have been developed as Java applets or applications, which can accessed by simply downloading them from various repositories. It is often necessary to share these resources in real time, for instance when an instructor teaches remote students how to use a certain resource explains the theory behind it. We have developed some tools for this purpose that emulate a virtual classroom, and are primarily designed for synchronous sharing of resources. They enable participants to share Java objects in real time and also allow the instructor to dynamically manage the telebearing session.
Shervin Shirmohammadi, Abdulmotaleb El Saddik, Nicolas D. Georganas, Ralf Steinmetz
ACM J. Educ. Resour. Comput.2
2000 Multibook's test environment
abstract
Well engineered Web based courseware and exercises provide flexibility and added value to the students, which goes beyond the traditional text book or CD-ROM based courses. The Multibook project explores the boundaries of customized learning materials by composing learning trails dynamically as learners have set their profile to access a course. In this paper we first give an overview of the core project ideas and illustrate them along our Software Engineering course. Then we present a novel extension to the project's exercise environment with a graph editing component that particularly fits the needs of structure-related assignments.
Nathalie Poerwantoro, Abdulmotaleb El Saddik, Bernd J. Krämer, Ralf Steinmetz
ICSE2